AI General Thread

Started by Legend, Dec 05, 2022, 04:35 AM

Legend and 2 Guests are viewing this topic.

the-pi-guy

Quote from: Legend on Jul 21, 2026, 08:21 PMTokens too expensive for you too?  :-[
Nah I just use the free stuff. 

Just haven't felt like it's been as needed lately. 

the-pi-guy

Me: "I got a 90% reduction in file size."

Copilot: "Oh that's really great. It's shaving off a lot of unnecessary data."

Me: "Oh I was looking at the wrong line, it was actually an 80% reduction."

Copilot: "Yeah that makes more sense. It would have been a red flag if it was a 90% reduction."


Legend

Quote from: the-Pi-guy on Jul 22, 2026, 02:55 PMMe: "I got a 90% reduction in file size."

Copilot: "Oh that's really great. It's shaving off a lot of unnecessary data."

Me: "Oh I was looking at the wrong line, it was actually an 80% reduction."

Copilot: "Yeah that makes more sense. It would have been a red flag if it was a 90% reduction."


I hate these types of convo quirks in ai so much lol.

There is zero reason to trust anything they say cause they will happily say the opposite ten seconds later!

Legend

OK this AI world on instagram is just too much.

Every ukulele girl has the same background.
 

These are just two examples. And the top one I know is a real girl cause I have previously watched her on youtube: https://www.youtube.com/watch?v=Vg1rl65u19Y

But is this a real video of her in this room? Who is the actual owner of this room!?!?!?!?

the-pi-guy

Quote from: Legend on Jul 25, 2026, 06:14 PMOK this AI world on instagram is just too much.

Every ukulele girl has the same background.
 

These are just two examples. And the top one I know is a real girl cause I have previously watched her on youtube: https://www.youtube.com/watch?v=Vg1rl65u19Y

But is this a real video of her in this room? Who is the actual owner of this room!?!?!?!?
That's just Green Screen Gary's room. 

the-pi-guy

Apparently there is an all in one model MiniMax-H3. Does video, text, image, audio apparently.  

Apparently it's incredible, even uncensored and all that jazz.


And apparently it's against the license agreement to use it in the US/Europe....  

Legend

Quote from: the-Pi-guy on Aug 03, 2026, 08:27 PMApparently there is an all in one model MiniMax-H3. Does video, text, image, audio apparently. 

Apparently it's incredible, even uncensored and all that jazz.


And apparently it's against the license agreement to use it in the US/Europe.... 
How big is it?

the-pi-guy

Quote from: Legend on Aug 03, 2026, 09:53 PMHow big is it?

I stole someone's answer here:
Quoteint8convrot models are 34gb, bf16 models are 66.3gb.

there's also pruned variants that are compacted to 21gb.

the text encoder: bf16 at 51.5gb, int8convrot at 27.1gb, nvfp4_awq at 15.7gb

the vae's are also massive, 605mb for audio and 5.21gb for video.

Legend

Quote from: the-Pi-guy on Aug 03, 2026, 09:55 PMI stole someone's answer here:
Wow that's pretty small!!!

Well, compared to kimi and deepseek lol.

What's the biggest one you've downloaded yourself?

the-pi-guy

Quote from: Legend on Aug 03, 2026, 11:21 PMWow that's pretty small!!!

Well, compared to kimi and deepseek lol.

What's the biggest one you've downloaded yourself?
Like ~30 GB.  

Legend

Quote from: the-Pi-guy on Aug 04, 2026, 12:55 AMLike ~30 GB. 

Not bad!

the-pi-guy

#506
Quote from: Legend on Aug 04, 2026, 09:23 AM
Not bad!

It seems to do wildly better with scene transitions than any other local model I've seen.

Legend

Google is having issues. Deepmind CEO just stepped down to chair, 3.5 pro is missing in action.

the-pi-guy

#508
I was able to get the model running on my computer. 

It's quite a bit harder to run.
With LTX-2.3 I was making 10 seconds videos in about 10 minutes, at 720p.

Here I'm making 7 second videos in like 13 minutes at 480p. 

You still see the funky AI stuff. Like I generated a video of someone going under a blanket, and the blanket very obviously phased through them.

I don't think LTX-2.3 was as capable of getting completely different views of the same scene. It could do smooth turning the camera around for sure. But I don't think it could swap back and forth on two characters. I had it generate in a single go, where it swaps back and forth between two characters twice, and it does really well. 

I also got my pipeline working for making extended videos. And I had it basically make 3 7 second videos that could be stitched together as a seamless video. Although I didn't listen to the audio back and forth to see how close that was. I just had it make more from the image.

There is a reference pipeline that would probably make it even better.

Legend

Quote from: the-Pi-guy on Today at 03:00 AMI was able to get the model running on my computer. 

It's quite a bit harder to run.
With LTX-2.3 I was making 10 seconds videos in about 10 minutes, at 720p.

Here I'm making 7 second videos in like 13 minutes at 480p. 

You still see the funky AI stuff. Like I generated a video of someone going under a blanket, and the blanket very obviously phased through them.

I don't think LTX-2.3 was as capable of getting completely different views of the same scene. It could do smooth turning the camera around for sure. But I don't think it could swap back and forth on two characters. I had it generate in a single go, where it swaps back and forth between two characters twice, and it does really well. 

I also got my pipeline working for making extended videos. And I had it basically make 3 7 second videos that could be stitched together as a seamless video. Although I didn't listen to the audio back and forth to see how close that was. I just had it make more from the image.

There is a reference pipeline that would probably make it even better.
Mind sharing your best vs worst results so far?

13 minutes for 7 seconds isn't bad if most generations are useful. Could let it render overnight