Index · Speaking

Intro to AI Art — Lesson 4: Tools & Workflows (Part 2)

Speaker 0 As we enter the second quarter of the twenty first century, human civilization is experiencing an accelerated pace of evolution. A technological and socioeconomic singularity is ahead of us, and it's driven by artificial intelligence and blockchain technology. NPEAK equips entrepreneurs and business professionals with practical knowledge to keep them ahead of the curve in the exponential age. Our live mentoring sessions, on demand training, and exclusive networking opportunities keep you at the cutting edge of Web three, NFTs, the metaverse, decentralized finance, automation, and so much more. Meet top industry leaders during our live mentoring sessions to ask your questions directly or simply follow the recorded sessions in your own time. And PEEP, inclusive, inspired, in the know, in this together.

Ben Helps if I'm off mute. Good morning, everyone. Alright. So let me share the presentation for today. Here we go. Alright. So we are on our fourth lesson in AI art, and I think this one today is going to be a fun one because we're gonna deal with some more advanced topics, some more cutting edge experimental topics. And so, you know, maybe slightly less practical and a little more experimental, but I think it's really gonna be fun. Although you could argue that every single thing in AI are just kind of experimental and new at the moment. So as they said, head on over to the polls so that we can look through those in a few minutes and see what's going on. Let's see. What's our first slide? Alright. So this is normally where I talk a little bit about myself. I'll give you the really short version.

If you wanna hear lots about me, you can go back to the first or second lesson. I'm a business executive. I've I've run a few different, companies, over the years. Been in blockchain for over a decade and then been working with AI, and art for about the last nine to twelve months now. But I have some art background going back twenty, thirty years. So that's about enough about me. I always put up an interesting AI photo before we get started every time. This one here is mid journey version five. We have in the top left Harry Potter as imagined by Tim Burton. We have in the bottom left, Gollum from Lord of the Rings played by Danny DeVito. And on the right, we have The Wizard of Oz mixed with The Shining. So what I really enjoyed about this is just how far AI has come in the last several months alone.

These are really stunning and relatively simple prompts that were put together. So alright. Let's see. Let's dive in. So what are we talking about today? So we're gonna go over some new developments in mid journey that have happened in the last two weeks, which is pretty crazy. We're gonna talk a little bit more. We didn't really get to it last time. We sort of ran out of time. Stable diffusion, we're gonna talk a little bit about prompt waiting and blending. We talked about how to wait prompts in mid journey last time. We didn't really get to it in stable diffusion other than to talk a little bit about negative prompting. And then we're gonna talk a little bit about prompt blending and stable diffusion, which is very similar to multi prompting in mid journey, which we dealt with last time.

We're also gonna go over control net in a little more detail. We talked about it back in lesson two as, hey. This is a new thing, and it's really cool. But we're gonna dive into a little bit more. We're gonna talk a little bit about how you can use ChatGPT for prompt engineering. We're gonna talk about fine tuning stable diffusion. We've talked a little bit about how there's custom models and checkpoints. We're gonna dive into that a little bit deeper, talk about textual inversion, LoRa modeling, and then we're gonna talk a little bit about how you can train your own model. And then we'll touch on a couple other things, but that's sort of the basics of what we're gonna go over in the next hour. So what's next? So we started out with a very broad intro to AIR. What is it?

Where does it come from? Then we went into what are the basic tools, and we talked about the big three diffusion models that are popular right now. We went over DALL E, mid journey and stable diffusion. Lesson three, we really dove into how do you use midjourney and stable diffusion at a very practical level. And then today, we're gonna go into some more advanced topics. And then next time, we'll be talking about GAN and how is GAN still used even though it's a little bit older technology than the latent diffusion models that we're talking about right now. It still has a lot of uses, and so we're gonna talk about that next time in a couple of weeks.

Speaker 2 Alright.

Ben So mid journey, new developments. So this has all happened in the last two weeks. We're gonna look at permutation prompts and what they are. There's a new described command. There's a new repeat parameter. There's a new version of the sort of custom, mid journey model Niji. So there's now a version five of that, and there is a mood board command, that's in alpha, that's coming up. I've I've got some, info about it. I don't have alpha access yet. I've asked them for it, but I can show you what's on the horizon. So let's dive right into it. What's a permutation prompt? So a permutation prompt allows you to quickly generate versions of a prompt with a single imagine command. However, you can use curly braces and commas to get multiple versions of the same prompt. What do I mean? Here's an example.

Imagine a bird, but normally, I would have to say a red bird and a green bird and a yellow bird. Well, now I can say, give me a red, green, yellow bird, and it will create and process all three jobs separately. Let's dive in and look at what that looks like. So here's imagine a naturalist illustration of and then here's where our permutation prompt is. A pineapple, a blueberry. I can't read what that is. A banana bird. And so it now creates some processes for different jobs with four different image grids, and you can see those on the screen. And you can do more, with this permutation, prompt than just that. You can also use it, to give you different variations, for instance, and, aspect ratios. So here we have, you know, a three two or one one or one two aspect ratio of a natural illustration of a fruit salad bird.

You can also use it to give you different models. So here we have the same thing, but our permutation is use version five, Niji, and the test model, and and it will do this. Big thing to note about this is that this is consuming more of your fast hours. So if you're on a limited plan and you're worried about burning through your hours, this will burn through your hours quickly, but it'll also save you some time on the front end. Before we dive into the next mid journey thing, let's look at the polls real quick and see see where we're at. So our first poll question is, did you attend the previous sessions? 40% of people said they were at session 20% were at session 30% were in session 10% said what sessions. If you said what sessions, I will say today will be interesting, will be fun, but you're probably missing some of the basics.

You're going, why are we talking about permutation prompts? You know, I just wanna make an image. So I would encourage you to go back to some of the earlier sessions for some more context. Have you ever created a piece of AI art? 60% to 40% said, yes. They have. So that's pretty cool. And then have you ever used the following to create an AI output? Mid journey and stable diffusion are currently in the lead with about a third each, Dolly, twenty percent one in five, and none of the above one in five. So that is our polls. Alright. So back back into things. So permutation prompts. I think everybody probably understands kinda how they work. It just lets you run multiple variations on a prompt at the same time without having to sit and enter in each variation as a separate command.

So next, we're gonna talk about the describe command. So what does the describe command do? It does the opposite of what an imagine command do. It's probably the best way to describe the describe command. So you would use the describe command with an uploaded image, and it will give you back four text prompts that try and describe the image in the way, you know, mid term would process this. So here's a good here's an example. And Nick Saint Pierre on Twitter did a great write up on this. I would recommend following him. He's got a lot of good stuff on AI. But let's look and see how this worked because what he did was he took an image, said describe it. It gave him the description, and then he put that back in as a prompt to see how close it would be.

So here on the left, he input the photo of the man on the left and asked it to describe it. Midjourney described it as a father reading two daughter for a home advertisement, Adobe, and the style of soft atmospheric scenes, etcetera. He put that in, and it came out with the image on the right. It's not an exact copy, but it's pretty close. It definitely gets the feel of it. The one on the right, same thing. He asked it to describe the image of the woman standing in the street on the left, and it created with that description being re input into Midjourney, it created the image on the right. So pretty powerful, certainly more accurate. There were some sort of open source web browser based versions of this that I've used in the past. They were not accurate enough that I wanted to put them in the references or recommend them.

They existed, but this seems to be much more accurate. Let's see. There's also now a repeat parameter. Say it's a parameter, not a command because you append it to the end of a command. So for instance, if you want mid journey to imagine cats, but instead of wanting one four by four grid, you want lots of cats, you would add dash dash repeat, and if you put five, it would create five two by two grids of cats. Now if you were to up the chaos level, which we talked about last time, this is going to give you exponentially more variation much quicker. You're still gonna burn through, you know, processing time and hours. But instead of saying, imagine cats, imagine cats, imagine cats, imagine cats, imagine cats five times, you're gonna say it once, and it's gonna give you the same number of outputs.

And if you were to append that chaos command, it's now gonna give you a bigger variety of outputs. So this is just speeding up your ideating, generally, on the front end of the process. Alright. So there's also a new model out. We talked a little bit last time about the Niji model, which is specifically tuned on anime aesthetics. They have advanced that model just like the main model, and I will show some examples of that here. So this is a pack of lions in the style of Tintoretto's The Last Supper. Happens to be one of my favorite paintings of all time. And so I wanted to see what mid journey would do version five with Tintoretto's The Last Supper with lions. I want it in the style of The Last Supper. It was a little more literal with it than I intended, but that's fine.

You can see, especially, like, in the top right, very nice looking image, nice composition. Nothing looks too strange. So then I input the same thing into Nidji, the exact same prompt, the pack of lines and the style of Tintoretto's The Last Supper. Remember that Nidji is trained on anime aesthetics, and so it's gonna vary from Tintoretto a bit. But we can see Nidji version four on the left. It has sort of a video game anime aesthetic, but it is lacking some detail. There's pixel art down on the bottom left. You look at Neji version five on the right, much more detailed, more coherent compositions, less weird artifacts. So just like mid journey version five is a significant upgrade over version four, Nidji version five appears to be a significant upgrade over version four. I haven't had a ton of time to play with it and test with test it out, but looks like it's a significant advancement.

Alright. So this is the alpha feature that I have not been able to access yet, but I have seen. So there's a mood board feature that is coming down the pike in mid journey. And what it allows you to do is it allows you to reference an entire collection of images in your prop. So instead of you know, we talked before that image prompting is limited in the amount of images you can put in, and and it's sort of a pain in the butt to get them all in to Discord. And then vice versa, blending images lacks some of the same parameters. Even though it's easier to get the images in there, you don't have as many parameters. So what this does is this lets you upload a bunch of images and then use that for prompting. So let me let me show you an example.

On the left, we have maybe 15 or 20 images, all sort of Gucci inspired images. And then with the prompt, the man standing by a horse, it outputs a very consistent style and look. On the right, we can see this is that Gucci mood board plus a man riding a black stallion through the country. Very coherent, very consistent styling. I think this is gonna be pretty powerful and pretty cool once they open it up to everyone. Alright. So that's kinda what's new in mid journey in the last two weeks, believe it or not. Next, we're gonna look at something that we talked about back in lesson two. So back in lesson two, so that's only been, what, maybe a month ago, we talked about Automancer, which is a platform for creating PFP collections, had developed a new AI generator fermenting PFP collections, large numbered PFP collections.

So you could say a portrait of a robot dog. It would give you four different base images. You would tell it to generate artwork, in this case, say, 3,500 NFTs based on one of those base images. In this case, the person picked the one on the far left, sort sort of the robotic mech robot dog, and then it would output a grid of those with different traits and characteristics. So they just had the launch of their first collection officially now. So four or five weeks after I talked about it coming up in beta, there's now a collection you can mount. It's the withdrawals. Full disclosure, I'm a member of the EV Mavericks DAO that put this out. One of them is named the Ben Medals after me. All the proceeds go to charity. It's almost minted out. It may already be minted out by the time we are having this discussion.

But pretty amazing that, I see the new poll. So there's a new poll, everyone. What are you using AI art for most? Creative projects for work, social posts, building a business. So make sure you do the poll. So, anyway, just shows the pace at which things move. Five weeks after this was announced in beta, it's still in beta, and there's already a collection you can mint out there. Alright. So let's talk about stable diffusion, some advanced prompting a little bit. So let's talk about weighted prompts and blending keywords. Last time, we we talked about unstable diffusion. It's easy to do negative prompting to say what you don't want, but waiting prompts is a little bit more difficult. And, mostly, we just ran out of time to go over it. So waiting prompts and stable diffusion. So you can adjust the keyword strength with parentheses and brackets and stable diffusion.

So you use the parentheses to increase the weight, and you can use the brackets to decrease the weight. So here's an example from I think it's automatic's automatic one one one one's WebView as feature showcase. So the original image is a photo of eggs and bacon on a frying pan, and we see that there's about equal amount of eggs and bacon on this frying pan. But then you could weight it. So a photo of eggs, right, weighted with parentheses and bacon. So now we've got the same number of eggs, but slightly less bacon. Now we have it's weighted even harder, a photo of eggs, and then we have it weighted really hard, a photo of eggs, and and there's almost no bacon on the picture. Likewise, you could do the same for the bacon and weight the bacon differently from the eggs.

Now like I mentioned just a moment ago, there's three different ways to weight prompts in stable diffusion, and that's sort of one of the themes of stable diffusion. I I like to say that it's the Linux of AI art generators. There's a million different ways to do things. And so you can weight things with brackets and parentheses. Right? And as you weight these, it actually assigns a numeric weight. So one parentheses would be a weight of 1.1, two would be a weight of 1.21, three would be a weight of 1.33, and it works the opposite way with brackets in terms of sort of negatively weighting. You can also so that's way number one, parentheses and brackets. Way number two is adjusting the keyword strength using a numerical weight. This is more like how Midjourney does it. So in this case, you enter the keyword, then a colon, and then the weight.

Relatively self explanatory. You can also, like we discussed, use negative prompting. And when you use negative prompting, you can also wait the negative prompts. So in this case, because these are negative prompts, we're actually using a positive weight on the negative prompts. So if I don't want ugly in there, I could put four parentheses and wait at, like, a 1.44. So I'm telling you I really don't want that. Anyway, you can you can also so three different ways of weighting your Prompts in stable diffusion. You can also blend your keywords, in stable diffusion. So here's a great example. Person one colon, person two colon, the amount of the blending. So we have on the left, this is Joe Biden, Donald Trump, with a weight of point one, so it's weighted much more towards Donald Trump.

In the middle, we have an even weighting, and on the right, we have it weighted much more towards Joe Biden. It looks much more like Joe Biden. So the last number ranges from a zero to a one, and it essentially tells us how much weight is given through how many sampling steps to which word. We can look at that a little more over here. Here's Emma Watson, Amber Heard with a weight of point eight five, 40 steps, and we can see how that blending would look. And you can even combine this and blend more than just two words. You could blend four different words. So here's a weighting between the actresses, Evan Rachel Wood, Jennifer Lawrence, Jennifer Aniston, and Jennifer Connolly, and you can see the image incorporates parts of all four of those actors.

So kinda interesting, little bit different than the way mid journey works, but something to be aware of when you're looking at more advanced prompting. So I wanna talk a little bit about ControlNet because this is sort of I think it's kinda mind blowing personally, what they've been able to do with ControlNet and stable diffusion. So ControlNet is a neural network structure that can control diffusion models by adding extra conditions. It works by modifying the text prompt embedding with another embedding derived from an additional input. So I'm gonna go from sort of the most technical definition, and we'll go down to something a little more palatable in a second. But what is that additional input? The additional input can be anything that provides more information or guidance for the image generation. So it could be an image, a sketch, a pose, a depth map, or even a scribble.

So why is that interesting? The revolutionary thing about ControlNet is its solution to the problem of spatial consistency. Whereas previously, there was simply no efficient way to tell an AI model which parts of an input image to keep. ControlNet changes this by introducing a method to enable stable diffusion models to use additional input conditions that tell the model exactly what to do. So you could inpaint, but you kind of lost everything that was being inpainted within that masked area. This is different because it maintains the spatial consistency of the entire image while letting you change things. So really short version, TLDR, this is pics to pics or image to image, but god mode. You can do so much more with this than than a traditional image to image. You've got so much more control. So here we're gonna take and so I did this practically because I love a good practical demonstration.

So I took an image that I had made for the MPEG gold getters. This was one of the images that didn't make the cut. It's actually, the basis for my first on chain piece, significantly simplified and changed. But I took this relatively raw image, and I said, let's run this through ControlNet, and let's see what happens. So here's ControlNet. If you go to the references page, there are a couple of demos of this that you can use and play with, but I used one of them online on Hugging Face. And I took this image, put it in, and my prompt was three people looking at metal pyramid. And so this is sort of the default mode for ControlNet. It's a canny edge model. It uses an edge detection algorithm, and that's what you see on the top right.

It essentially pulled out all the edges from the original image, and then it created that image down on the bottom right. Now you can you can play with some, you know, different options in terms of how many images, what resolution, and then thresholds and things just like always with stable diffusion. I left everything alone for the purposes of this, but we're gonna run through the different types of control net models real quickly. So this is sort of the default canny edge map model that we see right here. Next model, I do not know how to pronounce that word, but it's a line map, and you can see that it it you know the only thing it really saw was this very rudimentary top of the mountain and then these couple lines on the left, but it output this image on the right.

And it does have spatial consistency with the original image. May not look exactly like it, but it certainly has spatial consistency with those lines that it did derive from the image. So very cool. Next model is an HED map. So you can see this differs from the canny map, right, a little bit, and it's using a g d boundary detection, which is a different method of boundary detection than CANNY, and it came up with this image on the right. Once again, spatial consistency is there with the the original image. Next model is a scribble model. Now this didn't work as good because this isn't a scribble. Right? This is a relatively fully formed image, not a scribble, but it still detected, lines, and it put out this, once again, spatially consistent work on the right. Now as with all things AI, it doesn't always listen.

Instead of three people, we get one person, but still very interesting. There's also a fake scribble model. So recognizing that this is not really a scribble, this fake scribble model can take a fully formed image and sort of create a scribble off of it, which is what we did here. Now if you look so I ran this through a couple different stable diffusion models. Top right is stable diffusion 1.5, sort of the default model. Bottom right is using I think it's anything. It's a custom checkpoint, and, obviously, it gives us some more aesthetically pleasing look. Here's what's interesting. So if you look at the scribble model on the left, the fake scribble, it interpolated the space between the three mountain climbers as two smaller climbers, maybe children. And so in both of the images that resulted from this particular ControlNet mapping, we ended up with with five people.

There's two children in between the three larger mountain climbers. So kinda interesting to see, you the intricacies between the different mappings here. So I wanted to mix things up a little bit. This is a pose based model. Same exact input, except I said three robots staring at a metal pyramid. And what you'll see here is the only consistency the model is really going for is the posing. And so we have the three robots looking off in the distance in the same way. Our pyramid, our mountain is in the same spot, but it has a lot more artistic freedom outside of posing. All it's trying to do is maintain the skeletal posing of the figures in the original image. Very cool stuff. Then we also have segmentation mapping. So instead of mapping lines and boundaries, we're gonna map segments and areas.

So if you look on the left, we've sort of segmented into in the color blobs. And then on the right, same thing. We have the top right is stable diffusion standard 1.5 model. Down on the bottom right, the anything custom checkpoint. But you can see this gives us a very different sort of interpretation than, say, our fake scribble does. Then we have depth mapping. So the control net model is trying to map the depth of the image. What's in the foreground? What's in the background? And it's going to then reinterpret based on the depth. And it does a pretty good job. Top right, once again, is the default model. Bottom right is a custom model, but you can see it really gets the concept of depth here while still maintaining most of the edges. Here is a normal mapping, so it almost tries to do, like, a three d mapping.

It's a combination of edge detection, depth mapping, and you can see it really interpreted it pretty wildly on the top right, but also another form of mapping. So ControlNet's got huge amount of granularity, different options, but it's really wild. So I actually used ControlNet not too long ago and didn't tell anybody. So when I did the art for the goal getters, the top left was our final image that we came out with and then made variations based on that. Most of the variations that I made were very sort of manual variations. It was me in Photoshop or me using a glitch algorithm. However, three of them were made with ControlNet. And remember, this was a month or two ago, so ControlNet was newer. But if you look at the top right, bottom left, and bottom right, I tried to keep things pretty close.

They maintained a lot more consistency than some of the examples I just showed you, but it's pretty wild what it could do with a very simple command. So I love it. These are three of my favorites in the collection. I think ControlNet's really awesome, and, obviously, we're gonna see huge advances in that. Midjourney is gonna have something similar, but definitely a very powerful tool you should be aware of. And it's something that's pretty easy to use. You can go on a web page and use it. You can also install it with stable diffusion, but it's a lot more manual. Takes a little bit more time. If you wanna just do a one off or play around with it, go into the references, and there's a couple demos available. Alright. So next thing. I was reading on Twitter the other day, and this fellow, Anonymouse, said cameras and lenses for med journey.

I've compiled a list of some of the best professional cameras and lenses for various scenarios, and he goes through. And I know you probably can't see it because it's really tiny. It's a very, very long Twitter thread. But he goes through and says, you know, if you want a dark and moody ambiance, you should use this Sony camera with this lens, with these settings. And he gives this very exhaustive list, which is super helpful. We talked about a little bit last time sort of understanding different tables and different options for, you know, photographic AI photos or artistic ones. And somebody responds to this Twitter thread with, thanks for that. I'm creating tables in chat GPT that describe the key elements in a photo and then different options for each. And then he says, you can build your prompts using this. So he kinda says, hey.

You've done all this manual work, and I'm just gonna dump it all in the chat GPT and make it a lot easier. So here he's got a table in chat GPT doing something very similar. You know, here's a subject of a young woman with a camera angle, low angle, studio location, DSLR, and you can go through this and essentially do the same

Speaker 2 thing.

Ben And here's another one for landscape photography. Grand Canyon, sunset, foggy, soft, and we talked about doing this manually last time a little bit that, you know, you can look up these things and think, what do I want, where, who, how, what would the final result look like? So this is just automating that a little bit. And I've seen people use ChatGPT to write prompts before. I've even done it before, but it was much more open ended. It was, know, help me come up with a really cool, fantastical alien landscape or something. And it does a good job, But GPT doesn't always understand the constraints of good prompt engineering. Right? So maybe it'll give you a 100 word paragraph. Well, as we looked at last time, stable diffusion can only handle 70 some odd words. So this is what I found.

So you could and this isn't a chat GPT class, so we're gonna go through this pretty quickly. However, you could take a prompt blueprint. So in this case, image type, macro close-up, genre, emotion, scene. You So essentially write your prompt with good sort of prompt engineering grammar. Right? And so we've got a couple here, and we've tested them out, these prompts. Right? The one on the right is an aerial drone shot of a fantasy place. So then we take these two prompts knowing what outputs they give us and knowing that they're aesthetically pleasing, and then we put them in the ChatGPT. So on the left, we say ChatGPT create five prompts with random parameters by using the same construct as these two example prompts. And then ChatGPT comes up with the ones on the right, so it comes up with a macro close-up of a sci fi device or an aerial drone shot.

Now here's where things get interesting from an ideating standpoint. So we're looking at these links. You know, I I really love that macro close-up. So then you say, chat GPT, please provide me five more random prompts focusing on macro close ups only and using very bizarre and unusual scenes, and it then generates here's two of the examples on the right. So very interesting way we can use ChatGPT to help us ideate, but in a in a constructive, you know, sort of parameterized way. Alright. So let's talk a little bit about fine tuning stable diffusion. I think everybody will be kinda interested in this. Right? Because we've got we've talked about we've got the base model, which sometimes comes up with not the most aesthetically pleasing results. Right? It's fine tuned on thousands, millions of images, very broad based, but it can take a lot of work to kinda get what you want out of it.

But what if I want pictures of my dog? Let's say that's a dachshund on the left, and I really want him. I don't just want a random dachshund. I want one that looks like mine, a long haired brown dachshund in my photos. How do I do that? So here's an example where these have been used as input images to fine tune stable diffusion, then I can use those images to inform generations. So let's talk about how we can do that. So there's a few different options. There's Dreambooth or checkpoints. There's textual inversion or embeddings. There's LoRa models, and there's hyper networks. There's also aesthetic embeddings. But from what I've seen, aesthetic embeddings don't give us good even though it's called aesthetic, the aesthetics of aesthetic embeddings aren't as good as the other options, so we're not gonna deal with that. So what's the difference?

There's a really good twenty something long minute YouTube video on the difference. I've linked it in the references, but we'll go through very quickly so that we don't have to spend twenty five minutes talking about this, and we'll get to sort of the TLDR version. There's DreamBooth custom models. This is probably the most effective aesthetically as a whole. However, it's very storage inefficient. It uses an entirely new model. It's not compatible with other models by nature if you're using that model. That's all you can do is use that particular model, and it creates very large file sizes. I'm talking two, four, six gigabytes. You gotta have a lot of room for these custom models. Then you've got textual inversion. So your output is a very, very tiny embedding. It usually uses less photos to train it, very sort of narrow in terms of, you know, what you're looking for for an input and an output, but very, very tiny.

I'm talking 50, a 100 kilobytes for the file size. And you can use different textual inversion embeddings with different models, so there's a lot of flexibility. There's Lora's. Lora's are kind of an in between between a textual embedding and a and a full checkpoint. Their big advantage is they're very quick to train, and they kinda use some tricks that make them almost like a custom diffusion model and almost like a textual embedding, and we'll talk about that in a second. And then there's hyper networks. Short version on hyper networks is they're very similar to Alora, same size file that you get, but I would say they generally have worse results. They use a hyper network instead of some of the tricks that Laura uses with the diffusion model. So for all intents and purposes, at least right now today, it's really not worth using.

We'll focus on textual embeddings, Laura's full models, checkpoints. So speaking of, let's talk about kind of pros and cons. So a checkpoint, usually trained through DreamBooth, offers the best quality, but it also, like I said, has the largest file size. We're talking gigabytes, and it cannot be combined with other models, so it's very inflexible. Textual inversion or embedding has a lot of flexibility. You can use it with most custom models. I find that it produces better quality than a LoRa, but with a more limited scope. And we'll look at some examples of that in a minute. And like I said, smallest size, you're talking a 100 kilobytes, maybe. The size isn't even really worth talking about. And then you've got LoRa, so low rank adaptations. Also has a lot of flexibility, like textual embeddings, works with multiple different models.

It's faster to train than a textual inversion and requires less v RAM, which is one of the big advantages, and medium sized files, maybe a hundred, hundred and fifty megabytes. So much, much smaller than a full checkpoint, much, much larger than a textual inversion, but probably relatively, you know, not a big deal nowadays. Everybody's got plenty of storage. Albert, I see your question. Where can you find the references? They will be at the end, but I will also throw them in the chat here real quick. Give you the the short link. It's just bmeadows.xyz/npeak has the references and the PDFs of these presentations. Alright. So I wanna show an example of a LoRa and what you can do with it that's a little bit outside the box, and then we're gonna talk about how to train your own model, maybe with your face or your animal or but this is this is one of the really interesting things.

You know, we think about styles, custom models. We looked last time at some custom models that could make things look like Pixar or modern Disney movies or anime. But one of the things I don't think that we think about is so here's something interesting. So offset noise is essentially the ability to darken an image. And one of the one of the ways in which let me think about what the best way to say this. One of the disadvantages of the diffusion model that's used to create image Is is that it tends to create light images. The model has trouble with really dark or really contrasted images because of the way creating an image through diffusion works. And so you often end up with images that are simply too light or lack the contrast that a real photograph or artistic image would have.

But you can use a LoRa model that's been trained on this to essentially add darkness to an image and add contrast. So here we see on the left a picture of a corgi at night just coming out of a standard stable diffusion. If you look to the right a little bit, this has some noise offset. If you look all the way on the right, it has a much more significant noise offset using this LoRa model, and that looks much more like what a photograph would probably look like. Here's another example. On the left is sort of a basic level output. On the right is one that's been run through a LoRa with offset diffusion for darkness. There's references in there for how to use this and add this in. Here's another example. You often get photos of the night sky that look kinda blown out and unrealistic.

So here's this particular offset noise Laura applied to a night sky photo, and that looks way more accurate and aesthetically pleasing. Here's another one. This artwork kinda looks blown out on the left. It lacks the moodiness that the author intended, but here we've run that same prompt through an offset noise, Laura, and we get a much more dark sort of moody atmosphere, which is what the author is looking for. So training your own model. And I'm gonna stop and look at the poll question real quick. Let's see. So what are you using AI art foremost? 60% of respondents said for creative projects, 20% said for social media posts, 10% said for work, and 10% said for building a business. So let's see. Just checking to see if we had any questions. Looking at our chat. Alright. So training your own model.

So I really wanted to do this, and I ran out of time in the last couple weeks. It it only takes, I would say, maybe two to three hours. So maybe by next time, I'll have done it. But here's an example of a fellow. It's on the references, this article by Eric Richards. He trained his own model on his face at the very top, and he trained it a couple different ways through textual inversion and through Dreambooth to kinda see what the results would be. You can see on the bottom left, textual inversion, on the bottom right, Dreambooth. And he said that textual inversion consistently got his face correct more often. However, Dreambooth, while it didn't get his face right as often, when it did, it looked more aesthetically pleasing. There were more subtleties there, and it looked better. So I've included his article with his experiences, how he did it.

I've also included two articles. So one is how to train a custom model using Dreambooth. You can do that as long as you have a Google Drive with about nine, ten gigabytes free, about 10 Google Compute credits you'll have to use. And I would say you need six to 20 good images of yourself in about two to three hours. You can create your own custom DreamBooth model. I also included another article on how to create a custom textual inversion embedding on a local Mac computer. It's a little more technical. I kinda wanna do that now too, but I included a couple guides for both of those that are pretty recent so you could create your own and train your own model. Now I also something I love to do is always talk about new developments before we end because there's always something new coming out that's sort of mind blowing.

So today's new thing that I wanna look at is Meta, Facebook, Meta AI, came out with a project called SAM, segment anything by Meta, and I included the reference. It's segment-anything.com. And, essentially, what this does, the segment anything model, is a promptable segmentation system. So it can cut out any object in any image with a single click. You can track masks and videos, enable image editing apps, and even lift things out to three d. I'm gonna show you a couple examples, but here's here's a picture of some horses. You can click on any single horse, and it will accurately select the single horse, or you can tell it to select everything in the photo, and it will properly mask and select everything. Now what's really cool about this is this has been trained on about 11,000,000 images with over a billion different masks.

And because of that, you can you can do everything from I think I have a photo of it here. Yeah. Output mask can be used as inputs to other AI systems. You could take a mask of a chair and then use that as an input to a three d modeling AI system and come up with a three d model of that same chair. It also has the ability so it's learned a general notion of what objects are, and so you can actually use this masking on pictures it wasn't trained on and objects it hasn't seen before, and it's very accurate. So this is coming down. It's also something very cool that uses AI, and that I think we're gonna see a lot more of. Like I said, play around with it. They've got a demo. Really neat. Let's see. Do we have any questions?

As always, I left some quotes about art on the right as an AI as a picture of an AI piece that I'm still working on right now. But do we have any questions? Anybody wanna come on stage, ask anything, talk about anything?

Ben Alright. Well, I don't see any questions or comments. So let's talk briefly about next time. So next time, April 27, two weeks from today, we're gonna talk about GANs, generative adversarial networks, and how they're used for face restoration, upscaling, artistic, you know, generation of photos and art still. So we're gonna we're gonna deal with that next time. So Wu asked a question. He said, are there any ways to generate AI text and images? Not that I have seen yet. I have seen some sort of alpha and beta ideas, but nothing that's been released. It still is a big problem. Even if you negative prompt sometimes, you know, the generator not to include text and images, it will still include some gobbledygook. You know, I've seen I've seen you know, if images were used to train the AI that had a bunch of watermarks in it, I've seen it produce generated images that have fake watermarks, all sorts of weird things.

So, no, I I haven't seen anybody solve that problem yet. I think it's getting better, but it's not solved. And, well, apologies for continually pronouncing your name wrong. So I will I will try and get that right in the future. Alright. Well, if that is it, if anybody doesn't have any other questions, that's it for today. I'll see everybody in two weeks, hopefully, and, everybody have a great day. Thank you for coming.


All transcripts