Index · Speaking

Intro to AI Art — Lesson 3: Tools & Workflows (Part 1)

Speaker 1 As we enter the second quarter of the twenty first century, human civilization is experiencing an accelerated pace of evolution. A technological and socioeconomic singularity is ahead of us, and it's driven by artificial intelligence and blockchain technology. NPEAK equips entrepreneurs and business professionals with practical knowledge to keep them ahead of the curve in the exponential age. Our live mentoring sessions, on demand training, and exclusive networking opportunities keep you at the cutting edge of web three, NFTs, the metaverse, decentralized finance, automation, and so much more. Meet top industry leaders during our live mentoring sessions to ask your questions directly or simply follow the recorded sessions in your own time. And PEEP, inclusive, inspired, in the know, in this together.

Ben Good morning, everyone. So we are here for our third lesson on AI tools. So I would say, let me just dive into it. Let's see. Here we go. There's our PDF. So let's just dive into it. I would say at this point, hopefully, you know who I am. Our photo for today, is this photo that as the news described it, immaculate AI images of Pope Francis trick the masses. I actually saw this on Twitter and had to had to take a look at it and figure out if this was real or not. There are a few tells. If you look at, Pope Francis' hand down in the bottom left of the screen, you can tell it's a little bit misformed with the coffee there. There's also a couple of funny, things going on with the pectoral cross. Some of the lines aren't exactly right, but AI is getting really, really good.

This, was mid journey version five. Thought it'd be a fun, AI photo to look at today. You may also have seen some of the AI generated photos of Donald Trump being arrested, but I think this one's actually more realistic looking. So, anyway, scary how good this stuff's getting, but let's dive into making AI art. So what are we gonna go over today? I think this will probably be a little bit more succinct than our previous two sessions because we're really focusing on how to use mid journey and stable diffusion to make AI art. And let me I'm gonna jump forward and then jump backwards. So our first lesson, we really went over kind of conceptually what is AI art. You know? Where did it come from? What is it? Where is it going? What's some of the controversies around it?

Second lesson, we talked about what are people using in 2023 to actually make AI art. We looked at DALL E two, mid journey, and stable diffusion, but we really looked at them also kind of from a conceptual standpoint. How do you access them? How do you use them big picture? What do they cost? What are their differences? How are they similar? Today, we're gonna look at, you know if we looked at how do you use them, like, how do you access it? Do you do you access it through Discord or a website? Today, we're gonna look at, like, really nitty gritty, how do you construct a prompt and how do you make the art? Next time, we're gonna talk about some slightly more advanced topics. And then the final lesson, we're gonna talk about ways in which, again, is still used to edit AI art and in art creation.

So to go back to what are we gonna go over today, I I did change around the scope a little bit, as I kinda dove into things. So we're gonna talk about prompt engineering. Literally, how do you engineer a prompt, to get these models to create what you want? We're gonna talk about image prompting, you know, putting an image plus text in two images. We're gonna go over in painting, out painting. We're gonna talk a little bit about upscaling, mostly native resolutions, what resolution, is originally output versus how do you get to a resolution you want. We'll probably save a little bit of the upscaling talk for later on. And then we're gonna talk about different versions and custom models in mid journey and stable diffusion. I moved ControlNet to next week because it didn't really fit in with the mid journey and stable diffusion talk.

It's kind of its own thing right now. And then instead of making a piece live, I I originally wanted us to make a piece of art live. There have been some issues with mid journey being slow in the last week, And so I'm gonna move that a little bit, but some of the poll questions do relate to that. So make sure you do the poll questions. So alright. Let's move on. Actually, let's stop. Let's see. Let's stop and look at the polls real quick and see what we've got, and then we will move on. So let's see. Our questions are, have you ever used the following to create, an AI output? 45% said Midjourney, 36% said DALL E, only 9% have used stable diffusion, and about one in 10, now one in five has not. Have you ever created a piece of AI art?

Two thirds said, yes. They have. One third said, no. They have not. Did you attend the first or second session? Two thirds of people have been to one of the two sessions previous. About a third have not, and one out of, 10 said, what other sessions? I would say it is not necessary, to have attended the first or second session for this session. I would say the second lesson is probably beneficial for this one, but not necessary. And then we have our pick one. So pick one, number one, landscape person or an animal, landscape is winning. Photograph or a painting, photograph is winning, and urban is winning over rural. And we'll talk a little bit about that later, why those are there. So let's see. Alright. So prompt engineering. What is it? So I enjoy doing this.

I asked Google Bard what is prompt engineering because I now have access to Google's large language model, and it gave what I think is a pretty fair definition. It says prompt engineering is the art of crafting prompts that will produce the best possible results from a text to image generator. And so then I also asked it to give me some basic tips, and I kinda trimmed them a little bit. But this is generally what Google Bard and Bing AI gave me for basic tips in prompt engineering. Be specific. Instead of a city, say a large industrial city in the style of Monet. Use keywords to indicate style, mood, or subject matter. Don't be afraid to experiment, and be patient and persistent. And I'm gonna take a brief aside on the patient and

Speaker 3 persist

Speaker 3 Hi, everyone.

Speaker 2 GM,

Speaker 3 let's give him

Speaker 2 a seconds. Oh, hi, Fomi. I was just gonna say let's give him a few seconds. He's back now. It's probably something with his Internet.

Ben Sorry about that. I I there is something going on with our router and modem here, and I've I've been troubleshooting it for, a week now, and it works fine until I'm doing this. So, anyway, sorry about that. Alright. So I wanna give y'all a slight behind the scenes look at the NPEAK NFT we did, for the goal getters and walk you through, sort of this be patient and persistent idea. So, originally, Woot sent me this image up in the top left, and he said, this is our POAP or POAP as they say. This is this is you know, we're gonna be giving to people, and this is sort of the inspiration. We want something with a mountain and some people in it and a celebratory idea. So I asked Midjourney to give us some celebratory people, and it gave us this one in the top middle.

And if you look, you know, it's not a diverse group. It's all men. There's some bizarre things going on in some of these photos. There's some, you know, misformed faces, hands, people. It looks like somebody's sitting on top of a cake on top of a mountain in one of these. So definitely not exactly what we were looking for. So I refined the prompts a little bit, and we got the image in the top right. So now we've got a mountain in the background. We've got some people. It still looks like a all male cast. There's some weird malformed people. The color is not what I wanted. I wanted more purples for in peak, but we're starting to get there. So then I switched over to stable diffusion, bottom left, and it gave us this more stylized image, still some weird artifacts.

It wasn't exactly the style I wanted, but it gave us a starting point. So then I took this image and did an image to image, which we'll talk about in a little bit, back in the mid journey, which is what you see in the bottom middle there. With that image, we're getting closer. It it did some strange collage effects, did some weird things with the mountain peaks. The people are kind of on top of each other, but we were getting closer. And then through through image prompting and prompt engineering, we eventually got to our final output, which was the bottom right output, which was number one in the NFT series, and then all the others are variations on that one. But what I'm saying is patience is key.

It took about 300 plus iterations to get from, you know, these pictures of guys sitting on a cake in the top middle to our final result down on the bottom right, and it took a lot of patience and a lot of iterations. So we'll talk about that a little bit more. So prompt engineering. Basics one zero one, how do you start? And I would say this is probably the best way I've seen to start. You start by asking a list of questions. So do you want a photo or a painting? You know, are you looking for realism? Are you looking for, you know, artistically more of a painterly style? It's probably your very first question. Next is what's the subject. Right? Is it a person? Is it an animal? Is it a landscape? Is it an object? Next, what details do you want to add?

Right? So are you looking for special lighting? Are you looking for studio lighting or outdoor lighting? What's your environment? Is it indoor, outdoor, underwater, in space? What's your color scheme? Do you want a bright, neon, vibrant color scheme, maybe a dark, moody color scheme, a pastel color? What's your point of view? This was one of the hard things when we were making the NFT for MPEAK, right, is is the point of view face on? Is it over the shoulder, overhead, looking at it from the side? What's your background? You know? Are you are you looking at a mountain peak in the background or a valley? Or you need to you need to think about in your mind's eye, you know, what am I trying to create here? Are you looking for a specific art style?

Are you looking for something that looks like it came out of the Unreal Engine or a three d render or a Studio Ghibli or a movie poster? Are you looking for an a national park poster or a National Geographic cover? What what specific art style are you looking for? And are you looking for a specific photo type, a a macro or a telephoto or a wide angle? So I would start by just kinda, you know, asking yourself these questions. What am I trying to create? And here's a great example of this. So answer the following questions. Do we want a photo or a painting? We want a painting of a goat a goldendoodle wearing a suit, natural light, in the sky with bright colors, by Studio Ghibli. And so we take those questions and we put it in a sentence, a painting of a cute goldendoodle wearing a suit, natural light in the sky with bright colors by Studio Ghibli.

And we get this. So it is a painting of a goldendoodle wearing a suit, but he's not in the sky. Right? So we can move the word order around to to sort of tell the AI different things. So in this case, we changed the word order. And if you look, originally, it said, Goldendoodle wearing a suit, natural light in the sky was later on in the prompt. If we move the words in the sky earlier in the prompt, it's gonna change the way the AI interprets it. And so in this case, now we have a goldendoodle in the sky, very first descriptive thing we give it, and it is now the dog is now in the sky. So this teaches us that word order is important when we're crafting these prompts, and here is a great example. Right?

So the order and presentation of the desired output is almost as important as the actual words we're using. So we have a cat sitting on a Martian table. It's just a cat sitting on a strange looking red table. If we say a cat that is sitting on a table on Mars, it's a cat sitting on what looks like a standard, maybe, outdoor table on Mars. If we say a cat that is sitting on a table, the table is on Mars, it's now got a cat sitting on a table like surface on the planet of Mars. So where you put the words can be as important as the words you use when you are prompting. Additionally, just to sort of reiterate this, sometimes the weirder the prompt gets, the harder the AI time it has, you know, sort of determining what you're looking for.

And so in this case, the prompt is pink ice cream truck with machine gun mounted on it, technical, and there is no machine gun. However, if you change it to machine gun mounted on top of an ice cream truck, technical, now all of a sudden, the machine gun appears. So depending on where you put the words, sometimes, you know, the the AI will actually totally ignore the prompt. So just something to be aware of, kinda back to that patience and persistence. Sometimes it takes a lot of trial and error to get what you want. In this case, Woot, I see your question. Why the technical prompt? I think they're looking for something that looks less artistic. I haven't tried this exact one. I stole this from one of the guides that's in the reference because it was so good.

But I think they were simply saying they want something that looks very realistic. Alright. So prompt engineering. Couple other things to talk about before we start diving into, you know, more of the nitty gritty is is types of, prompt modifiers. So there are style modifiers, which are descriptors, which consistently produce certain styles. So the way to think about this is if I say photorealistic, I will consistently get a certain style. If I say by Norman Rockwell, it will consistently look a certain way. No matter whether I say a Goldendoodle or an ice cream truck, if I say by Norman Rockwell or with film grain, it will consistently style, the output in a in a similar way versus another type of modifier as a quality booster. So these are terms that are added to improve nonspecific quality. So high resolution, eight k, clear, beautiful, realistic.

These won't consistently style it in the same way, but they will consistently enhance the quality of the output. So kind of a fine def, you know, difference, but it does make a difference, as we'll see. So let's start with photography modifiers. When you're crafting a prompt, photography, you really need to think about sort of the technical details of how you're creating this prompt if you're trying to get a photograph. So you wanna talk about shot type. Is it close? Is it extremely close? Medium, long? You know, where is this shot being taken? What style? Is it a Polaroid? Is it black and white? Is it tilt shifted or blur shifted? Long exposures. You're gonna have light trails. Who's the subject? Right? Is it a woman, an old man, a cat? What's the lighting? Is it soft lighting, ring lighting, cinematic lighting? What's the context?

Indoor, outdoor, at night, studio. What kind of lens is being used? So are you looking for a wide angle or a telephoto or a a portrait lens? Do you want bokeh in the background? And what sort of device was being used? Right? Are you using a, you know, Fuji x? Are you using a Nikon z? An iPhone will also give you a different result. So I would say photography is particularly specific if you were looking for something, and it's good to reference a table. Or one of the references in the reference sheet is which we'll get to a little bit later is let's see. The stable diffusion prompt book resource is great about this. It's got some really good tables like this. This is where I took this from. So let's see. Albert, I see your question. Do you have a prompt generator where I can choose?

Yes. There are entire websites of prompt generation. I didn't put it in the references yet, but I will. There's dozens of websites that will help you construct a prompt. And one of the things we will talk about next lesson is you can even use large language models like ChatGPT, thank you, Rachel, to help you write a prompt. Now they have some limitations when you're doing that. They don't always understand word order or character limits, but they can certainly help you with ideas and ideating different styles, more descriptive words, vocabulary. So just wanna run through some other modifiers that you can that you can use. So we we kinda went through the photographic modifiers at a at a very high level. Let's talk about artistic modifiers a little bit. So you can talk about the medium you want. Do you want a chalk painting, wall graffiti art, Watercolors, oil paintings.

You can you can give artist styles that you're looking for. Maybe I want, you know, a chalk painting in the style of Ralph McCrory who did some of the Star Wars concept art. You can also give it three d illustrations. I want a low poly three d isometric render or an isometric asset or a pixel render, and it can understand that. You can ask for even more illustrative styles, vector illustrations, scientific illustrations, comic style artwork. Like I said, one of my favorites is to tell it to do, a portrait in the style of a National Geographic magazine cover. You can also use emotions as a style modifier. So I gay I I use positive emotions for our examples here, cozy, romantic, joyful, energetic, but you could also say, I want a dark image in the style of, Geiger, and it's gonna give you something, you know, grim or dark, or morose.

So you can also use emotional modifiers to give a feel, to the output. You can also give sort of color modifiers. So here's an example of some vibrant ones like we were talking about. Weird core image of a zoo or an assive wave aesthetic or a dream core style or a vaporware pool, and it'll also interpret those. You can also ask for historic styles. So here's a painting of Danny DeVito in Baroque cloth, you know, Soviet way, Wild West. So you can you can also give historic styles. Say, you know, I want a renaissance painting, or you can say, I want, you know, a painting in the style of Michelangelo, or you can be even more specific. I want a painting of Danny DeVito in the style of Michelangelo's last judgment. So, you can get as specific as you want.

Sometimes you can get too specific, and we'll talk about that in a minute. So here's sort of quality modifiers like we were talking about quality boosters. So you can on the left, we have a landscape. On the right, we have a landscape, HDR, UHD 64 k, and you can see there's a lot more detail there. So using these quality boosters can help you with specificity of your outputs. Also, you can use words as simple as highly detailed. So on the left is the prompt is Joan of Arc portrayed by Jennifer Lawrence concept, and you can see that it's very sort of cartoony. The detail's not there. The nose is not exactly correct. But on the right, the prompt is exactly the same with the word highly detailed added. Now my personal opinion is this is gonna get less important as, these models evolve and get better.

But right now, just using words like very, very, very highly detailed will give you a more detailed image than not using a term or simply using highly detailed. So there are some strange idiosyncrasies right now that I think they'll get this worked out, and we won't be talking about this in a year. But right now, these quality boosters can make a big difference in terms of your outputs. So that's sort of the high level how do you make a prompt. Now we're gonna dive into mid journey. We're gonna go through mid journey best practices, and then we're gonna go through stable diffusion best practices. We're not gonna touch on DALL two simply because I haven't used it as much. I've produced thousands and thousands of images with mid journey and stable diffusion, and they're they're cheaper, generally. Mid journey, you can get a subscription for $10 a month.

Stable diffusion's free if you use it on your own computer. So it's much more cost effective to produce a ton of pieces with Midjourney and Stable Diffusion right now. So that's partly why we're gonna focus on those. So Midjourney. This is the very basic structure to a prompt in Midjourney. So Midjourney, remember, you access through a Discord bot. So every prompt begins with slash imagine prompt. And then if there is an image, the image comes first, and we'll talk about that in a little while. If there is not an image, the text prompt comes first or next. And then there's parameters after the text prompt, and we'll talk about that in a minute. But this is the basic structure of a mid journey prompt. Alright. So basic usage. The basic prompt anatomy like we just talked about. So here's an example.

So image prompt astronaut on a horse, and we get four generations. MedJourney always gives us four generations. Now here we have an example with some parameters behind it. So imagine prompt astronaut on a horse, aspect ratio, three two, chaos, 70, quality, two, seed, 1,000. We're gonna talk about what all those parameters mean in just a moment. Alright. So let's start with aspect ratio because we already talked about sort of how to construct the text part. Right? Astronaut on a horse, and we'll talk a little bit more about that. But I think we've gone over the basics. So now let's dive into the parameters a little bit. So aspect ratio. Now this is something that's changing all the time. So mid journey version five can handle any aspect ratio. Mid journey version four, and earlier, the aspect ratios are more constrained.

But if you want a specific aspect ratio, the default is one one. It's a square. You have to define that with a dash dash a r for aspect ratio, and then you define your aspect ratio. Whether you want a two three portrait or a three two, landscape aspect ratio, this is where you define that. So chaos would be the next thing. That was the dash dash c, and that influences now this is interesting because there's a difference between, we'll talk about it in a little bit, stylized dash s and dash c. They sound very similar, but they do slightly different things. So chaos influences how varied the initial image grid is. So remember, we get four images when we first put a prompt in, and the chaos rating will tell tell how different they are. So if you look on the bottom left, this is a low chaos value.

So we've got a prompt watermelon, owl, hybrid. And if you look on the bottom left, all of those grids of four are relatively similar. As we increase the chaos value and the default value is zero. But as we increase that value towards a 100, that grid of four, you know, initial images is gonna get the four images will be very different from each other. So if we look at the bottom right, that set of images, the stylization of the different images is very, very different. So if you don't have a specific idea in mind and you're sort of experimenting and and, you know, playing around with different styles, images, compositions, this is a great way to do that quickly. Increase the chaos rating, and you'll get four very different images each time off the same prompt. Quality settings. So this is this is not resolution.

This is sort of detail. So and I'm gonna read a little bit of this because it's interesting. Higher quality settings aren't always better. Sometimes a lower quality setting can actually produce better results depending on the image you're trying to create. So if you're looking for a very abstract work, you actually might want the quality, which is sort of the amount of detail in the image, to be lower. If you're creating, you know, a very detailed architectural drawing, you may want the quality setting to be very high because you want a lot of detail in

Speaker 3 there.

Ben Now this is sort of one of the idiosyncrasies with Midjourney. Your subscription is based on how many GPU minutes and hours you're using. The higher the quality, the more of your GPU minutes and hours you're using, essentially. It will generate the images slower, and you're using more of your, you know, subscribed time. So something interesting, certainly something to play around with depending on the style you're looking for. Seeds. Now this is sort of interesting. This is this is sort of reproducibility when you're generating AI images, and we'll we'll talk about this also in stable diffusion in a few minutes. So there's a seed number that creates sort of the visual noise. We talked about how these are diffusion models. Right? So they're diffusing an image out of noise. So Albert said, would you recommend to only use low quality and then use high quality details for the final one?

I wouldn't because they actually give you a different style sometimes. So, like, if you look at this woodcut birch forest, point two five is not necessarily a worse image than a quality of one. They're they're very different. So I wouldn't I I would I would if you wanted to go quickly, I would put my chaos high, my quality low, and then fine tune it. But it really depends on what your workflow is. If you're looking for a specific image or if you're experimenting more conceptually, it really just depends on what you're looking to do. But, yes, to save GPU, absolutely. But same thing, it depends on what subscription you're using. If you're using the $10 and you're trying to save hours or if you're doing the $30 a month and you have unlimited relaxed hours and you just wanna put out a whole bunch of things or you're looking to do very high highly detailed architectural work, you could just do that under the relaxed mode.

You'd wait a little longer, but you wouldn't use your fast hours. So there's there's some decisions to be made there. So seeds. So, essentially, the seed number influences this random static that the image is being made out of. Normally, this is random, and you don't really touch it. However, if you have an image and you want that specific original seed to produce very similar images, you can go into Midjourney. You can ask it for the seed that was used, and you can reuse that seed to create more images, and they're gonna be very, very similar. So here here's an example. If we look at on the bottom left, this is a Celadon Al pitcher. You can see there's a fair amount of difference between those four grids. However, if we look at the grids on the right, these were all run with the same seed, one, two, three.

And you can see that they are very well, these are exactly the same because they were run with exactly the same seed. So this gives you some reproducibility if you wanna go back and adjust an image so that it's not as random, if that if that makes some sense. Now stylize. So like I said, this is similar to chaos, but it's different. So in this particular instance, the mid journey bot's been trained to produce images that favor artistic color, composition, and form. Stylize influences how strongly this training is applied. So low stylization produces images that are very close to your prompt, but essentially, you're giving it less artistic freedom. High stylization is going to give you images that are very artistic, but they're gonna be less connected to your prompt. So here's a prompt example of illustrated figs.

And if you look on the bottom left, this is with very low stylized, which which is the default. Right? A hun 50 and a 100 on the far left, and we can look on the right too. It's gonna look like what you asked it to, but it's not it's not gonna create something interesting and artistic. And so this really depends on what you're looking to do, you know, how much you wanna experiment. But I find this particular setting to be very helpful once I've kinda come up with what I'm looking for. It allows you to experiment a little bit. We'll talk about this with stable diffusion a little while. You can experiment so far that it's really not giving you what you want at all. So there's some trial and error with this as well. But really, really an interesting, you know, setting to play around with, I find.

Tiling, this is something that was not available in version four, but it came back in version five. So the tile parameter generates images that can be used as repeating tiles to create seamless patterns. So you can use this for, you know let's say, you want to combine a background with a a character. You could use this to create, you know, the background if it was a repeating background or a wallpaper or a fabric, or you could use this one layering. Yes, Albert. The presentation is available in the references, which we'll get to at the end of the presentation. But so this tiling can be very useful. Very specific use case, but it can be really useful, and I would say most people don't even realize this is available in mid journey. Alright. Version. So this is sort of you know, with stable diffusion, we're gonna talk about models.

I would say this is mid journey's equivalent to models. Right? Mid journey itself is a custom model, maybe based on stable diffusion, but this is sort of mid journey's version of custom models. It's something that you can you can play with, change, and you will get very it changes everything about Midjourney. So the current version is version five. That's on the bottom left. It's the newest model. It's their most advanced model. It was just released two weeks ago, March 15. And I've been playing around with it. You can go on Twitter and look it up, and lots of people are doing some very interesting things. I would say one of the most unique things about version five is it can create some of the most realistic images that Midjourney's been able to create yet.

Now version four, which was the previous version and what we've all been using for the last few months, if we've been using Midjourney, was a little more constrained than version five. I would say it's a little more artistic while still being more realistic than the previous three versions. Something most people probably don't realize is that there were actually three different flavors of version four. Each one had a slightly different stylistic tuning. So there's a four a, a four b, and a four c, and you could select which of those you wanted. Here's a couple examples on the right side of the screen. So if version five is giving you something too realistic, right, on the bottom left, we're looking at vibrant California poppies, You can see that these version four models will still retain, you know, the realisticness. They're not quite so abstract as version one, two, and three.

However, they have different styles within themselves, and you can select those and experiment with those. I'm not gonna go back to version one, two, and three because as we talked about in lesson two, I think they had too many artifacts and too much of a waxy look to be as useful as version four and five, but I think right now there's still a place for using version four and five together depending on your desired output. Now this is something that I don't think a lot of people are aware of. There is a Nidji model, which is a custom trained it's it's essentially version four that was custom trained on anime, anime styles, and anime aesthetics. So it's not necessarily better than four. It's better in some situations, especially if you're looking for something cartoony, very vibrant.

So in this case, we have version four's fancy peacock, and on the right, we have Nidji's fancy peacock, which, you know, you could argue whether that's more aesthetically pleasing or or or less, but it's certainly different. It focuses it has a different strength. So, also, something very interesting that you can play around with and get a very different output. And you can also you could create an image and say version four or five, and then you could do a variation using this model to sort of fine tune that. So let's talk about variations a little bit. There's a remix mode that influences how variations work. I just want to briefly go over it. But, essentially, once you get your image, there will usually be a button for make variations. You can click on that button. So in this case, we have a line art stack of pumpkins on the top left, step one.

We click that make variation button, and then we can modify or enter a new prompt. And this is where it's interesting. We can also apply different versions to this one single image. So in this case, the first thing that person did was said, change this stack of pumpkins into a pile of cartoon owls, and it did that. But we could also say, you know, use the you we could change the model. So if you look at the bottom left, this used a test model, which we're not gonna go over, but there are occasionally beta test models that Midjourney releases. So you could use a different model. You could change the subject. You could change the medium. So this is sort of a this is almost mid journey's version of ControlNet, right, where it's trying to keep things very, very similar, maybe at a line level or an outline level and change styles or models or versions.

So very interesting, but very fine tuned. Right? You're down to your you're down to one image at this point, and then you're making small changes. Multi prompts. So this is something that's interesting and worth talking about. So you can write cupcake illustration, or you can write cup and then two colons, cake illustration, and then you could write cup, two colons, cake, two colons, illustration. And this will actually cause the AI to interpret these prompts differently. So cupcake illustration is considered together. Right? So it's producing you illustrated images of cupcakes. That's what we see on the left. In the middle, cup is considered separately from cake illustration. So it's producing images of cakes and cups. And then if you look on the right, cupcake and illustration are all considered separately, producing cake and a cup with common illustration elements like flowers and butterflies.

So you can actually force the AI to consider phrases fully separately as as sort of separate but interconnected prompts. Not gonna dive too deep into this so we don't go for two hours, but multi prompts are very, very powerful in mid journey. There's also prompt weights. So I moved this to this lesson from next lesson, because I felt like it flowed better, in this section, but you can you can weight a prompt. So when a double colon is used to separate a prompt into different parts, you can then add a number immediately after the double colon to assign relative importance to that part of the prompt. Quick example, bottom left, this is a hot dog. Right? It's not a hot dog because it's been separated. It's hot and it's a dog.

However, on the image to the right, hot, double colon, to dog, we've given hot twice as much important as dog, and we can see now the dog is on the left. Flames are coming out of it. So we waited the word hot in that section of the prompt. Heavier than, dog. On you can also do the negative of that. Right? So on the right, we have a vibrant tulip field. And on the far right, we have a vibrant tulip field, but we have negatively weighted the word red. So we

Speaker 3 we

Ben so the field is less likely to contain the color red. We could also use the no parameter, which essentially just weights it at a weight of negative half. So we can we can weight some sections heavier or less. So if we wanted a picture of a mountain with lots of trees on it, we could weight it like that. If we wanted a picture of a mountain without any trees on it, we could weight down trees or weight up mountain. There's different ways of doing it. But prompt waiting along with multi prompts is super powerful in terms of getting you the specificity that you want. Alright. Let's see what we've got next. Alright. Upscaling. So mid journey and stable diffusion both handle this slightly differently, but they both have some upscaling capabilities built into them. So mid journey starts by generating a grid of low resolution images.

Right? So you generally get five twelve by five twelve. If we're if our aspect ratio is one one, we're gonna get five twelve by five twelve images in that grid. Then when we select an image, it's gonna upscale that to, generally, twenty forty eight by twenty forty eight or ten twenty four. We'll talk about that in just a second. But your original image in that grid is realistically probably gonna be lower resolution than you want to use, and then you're gonna upscale from there when you and it'll do this automatically when you select an individual image. So let's talk really briefly about upscaling. So there's the regular upscaler on the left, which really is just increasing the image size while kinda smoothing out or refining details. Then there's a light upscaler that handles, you know, moderate amounts of details and textures. There's a detailed upscaler and a beta upscaler.

And if you look, you can probably see the differences. If you're really interested in this, let's say download the presentation and look at them. But probably the easiest difference to see is the detailed upscaler on the left. There's a ton of detail. Beta, there's a little bit more. And then on some of these light ones and the regular one, you can tell there's less detail. Sometimes you want less detail. You know? Say I have a oil painting and I try and do a detailed upscale of the oil painting, it may add too many brush strokes and too much detail, and it won't look aesthetically pleasing. So you can play around with this too, and you'll kinda find what upscaling you want for the style of image that you're looking for. That's something to understand. Alright. Blending. So let's talk a little bit about image prompting.

There's two ways to do image prompting in Midjourney as of today. So the first option is the blend option. The second opt in option is true image prompting. The blend option will allow you to upload two to five images quickly, then it kind of looks at the images and mashes them together into a new image. There's really no text options with blending in the sense that you can't add text after that. You just pick the images, and the AI decides how it's gonna blend them together. But it's quick and it's easy. You can do it from your phone. You can select local images from your computer, throw them together, and kinda see what happens. So in this case, we have this strange illustration of, like, a mask or some sort of, you know, mythical forest god or something.

And then on the right, we have this blue bird, and it mixes them together in some different ways. Image prompting, which is sort of the more robust workflow, allows you to upload one or two images and then add the text prompt and then add the parameters after it. You could do the same thing as blend and simply dump two images together. What you can't do is put in an image prompt with no words. It just won't process that. So you do have to either give it two images or an image plus a text prompt for it to process that. And version five's come a long way from version four. If you look, here's some examples, statue and flowers, statue and jellyfish, etcetera, but image prompts are extremely powerful. One thing to note is that you have to upload the image to Discord or link to a publicly available URL of the image.

So it's a little bit more of a pain in the butt than Blend. Blend, I can just grab a couple images off my phone or my computer and put them right in the prompt. Image prompting, this traditional more powerful way, I have to either upload them to Discord or upload them somewhere publicly on the Internet to use them. So while it's more powerful, it's not as as fast, not quite as easy to use. Alright. Let's jump into stable diffusion. So we're gonna talk a little bit more generically about stable diffusion because, like I mentioned before, stable diffusion is the Linux of AI, image generation. Right? There's lots of different ways to run it, whether you have Windows, Mac, Linux, whether you access it through a GUI like this. This is Mochi diffusion on the Mac.

Whether you access it through a command line, It's hard to get too specific because it can look different and behave slightly different depending on how you installed it, where you installed it, how powerful your computer is. In this instance, though, we can see sort of the basics. We have our text prompt in the top left. We have exclude from image. We have HD. We have starting image, strength, number of images, steps, guidance scale, seed, and model, and we'll kinda go through these and talk about how this works. So very first thing to note as we walk through this is the default resolution is five twelve by five twelve, one one aspect ratio, same as mid journey. However, since you're generally running stable diffusion on your own computer, this is where things get a little different. You can change the resolution. You can do five seventy six by four forty eight.

You could do, you know, an aspect ratio of three two. However, this is gonna increase the required RAM and the time because you're running this on your computer. There's not some giant server farm where you're doing this, and so this will impact how long it takes for you to get images. So, generally, if I was gonna create a number of images, I would probably do them in five twelve by five twelve on stable diffusion. The other difference is you can tell stable diffusion how many images you wanna create and how many you wanna create at once. So maybe instead of a grid of four, I want a grid of 16, but I know that it's gonna, you know, kill my computer to do all 16 at once. So I have to tell it to do two at a time, and it takes 10 to generate them.

So something to be aware of. There is an upscaler. I don't wanna say it's built in to stable diffusion, but it's it is produced by Stability AI, and it is generally a part of a stable diffusion install that automatically upscales times four. So same as mid journey up to twenty forty eight. Obviously, they gave us a little more dis detail on how it works than mid journey does. The model card focuses on the model associated with stable diffusion upscaler. The model is trained for 1,250,000 steps on a 10,000,000 image subset containing images that are over twenty forty eight. So, anyway, it essentially uses AI to upscale the images, and you can see on the left is a five twelve, and on the right is is a 20 by twenty forty eight image. Alright. Seed. We talked about seed with mid journey a little bit.

We'll talk about it with stable diffusion. Works very similar. Seed is a number that controls the initial noise that's you know, the diffusion model's using, and it's why you get a different image each time you generate when all other parameters are fixed. Normally, this is random, but same as mid journey. You could go back and you could find the same seed, and you could get similar generation. So and this is true too. Some seeds are better for a prompt. For whatever reason, the way that randomness and that seed with that noise works, the prompts come out more aesthetically pleasing or closer to what you're looking for. So by reusing those seeds, this is a great, sentence here. You can use that to test the effect of different modifiers. So by keeping the seed the same, you can really start fine tuning what the different modifiers do.

Monika, to answer your question, yes. Default stable diffusion model often focuses more on a realistic look than mid journey. I would say it's it's less artistic by default. Obviously, you can change that very, very quickly and easily, you know, mid journey four versus five or what type of custom model you're using on stable diffusion, which we'll talk about in a moment. But, yes, it generally does tend towards producing more realistic default outputs than mid journey does. Okay. So this is something interesting in stable diffusion that is slightly different from mid journey. Mid journey, we talked about chaos and stylization and stable diffusion, they call it guidance. So if you look at the GUI here, we have a guidance scale down on the bottom left, which is eight on this one. On the over here, the full name is the classifier free guidance, but it's usually in a GUI.

You'll see it as guidance. So this parameter is, the creativity versus prompt scale. Right? Lower numbers give the AI more freedom to be creative while higher numbers force it to stick to the prompt. So if we were to give something a guidance scale of zero, it's essentially going to completely ignore the prompt and just give us really gobbledygook. If we go too high, it tries to stick too closely to the prompt, and it will actually do some strange things. So the default is seven. Four may be more creative, but it also might miss things from your prompt, and you go too high, and it'll it'll start looking too strange and too structured. We'll look at another example of this. So why would you use these values? So here's a prompt, an anime illustration of a blue rabbit riding a scooter near a lake with the sun in the sky.

So if we go a guidance scale of two, it's very strange looking. It's not really a coherent image based on that prompt. Seven, which is our default, we do have a rabbit. We do have light in the background, but we don't really have a sun, and it the scooter is malformed. 12, we get everything we're looking for. We've got our rabbit on a scooter near a lake, sun in the sky, 19 you know, I I think this is true. Composition seems too forced. It can it can tend to be a little too literal and do some strange things. So just something interesting to know, kind of stylized as the same as guidance scale and stable diffusion. Now step count, this is something else that is different. And stable diffusion from mid journey, you've got some more control over the diffusion process.

So stable diffusion creates an image by starting with the canvas full of noise and denoising it. Right? We've talked about how diffusion models work. So this controls the number of denoising steps. So usually higher is better, but some strange things can happen if you go too high or too low on the steps too. So step one, it's just noise. Step two, it's still just noise. You've probably seen this when you create an image in MIDjourney, and it sort of goes from blurry to the actual image. They're showing you this process. Stable diffusion generally doesn't show it. It does this on the background, but it does take more time to generate an image with more steps, which makes sense. It's processing more steps. And this, you can see in the GUI from earlier on, is right here in the steps section.

And this I've got it set at 40 for this one. Let's see. So default is 50. I forget if you can go to 75 or 80, but like I said, you can actually go too high, and it doesn't improve the image. Sometimes it actually makes it worse. So step count. Token efficiency. So this is something different also in stable diffusion. So if we go back to the image of the GUI, you can see I put I pasted a very long prompt from mid journey version five that I found on Twitter, and it's a 108 characters. And this says the description is too long, maximum of 75. So one of the things that is currently different in stable diffusion is your prompt is limited to 75 tokens, essentially 75 words. So if you're working with a long prompt, you do have to be more considerate of your efficiency in stable diffusion than you do in mid journey.

So and here's some examples. Right? You could say horse van Gogh. You could say a horse in the style of Vincent van Gogh, but that takes more than twice as many tokens. So it's something to be aware of when you're creating in stable diffusion with a very long complex prompt. You can actually go over the limit, essentially. Negative prompting. So this is sort of stable diffusion's version of prompt waiting. It's slightly different the way it's done. It's usually done in an in a GUI, it's done in an entirely different box. So you don't you don't wait in the same way in stable diffusion as you do in mid journey, but it's still possible. So in this case, we have a portrait photo of Genghis Khan. Second image, we have a portrait photo of Genghis Khan, but we added the negative prompt, black and white and monochrome, telling it that we don't want a black and white or monochrome image.

We want a covered image. So you could do that in the original prompt, you know, portrait photo of Genghis Khan color, or you could do it with a negative prompt. Let's see. Alright. So here we're looking at stable diffusion's version of image to image, and it's it's slightly different, but not so different from Midjourney. You could either start with, say, a cartoon image, give it a text prompt, and it'll give you, this generated image on the right, or you could, give a more detailed image and put it through. You could even take an image out of mid journey, put it through stable diffusion, and you will get a slightly different second image. So very similar in terms of sort of image prompting in mid journey or stable stable diffusion. Diffusion.

Prompting You can also, when you are doing image to image, adjust the steps, and it will give you and this is where you have some more fine tuned control and in stable diffusion than mid mid journey because you can adjust the steps on your image to image and get a very different variety based on how many diffusion steps you're using for your image to image prompting in stable diffusion. So very interesting. Inpainting is something that you can do in stable diffusion that you also can't do in mid journey, and I'd say that's sort of one of the themes here. Right? Is mid journeys a little bit easier and a little more polished, but you have a little less control over the process. If you really wanna get into the nitty gritty fine tuning details, you've got way more options with stable diffusion.

So in painting could be used to remove something. So in this case, we wanna remove Homer Simpson from the couch, so we mask off Homer, and we give it a prompt, orange couch in the living room, cartoon show, and Homer's gone. In the next one, we want to change Homer into a cartoon boy eating pizza on a couch. And so we mask off Homer, and then we end up with a cartoon boy eating pizza on a couch. We could also use this to fix the photograph, you know, almost like a advanced AI version of a repair tool in Photoshop. So we wanna remove these hot air balloons. So we mask off the hot air balloons, we give it a prompt photo of a sunset, and it removes them. Outpainting is kind of the opposite. So stable diffusion will give us a cropped image oftentimes.

So in this case, there's a a very up close portrait image of this woman, but maybe we want to see, her full head or her shoulders, so we out paint. We essentially out mask the image. I would say both in painting and out painting painting take a lot of trial and error. In my experience, it they're not all there yet. The results can be amazing, when you get what you want, but it can take a lot of trial and error to get there. Custom models. So this is this is probably the most powerful thing in stable diffusion that Midjourney doesn't have. Midjourney has their versions. Right? They've got version one, two, three, four, a, b, c, Niji, and version five. But with stable diffusion, you have an almost unlimited number of custom models. So custom stable diffusion models are fine tuned. So so what are they?

Right? How are they different? Essentially, they're fine tuned on different, usually smaller data So if I want something that looks like a Pixar movie, or I want stable diffusion outputs that look like mid journey, or I want stable diffusion outputs that are gonna look like a Studio Ghibli movie, this is how I can accomplish it. They're available on several different websites. I've linked a couple in, the references, everything from, Civic AI to I forget what one of the others is. What is it? StableRes. You can get them directly off of Hugging Face. I will say it's a little bit technical. Right? You have to manually install these models. Oftentimes, the names are not intuitive as you can see on the right here. Some of the basic ones I have in one of my stable diffusion installs on one of my computers, and you have to make sure the model is compatible with that particular build of stable diffusion.

Is it 1.4, 1.5, two, two point one, and with your processor type. Am I running a Intel processor with, you know, an NVIDIA GPU? Am I trying to do a CoreML model on an Apple neural engine? So you have to know, you know, what processor type, what build I'm using, where to find them, what you're looking for. So it takes a bit of research, and it's a little bit technical. Here's three examples of different models. So we can see on the left, Open Journey. This is this is a stable diffusion model that is trained on mid journey, so it produces results that look more like mid journey. In the middle, we have one that's trained on Studio Ghibli movies, and so it will produce images in that style. And on the right is sort of a, modern Disney aesthetic. So I believe that's Barack Obama in the middle there.

So this is going to give you it's trained on images from modern Disney movies, so it will give you Images that look like they came from modern Disney movies. So really powerful, but it's not just plug and go, you know, like mid journey is. You can't just hop in and then five minutes be producing images. You have to find this and install it and make sure it's the right thing and test it. So we will stop there. I thought this was gonna be a much shorter session, but we're already at an hour. So do we have any questions? So, Monica, yes, there are installs for Mac. If you go which we'll get to in just a minute. If you go to the references page, the two GUI versions that I like to use are diffusion b and mochi diffusion. Both of those are referenced, and you can download those.

So, JJ, let's see. Here we go. Will there be any discussion on legal copyright or trademark? No. But we did deal with that in lesson one and lesson two. And, obviously, those things are still changing sort of on a weekly basis, but we did touch on them some. And I think there's some articles in the references references that discuss that a little bit. Let's see. I just turned the chat off by accident. So let me load this up. Let's see. Where do you find the seed in mid journey? And we have an upcoming session on web three law. So where can you find the seed in mid journey? You have to go into the let's see. You have to go into the web interface, and then in the web interface under that image, it will allow you to find the seed number.

The fastest way to find that, Lucas, would be in this references here. There is under lesson three, mid journey docs, and just search for seed there, and it will take you right to where that is. So our next lesson is coming up in two weeks on back to our sort of regularly scheduled time. We'll be on a Thursday, a little bit later in the day, and we're gonna talk about tools and workflows part two. We're going to go over some more advanced topics. Next time, we're going to talk about textual inversion, prompt blending. We will go over briefly using GPT to write prompts, both pros and cons of that. I wanna talk a little bit about how to train your own model. You've probably seen in the news, you know, people training a model on their face, and then they can produce portraits or them as Spider Man or Batman or something.

We'll talk briefly about how do you train your own model. We will go over ControlNet. I wanted to get to it today, but I didn't wanna go too long. So we'll briefly go over ControlNet, and then we'll also talk about offset noise and Laura's a little bit, for sort of our advanced session, and then we'll talk about GANs, at the end. Let's see. And then once again, there are references available. Bmeadows.xyz/npeak has, references for all the lessons as well as, the PDFs for this lesson and the previous two. So that's it. Everybody have a great day. Thank you.


All transcripts