Art made with artificial intelligence

Explorations with VQGAN + CLIP, prompt engineering and the Rusty Roboz visual universe.

6/1/20225 min

Art made with artificial intelligence

In the summer of 2021, a Google Colab notebook came into my hands, which using VQGAN + CLIP (paper) was capable of generating images from a text sentence. The results were quite impressive, and there was something magical in writing a sentence and seeing how the AI interpreted it and generated what for me was undoubtedly a work of art. I was hooked.

I kept testing for months, using initial images that the AI used as a starting point, mixing text inputs so that the image was a combination of two sentences, and experimenting with a new concept for me: prompt engineering. Some typical examples were ending the sentence with "top of artstation" to give it a professional drawing style or "rendered by Unreal Engine 4" so that the image looks like it was generated by that engine. Some of these images can be seen in the Instagram account where I uploaded some of my first tests, using titles of books, movies, comics and video games as prompts.

At one point I started using as a starting image a sketch I made some time ago of Jax, the little robot that appears in Ted Chiang's The Lifecycle of Software Objects, a book I recommend to everyone, especially if you are interested in Artificial Intelligence and you like science fiction. I also began including "Rusty Roboz" in all the prompts. The idea was to see if interesting variations could be created from the same initial image and the words "Rusty Roboz".

In this way I started to create a lot of Rusty Robozs, which took a long time since running these models requires a good graphics card, which I did not have, so I had to use Google Colab. It is an excellent service that provides free access to a notebook development environment with graphics cards, even with its time constraints and limited hardware. Sometimes I let the model generate for only a few hundred iterations and other times I left it for hours, reaching several thousand iterations. I tried variations of the same prompt, colors, objects and more. Little by little you start to get a kind of intuition about which prompt will generate the most interesting output.

Original sketch of Jax
Original sketch of Jax

Generated pieces

Here are some generated images and videos from those experiments.

From text: Rusty roboz with a lightsaber in a volcano

From text: Red roboz surfing in the sea

Ciborg with neons
Ciborg with neons
Luke Skywalker generated with VQGAN and CLIP
Luke Skywalker
Blue Rusty Roboz space pirate
Blue rusty roboz space pirate
Dune generated with VQGAN and CLIP
Dune
Assassins Creed generated with VQGAN and CLIP
Assassins Creed
Uncharted 2 generated with VQGAN and CLIP
Uncharted 2
GTA VI generated with VQGAN and CLIP
GTA VI
Batman The Dark Knight Returns generated with VQGAN and CLIP
Batman: The Dark Knight Returns
Superman Red Son generated with VQGAN and CLIP
Superman: Red Son
Brave little Rusty Roboz after 500 iterations
Brave little rusty roboz (500 iterations)
Brave little Rusty Roboz after 2000 iterations
Brave little rusty roboz (2000 iterations)
Brave little Rusty Roboz after 4000 iterations
Brave little rusty roboz (4000 iterations)

Brave red roboz on a hill

Brave little rusty roboz

After a while I had a considerable number of images and videos of different Rusty Robozs and inevitably it occurred to me to make a collection of NFTs.

Bonus

Some other generative models that I tried in that period were:

Rusty Roboz experiment made with StyleGAN-NADA
StyleGAN-NADA
Second Rusty Roboz experiment made with StyleGAN-NADA
StyleGAN-NADA
Rusty Roboz experiment made with GauGAN 2
GauGAN 2 (from Nvidia)
Rusty Roboz experiment made with GLIDE
GLIDE (from OpenAI)