Andrey Palaev
LLM-guided instance-level image manipulation with diffusion u-net cross-attention maps
Palaev, Andrey; Khan, Adil; Kazmi, Ahsan
Abstract
The advancement of text-to-image synthesis has introduced powerful generative models capable of creating realistic images from textual prompts. However, precise control over image attributes remains challenging, especially at the instance level. While existing methods offer some control through fine-tuning or auxiliary information, they often face limitations in flexibility and accuracy. To address these challenges, we propose a pipeline leveraging Large Language Models (LLMs), open-vocabulary detectors and cross-attention maps and intermediate activations of diffusion U-Net for instance-level image manipulation. Our method detects objects mentioned in the prompt and present in the generated image, enabling precise manipulation without extensive training or input masks. By incorporating cross-attention maps, our approach ensures coherence in manipulated images while controlling object positions. Our approach enables precise manipulations at the instance level without fine-tuning or auxiliary information such as masks or bounding boxes.
Presentation Conference Type | Conference Paper (unpublished) |
---|---|
Conference Name | British Machine Vision Conference |
Start Date | Nov 25, 2024 |
End Date | Nov 28, 2024 |
Acceptance Date | Jul 20, 2024 |
Deposit Date | Oct 3, 2024 |
Publicly Available Date | Oct 3, 2024 |
Peer Reviewed | Peer Reviewed |
Public URL | https://uwe-repository.worktribe.com/output/13263543 |
Files
LLM-guided instance-level image manipulation with diffusion u-net cross-attention maps
(5.1 Mb)
PDF
You might also like
Cache sharing in UAV-enabled cellular network: A deep reinforcement learning-based approach
(2024)
Journal Article
Multiple adversarial domains adaptation approach for mitigating adversarial attacks effects
(2022)
Journal Article
PbCP: A profit-based cache placement scheme for next-generation IoT-based ICN networks
(2022)
Journal Article
Downloadable Citations
About UWE Bristol Research Repository
Administrator e-mail: repository@uwe.ac.uk
This application uses the following open-source libraries:
SheetJS Community Edition
Apache License Version 2.0 (http://www.apache.org/licenses/)
PDF.js
Apache License Version 2.0 (http://www.apache.org/licenses/)
Font Awesome
SIL OFL 1.1 (http://scripts.sil.org/OFL)
MIT License (http://opensource.org/licenses/mit-license.html)
CC BY 3.0 ( http://creativecommons.org/licenses/by/3.0/)
Powered by Worktribe © 2025
Advanced Search