System and method for the automated generation of an optimised ai prompt for image and video analysis, computer program and computer-readable storage medium
Patent Information
- Application Number
- PCT/EP2026/058945
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2025-12-05
- Filing Date
- 2026-03-27
- Publication Date
- 2026-10-01
Smart Images

Figure EP2026058945_01102026_PF_FP_ABST
Abstract
Description
[0001] System and method for the automated generation of an optimized AI prompt for image and video analysis, computer program and computer-readable storage medium
[0002] Description
[0003] The present disclosure relates to a system and a method for the automated generation of at least one optimized AI prompt for image and video analysis. Furthermore, the present disclosure relates to a computer program and a computer-readable storage medium.
[0004] AI-based analysis of image and video material is becoming an increasingly important component in security, surveillance, and automation systems. Currently, AI video systems increasingly use language models in the form of conventional neural networks optimized for specific classes to analyze image and video material. To reliably recognize specific objects or actions, these language models require precisely formulated prompts. Creating these prompts requires technical expertise and iterative refinement, making it challenging for many system integrators and users. Existing systems either offer static, predefined prompts or require technical expertise for their creation.
[0005] US patent 2024296315 discloses a system for storing and processing generative AI prompts. US patent 2025053799 discloses dynamic UL visualizations. US patent 2025086467 discloses the use of metadata to improve prompts.
[0006] The purpose of this disclosure is to provide a technique for the automated generation of optimized AI prompts for image and video analysis.
[0007] This task is accomplished by a system, a procedure, a computer program and a computer-readable storage medium having the characteristics specified in the independent KF2009P-WG-0001
[0008] 2 / 18
[0009] Claims resolved. Advantageous embodiments are the subject of dependent claims.
[0010] A disclosed system for the automated generation of at least one optimized AI prompt for image and video analysis. The system is particularly well-suited for use in security, surveillance, and automation applications. In these areas, images and videos are to be monitored for the potential recurrence of specific events. The videos may contain both visual and audio content. The system features a user interface configured to receive single-modal and / or multi-modal input, in particular text-based, speech-based, image-based, audio-based, and / or video-based input. Furthermore, the user interface is configured to create and transmit a prompt generation request based on the single-modal or multi-modal input.The disclosed system further comprises a prompt engine configured to receive the prompt generation request from the user interface and to generate and transmit at least one prompt in response to the request. The user interface is further configured to receive and output the at least one prompt from the prompt engine, as well as to receive an acknowledgment regarding the at least one prompt and, upon receiving the acknowledgment, to store and / or transmit the at least one prompt. The disclosed system enables the simple generation of prompts for image and video analysis without requiring detailed technical knowledge. Therefore, the system represents a helpful tool for...
[0011] System integrators.
[0012] Single-modal input can take the form of text, speech, image(s), audio, or video. Multi-modal input can take the form of a combination of text, speech, image(s), audio, and / or video. Both single-modal and multi-modal input correspond to a user request for image and / or video analysis. This user request defines which events in the image and / or video material are to be detected. Data transmission, such as the prompt generation request or at least one prompt, can occur in the disclosed system via wired connections such as LAN, Ethernet cable, fiber optic connection, powerline, direct USB connection, and / or SD card transfer. Furthermore, wireless transmission is possible via WLAN, mobile networks (LTE, 5G), Bluetooth, radio connections (433 MHz, 868 MHz), satellite connections, and / or mesh networks.A combination of wired and wireless transmission is also conceivable. Furthermore, the following transmission methods can be used for wired and / or wireless transmission: VPN, encrypted TCP / IP connections, cloud services, edge computing with local forwarding, HTTP / HTTPS protocols, RTSP, FTP / SFTP, MQTT, and WebSocket connections.
[0013] The output of at least one prompt through the user interface can, for example, take the form of text and / or sound and / or speech output.
[0014] According to one advantageous aspect, the user interface can be trained to receive an adaptation request instead of an acknowledgment and to create and transmit a prompt adaptation request based on that request. The prompt engine can be trained to receive the prompt adaptation request from the user interface and to adapt or regenerate the at least one prompt accordingly. This allows the at least one generated prompt and the associated precision (recognition rate) and relevance of the image and video analysis to be iteratively improved in a dynamic optimization loop.
[0015] According to another aspect, the prompt engine can be trained to receive annotated sample images / videos, apply the generated prompt to these annotated sample images / videos, and adjust the prompt to increase the recognition rate. The prompt can be adjusted or regenerated until a predefined recognition rate for the annotated sample images / videos is achieved. Thus, evaluation can be automated by automatically testing and adjusting the generated prompt based on (real-time) evaluation metrics (such as sample images / videos), thereby further increasing the recognition rate.
[0016] According to another aspect, the user interface can also be configured to output and receive a selection of example prompts. The user interface can then be configured to create the prompt generation request based on this selection of example prompts. This provides users, especially those with little experience, with a basis for decision-making when creating prompts, thereby enabling, for example, faster learning, fewer errors, and greater usability.
[0017] According to another advantageous aspect, the Prompt Engine can be a single-modal and / or multi-modal generative AI model, such as ChatGPT, DALL E, Midjourney, Stable Diffusion, Claude, Gemini, LLaMA, Mistral, Jurassic, BLOOM, MusicLM, VALL-E, Codex, Whisper, StyleGAN, BigGAN, Imagen, Synthesia, RunwayML and Sora.
[0018] Another advantage is that the prompt engine can possess pre-contextualized knowledge about the images and / or videos to be analyzed. For example, that the generated prompt is intended for the analysis of security camera videos. This pre-contextualized knowledge allows redundant general elements to be avoided and relevant parameters to be specifically incorporated into the prompt.
[0019] The system can, for example, capture the image and / or video material to be analyzed using cameras, lenses, microphones, infrared sensors, thermal imaging sensors, motion sensors, a network connection (LAN / WLAN), lighting units, automatic exposure control, noise reduction, a night vision function, time-based or event-based recording, mobile capture units, synchronization with time servers, and real-time data transmission. Furthermore, the system KF2009P-WG-0001
[0020] 5 / 18
[0021] For example, it may include motion detection, facial recognition, license plate recognition, behavioral analysis, object detection, person tracking, event triggering, automatic alerting, data storage, data compression, data encryption, metadata generation, time stamping, event categorization, relevant content filtering, report generation, export functions, access control, user logging, forwarding to security services, live monitoring, dashboard visualization, integration with other security systems, automatic notifications (e.g., via email or app), deletion period management, and historical data analysis.
[0022] According to another aspect, the system can include an image / video analysis unit configured to receive and store at least one prompt from the user interface. This image / video analysis unit can further be configured to receive an image and / or video and analyze it according to the at least one prompt. The image / video analysis unit can, for example, be integrated into a camera and / or video processor or connected via an interface (e.g., cloud-based).
[0023] According to one advantageous aspect, the image / video analysis unit can be configured to perform a time-based analysis of the image and / or video according to at least one prompt. For example, an analysis of the image and / or video is performed at one or more predefined times. It is also conceivable that an analysis is performed after a predetermined duration. In this way, temporal relationships, such as movements, processes, behavioral patterns, and / or causal relationships, can be recognized that would not be visible in a single image / video.
[0024] According to another advantageous aspect, the image / video analysis unit can be configured to perform an event-based analysis of the image and / or video based on at least one prompt. Thus, an analysis of the image and / or video only occurs when one or more triggers are present, and not [KF2009P-WG-0001].
[0025] 6 / 18
[0026] permanent. As a result, reduced data requirements, lower computing costs, energy efficiency, real-time response to relevant events, a focus on safety-critical or relevant situations, higher interpretability, a lower false alarm rate, better scalability with many cameras or sensors, easy integration into alarm systems, and better storage and bandwidth utilization can be achieved.
[0027] According to another aspect, the image / video analysis unit can be configured to receive an event signal from an external sensor. The image / video analysis unit can itself incorporate one or more sensors or be connected to one or more sensors via cable or wirelessly to receive the event signal. The sensor can be a camera sensor, a motion sensor, a door contact, a fire or smoke detector, and / or a microphone. The event signal can indicate an event that has occurred, such as movement (detected by a motion sensor), the opening of a door (detected by a door contact), a fire or smoke (detected by a fire or smoke detector), and / or a sound (detected by a microphone).
[0028] A disclosed method for the automated generation of at least one optimized AI prompt for image and video analysis comprises the following steps: inputting a single-modal and / or multi-modal input, in particular a text- and / or speech- and / or image- and / or video-based input, creating a prompt generation request based on the single-modal or multi-modal input and transmitting the prompt generation request, receiving the prompt generation request, generating and transmitting at least one prompt, receiving and outputting the at least one prompt, receiving an acknowledgment of the at least one prompt, and storing and / or transmitting the at least one prompt upon receipt of the acknowledgment.
[0029] The input of single-modal or multi-modal input, as well as the creation and transmission of the prompt generation request, can be performed via a user interface as disclosed. Receiving the prompt KF2009P-WG-0001
[0030] 7 / 18
[0031] A prompt engine can be used to initiate a generation request and to generate and transmit at least one prompt. Receiving and outputting the at least one prompt, as well as receiving an acknowledgment of the at least one prompt and storing and / or transmitting the at least one prompt upon receiving the acknowledgment, can be done through a user interface according to the disclosure. Thus, the disclosed method allows for the simple generation of prompts for image and video analysis without requiring detailed technical knowledge. Therefore, the method represents a helpful tool for system integrators.
[0032] According to one aspect, the disclosed procedure may further include the following steps: receiving an adaptation request instead of an acknowledgment and creating a prompt adaptation request based on the adaptation request and transmitting the prompt adaptation request, receiving the prompt adaptation request, and adapting or regenerating and transmitting the at least one prompt.
[0033] Receiving the adaptation request instead of the confirmation, and creating and transmitting the prompt adaptation request, can be done through a user interface as disclosed. Receiving the prompt adaptation request, as well as adapting or regenerating and transmitting the at least one prompt according to the adaptation request, can be done through a prompt engine as disclosed. This allows the at least one generated prompt and the associated precision (recognition rate) and relevance of the image and video analysis to be iteratively improved.
[0034] According to another aspect, the disclosed procedure may further include the following steps: outputting sample prompts and receiving a selection of sample prompts, as well as creating the prompt generation request based on the selection of sample prompts and transmitting the prompt generation request. KF2009P-WG-0001
[0035] 8 / 18
[0036] The output of sample prompts and the receipt of a selection of sample prompts can be accomplished through a user interface in accordance with the revelation. The creation of the prompt generation request based on the selection of sample prompts can be performed by a prompt engine in accordance with the revelation. Consequently, users, especially those with little experience, are provided with a basis for decision-making when creating prompts, thereby enabling, for example, faster learning, fewer errors, and greater usability.
[0037] Furthermore, the procedure as disclosed can be further developed in accordance with the aspects described above for the system.
[0038] A disclosed computer program comprises instructions that, when executed by a computer, cause it to perform the disclosed procedure. The disclosed computer program can be implemented in any possible programming language and on various technological levels and platforms, such as a desktop application (Windows, Linux, macOS), a smartphone app (Android, iOS), an embedded system (e.g., Raspberry Pi, Jetson Nano), an edge device (Edge TPU, NVIDIA Jetson), a cloud platform (e.g., AWS, Azure, Google Cloud), a web application (browser-based, e.g., with WebRTC or WebGL), a server-based application (with GPU acceleration), a Docker container (for portable AI applications), an integrated circuit (e.g., FPGA), and / or an IoT system with a camera and processing unit.Therefore, the technical effects and advantages described for the disclosed process can also be achieved by the computer program as disclosed.
[0039] The disclosed computer program is stored on a disclosed, computer-readable storage medium. This storage medium can be a hard disk drive (HDD), a solid-state drive (SSD), a USB flash drive, a CD, a DVD, cloud storage, network-attached storage (NAS), or the like. Accordingly, the KF2009P-WG-0001 described for the disclosed method can
[0040] 9 / 18
[0041] The technical effects and advantages can also be achieved through the computer-readable storage medium as disclosed.
[0042] The items described above are primarily used in security monitoring and automation. Furthermore, they can also be used in other areas for generating optimized AI prompts and for any type of image and video analysis, such as image and video analysis of (industrial) automation systems, traffic, customer behavior in retail, access control, smart cities, health, sports, animals, agriculture, education and distance learning, logistics and warehouse monitoring, environmental monitoring, construction sites, fall detection, facial recognition, emotion recognition, production quality, and driver monitoring in vehicles.
[0043] An embodiment of the present disclosure is described below with reference to a figure. It should be noted that the following description of the embodiment is only exemplary and is not intended to limit the scope of the disclosure. It shows:
[0044] Fig. 1 is a schematic representation of an embodiment of the system and method disclosed.
[0045] Fig. 1 shows a system 1 and a method 100 according to a disclosed embodiment. The system 1 comprises a user interface 4, a prompt engine 6, and an image / video analysis unit 8. These units work together to execute the steps of the method 100 disclosed. The user interface 4 is configured to receive multimodal input from a user 2 in step S1. The multimodal input can be, for example, text-based, speech-based, image-based, audio-based, and / or video-based. The multimodal input corresponds to a user request for image and / or video analysis. Furthermore, the user interface 4 is configured to create a prompt generation request based on the multimodal input in step S2 and transmit it to the prompt engine 6. The user interface 4 can also be configured to output example prompts to the user 2 in step S10.User 2 is thus provided with a basis for decision-making and can further refine the multimodal input by selecting from the example prompts. User interface 4 can also be configured to receive a selection of example prompts in step S11 and, based on this selection, to create and transmit the prompt generation request in step S12.
[0046] The Prompt Engine 6 is configured to receive the prompt generation request from the user interface 4 in step S3 and to generate and transmit at least one prompt in response to the prompt generation request. In the illustrated embodiment, the Prompt Engine 6 is a multimodal generative AI model with precontextualized knowledge, which avoids redundant general elements and allows relevant parameters to be specifically incorporated into the prompt. Furthermore, the user interface 4 is configured to receive the at least one prompt from the Prompt Engine 6 in step S4 and output it to the user 2. In step S5, the user interface 4 receives an acknowledgment regarding the at least one prompt from the user 2. Subsequently, in step S6, the user interface 4 stores and / or transmits the at least one prompt upon receiving the acknowledgment.In step S4, user interface 4 can output at least one prompt to user 2 in the form of text and / or speech. If user 2 is satisfied with the generated prompt, they enter their confirmation into user interface 4.
[0047] If user 2 is not satisfied with the generated prompt, user 2 can submit an adjustment request to user interface 4. In the embodiment shown, the system 1 according to the disclosure can be configured to execute an optimization loop 20. This allows the at least one generated prompt and the associated precision (recognition rate) and relevance of the image and video analysis to be iteratively improved. User interface 4 can further be configured to receive the adjustment request in step S20 and, in step S21, to create a prompt adjustment request based on the adjustment request and transmit it to prompt engine 6. Prompt engine 6 can further be configured to receive the prompt adjustment request from user interface 4 and to adjust or regenerate the at least one prompt according to the prompt adjustment request.After User 2 submits the customization request to User Interface 4, User Interface 4 receives the customization request in step S20, and in step S21 creates the prompt customization request based on the customization request and transmits it to Prompt Engine 6, Prompt Engine 6 can output a customized or newly generated prompt to User Interface 4 in step S22. User Interface 4 then outputs the customized or newly generated prompt to User 2 in step S4 to check whether User 2 is satisfied with the customized or newly generated prompt. User 2 then either enters confirmation into User Interface 4. Alternatively, User 2 can submit another customization request to User Interface 4, and the system 1, as revealed, executes another optimization loop 20.
[0048] After receiving confirmation regarding the at least one prompt in step S5, the user interface 4 stores the at least one prompt in step S6 and / or transmits the at least one prompt to the image / video analysis unit 8. The image / video analysis unit 8 is configured to receive and store the at least one prompt from the user interface 4. Furthermore, the image / video analysis unit 8 is configured to receive an image and / or video and analyze the image and / or video according to the at least one prompt. The image and / or video preferably originates from a surveillance camera. The image / video analysis unit 8 can be integrated into a camera and / or video processor or connected via an interface.The analysis, according to at least one prompt, is performed using AI and / or based on feature extraction, model applications, result interpretation, video-specific analyses, object detection, classifications, segmentations, face detection, emotion detection, or using tools. The image / video analysis unit 8 is further configured to perform the analysis of the image and / or video, according to at least one prompt, in a time-based and / or event-based manner. The image / video analysis unit 8 is also configured to receive an event signal from an external sensor. This event signal indicates that an event has occurred. KF2009P-WG-0001.
[0049] 12 / 18
[0050] The transmission of data, such as the prompt generation request or the at least one prompt, in the system 1 disclosed is carried out via wired and / or wireless and / or both wired and wireless means.
[0051] According to another embodiment, the Prompt Engine 4 can be configured to receive annotated sample images / videos, apply the generated prompt to the annotated sample images / videos, and adjust the prompt to increase the recognition rate. The prompt can be adjusted or regenerated until a predetermined recognition rate for the annotated sample images / videos is achieved. Consequently, the image / video analysis unit 8 can achieve higher recognition accuracy when analyzing the image and / or video according to the at least one prompt.
[0052] system
[0053] users
[0054] User interface
[0055] Prompt Engine
[0056] Image / video analysis unit
[0057] Optimization loop
[0058] Proceedings
[0059] Entering a multi-modal input
[0060] Creating and submitting a prompt generation request. Generating and submitting at least one prompt.
[0061] Receiving and outputting at least one prompt
[0062] Receiving confirmation of at least one prompt
[0063] Saving and / or transferring at least one prompt; outputting example prompts
[0064] Receiving a selection of sample prompts
[0065] Creating and submitting the prompt generation request based on the selection of sample prompts
[0066] Receiving an adjustment request
[0067] Creating and submitting a prompt customization request based on the customization request
[0068] Adapt or regenerate and transfer at least one prompt
Claims
KF2009P-WG-0001 14 / 18 Claims 1. System (1) for the automated generation of at least one optimized AI prompt for image and video analysis, comprising: a user interface (4) that is designed, to receive a single-modal and / or multi-modal input, in particular a text- and / or speech- and / or image- and / or sound- and / or video-based input, and to create and transmit a prompt generation request based on single-modal or multi-modal input, a Prompt Engine (6) that is trained, to receive the prompt generation request from the user interface (4) and to generate and transmit at least one prompt in response to the prompt generation request, wherein the user interface (4) is further designed, to receive and output at least one prompt from the Prompt Engine (6), to receive confirmation regarding at least one prompt, and to save and / or transmit at least one prompt upon receipt of the confirmation.
2. System (1) according to claim 1, wherein the user interface (4) is configured to receive an adaptation request instead of confirmation, to create a prompt customization request based on the customization request and to transmit the prompt customization request, whereby The Prompt Engine (6) is configured to receive the prompt adaptation request from the user interface (4), adapt or regenerate at least one prompt according to the prompt adaptation request, and transmit the adapted or regenerated prompt. KF2009P-WG-0001 15 / 18 3. System (1) according to claim 1 or 2, wherein the Prompt Engine (4) is configured to obtain annotated example imageZ-videos, to apply the generated prompt to the annotated example imageZ-videos and to adapt or regenerate the prompt to increase a recognition rate.
4. System (1) according to one of claims 1 to 3, wherein the user interface (4) is further configured to output example prompts and to receive a selection of example prompts and to create and transmit the prompt generation request also based on the selection of example prompts.
5. System (1) according to any one of claims 1 to 4, wherein the Prompt Engine (6) is a single-modal and / or multi-modal generative AI model.
6. System (1) according to any one of claims 1 to 5, wherein the Prompt Engine has precontextualized knowledge about the images and / or videos to be analyzed.
7. System (1) according to any one of claims 1 to 6, further comprising: an image / video analysis unit (8) configured to receive and store the at least one prompt from the user interface (4), wherein the image / video analysis unit (8) is further configured to receive an image and / or video and to analyze the image and / or video according to the at least one prompt.
8. System (1) according to claim 7, wherein the image-Z video analysis unit (8) is configured to perform an analysis of the image and / or the video according to the at least one prompt on a time-based basis.
9. System (1) according to claim 7 or 8, wherein the image / video analysis unit (8) is configured to perform an event-based analysis of the image and / or video according to at least one prompt. KF2009P-WG-0001 16 / 18 10. System (1) according to claim 9, wherein the image / video analysis unit (8) is configured to receive an event signal from an external sensor.
11. Method (100) for the automated generation of at least one optimized AI prompt for image and video analysis, comprising the following steps: Input (S1) of a single-modal and / or multi-modal input, in particular a text- and / or speech- and / or image- and / or video-based input, Creating a prompt generation request based on single-modal or multi-modal input and transmitting (S1) the prompt generation request, receiving the prompt generation request, Generate and transmit (S3) at least one prompt, Receiving and outputting (S4) the at least one prompt, Receiving (S5) an acknowledgment regarding the at least one prompt, storing and / or transmitting (S6) the at least one prompt upon receipt of the acknowledgment.
12. Method (100) according to claim 11, further comprising the following steps: receiving (S20) an adjustment request instead of the acknowledgment, creating a prompt adjustment request based on the adjustment request, and transmitting (S21) the prompt adjustment request. Receiving the prompt customization request, and Adapting or regenerating the at least one prompt according to the prompt adaptation request and transmitting (S22) the adapted or regenerated prompt.
13. Method (100) according to claim 11 or 12, further comprising the following steps: Output (S10) of example prompts, Receiving (S11) a selection of sample prompts, and Creating the prompt generation request based on the selection of sample prompts and submitting (S12) the prompt generation request. KF2009P-WC-0001 17 / 18 14. Computer program comprising instructions which, when executed by a computer, cause it to execute the method (100) according to any one of claims 11 to 13.
15. Computer-readable storage medium on which the computer program according to claim 14 is stored.