Information processing system
Patent Information
- Application Number
- CN202610284859.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2025-03-19
- Filing Date
- 2026-03-10
- Publication Date
- 2026-09-22
AI Technical Summary
一、提示文本的编写依赖人工经验,缺乏体系化方法,普通用户难以高效地利用生成式人工智能模型生成符合需求的视觉显示物
[0009]又进一步地,为了优先利用效果更佳的构成要素,所述处理器还被配置为:根据选择率较高的视觉显示物的构成要素对提示文本进行调整,以优先采用所述选择率较高的构成要素;并基于经调整的提示文本指示所述生成人工智能模型对所述视觉显示物进行重新生成。通过将统计得到的高选择率视觉显示物中的文案、图像、色彩、布局等构成要素优先写入提示文本,并向生成人工智能模型发出重新生成指令,系统能够在生成—展示—统计—优化的闭环中持续提升视觉显示物的表现效果,从而有效解决现有技术中难以及时、自动地基于用户行为优化生成结果的问题。
Smart Images

Figure CN122802469A_ABST
Abstract
Description
Technical Field
[0001] The technology disclosed herein relates to an information processing system. Background Technology
[0002] Japanese Patent Application Publication No. 2022-180282 discloses a method for controlling a role-based chatbot executed by at least one processor. The method includes the following steps: receiving a user's speech; adding the user's speech to a prompt word, the prompt word containing instruction statements associated with an explanation of the chatbot's role; encoding the prompt word; and inputting the encoded prompt word into a language model to generate a chatbot response to the user's speech.
[0003] In existing technologies, the process of automatically generating visual displays such as banner ads and promotional images based on generative artificial intelligence models typically requires professional designers or operations personnel to manually write the generated prompts and repeatedly adjust the prompts and constituent elements (such as copywriting, color scheme, and layout) through trial and error to obtain visual displays with high click-through or selection rates. This approach has the following problems: First, the creation of prompt text relies on human experience and lacks a systematic approach, making it difficult for ordinary users to efficiently utilize generative artificial intelligence models to generate visual displays that meet their needs.
[0004] Second, existing systems often cannot automatically optimize the components of visual displays based on users' actual selection behavior and statistically obtained selection rates. They can only manually modify the data after observation, which makes the optimization process time-consuming and labor-intensive, and difficult to respond to changes in user preferences in a timely manner.
[0005] Third, even if some systems have the function of performing simple statistics on the generated results, they lack a mechanism to automatically feed the statistical results back to the process of generating and adjusting the prompt text, which makes it impossible for generative artificial intelligence models to continuously improve the quality and effect of the generated results in a closed loop.
[0006] Therefore, it is necessary to provide a way to: 1) Users can automatically generate prompts and obtain visual displays to drive generative artificial intelligence models simply by inputting dimensions and constituent elements; 2) Based on the user's selection operations and selection rate, automatically replace and optimize the constituent elements of the visual display; 3) The optimization results are reflected in the new prompt text, and the generative AI model is automatically instructed to regenerate the visual display. This lowers the barrier to entry for users, reduces the cost of manual trial and error, and improves the selection rate and overall effect of visual displays. Summary of the Invention
[0007] To address the aforementioned issues, this invention provides an information processing system comprising a processor. The processor is configured to: receive input from a user regarding dimensions and constituent elements; generate prompt text based on the received input to instruct a generative artificial intelligence model to generate a visual display; retrieve the visual display from the generative artificial intelligence model using the generated prompt text; and display the visual display on the user's terminal. With this configuration, the user only needs to provide dimension and constituent element information, and the processor can automatically generate prompt text suitable for the generative artificial intelligence model, thereby automatically obtaining a visual display matching the user's needs and reducing the user's reliance on the ability to write prompt text.
[0008] Furthermore, to achieve automatic optimization of visual displays based on user behavior, the processor is also configured to: monitor user selection operations; measure the selection rate based on the monitored selection operations; and generate new prompt text for automatically replacing the constituent elements of the visual displays. By monitoring user operations such as selection, clicking, or confirmation on multiple visual displays on the terminal, the processor can calculate the selection rate of each visual display and automatically generate new prompt text containing new combinations of constituent elements, thereby achieving automatic replacement and version iteration of constituent elements without manual intervention.
[0009] Furthermore, to prioritize the use of better-performing constituent elements, the processor is configured to: adjust the prompt text based on the constituent elements of visual displays with higher selection rates, prioritizing the use of these higher-selection-rate constituent elements; and instruct the generative AI model to regenerate the visual displays based on the adjusted prompt text. By prioritizing the inclusion of constituent elements such as text, images, colors, and layouts from statistically high-selection-rate visual displays in the prompt text and issuing regeneration instructions to the generative AI model, the system can continuously improve the performance of visual displays within a closed loop of generation-display-statistics-optimization, thereby effectively solving the problem in existing technologies of difficulty in timely and automatically optimizing generation results based on user behavior.
[0010] "System" refers to an overall device or combination of devices consisting of at least one processor and optional hardware and / or software modules such as memory and communication interfaces, used to perform the various functional processes described in this invention. It can be deployed on a single physical device or distributed across multiple network nodes to work collaboratively.
[0011] A “processor” is a hardware unit used to execute program instructions to perform functions such as data reception, processing, generation and control, including but not limited to a central processing unit (CPU), a graphics processing unit (GPU), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or any combination thereof.
[0012] "User" refers to the entity that uses the system and provides input to the system and / or obtains output from the system through a terminal, and may be a natural person, legal person or other organization.
[0013] "Terminal" refers to a device that allows users to operate and interact with it, and is a visual display used to send user input to the system and display system output, including but not limited to personal computers, laptops, tablets, smartphones, smart displays, or other electronic devices with display and / or input functions.
[0014] "Size" refers to the spatial range of a visual display object on the display plane as specified by the user. It usually includes numerical parameters of width and height and can be expressed in pixels, physical length or other predetermined units.
[0015] "Component elements" refer to the various contents and visual elements used to make up a visual display, including but not limited to text content (titles, subtitles, explanatory text, etc.), images, icons, backgrounds, buttons, color schemes, font styles, and one or more layout-related parameters.
[0016] "Visual display" refers to graphic or image content generated by a generative artificial intelligence model and displayed on a terminal for users to view and select, including but not limited to banner ads, promotional images, posters, interface elements, illustrations, or other image-based display content.
[0017] "Generative artificial intelligence models" refer to artificial intelligence models that can automatically generate visual displays based on input prompt text, including but not limited to deep learning-based text-to-image models, diffusion models, generative adversarial network (GAN) models or their variants.
[0018] "Prompt text" refers to textual information generated by a processor based on the size and composition of user input, used as input to a generative artificial intelligence model to instruct or constrain the model to generate visual displays that conform to the expected content and style.
[0019] "Selection action" refers to the interactive behavior of a user on a terminal towards at least one visual display object, used to express preferences or selection intentions, including but not limited to one or more of the following actions: clicking, tapping, long pressing, hovering to confirm, checking, swiping to confirm, and touch to confirm.
[0020] "Selection rate" is a statistical metric used to represent the frequency with which a user selects a visual display relative to the number of times it is shown or selectable. It can be calculated as the ratio of the number of times the visual display is selected to the number of times it is shown or selectable.
[0021] "Automatic replacement" refers to the process by which the processor modifies, replaces, or reorganizes the constituent elements used in a visual display based on statistical results such as selection rate, without requiring the user to manually specify them one by one, thereby forming a new combination of constituent elements.
[0022] "Regeneration" refers to the process where, after the prompt text has been adjusted, the processor sends a generation instruction to the generative AI model again, causing the generative AI model to generate a new visual display based on the adjusted prompt text. Attached Figure Description
[0023] Figure 1 This is a conceptual diagram illustrating an example of the configuration of the data processing system according to the first embodiment.
[0024] Figure 2 This is a conceptual diagram illustrating an example of the main functions of the data processing apparatus and smart device according to the first embodiment.
[0025] Figure 3 This is a conceptual diagram illustrating an example of the configuration of the data processing system according to the second embodiment.
[0026] Figure 4 This is a conceptual diagram illustrating an example of the main functions of the data processing device and smart glasses according to the second embodiment.
[0027] Figure 5 This is a conceptual diagram illustrating an example of the configuration of the data processing system according to the third embodiment.
[0028] Figure 6 This is a conceptual diagram illustrating an example of the main functions of the data processing device and head-mounted terminal according to the third embodiment.
[0029] Figure 7 This is a conceptual diagram illustrating an example of the configuration of the data processing system according to the fourth embodiment.
[0030] Figure 8 This is a conceptual diagram illustrating an example of the main functions of the data processing device and robot according to the fourth embodiment.
[0031] Figure 9 This represents an emotion map that maps multiple emotions.
[0032] Figure 10 This represents an emotion map that maps multiple emotions.
[0033] Figure 11 This is a sequence diagram illustrating the processing flow of the data processing system of the first embodiment.
[0034] Figure 12 This is a sequence diagram illustrating the processing flow of the data processing system in Application Example 1.
[0035] Figure 13 This is a sequence diagram illustrating the processing flow of the data processing system of the second embodiment.
[0036] Figure 14 This is a sequence diagram illustrating the processing flow of the data processing system in Application Example 2. Detailed Implementation
[0037] Hereinafter, an example of an implementation of the system according to the present disclosure will be described with reference to the accompanying drawings.
[0038] First, let me explain the terminology used in the following instructions.
[0039] In the following embodiments, the processor (hereinafter referred to as "processor") with reference numerals may be a single computing device or a combination of multiple computing devices. Furthermore, the processor may be a single computing device or a combination of multiple computing devices. Examples of computing devices include CPU (Central Processing Unit), GPU (Graphics Processing Unit), GPGPU (General-Purpose computing on Graphics Processing Units), APU (Accelerated Processing Unit), etc.
[0040] In the following embodiments, RAM (Random Access Memory), as indicated in the figures, is a memory that temporarily stores information and is used as working memory by the processor.
[0041] In the following embodiments, the memory, as indicated by the reference numerals, is one or more non-volatile storage devices that store various programs and parameters. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), disks (e.g., hard disks), or magnetic tapes.
[0042] In the following embodiments, the communication I / F (Interface) with reference numerals is an interface that includes a communication processor and an antenna, etc. The communication I / F is responsible for communication between multiple computers. As an example of a communication specification applicable to the communication I / F, wireless communication specifications such as 5G (5th Generation Mobile Communication System), Wi-Fi (wireless fidelity) (registered trademark), or Bluetooth (registered trademark) can be listed.
[0043] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B". That is, "A and / or B" means that it can be only A, only B, or a combination of A and B. Furthermore, in this specification, when "and / or" connects to express more than three items, the same interpretation as "A and / or B" applies.
[0044] First Implementation Method Figure 1 An example of the configuration of the data processing system 10 according to the first embodiment is shown.
[0045] like Figure 1 As shown, the data processing system 10 includes a data processing device 12 and an intelligent device 14. A server can be cited as an example of the data processing device 12.
[0046] The data processing apparatus 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" as understood in this disclosure. The computer 22 includes a processor 28, RAM 30, and memory 32. The processor 28, RAM 30, and memory 32 are connected to a bus 34. Furthermore, the database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0047] The smart device 14 includes a computer 36, a receiving device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, RAM 48, and memory 50. The processor 46, RAM 48, and memory 50 are connected to a bus 52. In addition, the receiving device 38, output device 40, camera 42, and communication I / F 44 are also connected to the bus 52.
[0048] The receiving device 38 includes a touchscreen 38A and a microphone 38B, and receives user input. The touchscreen 38A receives user input via touch by detecting contact with an indicator (e.g., a pen or finger). The microphone 38B receives user input via sound by detecting the user's voice. The control unit 46A in the processor 46 sends data representing the user input received by the touchscreen 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data representing the user input.
[0049] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting data in a form perceptible to the user 20 (e.g., sound and / or text). The display 40A displays visual information such as text and images according to instructions from the processor 46. The speaker 40B outputs sound according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0050] Communication I / F44 is connected to network 54. Communication I / F44 and 26 are responsible for sending and receiving various information between processor 46 and processor 28 via network 54.
[0051] Figure 2 The diagram shows an example of the main functions of the data processing device 12 and the smart device 14.
[0052] like Figure 2 As shown, in the data processing apparatus 12, specific processing is performed by the processor 28. A specific processing program 56 is stored in the memory 32. The specific processing program 56 is an example of a "program" as understood in this disclosure. The processor 28 reads the specific processing program 56 from the memory 32 and executes the read specific processing program 56 on the RAM 30. Specific processing is implemented by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0053] The memory 32 stores a data generation model 58 and an emotion-specific model 59. The data generation model 58 and the emotion-specific model 59 are used by the specific processing unit 290. The specific processing unit 290 can use the emotion-specific model 59 to infer the user's emotions and perform specific processing based on the user's emotions. In the emotion inference function (emotion-specific function) using the emotion-specific model 59, various inferences and predictions related to the user's emotions are performed, including inferences and predictions of the user's emotions, but this is not limited to this example. Furthermore, emotion inference and prediction may also include, for example, emotion analysis (parsing).
[0054] In the smart device 14, the processor 46 performs the acceptance output processing. The memory 50 stores the acceptance output program 60. The acceptance output program 60 is used in conjunction with the data processing system 10 and the specific processing program 56. The processor 46 reads the acceptance output program 60 from the memory 50 and executes the read acceptance output program 60 on the RAM 48. The acceptance output processing is implemented by the processor 46 acting as the control unit 46A according to the acceptance output program 60 executed on the RAM 48. Furthermore, the smart device 14 has the same data generation model and emotion-specific model as the data generation model 58 and the emotion-specific model 59, and these models can also be used to perform the same processing as the specific processing unit 290. The acceptance output processing is implemented by the processor 46 acting as the control unit 46A according to the acceptance output program 60 executed on the RAM 48.
[0055] Alternatively, other devices besides the data processing device 12 may also have the data generation model 58. For example, a server device (e.g., a generation server) may have the data generation model 58. In this case, the data processing device 12 obtains the processing results (prediction results, etc.) using the data generation model 58 by communicating with the server device that has the data generation model 58. Furthermore, the data processing device 12 may be a server device or a user-held terminal device (e.g., a mobile phone, robot, home appliance, etc.). Next, an example of the processing of the data processing system 10 of the first embodiment will be described.
[0056] Example 1 The flow of a specific process in Example 1 will be described. Each part of the system described below is implemented by the data processing device 12 and the smart device 14. Furthermore, the data processing device 12 is referred to as the "server," and the smart device 14 is referred to as the "terminal."
[0057] With the widespread application of generative artificial intelligence models in the field of visual content generation, technologies that automatically generate visual displays such as images and banners based on natural language prompts are gradually becoming more common. However, existing technologies still have the following shortcomings at the computer technology level.
[0058] First, servers typically receive descriptions input by users in natural language and forward them to generative AI models. There is a lack of a mechanism to programmatically convert users' structured needs (such as size information and composition information) into high-quality prompts. This results in prompts that are not precise enough and have poor consistency. As a result, generative AI models have difficulty accurately constraining parameters such as size, layout, and color during the reasoning process, thereby reducing the controllability and stability of the generated results and increasing the interaction cost and computational resource consumption of users through multiple trials and errors.
[0059] Second, most existing systems treat generative AI models as black boxes, performing only a one-time generation and lacking the ability to provide closed-loop feedback and automatic optimization on the server side based on user interaction behavior (such as selection actions). Specifically, servers typically do not collect user preferences for multiple generated results from the terminal, nor do they quantitatively evaluate visual displays internally using statistical indicators such as selection rates and adjust prompts accordingly. This results in subsequent generation failing to fully utilize historical interaction data for adaptive optimization, leading to low efficiency in computing resource utilization.
[0060] Third, in existing technologies, server management of prompt statements is often static and decentralized. Servers neither explicitly provide the actual prompt statements used to invoke generative AI models to the terminal, nor do they possess a general processing flow for dynamically rewriting prompt statements based on selectivity. This makes it difficult to debug and reuse the prompt statement generation logic, and hinders the server from developing an internally evolving prompt statement generation strategy, thus limiting the scalability and maintainability of the entire platform in terms of multi-round generation and automated optimization.
[0061] Fourth, from a system architecture perspective, traditional solutions are mostly simple pipelines of "terminal input—server forwarding—model generation—terminal display," where the server mainly acts as a data channel without organically integrating server-side processing such as structured data construction, prompt generation, result storage, selection rate statistics, and regeneration control. This architecture cannot fully utilize the server's computing and storage capabilities to improve the end-to-end generation process, making it difficult to achieve a systematic improvement in overall generation quality and efficiency.
[0062] Therefore, it is necessary to provide a new system that introduces an automatic prompt generation mechanism based on structured data, a dynamic adjustment mechanism for compositional information based on selection rate, and an automatic optimization mechanism for multi-round regeneration into the server. This will improve the prompt generation process and the calling process of generative artificial intelligence models in a programmatic way within the computer, thereby achieving a comprehensive improvement in the quality of visual display generation, user interaction efficiency, and computing resource utilization efficiency.
[0063] The specific processing performed by the specific processing unit 290 of the data processing apparatus 12 in Embodiment 1 is achieved by the following means.
[0064] In this invention, the server includes means for obtaining size and composition information from a user via an input device; means for generating structured data containing the size and composition information and automatically generating prompt statements for a generative artificial intelligence model based on the structured data; means for sending generation instruction information containing the prompt statements and size information to the generative artificial intelligence model via a communication device and obtaining visual information generated in response to the prompt statements from the generative artificial intelligence model; means for storing the visual information in a storage device or generating reference information about the visual information; means for sending the visual information or the reference information to a terminal device via a communication device and outputting display control information for displaying the visual information on a display device of the terminal device; and means for sending the prompt statements to the terminal device and outputting information for displaying the prompt statements on the terminal device. This allows for structured management of size and composition information from the user on the server side, automatic construction of high-quality prompt statements adapted to the generative artificial intelligence model, and provision of the generation results and prompt statements to the terminal, thereby improving the consistency and controllability of the prompt statement expression, reducing the number of trial and error attempts by the user, and enhancing the interpretability and maintainability of the visual information generation process.
[0065] In this invention, the server further includes means for acquiring user selection operation information for multiple visual information displayed on a terminal device via a communication device and calculating the selection rate of each visual information based on the selection operation information; means for generating modification information to change the constituent information based on the selection rate and automatically generating new prompt statements based on the modified constituent information and the size information; and means for issuing visual information regeneration instructions to the generative artificial intelligence model using the new prompt statements. Thus, the performance of each visual information can be automatically quantified within the server based on user selection behavior, the programmatic adjustment of constituent information can be driven by the selection rate, and the prompt statements can be dynamically rewritten accordingly, achieving closed-loop control of the generative artificial intelligence model. This allows for continuous optimization of the generation results through server-side software logic without changing the terminal hardware structure.
[0066] In this invention, the server further includes means for adjusting the content of the prompt statement based on the compositional information contained in visual information where a specific selection rate meets predetermined conditions, so as to preferentially use the specific compositional information; and means for inputting the adjusted prompt statement and the size information into the generative artificial intelligence model and repeatedly performing regeneration processing to automatically optimize the compositional information of the visual information. This allows the server to form a compositional information optimization mechanism based on statistical indicators, continuously strengthening the weight of high-selection-rate compositional information in the prompt statement through multiple rounds of regeneration, achieving automated iterative optimization of the composition of visual information, and improving the adaptability, generation quality, and resource utilization efficiency of the visual content generation system from a computer technology perspective.
[0067] A "system" refers to an entire system consisting of one or more devices, components, or programs that performs information processing electronically to achieve the automatic generation and optimization of visual information.
[0068] A "server" refers to an electronic device or collection of electronic devices equipped with a processor, storage device, and communication device, used to perform tasks such as generating prompts, calling generative artificial intelligence models, managing visual information, and interacting with terminal devices.
[0069] "Terminal device" refers to an electronic device used for user input and displaying visual information and prompts returned by a server, including but not limited to smartphones, tablets, portable computing devices or desktop computing devices.
[0070] "Input device" refers to a hardware or software component used to receive size and composition information input by a user, including but not limited to keyboards, touch screens, mice, and input controls based on graphical user interfaces.
[0071] "Communication device" refers to a network interface or communication module used for sending and receiving data between a server and a terminal device, or between a server and a generative artificial intelligence model, including but not limited to wired or wireless network interface devices.
[0072] A "processor" is a computing unit that can execute program instructions to perform information processing functions, including but not limited to a central processing unit, a graphics processing unit, or other logical operation units.
[0073] "Storage device" means a storage medium used to store programs, structured data, prompts, visual information and references, including but not limited to semiconductor memory, magnetic storage medium or optical storage medium.
[0074] "Display device" refers to an output device used to display visual information and prompts to users, including but not limited to liquid crystal displays, organic light-emitting displays, or other image display devices.
[0075] "Size information" refers to a set of parameters used to define the spatial extent of visual information during generation, including at least width and height information, and may include information such as resolution or scale.
[0076] "Compositional information" refers to a set of parameters used to describe the internal components and attributes of visual information, including but not limited to background color, text content, text color, font size, layout position, image area or other visual elements.
[0077] "Structured data" refers to a combination of size and composition information represented by a predetermined data structure, which is suitable for parsing, processing and transformation by a program.
[0078] "Generative artificial intelligence models" refer to machine learning models that automatically generate visual information based on input prompts and optional parameters, including but not limited to deep learning-based image generation models or diffusion models.
[0079] "Prompt statements" refer to natural language or formal text information used to describe the content, style, size, and composition requirements of the visual information to be generated to a generative artificial intelligence model.
[0080] "Generation instruction information" refers to a data set containing prompts and parameters (including size information and / or other control parameters) related to the generation process, used to instruct generative artificial intelligence models to perform visual information generation processing.
[0081] "Visual information" refers to image-based data generated by generative artificial intelligence models based on prompts, including but not limited to advertising banners, icons, illustrations, or other digital visual content.
[0082] "Referencing information" refers to identification data used to uniquely identify and access visual information, including but not limited to file paths, Uniform Resource Locators, or identifiers.
[0083] "Display control information" refers to parameters or instructions used to control the display device of the terminal device to display visual information, including but not limited to information such as display position, display size or display mode.
[0084] "User selection operation information" refers to interactive data that reflects the user's selection behavior in response to multiple visual information displayed on the terminal device, including but not limited to information such as clicks, touches, selection marks, or confirmation operations.
[0085] "Selection rate" refers to a statistical indicator calculated based on user selection operation information, representing the frequency or proportion of a certain visual information being selected by a user among multiple visual information.
[0086] "Change information" refers to control data used to indicate adjustments such as modification, replacement, addition, and deletion of existing constituent information.
[0087] "New prompt statements" refer to prompt statements that are regenerated based on the changed composition and size information, and are used to call the generative artificial intelligence model again to generate visual information.
[0088] "Regeneration instruction" refers to issuing a request to a generative artificial intelligence model to regenerate visual information using new prompts.
[0089] "Predetermined conditions" refer to the thresholds or rules used to determine whether the selection rate has reached the priority adoption standard, including but not limited to the lower limit of the selection rate, ranking conditions, or statistical stability conditions.
[0090] "Automatic optimization" refers to the process by which the server, based on the relationship between selectivity, constituent information, and prompt statements, repeatedly adjusts and regenerates visual information to improve the overall performance of visual information without requiring manual intervention.
[0091] This invention, in conjunction with the technical solutions defined in the appendix, describes a visual information generation and optimization system based on the collaborative work of a server, a terminal, and a generative artificial intelligence model, and its specific implementation. This description of the embodiments does not limit the scope of protection of this invention, but rather serves to illustrate how this invention can be implemented in specific hardware and software environments.
[0092] I. Overall System Composition The server, as the main processing unit of this invention, includes: a processor (e.g., a multi-core central processing unit and an optional graphics processing unit), a storage device (e.g., solid-state memory, random access memory), a communication device (e.g., an Ethernet interface, a wireless network interface), and a program that executes the various functional modules of this invention. The server can run a general-purpose operating system, such as a server operating system, and deploy a network service framework on it, such as a network framework based on an interpreted language.
[0093] As the primary means of user interaction, a terminal can be a smartphone, tablet, portable computing device, or desktop computing device. A terminal typically includes a display device (such as an LCD screen), an input device (such as a touchscreen, keyboard, or mouse), and a network communication module. The terminal can run a web browser or a dedicated client application.
[0094] Users, as the main users of the system, provide visual information requests to the server through the input device on the terminal, and view the visual information generated by the server based on the generative artificial intelligence model through the display device on the terminal.
[0095] Generative AI models are deployed in a computing environment with servers or provided by external inference services over a network. Servers can use diffusion-based image generation models based on deep learning frameworks (such as tensor computation frameworks), or image generation models implemented using structures such as generative adversarial networks and variational autoencoders.
[0096] II. Program Module Composition and Data Structure The server stores programs for implementing multiple functional modules of the present invention in a storage device. The server executes these programs through a processor to implement the following functional modules (the following names are for illustrative purposes; in actual implementation, different names or combined / split structures may be used): 1. Size and Composition Information Receiving Module The server receives user input data sent by the terminal via a communication device. The server parses the received data into size and composition information. The server can use the following data structure for internal representation: The server represents the size information as a set of parameters including width, height, and optional resolution; the server represents the composition information as a set of parameters including fields such as background color, text content, text color, font size, text position, layout style, and image area description. The server combines the above information into structured data, such as a key-value pair data object.
[0097] 2. Prompt Statement Generation Module The server reads size and composition information from structured data and automatically generates prompts for calling generative artificial intelligence models based on predefined template rules. Instead of directly relying on user natural language input, the server uses a programmatic generation method to concatenate and convert structured parameters into natural language descriptions, thereby improving the consistency and accuracy of the prompts.
[0098] The server can use various templates, such as: "Please generate a visual display with dimensions of 300x250 pixels, a red background, and the words 'On Sale!' displayed in white text in the center of the image. The font size should be approximately 16 pixels. The overall style should be simple and clear, suitable for use as a web advertisement." "Please generate a visual display with a size of 1080x1920 pixels to be used as a pop-up in the mobile application. The background should be a blue gradient, and the text 'New version is online!' should be displayed in large white text centered at the top. Place a simple gift box icon below the text. The overall style should be modern and clean." "Please generate a visual display with a size of 1080x1080 pixels for social media promotion. The background should be light green, and the text '20% off for a limited time today' should be displayed in bold white text in the upper center. Leave space below to display product images. The overall layout should be simple and have an e-commerce promotional atmosphere." The server generates prompts by filling in parameters such as size, color, text, and position into templates, thereby achieving fine-grained control over generative artificial intelligence models. This structured-to-natural-language mapping process is procedural, unlike the traditional method where users write natural-language prompts themselves. This reduces semantic ambiguity and improves the predictability of the generated results.
[0099] 3. Command Information Generation Module After generating the prompt statement, the server further constructs the generation instruction information. In addition to the prompt statement, the server's generation instruction information also includes control parameters such as image width, height, sampling steps, and random seed. These parameters are used to directly control the inference behavior of the generative artificial intelligence model, for example: The server can specify the number of denoising steps for the generative diffusion model (e.g., 20 steps, 50 steps), the classification degrees of freedom parameter (a numerical parameter that controls the strength of the text constraint), and a random seed (so that the same results can be reproduced when needed).
[0100] The server generates instruction information and passes it to the generative artificial intelligence model interface in the form of data objects.
[0101] 4. Generative Artificial Intelligence Model Calling Module The server uses a processor and graphics processing unit to load and execute generative artificial intelligence models. In a typical implementation, the server uses a diffusion-based image generation network, which includes: The server uses a text encoding subnetwork to input the prompt statement into a text encoder (e.g., a text embedding model based on a multi-layer self-attention structure) to generate a fixed-length text feature vector. The server uses an image decoding subnetwork to perform multi-step iterative denoising on the initial random noise image. At each step, the feature distribution of the noise image is adjusted according to the text feature vector, so that the image gradually approaches the content described by the prompt statement.
[0102] The server uses a large amount of image data with text descriptions when training this generative AI model. During training, the server defines loss functions, such as noise prediction loss based on mean squared error and text-image consistency loss, and updates the model weights using gradient descent algorithms (e.g., adaptive learning rate optimization algorithms). The server can employ data augmentation methods (image scaling, cropping, color dithering, etc.) to improve the model's adaptability to different sizes and styles.
[0103] During the inference phase, the server sets the model input based on the prompts and size information in the generated instructions: the server sets the model output resolution according to the width and height parameters, obtains text features based on the prompts, and generates initial noise based on a random seed. The server performs multiple forward propagation operations on the graphics processing unit, performing convolution operations, nonlinear transformations, and attention calculations on the feature map in each denoising step, ultimately outputting visual information that meets the size and composition requirements.
[0104] This iterative denoising process based on a diffusion model allows the generation process to gradually approximate a high-quality solution numerically, rather than simply relying on a single sampling. Therefore, it can achieve more stable image quality with the same computing resources. By using size information as a mandatory constraint input, the server can reduce the extra computations required for cropping and scaling after generation, thereby improving overall processing efficiency.
[0105] 5. Visual information storage and reference information generation module After receiving visual information from the generative artificial intelligence model, the server temporarily stores the image data in memory and may selectively store it in persistent storage. The server can use image compression formats to save the images as bitmap or vector graphics files. The server assigns a unique identifier to each piece of visual information and constructs referencing information, such as a Uniform Resource Locator (URL), based on the storage path.
[0106] The server manages visual information uniformly in this way, which is beneficial for subsequent reuse, statistical analysis, and cache control. Through the reference information mechanism, the terminal does not need to transmit the entire image via data stream each time; instead, it can obtain the image on demand via network address, thereby reducing communication load in certain implementations.
[0107] 6. Visual Information and Prompt Output Module The server sends visual information or its referenced information to the terminal via a communication device. Simultaneously, the server also sends the actual prompts used to generate the message to the terminal. The server can add display control information to the response data, such as suggested display size, position, and scaling strategies, to guide the terminal's display control.
[0108] After receiving the server's response, the terminal displays visual information on the display device and can show corresponding prompts on the interface, allowing the user to clearly understand the correspondence between the generated result and the prompts. Users can then modify the prompts accordingly, and the server, in subsequent calls to the prompt generation module, updates the system based on structured data rather than user-defined text, thus maintaining internal logical consistency.
[0109] 7. User selection operation information collection and selection rate calculation module When the terminal displays multiple visual pieces of information on the interface, it accepts user selections, such as clicking on a specific image or selecting a preference from multiple images. The terminal converts these actions into user selection information and sends it to the server via a communication device.
[0110] After receiving the user's selection information, the server counts the number of times each visual element is selected and displayed, calculating the selection rate as a numerical metric. The server can maintain a selection rate statistics table in the storage device, associating the identifier of the visual element with the corresponding selection rate, timestamp, and other information.
[0111] The server uses this selection rate based on numerical statistics, rather than relying solely on individual user feedback, to objectively quantify the quality of visual information. This differs from traditional design adjustments based solely on editors' subjective judgment and is part of the computer's internal data processing and optimization.
[0112] 8. Module for generating new prompts and information changes After obtaining the selection rates from multiple visual information sets, the server analyzes which constituent elements significantly influence the selection rates. The server can employ a rule-driven processing approach, for example: The server can treat background color, text content, font size, layout position, etc. as different constituent variables. The server can generate change information according to preset rules (for example, if the selection rate of a certain color combination is consistently higher than a threshold, the probability of using that color in subsequent prompts will be increased; if the selection rate of a certain text length is low, the text will be shortened or the wording will be adjusted).
[0113] After generating change information, the server updates the composition information fields stored in the structured data. For example, it might change the background color from "red" to "blue," or the text content from "On Sale!" to "Limited Time Discount!". The server then calls the prompt generation module again to automatically generate a new prompt based on the updated composition information and existing size information.
[0114] For example, when the server finds that the selection rate of "red background, white text 'On Sale!'" is lower than that of "blue background, white text 'Limited-Time Discount!'", the server can use the latter's composition information as the preferred composition and prioritize the combination of "blue background" and "Limited-Time Discount!" in the new prompt statement. An example of the new prompt statement generated by the server is as follows: "Please generate a visual display with dimensions of 300x250 pixels, a blue background, and the words 'Limited-time discount!' displayed in white text in the center of the image. The font size should be approximately 16 pixels. The overall style should be eye-catching but not overly glaring, suitable for use as a web advertisement." This process demonstrates the technical features of the present invention: the server does not simply repeat the user's original input, but rather, based on the statistical indicator of selection rate, it programmatically adjusts the constituent information and automatically rewrites the prompt statements. This data-driven prompt statement rewriting forms a self-optimizing mechanism within the computer.
[0115] 9. Regeneration and Automatic Optimization Module After generating a new prompt, the server regenerates the message by using the instruction information building module and the generative artificial intelligence model calling module. The server can repeat this process periodically or under specific triggering conditions, allowing the compositional information of the visual information to gradually converge towards a high-selectivity pattern.
[0116] In multiple rounds of regeneration, the server can employ unconventional techniques, such as: The server can assign different random seeds for each round of generation to increase sample diversity and prevent overfitting; the server can fine-tune the tone and number of adjectives in the prompts to control the constraint of the text description on the model; the server can assign weights to different components based on the selection rate, and indirectly affect the generation results by increasing or decreasing the descriptive intensity of certain elements in the prompts.
[0117] The server automatically optimizes the composition of visual information through multiple rounds of iteration and weight adjustments. Because the server uses the selection rate as the optimization target and internally analyzes the relationship between the composition information and the selection rate through algorithmic logic, its processing method differs from traditional human-based graph design. The server uses a unified data structure and modular programs, enabling the optimization process to run stably in scenarios with large-scale data and high-frequency interactions, thereby improving generation efficiency and result quality at the computer technology level.
[0118] III. Explanation of Technical Effects and Causal Relationship By converting user input into structured data and then programmatically generating prompts based on this data, the server can significantly reduce semantic uncertainty and lower input noise during inference in generative AI models compared to the traditional method of users directly writing natural language prompts. Since size and composition information are explicitly passed to the model through structured fields, the server can obtain images that meet size requirements without relying on post-processing cropping or scaling, reducing additional image transformation operations and directly improving processing speed and resource utilization efficiency.
[0119] The server generates a selection rate metric by statistically analyzing user action choices over a long period. Based on this selection rate, it dynamically adjusts the constituent information and prompts, creating an optimization process within the computer similar to "online learning." Although the model's weights are no longer updated during the inference phase, the server achieves system-level adaptive performance improvements by altering the model inputs (prompts and parameters). This structured feedback and regeneration mechanism allows the system to continuously improve the relevance and appeal of the generated results through server-side software logic without changing the terminal or user's usage patterns.
[0120] During the rewriting of prompt statements, the server does not simply copy the user's language. Instead, it transforms the prompts using a set of rules based on structured information and statistical indicators. These rules include prioritizing key elements, weakening or eliminating inefficient elements, and standardizing the description order. These rules are implemented in the program through conditional statements, weight adjustments, and template selection, forming a non-inertial processing path that differs from manual design processes. This path leverages the computer's rapid statistical, combinatorial, and iterative capabilities to achieve higher stability and consistency in large-scale visual information generation tasks.
[0121] When the server employs a diffusion-based generative model, a balance can be struck between quality and computational cost by setting the number of sampling steps, text constraint strength, and random seed parameters. During multiple rounds of generation and optimization, the server can automatically adjust the number of sampling steps based on the selection rate (e.g., using fewer steps for well-performing combinations to accelerate generation, and more steps for exploratory combinations to improve quality), thereby further improving overall inference efficiency. This computational resource allocation strategy based on statistical feedback is also a key factor contributing to the technical effectiveness of this invention.
[0122] IV. Other Implementation Forms and Variations In one implementation, the server can deploy the generative AI model on a dedicated inference device and communicate with the main business server via a high-speed network interface. The server can package prompts and parameters into a remote call request, which the inference device then uses to perform forward computation on the model and returns the generated visual information to the server. This can improve inference throughput at the physical level and adapt to high-concurrency scenarios.
[0123] In another implementation, the server can introduce a multi-model integration mechanism. This involves calling multiple generative AI models with different structures (such as diffusion models and generative adversarial network models) for the same prompt, and then using the server's selection rate statistics to select the better-performing model or combination result. In subsequent generation, the server can prioritize using models with higher reliability, thereby further improving generation stability at the system level.
[0124] In another implementation, the server can extract features from visual information, such as using an image encoding network to extract visual feature vectors, and store these feature vectors together with the prompt statement feature vectors and the selection rate. The server can analyze which visual features are associated with high selection rates through clustering or similarity search, and based on this, actively guide the generative AI model to move closer to these feature spaces when generating new prompt statements, achieving more fine-grained automatic optimization.
[0125] The terminal, in different implementations, can be a browser, desktop application, or mobile application, but its core functions all include: sending size and composition information to the server, receiving and displaying visual information and prompts, and collecting user selection information. Users only need to perform simple input and selection operations through the terminal, and the server can complete complex statistical, analytical, and regeneration processes in the background, thereby transferring complex calculation and optimization tasks from the user side to the server side. This achieves the computer technology improvement effect described in this invention without increasing the user's burden.
[0126] use Figure 11 The processing flow is explained.
[0127] Step 1: Users input size and composition information on the terminal and submit a generation request.
[0128] In the graphical interface of the terminal, users input the size information (such as width and height) and composition information (such as background color, text content, text color, font size, text position, layout style, etc.) of the visual display object through input boxes, drop-down menus, color pickers, and other controls, and then click the "Generate" button.
[0129] Input: The size and composition information entered by the user in the terminal interface.
[0130] Output: Size and composition information displayed internally by the terminal.
[0131] Step 2: The terminal converts user input into structured data and sends it to the server.
[0132] The terminal reads various fields of user input from the interface controls, encapsulates width, height, background color, text content, etc., into structured data objects, and sends them to the server's designated interface in the form of request messages via the network communication module. Before sending, the terminal can perform basic data type validation (e.g., whether the width and height are numerical values).
[0133] Input: The size and composition information entered and confirmed by the user in step 1.
[0134] Output: A request message sent to the server over the network, with the message body containing structured size and composition information.
[0135] Step 3: The server parses the request message and performs parameter validity checks.
[0136] The server receives request messages from the terminal via a communication device, extracts size and composition information from the messages using a parsing module, and verifies whether the width and height are positive integers, whether the color values are valid, and whether the text length is within a preset range. Internally, the server performs data operations such as conditional checks and range comparisons to filter out illegal or abnormal parameters.
[0137] Input: A request message for structured size information and composition information from the terminal.
[0138] Output: Verified valid size and composition information, or error information generated in abnormal situations.
[0139] Step 4: The server converts the valid parameters into structured data objects, which are then used as input for generating the prompt statements.
[0140] After the parameters pass validation, the server stores fields such as width, height, background color, text content, font size, and position into an internal data structure, such as a key-value pair mapping or a record object. The server then normalizes these fields (e.g., standardizes units and color names) for subsequent template population. This normalization process includes string manipulation, unit conversion, and other data processing.
[0141] Input: Verified original size information and composition information fields.
[0142] Output: A normalized, structured data object suitable for use by the prompt statement generation module.
[0143] Step 5: The server automatically generates prompts based on structured data.
[0144] The server reads parameters such as width, height, background color, and text content from structured data objects based on predefined prompt templates, and inserts them into reserved positions within the template to form a natural language description. For example, after the server fills in "width=300, height=250, background=red, text=Promotional!" into the template, it generates the following prompt: "Please generate a visual display with a size of 300x250 pixels, a red background, and white text 'Promotional!' displayed in the center of the image. The font size is approximately 16 pixels, and the overall style is simple and clear, suitable for web page advertising." The server performs specific operations such as string concatenation and placeholder replacement during this process.
[0145] Input: Normalized structured data object (size information and composition information).
[0146] Output: Natural language prompts used to invoke generative artificial intelligence models.
[0147] Step 6: The server builds and generates instruction information and prepares to call the generative artificial intelligence model.
[0148] After receiving the prompt, the server packages the prompt along with control parameters such as width, height, number of sampling steps, random seed, and text constraint strength into a generation instruction. The server then limits the range of numerical parameters and fills them with default values, for example, setting them to preset values when the user does not specify the number of sampling steps. Through this data processing, the server forms a complete set of model input parameters.
[0149] Input: Prompt statements and internally preset model control parameters and dimensional information.
[0150] Output: A data structure containing the generation instructions that serve as input to the generative artificial intelligence model.
[0151] Step 7: The server invokes a generative artificial intelligence model and performs image generation inference.
[0152] The server transmits the generated instruction information to the generative AI model interface, and the model performs forward inference on the server's graphics processing unit. Internally, the server first calls a text encoding sub-network to convert the prompt into a text feature vector; then, it initializes a random noise image tensor based on width and height parameters; next, it performs multi-step iterative denoising through a diffusion network or other generative networks, performing convolution, normalization, and attention operations at each step, and adjusting image features according to the text feature vector. By controlling the number of sampling steps and the random seed, the server ensures that the inference process converges to a high-quality image within a finite time.
[0153] Input: Generate instruction information (prompt statements, dimensions, sampling parameters, etc.).
[0154] Output: Visual information data (image tensors or image file data) generated by a generative artificial intelligence model.
[0155] Step 8: The server encodes and stores the generated visual information and generates reference information.
[0156] After obtaining visual information data, the server performs necessary encoding and compression on the image, such as encoding it into an image file format. The server writes the encoded image data to a storage device, generating a corresponding storage path or unique identifier. Based on the storage path, the server constructs reference information, such as a network access address. During this process, the server performs specific operations such as file writing, path concatenation, and identifier generation.
[0157] Input: Raw visual information data from a generative artificial intelligence model.
[0158] Output: Visual information files stored in a storage device and their corresponding reference information (such as Uniform Resource Locators or identifiers).
[0159] Step 9: The server constructs a response message and sends visual information and prompts to the terminal.
[0160] Depending on the system configuration, the server chooses to send either the image data itself or only the reference information to the terminal. The server encapsulates the visual information or its reference information, prompts, and display control parameters (such as recommended display size) into a response message. The server then sends this response message to the terminal via a communication device. During this process, the server performs specific operations such as data packaging, serialization, and network transmission.
[0161] Input: Stored visual information or its references, along with corresponding prompts and display control information.
[0162] Output: Response messages sent to the terminal over the network (including visual information / reference information and prompts).
[0163] Step 10: The terminal receives the response message and parses the visual information and prompts.
[0164] The terminal receives response messages from the server via the network communication module, uses parsing logic to extract image data or reference information and prompts from the messages, and converts them into display data structures that the terminal can directly use. For example, when the terminal obtains that the reference information is a network address, it will arrange to load the image from that address; after obtaining the prompts, it will store them in the interface state data.
[0165] Input: Response messages from the server (visual information / reference information and prompts).
[0166] Output: Image resource references and prompt text that can be displayed internally in the terminal.
[0167] Step 11: The terminal presents visual information and displays prompts on the display device.
[0168] Based on the parsing results, the terminal creates or updates display controls on the interface, sets the display source of the image controls to the address corresponding to the image data or reference information, and sets their display size and layout. Simultaneously, the terminal displays the actual prompts used by the server in the text area, allowing the user to compare the results with the prompts. During this process, the terminal performs specific operations such as interface layout calculations, image loading and rendering, and text drawing.
[0169] Input: Image data or reference information for display, and prompt text.
[0170] Output: Visual information displayed on the terminal display device and corresponding prompts.
[0171] Step 12: Users can view and select from multiple visual information options on the terminal.
[0172] Users compare multiple visual pieces of information generated by the server on a terminal display device. Based on personal preferences and task objectives, they click or select one or more visual pieces of information as the "optimal result" or "candidate solution." Users can make selections using controls such as buttons, checkboxes, and radio buttons.
[0173] Input: Multiple visual information and prompts displayed on the terminal interface.
[0174] Output: Selection actions generated by user input (e.g., click events, selection markers).
[0175] Step 13: The terminal packages the user's selected operation information and sends it to the server.
[0176] After the terminal captures a user's selection event for one or more visual information items, it reads the identifier of the corresponding visual information from its internal state and combines it with the selection action type (e.g., selected) to form the user's selection operation information. The terminal organizes multiple selections and can count the images selected in each display; then it sends this selection operation information to the server via the network.
[0177] Input: User selection events on the terminal interface (associated with specific visual information).
[0178] Output: User selection information sent to the server (including visual information identifiers and selection results).
[0179] Step 14: The server receives the user's selected operation information and calculates the selection rate for each visual information.
[0180] After receiving the selection operation information sent by the terminal from the communication device, the server reads the identifier and corresponding selection result of each visual information, and updates the number of times the visual information was displayed and selected in the storage device. The server performs a selection rate calculation for each visual information, that is, divides the "number of times selected" by the "number of times displayed", and obtains a statistical value between 0 and 1. The server can smooth the selection rate according to the time window or the sample size to reduce the impact of random fluctuations.
[0181] Input: User selection operation information (visual information identifier and selection result).
[0182] Output: Selectivity data (numerical statistics) corresponding to each visual information.
[0183] Step 15: The server analyzes the composition information based on selectivity and generates change information.
[0184] The server correlates selectivity with corresponding component information, analyzing which background colors, text content, font sizes, or layouts are statistically associated with higher selectivity. The server can use rules or simple algorithms (such as comparing high and low thresholds, sorting the top few) to determine "high-selectivity components" and "low-selectivity components." Based on the analysis results, the server generates change information, such as instructing the replacement of certain colors with high-selectivity colors, or shortening lengthy text into concise copy.
[0185] Input: Selectivity data for each visual information and stored composition information.
[0186] Output: Change information for the adjustments made to the constituent information (including the fields to be modified and the target values).
[0187] Step 16: The server updates the configuration information based on the changed information and generates a new prompt statement.
[0188] The server applies the change information obtained in step 15 to update the existing structured information, such as modifying the background color field, replacing the text content field, or adjusting the font size field. After completing the update, the server re-enters the updated structured data into the prompt generation module, and regenerates a new prompt by filling in the template. For example, changing "red background, 'On Sale!'" to "blue background, 'Limited-Time Discount!'" generates a new prompt: "Please generate a visual display with a size of 300x250 pixels, a blue background, and white text displaying 'Limited-Time Discount!' in the center of the image, with a font size of approximately 16 pixels. The overall style is eye-catching but not overly glaring, suitable for web page advertising." Input: Existing structured information and changes.
[0189] Output: Updated composition information and new prompts generated based on the updated composition information.
[0190] Step 17: The server then invokes the generative artificial intelligence model again to regenerate based on the new prompt.
[0191] The server repeats steps 6 and 7, packaging the new prompts along with the existing or adjusted size information and model control parameters into new generation instructions, and then calls the generative AI model again. The server performs the same text encoding, noise initialization, and multi-step iterative denoising operations to obtain new visual information based on the optimized composition information.
[0192] Input: New prompts, size information, and control parameters.
[0193] Output: A new batch of visual information data generated based on the new prompt statement.
[0194] Step 18: The server provides the newly generated visual information to the terminal, which then displays it to the user and collects feedback, thus forming a multi-round optimization cycle.
[0195] The server stores the new visual information and constructs a response message according to steps 8 and 9, sending the new visual information and corresponding prompts to the terminal. The terminal parses and displays this new visual information, and the user performs another viewing and selection operation. The terminal then sends the new selection operation information to the server. The server continuously repeats the process from steps 14 to 17, constantly adjusting the constructed information and prompts based on the selection rate in multiple loops, gradually improving the performance of the visual information.
[0196] Input: The newly generated visual information and prompts, as well as subsequent user selection information.
[0197] Output: Optimized new visual information updated and displayed on the terminal display device, and statistical results that gradually approach the high-selectivity composition pattern within the server.
[0198] Application Example 1 The process flow corresponding to the specific processing in Use Case 1 will be described below. The various parts of the system described below are implemented by the data processing device 12 and the intelligent device 14. Furthermore, the data processing device 12 is referred to as the "server" and the intelligent device 14 is referred to as the "terminal".
[0199] With the widespread application of generative artificial intelligence models in image generation, the commonly used approach in existing technologies is for users to directly input abstract text descriptions, and the server to input this text as a single instruction into the generative artificial intelligence model to generate visual displays. However, this approach has the following technical problems: First, the natural language descriptions input by users often lack structure and parameterization. After receiving the description, the server cannot accurately split and encode different types of data such as size information, text information, color information, and image information into conditional vectors that the model can use. This makes it difficult for generative artificial intelligence models to stably output visual display data that meets specific size and layout constraints, thereby reducing the consistency and controllability of the generated results.
[0200] Second, existing systems typically treat generative AI models as "black boxes," outputting image results only once and lacking mechanisms to utilize subsequent editing and selection operations performed by users on the terminal. In other words, the server does not extract the arrangement parameters and preference information of text, color, and image elements from user interactions, and cannot accumulate, summarize, and constrain these parameters in the next round of prompt generation and model invocation. This makes it difficult for the system to form a "learnable" and "memorable" generation process. Each time a user generates an image, it's as if they are setting conditions from scratch, resulting in high usage costs and wasted interactive feedback.
[0201] Third, existing technologies commonly involve simple image overlay or annotation on the terminal side, while the generative AI model on the server side is independent of the editing logic on the terminal side. The server does not dynamically adjust prompts and conditional vectors based on the terminal's editing results. This loose coupling prevents the system from performing structured parsing and parameter extraction of the editing results on the server side, and further prevents the system from using these parameters to constrain and optimize the generative AI model. Consequently, it cannot achieve a continuous, multi-round, progressively optimized visual display generation process, limiting the computer system's adaptive capabilities and overall performance in image generation tasks.
[0202] Fourth, in traditional solutions, judging the quality of different generated results often relies on subjective human judgment. There is a lack of a system-level mechanism to utilize interactive data such as selection and editing operations to form evaluation indicators, thereby driving the server to automatically correct prompts and optimize the information. This prevents the computer system from effectively utilizing large amounts of user interaction data to perform closed-loop optimization of the generation process, making it difficult to improve generation efficiency and quality.
[0203] Therefore, an improved computer implementation is needed to enable the server to: (1) Transform the size and composition information from the user terminal into structured prompt statements, and further decompose them into condition vectors for model reasoning; (2) During the inference process, which includes matrix operations, convolution operations and nonlinear transformations, ensure that the generated visual display data is consistent with the target size and constraints; (3) Transform the selection and editing operations on the user terminal into measurable evaluation indicators and layout parameters, and continuously feed them back into the generation of prompt statements and model invocation; (4) This enables the controllable invocation and adaptive optimization of generative artificial intelligence models at the system level, thereby improving the overall technical performance of computers in image generation and interactive editing tasks.
[0204] The specific processing performed by the specific processing unit 290 of the data processing apparatus 12 in Application Example 1 is achieved by the following means.
[0205] In this invention, the server includes: a processing unit for receiving input information containing size information and composition information from a user terminal; a processing unit for generating a prompt statement in natural language, based on the size information and the composition information, for issuing instructions to a generative artificial intelligence model to generate a visual display; a processing unit for decomposing the prompt statement, extracting text information, color information, image information, and numerical information representing size information, and encoding the information into a condition vector of the generative artificial intelligence model; and a processing unit for performing inference processing, including matrix operations, convolution operations, and nonlinear transformations, on a computing device using the condition vector as input, thereby generating a visual display. The processing unit for visual display data in pixel array form corresponding to the size information; the processing unit for performing text-based descriptive processing and image-based synthetic processing on the visual display data using image processing software to generate editable visual display data and send it to the user terminal; and the processing unit for calculating evaluation indicators of the constituent information based on selection and editing operations received from the user terminal, extracting arrangement parameters of text elements, color elements, and image elements, reflecting the evaluation indicators and arrangement parameters in subsequent prompt statements and conditional vectors, thereby instructing the generative artificial intelligence model to regenerate through the updated prompt statements. This allows for a closed-loop processing flow on the server side, from input parsing, conditional encoding, model inference to result editing and interactive feedback, enabling the computer system to adaptively optimize prompt statements and conditional vectors based on user interaction behavior, improving the controllability, consistency, and efficiency of visual display generation, and thus significantly improving the overall performance of computer technology in image generation and editing tasks driven by generative artificial intelligence models.
[0206] A "system" refers to an integrated computer implementation scheme consisting of multiple interconnected information processing devices, storage devices, and communication devices, used to perform tasks such as prompt generation, generative artificial intelligence model reasoning, image processing, and data communication.
[0207] A "server" refers to a computer device or computing node in a system that is responsible for receiving input information, generating prompts, calling generative artificial intelligence models, and performing image generation and processing.
[0208] "User terminal" refers to an electronic device operated by a user for inputting size and composition information, receiving and displaying visual displays, and performing editing operations, including but not limited to mobile terminals, fixed terminals, or other devices with display and input functions.
[0209] The "processing unit" refers to the processor and its software program configured in the server, which is a functional unit used to parse input information, generate prompt statements, construct condition vectors, perform inference processing, and control image processing and data communication.
[0210] "Input information" refers to a set of data sent by the user terminal and received by the server to indicate the conditions for generating visual displays. It includes at least size information and composition information, and may also include additional parameters such as layout preferences and style instructions.
[0211] "Size information" refers to parameters that represent the spatial extent or resolution of a visual display, including but not limited to width, height, number of pixels, or equivalent numerical information.
[0212] "Construction information" refers to the collection of information about various elements that constitute the content and style of a visual display, including textual information, color information, image information, and descriptive information such as their layout, hierarchy, and style.
[0213] "Text information" refers to the text content and related attributes displayed in a visual display, including parameters used to control text display such as character sequence, font, font size, font weight, and alignment.
[0214] "Color information" refers to data used to represent the color characteristics of the background, foreground, and various elements of a visual display, including color names, color value codes (such as RGB, HSV, etc.), and color gradients or color schemes.
[0215] "Image information" refers to the data of static images or graphic elements used in visual displays, including logo images, product images, background images, and their attributes such as size, position, and transparency.
[0216] "Prompt statements" refer to instruction text generated by the server based on size and composition information, written in natural language or a form similar to natural language, used to explicitly describe the visual display generation goals and constraints to the generative artificial intelligence model.
[0217] "Generative artificial intelligence models" refer to artificial intelligence models trained through machine learning that can automatically generate visual display data based on prompts or conditional vectors, including but not limited to image generation models based on deep neural networks.
[0218] "Conditional vector" refers to a numerical feature vector that is generated by the server through parsing and encoding of prompt statements, transforming text information, color information, image information, and size information into a process that can be processed by generative artificial intelligence models.
[0219] "Inference processing" refers to the computational process in generative artificial intelligence models where, with a conditional vector as input, matrix operations, convolution operations, nonlinear transformations, and other computational steps are sequentially executed on a computing device to output visual display data.
[0220] "Matrix operations" refer to linear algebraic operations performed in generative artificial intelligence models, including multiplication and addition between vectors and matrices, and between matrices, used to achieve feature transformation and inter-layer propagation.
[0221] "Convolution operation" refers to the operation of locally weighted summation of input features in a neural network model, used to extract local spatial features and construct multi-layer feature representations.
[0222] "Nonlinear transformation" refers to the process of transforming the results of linear operations in generative artificial intelligence models through activation functions or other nonlinear functions, in order to improve the expressive power of the model.
[0223] "Visual display data" refers to image data represented in the form of a pixel array, output by a generative artificial intelligence model. Its size corresponds to size information and can be further processed for rendering and synthesis.
[0224] "Editable visual display data" refers to image data that, based on visual display data, is processed through methods such as drawing text and synthesizing images, and can be edited on the user terminal in terms of position, content, color, etc.
[0225] "Image processing software" refers to program modules or library components used for image processing operations such as drawing, compositing, cropping, scaling, and color adjustment of visual display data.
[0226] "Selection operation" refers to the selection behavior of a user on a user terminal regarding a visual display, including but not limited to clicking, long pressing, swiping, or other interactive operations used to express preferences or selections.
[0227] "Editing operations" refer to the actions of users on their terminals to modify the content or layout of visual displays, including interactive operations such as adjusting text, changing colors, moving or replacing image elements.
[0228] "Evaluation metrics" refer to numerical or statistical quantities calculated by the server based on selection and editing operations, used to quantify the quality of information or user preferences.
[0229] "Layout parameters" refers to a set of parameters that define the spatial and hierarchical layout characteristics of various text elements, color elements, and image elements in a visual display, such as their position, size, spacing, and alignment.
[0230] In embodiments of the present invention, the server, terminal, and user work collaboratively in a network environment to generate and optimize visual displays with high precision and efficiency through a generative artificial intelligence model. Various embodiments of the present invention can be implemented on general-purpose computer hardware, such as a computing device including a multi-core central processing unit, a graphics processing unit, a large-capacity random access memory, and a non-volatile storage device. The server runs software modules for implementing the functions of the present invention through an operating system and middleware.
[0231] In one implementation, the server utilizes a general-purpose scripting language and a web application framework to build backend services. The server employs a generative artificial intelligence model based on a multi-layer neural network as the core of image generation. At the hardware level, the server can use a rack-mounted computer equipped with a graphics processor, while the terminal can use a mobile or fixed information terminal with a display and touch input device. The server exchanges request data and generated image data with the terminal via network communication protocols.
[0232] The program executed by the server in this invention can be divided into multiple functional modules, including a prompt statement generation module, a prompt statement parsing module, a condition vector construction module, a generative artificial intelligence model inference module, an image post-processing module, an interactive feedback analysis module, and a parameter update module. The program executed by the terminal can be divided into a user interface module, an input acquisition module, a prompt statement assembly module, an image display module, and a local editing module. The user interacts with the server through the terminal's user interface module.
[0233] When the server receives size and composition information from the terminal, it stores this information as structured data objects in its memory. The server uses internal data structures to distinguish fields such as width, height, color, text content, and image identifiers. When generating prompts, the server inserts these fields into a predefined natural language template to form prompts for generative artificial intelligence models. For example, the server can generate the following Chinese prompts: "Please generate a 300x250 pixel visual display with a blue background and the text 'New Product Launch', and use the logo image I uploaded in the image." or: "Please generate a 300x250 pixel visual display with a dark blue background, the large text 'Limited-Time Special Offer' in the center, and the uploaded brand logo in the lower left corner. The overall style should be simple and clear." After generating the prompt statement, the server performs lexical and syntactic analysis on it using the prompt statement parsing module. The server uses word segmentation and sub-word splitting algorithms to break down the prompt statement into a sequence of tokens, assigning a unique identifier to each token. These identifiers are then input into the embedding layer, where matrix multiplication converts the discrete identifiers into a dense vector representation, thus obtaining the vectorized features of the text information. Simultaneously, the server constructs numerical features based on size information, such as width, height, and aspect ratio. These values are normalized and then concatenated with the text vectors or mapped to a space of the same dimension as the text vectors through linear transformation, forming a unified conditional vector.
[0234] When constructing conditional vectors, the server can also design specific encoding methods for color and image information. The server can encode RGB values or color category indices into numerical vectors, mapping them to a high-dimensional feature space through a linear layer. The server can encode image categories or image uses (e.g., logos, backgrounds, product images) into one-hot vectors, then pass them through a transformation layer to obtain image-related conditional vectors. The server concatenates text features, color features, size features, and image use features along the vector dimension, or fuses them through an attention mechanism, to obtain the final conditional vectors used by the generative AI model.
[0235] The server invokes a generative artificial intelligence model during the inference phase. In a preferred embodiment, the server employs a deep neural network with an encoder-decoder structure, including multiple layers of self-attention modules, multiple layers of convolutional modules, and upsampling modules. The server inputs conditional vectors into the text encoding branch of the model and uses random noise tensors as the initial input to the image generation branch. The server propagates features between multiple network layers through matrix multiplication and convolution operations. In each layer, the server performs non-linear activation function operations, such as rectified linear units or other non-linear functions, to enhance the model's expressive power. The server iteratively denoises through multiple iterations, converting random noise into visual display data with specific content and structure. The server uses decoder and upsampling layers to bring the feature map size to a specified width and height, thereby obtaining a pixel array in memory consistent with the size information.
[0236] During the model's internal training phase, the server can pre-learn from a large amount of training data. The server uses a dataset containing size labels, color labels, text descriptions, and image feature annotations during training. The server uses a loss function to measure the difference between the generated image and the target image; this loss function can include pixel-level loss, perceptual loss, and adversarial loss. The server calculates the gradient of the loss function with respect to the network weights using backpropagation and updates the weight parameters using gradient descent or its variants. During training, the server can perform data augmentation operations, such as random cropping, color perturbation, and affine transformations, to improve the model's robustness to diverse layouts and color schemes. Through this training process, the server obtains a generative artificial intelligence model that can efficiently generate visual displays based on prompts during the inference phase.
[0237] After generating the visual display data from the model, the server further processes this data using an image post-processing module. The server uses an image processing software library to convert the pixel array into an image object. Based on the text content, font style, and positional rules specified in the composition information, the server draws text on the image. The server can arrange text and image elements according to pre-defined layout strategies, such as centering the title, placing the subtitle at the bottom, and positioning the logo in a corner. During text drawing and image overlay, the server calculates pixel coordinates and overlays or blends pixels in the target area to combine visual elements. The server then encodes the composite image into a compressed format and prepares it for transmission to the terminal.
[0238] In this invention, the terminal is responsible for interacting with the user. The terminal displays multiple controls on the user interface, and the user inputs dimensions, background color, text content, and image selection on the terminal. After receiving user input, the terminal combines these fields into structured parameters and prompts in natural language. The terminal can directly send the structured parameters to the server for the server to generate the prompts, or it can generate the prompts locally before sending them to the server. After receiving the visual display from the server, the terminal decodes the image data and presents it on the display device, allowing the user to observe the generated result on the terminal.
[0239] When users edit visual displays on the terminal, they can drag text boxes, modify text content, adjust colors, and replace or move image elements. The terminal records these operations and sends the edited composition information back to the server. The terminal can perform some simple compositing locally for instant preview, but the key parameters used for the next round of generation are uniformly parsed and modeled by the server to ensure consistent updates of the condition vectors.
[0240] After receiving selection and editing operations from the terminal, the server transforms these operations into evaluation metrics and layout parameters through an interactive feedback analysis module. The server can count the number of clicks, saves, or deletions made by users on different generated results, normalizing these statistics into score values. Based on the repeated use of a particular layout or color during editing operations, the server can infer the user's preferred layout parameters and color scheme. The server stores the evaluation metrics and layout parameters in a structured data table, indexed by user identifier, task identifier, or scene identifier, enabling accumulation and reuse.
[0241] In new generation tasks, the server references historical evaluation metrics and layout parameters when constructing prompts and condition vectors. For example, if a particular color scheme and text layout historically received higher evaluation metrics, the server automatically adds descriptions such as "simple and clear overall style," "centered text layout," and "using contrasting dark blue and white" to the prompts and increases the weights of the corresponding features in the condition vectors. This makes the server more inclined to generate visual displays that closely resemble historically preferred schemes during new inference processing. Because these weight adjustments numerically alter the distribution of the condition vectors, the output obtained by the server through the generative AI model will be more structurally and stylistically aligned with long-term user preferences, thereby improving generation accuracy and user satisfaction.
[0242] The server goes beyond simply automating the human design process; it deeply reconstructs the traditional image generation workflow through structured encoding of conditional vectors, weight adjustments driven by evaluation metrics, and regularized representation of parameter arrangements. By adding multi-dimensional constraints to the model input, the server reduces the randomness and uncertainty of the generated results, thereby achieving more efficient feature utilization and inference path convergence within the computer. Due to the fine-tuning of the dimensionality and numerical values of the conditional vectors, the server can improve generation quality without significantly increasing the model size, thus completing more high-quality generation tasks per unit time and achieving a dual improvement in processing speed and quality.
[0243] In an alternative implementation, the server can employ different generative AI model structures. For example, the server can use a diffusion-based generative model, progressively denoising along a temporal dimension. By injecting conditional vectors at each time step, the model continuously approaches the target layout and color preference through multiple iterations. In this architecture, the server can adjust parameters such as the number of sampling steps, noise intensity, and conditional guidance weights to strike a balance between image detail and generation speed. The server can also employ a generative adversarial network (GAN) structure, consisting of a generator network and a discriminator network. During the training phase, the server uses the adversarial loss output from the discriminator network to optimize the generator network, resulting in more natural and visually appealing outputs.
[0244] In another implementation, the server can introduce a multi-head self-attention mechanism, using text content, size information, and color information as different attention sources. When calculating attention weights, the server uses matrix multiplication to calculate the relevance between the query vector and the key vector, then weights and aggregates the value vector, thus focusing on key information within the conditional vector. This attention mechanism allows the server to explicitly capture the joint constraints of "specific size + specific background color + specific text," making it easier for the network to meet complex layout requirements when generating images. This structured attention calculation improves the model's convergence speed and stability under multiple constraints, thereby enhancing the computer's technical performance in multi-constraint image generation tasks.
[0245] In this embodiment of the invention, the terminal not only performs input / output functions but also undertakes local pre-construction of prompt statements and fine-grained acquisition of editing parameters. The terminal can locally sample the user's drag trajectory of the text box position as a coordinate sequence and store the final position as layout parameters in normalized coordinate form. The terminal can also record the final value of the color slider when the user adjusts the color and send it to the server as part of the color information. Because the terminal performs necessary preprocessing and compression, the server can reduce parsing overhead, lower communication load, and improve overall system throughput when receiving data.
[0246] In this embodiment of the invention, users interact with the system through a combination of natural language description and visual editing. Upon initial request, users only need to provide general size and style requirements. After seeing the generated results, users gradually refine their needs through editing operations. The server transforms these editing operations into quantifiable parameters and continuously utilizes them in subsequent generation processes, thus forming a progressive, human-computer collaborative generation process. Because the server internally performs structured modeling and parameterized encoding of user behavior, this collaborative process not only reduces repetitive user operations but also enhances the server's understanding of the task context. From a computer technology perspective, this achieves an adaptive generation system based on interactive feedback.
[0247] Through the aforementioned embodiments, this invention enables the server to establish clear data flow and transformation rules between size information, composition information, prompt statements, conditional vectors, and interactive feedback internally, thereby achieving fine-grained feature management and efficient inference processes at the computer level. This invention can be applied not only to the automatic generation and optimization of online advertising images, but also to various scenarios such as digital billboard content generation, application startup screen generation, and e-commerce product display image generation, tightly integrating visual content production with display control on terminal devices, achieving technical control and optimization of display device content in the real world. Through conditional vector constraints, multi-round generation, and feedback learning, this invention improves the performance of computers in complex image generation tasks in terms of accuracy, speed, and stability, demonstrating a substantial improvement in computer technology itself.
[0248] use Figure 12 The processing flow is explained.
[0249] Step 1: The user enters the initial design conditions on the terminal. Users select and input relevant parameters for visual displays on the terminal's interface. Inputs include: size information (e.g., 300 pixels wide, 250 pixels high), background color (e.g., blue), text content (e.g., "New Product Launch"), image selection (e.g., a brand logo), and optional style preferences (e.g., "minimalist" or "tech-savvy"). The terminal combines this input data into an internal structured object, which serves as the raw data for subsequent processing. The output is a set of structured parameters, including size, color, text, image, and style fields.
[0250] Step 2: The terminal generates prompts based on structured parameters. The terminal takes the structured parameter set from step 1 as input and generates a prompt statement for the generative AI model using a preset natural language template. The terminal fills in fields such as width, height, background color, text content, and logo description into the string template to form a natural language description. For example, the prompt statement generated by the terminal could be: "Please generate a 300x250 pixel visual display with a blue background, the text 'New Product Launch,' and use the logo image I uploaded." The terminal's output includes the natural language prompt statement and the original structured parameters, which are prepared as a request payload to be sent to the server.
[0251] Step 3: The terminal sends request data to the server. The terminal constructs a network request using the prompt statement and structured parameters generated in step 2 as input. The terminal packages the prompt statement as a text field, dimensions, colors, and text as numeric or string fields, and images such as the logo as binary data into a request message. The terminal sends this request message to the designated interface of the server via the communication module. The data processing in this step involves serializing and encoding the local structured objects and prompt statement (e.g., converting them to a byte stream and adding header information). The output is a request data stream transmitted over the network, with the server as the input.
[0252] Step 4: The server receives and parses the terminal request. The server takes the request data stream sent by the terminal as input and parses the request through the network receiving module and the request parsing module. The server extracts the prompt text, size information, background color information, text content, image binary data, and other parameters from the request body. The server transforms this data into internal data structures, such as key-value maps or object instances. During parsing, the server performs data validation, including checking if the size is a positive integer, if the color format is correct, and if the image data is complete. The output is a set of validated internal representation data in the server's memory, used for subsequent construction of conditional vectors.
[0253] Step 5: The server parses the prompt statement into text characteristics. The server takes the prompt obtained in step 4 as input and calls the text parsing and encoding module. First, the server performs word segmentation or sub-word splitting on the prompt, breaking down a natural language sentence into several token sequences. Then, the server maps each token to an integer ID based on a pre-trained vocabulary, forming discrete ID vectors. Next, the server performs matrix multiplication on these ID vectors through an embedding layer, converting them into a dense floating-point vector sequence. This data processing encodes semantic information from the character or word level into a high-dimensional vector space. The output is a sequence of text feature vectors, which serves as part of the conditional vectors for the generative artificial intelligence model.
[0254] Step 6: Server-encoded size and color information are numerical features. The server takes the size and color information from step 4 as input and converts the width, height, and color values into numerical features usable by the model. The server normalizes the width and height (e.g., by dividing by a preset maximum size) to generate proportional values in the range of 0 to 1. For color, the server can use RGB encoding to normalize the integer values of the three channels to floating-point numbers, or map color categories to one-hot vectors. The server performs matrix multiplication and bias addition on these numerical features through a linear transformation layer, transforming them into a vector space of the same dimension as the text features. The output is a size feature vector and a color feature vector.
[0255] Step 7: Server-encoded image information and layout preferences are used as auxiliary features. The server takes the image information (e.g., logo purpose and type) and layout preferences (e.g., "logo in the bottom right corner," "text centered") from step 4 as input, and converts these discrete attributes into numerical representations using label encoding, one-hot vectors, or embedding. The server performs mapping operations, mapping each purpose and layout rule to a vector representation, and combines these vectors into a unified auxiliary feature vector through vector addition or concatenation. This data processing encodes the logical layout rules into a numerical form that can participate in neural network operations. The output is the auxiliary feature vector.
[0256] Step 8: The server builds a unified condition vector. The server takes the text feature vector from step 5, the size and color feature vectors from step 6, and the auxiliary feature vector from step 7 as input. It concatenates these vectors along their dimensional axes or fuses them using an attention module. The server performs matrix multiplication, weighted summation, and non-linear activation to combine features from multiple sources into a unified conditional vector or synthesize it into a feature matrix. This conditional vector reflects all constraints and semantic requirements during the generation of the visual display. The output is a unified conditional vector, which serves as one of the main inputs to the inference phase of the generative artificial intelligence model.
[0257] Step 9: The server invokes a generative artificial intelligence model to perform inference. The server takes the unified conditional vector and random noise tensor from step 8 as input and feeds them into the generative artificial intelligence model. Within the model, the server performs multi-layer matrix multiplication, convolution operations, self-attention computation, and non-linear activation, updating the feature representation layer by layer. In a diffusion-based or generative adversarial architecture, the server iteratively denoises or rewrites the noisy image multiple times, gradually approximating the target distribution constrained by the conditional vector. Input data processing includes fusing the conditional vector and noise features at specific layers (e.g., through conditional normalization or cross-attention), and the output is initially generated visual display data, i.e., a pixel array with a specified width and height.
[0258] Step 10: The server performs image post-processing on the visual display data. The server takes the visual display data generated in step 9 as input and calls the image processing module to perform post-processing. First, the server converts the pixel array into an image object, then performs necessary cropping or scaling based on the size information. Next, the server determines the text area and image overlay area based on the composition information and calculates the corresponding pixel coordinates. The server performs text rendering operations, rendering the text content to the image pixel buffer according to parameters such as font, font size, and color. The server then overlays image elements such as the logo onto the corresponding positions through pixel compositing operations. The output is post-processed, editable visual display data that meets the requirements of the composition information.
[0259] Step 11: The server encodes and generates the result, which is then returned to the terminal. The server takes the editable visual display data obtained in step 10 as input and encodes it into a compressed image format, such as bitmap compression. The server encapsulates the encoded binary data into a response message and sends it to the terminal over the network. In this data processing, the server compresses and buffers the image data to reduce network transmission load. The output is a response data stream sent to the terminal over the network, containing the final generated visual display.
[0260] Step 12: The terminal receives and decodes the visual display. The terminal receives the response data stream sent by the server as input, passes it to the decoding module for processing via the network module, and decodes the compressed image data into a local image object or pixel buffer. The terminal then displays the image on the display device, allowing the user to view the generated result. The output consists of a visual display on the terminal screen and image data in memory for subsequent editing.
[0261] Step 13: Users perform editing and selection operations on the terminal. Users interact with the visual display shown in step 12 within the terminal's editing interface. Users can drag text, modify text content, change background color, replace the logo, or select a version as a candidate. Each user's edit or selection is recorded as event data by the terminal, including the operation type, target element identifier, position coordinates, color value, and text content. The output is a series of structured edit and selection records.
[0262] Step 14: Terminal upload editing records and selection records The terminal takes the edit and selection records generated in step 13 as input and organizes this data into a structured log, such as a list or time series. The terminal performs necessary compression or simplification of the log, merging multiple minor drags into a final position parameter and multiple color adjustments into a final color value. The terminal sends this structured feedback data to the server via a network interface as the basis for subsequent optimization. The output is a feedback data stream containing evaluation criteria and layout change information.
[0263] Step 15: The server analyzes interactive feedback and calculates evaluation metrics. The server takes the feedback data received in step 14 as input and calculates evaluation metrics through the interactive feedback analysis module. The server scores different components based on factors such as the number of user selections, retention time, and editing frequency, using statistical operations (e.g., counting, averaging, weighted summing) to obtain evaluation values for each layout method, color scheme, and text style. The server also extracts final layout parameters from the editing records, such as the final text position coordinates, the relative position of the logo, and commonly used color values. The output is a set of evaluation metrics and layout parameters, which serve as the basis for updating the prompt statements and conditional vectors.
[0264] Step 16: The server updates the prompt statement and condition vector based on evaluation metrics. The server takes the evaluation metrics and layout parameters obtained in step 15, along with the original prompts and condition vectors, as input to update the descriptions and features used in subsequent generation. At the text level, the server incorporates highly rated elements into the new prompts, such as adding descriptions like "simple and clear overall style," "using a contrasting dark blue and white color scheme," and "the logo is placed in the lower left corner." At the feature level, the server adjusts the weights corresponding to these elements in the condition vectors and recalculates the feature values through linear transformation and normalization. This data processing makes the model more inclined to generate visual displays that match user preferences in the next inference. The output is the updated prompts and the updated condition vectors.
[0265] Step 17: The server performs a regeneration based on the update conditions. The server takes the updated prompt and condition vector from step 16 as input and repeats the text parsing, feature encoding, model inference, and image post-processing processes described in steps 5 to 11. However, at this point, the condition vector reflects the user's historical preferences and evaluation results. During the new inference process, the server performs matrix operations, convolution operations, and nonlinear transformations based on the updated feature distribution, making the output visual display more aligned with user needs in terms of layout, color scheme, and text style. The output is the newly generated visual display data, which the terminal receives and displays again, allowing the user to continue interacting and optimizing.
[0266] Alternatively, an emotion engine for inferring user emotions can be combined. That is, the specific processing unit 290 can also use the emotion-specific model 59 to infer user emotions and perform specific processing using user emotions.
[0267] Example 2 The flow of a specific process in Example 2 will be described. Each part of the system described below is implemented by the data processing device 12 and the smart device 14. The data processing device 12 will be referred to as the "server," and the smart device 14 as the "terminal."
[0268] In existing technologies, the automatic generation of visual presentations (such as advertising images and promotional banners) mainly relies on generative artificial intelligence models to output single or a small number of candidate results based on a single instruction given by the user. This type of technology typically suffers from the following problems: First, servers often simply concatenate static parameters such as size and color input by the user directly into the instruction, lacking the ability to uniformly model the user's natural language intent and structured generation conditions. This results in the generated visual presentations failing to reflect business needs in a timely and accurate manner. Second, while servers can record basic data such as click counts, they do not transform the user's display and selection behaviors on the terminal into resolvable signals that can be fed back into the generation process at the system level, making it impossible to perform fine-grained optimization of prompts on the input side of the generative artificial intelligence model. Furthermore, existing systems typically only provide a rough evaluation of the overall solution (the entire image), without breaking down the visual presentation into fine-grained constituent elements such as color elements, text elements, layout elements, and operational component elements within the server, and without conducting statistical analysis based on the correlation between these constituent elements and the selection rate. As a result, it is difficult to automatically discover "combinations of constituent elements with high selection rates" within the computer, and even more difficult to achieve continuous automated iterative optimization.
[0269] Furthermore, in existing systems, servers often treat generative AI models as "black box image generators," with the generation process and behavior log analysis process being disconnected. Behavioral data cannot be used in a structured way to influence the content, weight, and template structure of instruction statements. This results in two main consequences: firstly, a large amount of computing resources are consumed in repeatedly trying out low-conversion design solutions; secondly, the generation system lacks the ability to adapt to environmental changes (such as changes in user preferences or marketing scenarios), limiting the overall generation efficiency and effectiveness of the computer system.
[0270] Therefore, it is necessary to provide a system and method that centrally manages attribute information, instruction statements, behavior logs, and constituent element attributes within a computer via a server, and automatically constructs, adjusts, and iterates prompt statements on the server side, transforming user behavior data into structured feedback for generative artificial intelligence models. This would improve the quality of visual presentation generation while enhancing the overall technical performance of the server in terms of generation task scheduling, data processing, and model invocation.
[0271] The specific processing performed by the specific processing unit 290 of the data processing apparatus 12 in Embodiment 2 is achieved by the following means.
[0272] In this invention, the server includes a module for acquiring attribute information and natural language instructions from the user in an information processing device and generating prompt statements; a module for executing a generative artificial intelligence model on computing resources to generate multiple visual presentation data based on the prompt statements; a module for sending and displaying the visual presentations to the terminal and collecting display event and selection event behavior logs; a module for statistically analyzing the behavior logs to calculate the selection rate of each visual presentation; a module for performing statistical analysis based on the constituent element information and the selection rate to extract high-selection-rate constituent elements and automatically adjust and update the prompt statements accordingly; and a module for re-inputting the updated prompt statements into the generative artificial intelligence model to repeatedly regenerate and optimize the constituent elements of the visual presentations. This allows for a closed-loop computation process within the server, encompassing everything from user input to model generation, from behavioral data collection to element-level statistical analysis, and then to automatic reconstruction of prompts and model re-invocation. This enables the input parameters and prompts of the generative AI model to adaptively adjust based on real user interactions, thereby reducing invalid generation, increasing click-through rates and conversion rates, and improving server resource utilization efficiency and overall technical performance at the computer level in terms of task scheduling, data processing, and model invocation.
[0273] "Information processing device" refers to a data processing device used to receive, store, process, and output results from input data, including but not limited to computing devices with processors, memory, and communication interfaces.
[0274] "User terminal" refers to a terminal device that allows users to operate and interact with it. It is used to send input information to the server and receive and present visual presentations, including but not limited to mobile terminals, fixed terminals, and computing devices with display functions.
[0275] "Generative artificial intelligence models" refer to models built based on machine learning algorithms, especially deep learning algorithms, that can automatically generate output data such as image data based on input prompts, including but not limited to diffusion models, autoregressive models and their variations.
[0276] "Prompt statements" refer to textual information, primarily in natural language, used to describe the conditions and content requirements for generating visual representations. This textual information is input into generative artificial intelligence models to guide the models in generating visual representations that conform to the semantics of the text.
[0277] "Generation conditions" refer to the various parameters and requirements used to constrain and describe the generation results of visual representations, including but not limited to size information, color information, layout information, target purpose information, and content-related constraints.
[0278] "Attribute information" refers to the structured data elements that constitute the generation conditions, including size attributes, component attributes, color attributes, layout attributes, and other generation-related information expressed in parameter form.
[0279] "Component elements" refer to the various element units used to make up a visual presentation, including but not limited to color elements, text elements, layout elements, interactive elements, and other graphic or text elements.
[0280] "Color elements" refer to the constituent elements used to represent the color characteristics in a visual presentation, including background color, foreground color, button color, and text color, as well as other attributes related to color representation.
[0281] "Text elements" refer to the constituent elements used to present text information in a visual presentation, including attributes such as text content, font type, font size, font weight, line spacing, and text position.
[0282] "Layout elements" refer to the constituent elements used to define the spatial arrangement of various constituent elements in a visual presentation, including the relative positions of image areas and text areas, alignment methods, white space methods, and overall layout structure.
[0283] "Interactive elements" refer to the interface elements in a visual presentation that allow users to interact with them, including buttons, icons, link areas, and other clickable or actionable graphic elements.
[0284] "Visual presentations" refer to visual information carriers displayed on user terminals in the form of image data, including but not limited to advertising images, promotional banners, interface component preview images, and other static or dynamic images.
[0285] "Behavior logs" refer to data records collected and recorded by the server from the user terminal related to the display and interaction of visual presentations, including display events, click events, selection events, evaluation information, and regeneration requests.
[0286] A "display event" is a record of the act of a visual object being presented once in the display area of a user's terminal, used to indicate the number of times the visual object is actually shown to the user.
[0287] "Selection event" refers to the behavior record generated when a user clicks on a visual object on the terminal, confirms its use, or performs a similar selection operation, which is used to indicate the number of times the visual object is selected by the user.
[0288] "Selection rate" refers to the ratio of the number of times a visual presentation is selected to the number of times it is displayed within a statistical period. It is used to measure the degree to which a visual presentation is selected relative to display opportunities.
[0289] "Component information" refers to descriptive or parametric data related to each component in a visual presentation, including color, text, layout, operating parts, and their specific attributes, which are used for structured management and statistical analysis on the server.
[0290] "High-selection-rate constituent elements" refer to constituent elements or combinations of constituent elements that are identified as helpful in improving the overall selectability of visual presentations after statistical analysis of constituent element information and corresponding selection rates.
[0291] Metadata refers to auxiliary data that is attached to the visual presentation data and is used to describe the characteristics and generation conditions of the visual presentation. It includes primary color attributes, button color attributes, text size attributes, layout type attributes, and whether price information exists.
[0292] "Instruction information" refers to a set of data generated by the server based on attribute information and instruction statements, used to explicitly provide the generative artificial intelligence model with the content, parameters, and constraints of the generation task. This includes prompt statements and control parameters related to the generation process.
[0293] "Computing resources" refers to the hardware and its supporting environment used to perform inference operations on generative artificial intelligence models, including but not limited to central processing units, graphics processing units, accelerators, and corresponding storage devices and network resources.
[0294] "Storage device" refers to the storage medium and device used to store visual presentation data, behavior logs, prompts and related metadata, including local storage devices and network storage devices.
[0295] "Communication device" refers to the communication interface and communication module used for data transmission and reception between the server and the user terminal, including network interface, communication protocol stack, and related hardware and software that support wired or wireless communication.
[0296] "Statistical processing" refers to the process by which the server summarizes, groups, counts, calculates ratios, and performs correlation analysis on behavior logs, constituent element information, and metadata to obtain selection rates and statistical results related to constituent elements.
[0297] "Prompt statement template" refers to a sentence framework or format structure that is predefined or dynamically generated in the server and used to construct prompt statements. It is formed by filling in attribute information and optimized constituent element conditions.
[0298] "Self-learning update" refers to the process by which the server automatically adjusts the prompt statement template and instruction content parameters based on continuously collected behavior logs and statistical results, thereby gradually improving the input configuration of the generative artificial intelligence model without human intervention.
[0299] The embodiments of the present invention will be described in conjunction with the information processing system structure described in Appendices 1 to 3. In the following description, the subjects are limited to "server," "terminal," and "user," and the description focuses on the specific composition of the data structure, algorithm flow, and generative artificial intelligence model within the computer, to ensure that the present invention can be implemented by those skilled in the art and to highlight its improvement effect on computer technology itself.
[0300] I. Overall System Composition A server includes at least one processor, storage device, communication device, and computing resources. The computing resources may include a central processing unit and a graphics processing unit. The server can run a general-purpose operating system, such as a Linux-based server operating system, and middleware software, such as a web server and backend service framework. The server can also deploy deep learning frameworks, such as Python-based PyTorch or TensorFlow, for loading and executing generative artificial intelligence models.
[0301] A terminal includes a computing device with a display and input section, such as a smartphone, tablet, laptop, or desktop computing device. The terminal can run browser software or dedicated applications and communicate with the server using HTML, CSS, JavaScript, or component-based front-end frameworks.
[0302] Users interact with the server through the terminal, input the conditions for generating visual representations, view the visual representations generated by the server, and perform operations such as selection and evaluation.
[0303] II. Server-side modules and data structures The server consists of multiple functional modules, each of which executes a corresponding program in collaboration with the processor and storage device.
[0304] 1. Attribute and Indicator Statement Acquisition Module The server maintains data structures in the storage device to describe the generation conditions. The server can store information corresponding to each generation task as a record object, which includes: - Task Identifier: A unique identifier that identifies the task that was generated; - User ID: An identifier associated with the current user or session; - Attribute information: including size attributes (width, height), color attributes (primary color, secondary color), layout attributes (relative position of text and images), target purpose attributes (promotion type, product category), etc.; - Natural language instructions: Descriptive information entered by the user in free text format; - Prompt statements: Input text generated by the server combining attribute information and natural language instructions.
[0305] The server can use a relational database table or key-value storage structure in the storage device, mapping the above fields to columns or keys respectively. Through this structured storage, the server can efficiently group and retrieve data based on attributes during subsequent statistics and analysis, thereby reducing the amount of data scanned and improving the speed of querying and statistics.
[0306] 2. Prompt Statement Construction and Template Management Module The server maintains a collection of prompt statement templates in the storage device. Each template is a text structure with placeholders that describes the generation conditions. For example, the server might maintain the following templates: "Please generate a banner ad for {target purpose}, with dimensions of {width} × {height} pixels, using {primary color} as the main color scheme, and employing {layout description}. {User supplementary notes}." The server extracts fields such as target purpose, size, main color, and layout from the attribute information, replaces the placeholders in the template with these fields, and inserts the user's natural language instructions at the end to form a complete prompt. For example, when the user enters the following attributes: - Dimensions: 1200×628 - Main color: Blue - Layout: Image on the left, text on the right - Target use: New product promotion And instructions: "Create a promotional banner for the newly launched smartwatch, highlighting its health monitoring features and long battery life, with an overall high-tech feel." The prompt message generated by the server can be: "Please generate a banner ad for promoting a new product. The banner should be 1200×628 pixels, with blue as the main color, and use a layout of image on the left and text on the right. Create a promotional banner for a newly launched smartwatch that highlights its health monitoring functions and long battery life, with an overall high-tech feel." The server uses this template-based construction to give the prompts a uniform structure, which makes it easier to insert high-selectivity constituent conditions into some segments later, and avoids simply concatenating strings. This reduces ambiguity and improves the consistency of model input, thereby improving the controllability and stability of the generated results.
[0307] 3. Generative Artificial Intelligence Models and Their Structure The server loads a generative artificial intelligence model onto its computing resources. This model can be an image generation network based on a diffusion model. The server can adopt the following structure as an example: - Text Encoding Module: Employs a text encoding network based on a transformer structure to convert prompts into high-dimensional text embedding vectors. This text encoding network consists of multiple layers of self-attention layers, a feedforward network, and a normalization layer, with parameters obtained through pre-training.
[0308] - Latent Diffusion Module: This module employs a U-shaped network structure to perform diffusion denoising on the latent image representation. The U-shaped network can include multiple convolutional layers, downsampling, upsampling, and cross-layer connections. Conditional information from text embeddings is introduced into each layer, and through conditional attention or feature injection mechanisms, the image content is constrained by the prompting text.
[0309] - Decoding module: Employs a convolutional decoding network to restore the feature maps in the latent space to image data in the pixel space.
[0310] During the model training phase, the server can use a large number of images and their corresponding descriptive texts as training data. During training, the server performs the following operations for each training sample: - Input the description text into the text encoding module to obtain the text embedding; - Encode the real image into a latent space, and add Gaussian noise; - Use a U-shaped network to predict the noise or denoised potential representation; - Use a loss function (such as mean squared error) to compare predicted noise with actual noise, or to compare denoised results with the true latent representation; - Update network weights based on backpropagation algorithm and optimizer (such as adaptive learning rate optimization algorithm).
[0311] By repeating this process, the server enables the model to learn a conditional generation mapping from prompts to images and encodes the relationships between various components and semantics in its internal parameters.
[0312] III. Visual Presentation Generation and Post-processing After receiving the prompt, the server inputs it into the text encoding module to obtain a text embedding vector. The server then samples initial noise in the latent space and performs a multi-step iterative diffusion denoising operation. In each step, the server performs numerous matrix multiplications, convolutions, and activation function operations, gradually converging the latent representation from noise to a structure corresponding to the semantics of the prompt.
[0313] After the diffusion process concludes, the server inputs the latent representation into the decoding module to generate a pixel matrix. The server converts the generated image into a standard format file (such as PNG or JPEG) and uses an image processing library to perform size verification and compression quality adjustments. The server stores these image files in object storage or a file system and records the file paths, associated task identifiers, and constituent element attributes in a database.
[0314] The server returns the access address of the visual representation to the terminal via a communication device. The terminal loads the image from that address and displays it on the screen. Users can browse these images on the terminal and perform operations such as "use this scheme" or "generate another one".
[0315] IV. Behavioral Logs and Selection Rate Statistics The server maintains a behavior log data structure for each visual representation, which includes: - Visual signage; - Displays event count; - Select event count; - Timestamp information; - User ID or session ID (optional).
[0316] After receiving display and selection event information from the terminal, the server increments the corresponding counter field through atomic increment or transaction operations to avoid concurrent write errors. The server can employ batch write or log buffering techniques to merge multiple behavior records in memory and write them to the storage device at once, thereby reducing write frequency, reducing disk I / O overhead, and improving statistical processing efficiency.
[0317] During the statistics phase, the server reads these count fields and calculates the selection rate for each visual presentation, which is the selection event count divided by the display event count. The server can also group the attributes of each component element, such as grouping by button color, by primary color, or by layout type, and calculate indicators such as average selection rate, total number of displays, and variance for each group.
[0318] V. Component Attribute Management and High-Selectivity Component Extraction When generating visual representations, the server records metadata for each image. This metadata includes at least: - Primary color attribute; - Button color attribute; - Text size attributes; - Layout type attribute; - Whether it includes price information; - Whether it contains a specific visual symbol (such as a discount sign).
[0319] The server establishes a component attribute table in the database to map visual representation icons to these attributes. During statistical analysis, the server constructs a statistical view grouped by attribute by connecting the visual representation table, attribute table, and behavior log table. For example, the server can calculate the overall selection rate of all images with "button color red" and compare it with other color groups such as "button color blue".
[0320] The server can implement the following algorithms in the statistics module: - Calculate the weighted selection rate for each component attribute; - When the selection rate of a certain attribute is significantly higher than the overall average and the sample size exceeds a preset threshold, the attribute is marked as a high selection rate component. - Calculate the joint selection rate based on multidimensional attribute combinations (e.g., "red button + left image and right text layout + large font price information"), and select attribute combinations that are significantly better than other combinations as high-selection-rate constituent element combinations.
[0321] Through this fine-grained, attribute-level statistical analysis, the server enables the system to automatically discover effective combinations of constituent elements within the computer, without requiring manual comparison. This process, using specific data structures (multidimensional attribute indexes, aggregation tables) and computational procedures (grouping, aggregation, ratio calculation, saliency judgment), directly improves the server's analytical efficiency and decision-making capabilities under large-scale visual data.
[0322] VI. Automatic Adjustment and Self-Learning Update of Prompt Statements After obtaining the high-selection-rate components or combinations of components, the server writes these results into the prompt message template configuration. For example, the server can reserve insertion points in the template for "high-selection-rate button color" and "high-selection-rate emphasized content." When generating prompt messages for new tasks, the server automatically fills in these priority components based on the current statistical results.
[0323] For example, when the server detects that the "red button" has the highest selection rate in recent statistics, the server can set the corresponding part of the template as follows: "Please use a prominent red 'Buy Now' button in the bottom right corner of the screen to make it stand out." This generates the following example of a prompt statement: "Please generate a banner ad for a summer beverage promotion, measuring 1200×628 pixels, with blue as the main color. The left side should display an image of the iced beverage, and the right side should display 'Cool Summer, Limited Time 20% Off' in large white font. Please use a striking red 'Buy Now' button in the lower right corner of the image to make it stand out." The server automatically modifies the specific text of the prompt statement template, making the constraints on high-selection-rate elements in the model input more explicit. This avoids adjustments based solely on human experience, allowing the prompt statement generation logic to adaptively evolve with behavioral data. This self-learning update process transforms behavioral logs into parameter updates for template configurations, representing automatic maintenance of internal computer rules and parameters, distinct from simple automation that manually sets fixed rules.
[0324] VII. Technical Effects and Causal Relationships Through the aforementioned structured data storage, attribute-level statistical analysis, and templated prompt generation mechanisms, the server brings about the following effects and causal relationships at the computer technology level: 1. Improved processing efficiency The server indexes and aggregates behavior logs and constituent attribute tables, enabling access to only necessary fields during grouped statistics and reducing the number of full table scans. Since high-frequency calculations are centralized in server-side aggregation operations, the terminal only needs to report simple events, thus reducing terminal load and network bandwidth consumption.
[0325] 2. Improved generation accuracy and conversion effect The server adjusts the prompts based on statistical results at the component level, giving the generative AI model more explicit and constrained generation conditions during the input phase. Because the prompts explicitly indicate high-selection-rate components, the model generates images that better conform to historically successful design patterns, thereby increasing user selection rates. This increase in selection rate can be verified through statistical significance tests.
[0326] 3. Improved utilization of computing resources Guided by behavioral data, the server gradually reduces attempts at combinations of constituent elements with obviously low selection rates, and concentrates computing resources on combinations of constituent elements with high potential. At the same time, through batch statistics and template updates, it can converge to a better prompt statement configuration in a few iterations, thereby shortening the exploration time from initial design to efficient design and reducing the number of redundant generation.
[0327] 4. Improvements in data management and model input control The server integrates generation conditions, behavioral results, and constituent element attributes into a structured database for unified management, simplifying subsequent queries, backtracking, and version control. This data management approach directly improves the traceability and adjustability of model inputs, making it easier to use historical data as a training set to expand or fine-tune the model when needed.
[0328] 5. Automatic generation of non-human rules In the statistics module, the server automatically determines which components should be considered as priority components based on predetermined evaluation criteria (such as selection rate threshold and sample size threshold). These rules are not predefined manually but generated by the system based on a large amount of interaction data. In this process, the server executes specific algorithmic steps to automatically generate a set of rules for controlling the construction of prompt statements, achieving an optimization path that goes beyond traditional human experience.
[0329] VIII. Alternative Implementation Forms and Expansion Servers can employ different generative artificial intelligence model structures in different implementations, for example: - Uses an image generation model based on an autoregressive network, which predicts image content pixel-by-pixel or block-by-block; - Using a multimodal transformer model, text and images are uniformly mapped to a shared representation space, and the image generation process is controlled through a cross-attention mechanism.
[0330] Different statistical methods can be selected for servers in different implementation forms, for example: - Using decision tree-based or gradient boosting models, learn a predictive function for selection rate from the attributes and behavioral statistics of constituent elements, and inversely deduce the optimal combination of constituent elements based on this function; - Using clustering methods, visual presentations are classified into multiple groups with similar appearances, and constituent elements are extracted and prompts are adjusted at the group level.
[0331] In another implementation, the terminal can be equipped with a local caching module to cache frequently used visual representations, thereby reducing the amount of data that the server repeatedly transmits for the same images, and further reducing the network communication burden.
[0332] In different implementations, users can directly edit certain parts of the prompts. The server retains the user-edited parts and automatically adjusts only the remaining parts related to the elements with high selection rates, thus achieving a balance between user subjective creativity and automatic system optimization.
[0333] In summary, this invention transforms the invocation of generative artificial intelligence models from a traditional "single, static, black-box" invocation to a "multi-round, dynamic, data-feedback-closed-loop" invocation process by introducing structured attribute management, behavior log statistics, high-selectivity component extraction, and templated prompt statement self-learning and updating mechanisms on the server side. This achieves a comprehensive improvement in generation accuracy, processing efficiency, and resource utilization at the computer technology level.
[0334] use Figure 13 The processing flow is explained.
[0335] Step 1: Users input generation conditions and instructions using the terminal.
[0336] On the terminal interface, users can input attribute information such as the size of the visual presentation (e.g., "1200×628 pixels"), main color (e.g., "blue"), layout (e.g., "image on the left, text on the right"), and target purpose (e.g., "new product promotion") through a keyboard, touch screen, or pointing device. They can also enter natural language instructions in the text input box, such as: "Create a promotional banner for the newly launched smartwatch, highlighting its health monitoring function and long battery life, with an overall high-tech feel." Input: Attribute information (size, color, layout, purpose, etc.) and natural language instructions.
[0337] Output: Structured request data generated within the terminal.
[0338] Step 2: The terminal sends structured request data to the server.
[0339] The terminal encapsulates user-input attribute information and natural language instructions into structured data objects, such as request messages containing multiple key-value pairs, and sends them to the interface address provided by the server via a network communication protocol (such as HTTPS). Before sending, the terminal can perform simple encoding processing on the text (such as UTF-8 encoding) and attach a session identifier.
[0340] Input: Structured request data saved by the user in the terminal.
[0341] Output: The request message transmitted to the server via the network channel.
[0342] Step 3: The server parses the request and performs parameter validation and standardization.
[0343] After receiving the request message from the terminal, the server uses the backend application to parse the message, converting the JSON or other formatted data into internal data structures. The server checks according to preset rules whether the size is within the allowed range, whether the color value belongs to the supported list, and whether the layout identifiers are valid. It also performs security checks on the length and content of natural language instructions (e.g., filtering prohibited words). Simultaneously, the server converts natural language attributes such as "blue" into standard color values (e.g., "#0000FF") and converts "left image, right text" into internal enumeration values.
[0344] Input: The original request message from the terminal.
[0345] Output: A validated and standardized internal task data object (containing standardized dimensions, colors, layouts, and safety-filtered instruction statements).
[0346] Step 4: The server constructs a prompt statement based on the attribute information and the instruction statement.
[0347] The server reads attributes such as target purpose, size, main color, and layout from the internal task data object, and selects the corresponding prompt template from the template store. The server fills the attributes into the placeholder positions in the template and appends the user's natural language instruction to the end of the template or a specified position, thus automatically generating the complete prompt text. For example, the server generates: "Please generate a banner advertisement for promoting a new product, with a size of 1200×628 pixels, a blue main color, and a layout of image on the left and text on the right. Create a promotional banner for a newly launched smartwatch, highlighting its health monitoring functions and long battery life, with an overall high-tech feel." Input: Standardized attribute information, natural language instructions, and prompt templates.
[0348] Output: The complete prompt text for the generative artificial intelligence model can be directly input.
[0349] Step 5: The server records the generated tasks and the metadata of the prompt statements.
[0350] The server writes the task identifier, user identifier, timestamp, attribute information, and pre-constructed prompts to the database or other persistent storage. The server can also pre-define element attribute fields (such as primary color, button color, and layout type) in the data records for later statistical analysis. During the write process, the server performs index updates and transaction commits to ensure record consistency and retrieval.
[0351] Input: Task data object and generated prompt statement.
[0352] Output: One or more task records in the database, including prompts and related metadata.
[0353] Step 6: The server invokes a generative artificial intelligence model to perform text encoding.
[0354] The server passes the prompt as an input string to the generative AI model deployed on computing resources. The server first calls the text encoding module in the model (such as a transformer-based encoding network) to break down the prompt into a sequence of tokens, and then performs embedding mapping and multi-layer self-attention operations on each token, ultimately obtaining a fixed-length or variable-length high-dimensional text embedding vector. This process performs numerous matrix multiplications and normalization operations on the graphics processing unit.
[0355] Input: The complete prompt text.
[0356] Output: A high-dimensional text embedding vector representing the semantics of the prompt statement.
[0357] Step 7: The server performs diffusion generation operations in the latent space.
[0358] The server samples an initial noise vector in the latent space and inputs this noise along with the text embedding vector into the U-shaped network of the diffusion model. The server iteratively performs denoising operations at multiple time steps: at each step, it uses operators such as convolution, downsampling, upsampling, and cross-layer connections to predict the next latent representation based on the current noise estimate and text conditions. In each iteration, the server uses activation functions and residual connections to perform nonlinear transformations, gradually approximating the latent representation from random noise to structured features that conform to the semantics of the prompt.
[0359] Input: text embedding vector, initial random noise vector, number of diffusion steps, and other generation parameters.
[0360] Output: The converged latent image representation (latent spatial feature tensor).
[0361] Step 8: The server decodes the latent representation into a pixel image and performs post-processing.
[0362] The server inputs the obtained latent representation into the decoding network, generating an RGB pixel matrix through deconvolution or upsampling operations. The server then uses an image processing library to perform size checks and necessary scaling on the generated image, ensuring the image size matches the user-specified dimensions (e.g., 1200×628 pixels). The server also sets the compression quality and format encoding for the image, generating PNG or JPEG image files.
[0363] Input: Latent image representation.
[0364] Output: Image file data conforming to the specified size and format.
[0365] Step 9: The server stores the visual representations and generates access addresses.
[0366] The server writes the generated image file to object storage or the file system, and records the path or Uniform Resource Locator (URL) in the task record in the database. Simultaneously, the server updates the task status to "generated" and generates an address string for the image that can be accessed by the terminal.
[0367] Input: Image file data, task identifier.
[0368] Output: The location identifier and accessible address of the image file in the storage system.
[0369] Step 10: The server returns the generated result information to the terminal.
[0370] The server constructs a response message, encapsulating information such as the task identifier, image access address, and image size into structured response data, and sends it to the terminal over the network. Before returning the response, the server can attach some metadata (such as primary color and layout identifier) according to the configuration, so that the terminal can perform simple categorization and display on the interface.
[0371] Input: Image access address, relevant metadata from the task record.
[0372] Output: The response message returned to the terminal.
[0373] Step 11: The terminal receives the generated results and displays the visual representation.
[0374] The terminal parses the response message returned by the server, extracting the image access address and size information. The terminal requests the image file from the server or content distribution node via HTTP or other protocols, and then renders the image on the interface at a suitable size and layout. The terminal displays buttons such as "Use this scheme" and "Generate another image" around the interface for user interaction.
[0375] Input: Server response message, image access address.
[0376] Output: Visual presentations and related operation controls displayed on the terminal display.
[0377] Step 12: Users can select or rate visual presentations on the device.
[0378] Users can browse multiple candidate visual presentations on the terminal, click on an image or its associated "Use this scheme" button to indicate selection, or click the "Generate another image" button to request new candidate results, or rate and comment on the current image.
[0379] Input: Visual representations and control elements displayed on the terminal interface.
[0380] Output: Selection events, regeneration requests, or evaluation information generated by user actions.
[0381] Step 13: The terminal reports and displays event logs and selects event behavior logs.
[0382] Each time a visual object is drawn onto the interface, the terminal reports a display event; each time the user clicks or confirms use, it reports a selection event. The terminal encapsulates these events into behavior log records, including the visual object identifier, event type (display / selection), timestamp, and optional user identifier, and sends them to the server's statistics interface via the network.
[0383] Input: The user's actual browsing and clicking actions.
[0384] Output: Behavior log data messages sent to the server.
[0385] Step 14: The server records behavior logs and updates the count field.
[0386] After receiving the behavior logs uploaded by the terminal, the server locates the corresponding record in the database based on the visual representation identifier and performs an accumulation operation on the display count or selection count field. The server can use an accumulation buffer in memory to merge and accumulate multiple logs within a short period of time before writing them in batches, thereby reducing the number of disk accesses.
[0387] Input: Behavior log records reported by the terminal.
[0388] Output: The updated display count and selection count fields in the database.
[0389] Step 15: The server calculates the selectivity rate metric for each visual presentation.
[0390] The server periodically or under triggered conditions reads the display count and selection count fields, and performs a selection rate calculation for each visual, which is the number of selections divided by the number of displays. The server can use a floating-point arithmetic library to accurately calculate the ratio and store the result in a separate metrics table or update the original record. The server can also filter data based on time windows to calculate the recent selection rate.
[0391] Input: The number of times each visual element is displayed and selected.
[0392] Output: Selectivity value for each visual representation.
[0393] Step 16: The server aggregates and calculates selection rates based on the attributes of the constituent elements.
[0394] The server reads information such as the primary color attribute, button color attribute, text size attribute, and layout type attribute for each visual element from the component attribute table, and associates these attributes with the selection rate. The server groups and aggregates by attribute or attribute combination, and calculates statistics such as the average selection rate, sample size, and variance for each group. For example, it calculates the average selection rate for all visual elements with "button color = red".
[0395] Input: Component attribute data, selection rate of each visual representation.
[0396] Output: Statistical results (average selection rate, sample size, etc.) for each component attribute group and attribute combination.
[0397] Step 17: The server extracts high-selectivity components and preferred combinations.
[0398] The server evaluates the statistical results obtained in step 16 based on preset thresholds or algorithm rules. When the average selection rate of a component attribute or attribute combination is higher than the global average by a certain percentage and the sample size exceeds a set threshold, the server marks the component or combination as a high-selection-rate component. The server writes these marking results into the rule configuration table for later reference when constructing prompt statements.
[0399] Input: Statistical results aggregated by attribute, preset threshold, and judgment rules.
[0400] Output: A set of rules marked as high-selectivity components and combinations of high-selectivity components.
[0401] Step 18: The server automatically adjusts the prompt message template based on elements with high selectivity.
[0402] The server reads the elements marked as having a high selection rate from the rule configuration table and embeds the relevant conditions into the placeholder section of the prompt statement template. For example, when the "red button" is marked as a high selection rate element, the server adds the statement to the template: "Please use the eye-catching red 'Buy Now' button in the lower right corner of the screen to make it stand out." When updating the template, the server retains the original structure and only inserts or replaces content in specified segments to ensure the overall sentence structure is stable.
[0403] Input: High selectivity component rule set, existing prompt statement template.
[0404] Output: An updated suggestion statement template containing constraints on elements with high selectivity.
[0405] Step 19: The server generates an updated prompt statement based on the updated template and the new task input.
[0406] When the server receives a new generation request, it uses the updated template from step 18 and combines it with the attribute information of the new task and the user's instructions to generate the prompt message again. For example, in a summer beverage promotion scenario, the server generates: "Please generate a banner ad for a summer beverage promotion, with a size of 1200×628 pixels, using blue as the main color, displaying an image of iced drinks on the left, and displaying 'Cool Summer, Limited Time 20% Off' in large white font on the right. Please use a striking red 'Buy Now' button in the lower right corner of the screen to make it very prominent." Input: Updated template, attribute information of new tasks, and natural language instructions.
[0407] Output: Update suggestion text reflecting the components with historically high selection rates.
[0408] Step 20: The server uses the updated prompts to invoke the generative AI model again to generate new visual representations.
[0409] The server uses the update prompt generated in step 19 as input to the model for the next round, repeating the text encoding, diffusion generation, decoding, and post-processing process to generate a new set of visual representation candidates. Since the prompt explicitly includes high-selectivity components, the server assigns higher weights to these features during model execution, thereby guiding the model to generate images in the latent space that better conform to statistically preferred components.
[0410] Input: Updated prompt statement, generated model parameters.
[0411] Output: A new set of visual representation data that incorporates historical statistical experience.
[0412] Step 21: The terminal receives new visual presentations and guides the user to make a new round of choices.
[0413] After receiving the new image access address from the server, the terminal loads and displays the new visual presentation, while simultaneously displaying a message on the interface stating, "The system has automatically optimized the button colors and layout based on historical data." The user then makes another selection and evaluation on the terminal, which then reports the behavior log again.
[0414] Input: The new image access address returned by the server, and the new visual representation.
[0415] Output: A new round of user interaction behavior and behavior log data, for use in subsequent iterations.
[0416] Application Example 2 The process flow corresponding to the specific processing in Use Case 2 will be described below. The various parts of the system described below are implemented by the data processing device 12 and the intelligent device 14. In addition, the data processing device 12 is referred to as the "server" and the intelligent device 14 is referred to as the "terminal".
[0417] With the widespread adoption of online services and interactive interfaces, the generation and optimization of visual information, such as advertising images, interface banners, and recommendation tiles, now largely rely on computer systems for automated processing. However, existing technologies suffer from the following technical problems: First, most systems generate visual information using only static templates or a few preset rules, lacking a mechanism for dynamic optimization based on user behavior data and contextual data. This results in limited appeal of visual information to users, making it difficult to continuously improve click-through rates and conversion rates. Second, even when generative AI models are introduced, fixed generation instructions are often manually written. The system cannot automatically update these instructions based on historical performance data, making the generation effect highly dependent on human experience. The computer itself lacks adaptive optimization capabilities for the generation process. Third, existing systems typically only count simple click or display counts without structurally modeling the relationship between these behavioral data and visual elements (size, color scheme, layout, text structure, etc.). Furthermore, they do not automatically adjust the content of "prompt statements" at the model level, thus failing to form a closed loop of "behavioral feedback—element evaluation—prompt statement update—regeneration" within the computer. Fourth, in most cases, existing technologies do not incorporate the user's emotional state into the visual information generation and sorting process. The system cannot automatically change the visual content and presentation order based on real-time emotional states, resulting in insufficient adaptability to the user's current psychological state and low user experience and interaction efficiency.
[0418] Therefore, the technical challenge this invention aims to address is to provide a novel information processing mechanism at the computer system level that enables the server to: automatically acquire abstract attribute information input by the user, construct and adjust prompt statements for generative artificial intelligence models; quantitatively evaluate and store the effectiveness of visual components as evaluation information based on statistical results of display and selection events; further combine emotional state inference results with machine learning models to predict more favorable element combinations under different emotional and attribute conditions, and automatically generate or correct prompt statements accordingly; thereby achieving adaptive optimization of the visual information generation process and improving resource utilization efficiency within the server in a programmatic manner.
[0419] The specific processing performed by the specific processing unit 290 of the data processing apparatus 12 in Application Example 2 is achieved by the following means.
[0420] In this invention, the server includes means for acquiring abstract attribute information about the size and constituent elements of visual information from a user in an information processing device; means for constructing a prompt statement for instructing a generative artificial intelligence model to generate visual information based on the abstract attribute information and previous prompt history information; means for inputting the prompt statement into the generative artificial intelligence model to obtain multiple visual information candidates; means for presenting the visual information candidates to the user's terminal through a display device and acquiring display event information and selection event information related to the presentation; and means for calculating the index values corresponding to each visual information candidate based on the display event information and the selection event information, and comparing the index values with the constituent elements. The system includes a device for storing evaluation information on the correspondence between elements in a recording medium; a device for automatically generating a second prompt statement based on the evaluation information, adjusting at least a portion of the adoption order, color scheme, layout, and text structure of the constituent elements, and instructing the generative artificial intelligence model to regenerate using the second prompt statement; a device for presenting the visual information obtained through the regeneration back to the user terminal; and a device for inferring an emotional state based on facial expression or voice information obtained from the user terminal, using the emotional state and the evaluation information together as learning data for a machine learning model to predict favorable combinations of constituent elements under a given emotional state and abstract attribute information, and automatically generating or correcting the prompt statement accordingly. This allows for the formation of an adaptive generation control mechanism centered on prompt statements within the server, enabling the computer system to programmatically optimize the input of the generative artificial intelligence model using behavioral and emotional data, achieving automatic closed-loop optimization of the visual information generation process, improving the relevance and effectiveness of visual content generation, reducing the intensity of human intervention, and thus improving human-computer interaction efficiency and resource utilization efficiency at the computer technology level.
[0421] A "system" refers to an overall technical solution consisting of one or more information processing devices, storage devices, and user terminals connected in communication with them, used to perform visual information generation, presentation, and optimization processing.
[0422] "Information processing device" refers to an electronic device with a processor, memory and communication interface, used to execute program instructions to achieve functions such as data acquisition, computation and processing, model calling and result output.
[0423] "User entity" refers to an individual or organizational entity that provides input information to the system and receives visual information output through the use of a terminal, including but not limited to ordinary users, operators or content providers.
[0424] "User terminal" refers to an electronic device operated by a user for interacting with information processing devices and presenting visual information, including but not limited to mobile terminals, fixed terminals, or wearable terminals.
[0425] "Visual information" refers to content data that can be presented on a display device and perceived by the user through visual means, including but not limited to images, icons, banners, advertising images, or interface components.
[0426] "Size" refers to a metric used to define the spatial extent of visual information in a display area, including but not limited to width, height, number of pixels, or proportional relationship.
[0427] "Constituent elements" refer to the various basic attributes or components used to make up visual information, including but not limited to color, color scheme, image elements, text content, font style, button elements, and layout structure.
[0428] "Abstract attribute information" refers to high-level attribute data that is input by the user or derived by the system to describe the overall characteristics of visual information rather than specific pixel data. This includes, but is not limited to, theme type, target audience, style preference, color preference, and size category.
[0429] "Prompt statements" refer to textual instructions provided to generative artificial intelligence models by information processing devices. These instructions describe the target features and constraints of the visual information to be generated, thereby guiding the model to generate output data that meets the requirements.
[0430] "Generative artificial intelligence models" refer to models built based on machine learning methods that can automatically generate new visual information or related content data based on input prompts, including but not limited to deep learning-based image generation models or multimodal generation models.
[0431] "Visual information candidates" refer to the various visual information instances that are yet to be evaluated and selected from multiple different visual information results generated by a generative artificial intelligence model under a given prompt.
[0432] "Display device" refers to an output device used to present visual information for viewing by a user, including but not limited to a display screen, projection device, or head-mounted display.
[0433] "Display event information" refers to recorded data related to the presentation of visual information on the user terminal, including but not limited to visual information identifiers, presentation time, presentation location, and terminal environment parameters.
[0434] "Selection event information" refers to the recorded data related to the user's interactive behavior such as selecting, clicking or confirming visual information, including but not limited to the selected visual information identifier, operation time, operation type and session identifier.
[0435] "Indicator value" refers to a quantitative evaluation parameter calculated for each visual information candidate based on the display event information and the selection event information, including but not limited to click-through rate, selection rate, conversion rate, or overall score.
[0436] "Evaluation information" refers to structured data that represents the correspondence between the components of visual information and their corresponding index values, and is used to reflect the degree of influence of each component or combination of components on the performance results.
[0437] "Recording medium" refers to a non-temporary storage medium used to store evaluation information, prompts, model parameters or other related data, including but not limited to magnetic media, optical media, flash memory or database systems.
[0438] "Second prompt statement" refers to an updated prompt statement that is generated based on the initial prompt statement and by adjusting at least some of the elements, such as the order of use, color scheme, layout, and text structure, according to the evaluation information generated by the information processing device.
[0439] "Regeneration" refers to the process of using a second prompt statement to re-invoke the generative artificial intelligence model in order to generate new visual information that is different from and optimized from the previous visual information candidates.
[0440] "Facial expression information" refers to images or feature data that reflect the facial expression characteristics of the user after being acquired and processed by a camera device, and is used to infer the user's emotional state.
[0441] "Voice information" refers to audio or feature data that is acquired through an audio acquisition device and processed to reflect the speech content and acoustic characteristics of the user, and is used to infer the emotional state of the user.
[0442] "Emotional state" refers to the state information inferred by the system based on facial or vocal information, used to represent the user's psychological and emotional tendencies, including but not limited to types and intensities such as joy, calmness, tension, or depression.
[0443] "Machine learning model" refers to a computational model that obtains parameters by learning from historical data, and is used to predict the output result given input conditions. In this invention, it is used to predict the favorable combination of constituent elements under specific emotional states and abstract attribute information.
[0444] "Combination of constituent elements" refers to a specific combination scheme of a set of constituent elements that are used together when generating visual information, including but not limited to the combined configuration of specific color schemes, layout methods, image types, and text styles.
[0445] In the following embodiments, the subject is limited to "server," "terminal," and "user," and the description focuses on the system described in the claims. This invention is not limited to the specific embodiments described below; those skilled in the art can make various modifications and substitutions without departing from the spirit of the invention.
[0446] In one implementation, a server is deployed as an information processing device in a data center or cloud computing environment. The server includes a processor, main memory, non-volatile storage, and a network interface. The server runs an operating system (e.g., a general-purpose server operating system), a database management system (e.g., a relational database management system or a document-oriented database management system), application server software (e.g., an HTTP-based backend framework), and libraries for machine learning and generative artificial intelligence model inference (e.g., a deep learning framework based on tensor operations).
[0447] In one implementation, a terminal includes a smartphone, tablet, or computer terminal. The terminal's hardware includes at least a display device, a camera device, an audio acquisition device, a memory, and a communication module. Its software runs a browser or mobile application for bidirectional communication with a server, presenting visual information, and collecting user behavior and emotion-related data.
[0448] In one implementation, a user operates a graphical user interface (GUI) via a terminal. The user inputs visual information, including dimensions, constituent elements, and abstract attributes such as width, height, background color, theme type, and target audience. In some implementations, the user allows the terminal to access a camera and microphone, enabling the server to infer the user's emotional state based on facial or voice information.
[0449] After receiving user input, the server stores the size and component information as structured records. In one implementation, the server uses a relational database, representing each visual information request as a record with fields including request identifier, timestamp, user identifier, width, height, color preference, target audience tags, and topic tags. In another implementation, the server uses a document-oriented database, storing the above fields as key-value pairs in a document.
[0450] When generating prompt statements, the server reads abstract attribute information and historical prompt records from the database. The server maintains several prompt statement templates in the program; these templates describe the image generation target in natural language. Based on the user-inputted size, theme, and constituent elements, the server inserts variables into the templates to form a complete prompt statement. In one embodiment, the server generates the following example prompt statement: "Please generate a 300x250 pixel banner ad with a blue background, a product image in the center, and a prominent 'Buy Now' button below. The overall style should be bright and modern." In another embodiment, the server generates the following example of a prompt message for a specific group of people: "Please generate a 300x250 pixel fashion product advertisement banner for women in their 20s, with a light pink background, clear product images, and concise discount copy. The overall style should be light and youthful." In other implementations, the server uses historical performance data to enhance the prompt statements. For example, after analyzing that the combination of a light blue background and a large button has a high click-through rate, the server adds constraints to the prompt statements: "Please try to use light blue backgrounds and large buttons, as these elements have a higher click-through rate in historical ads." When the server invokes the generative AI model, it communicates with the model inference service via a network interface. In one implementation, the server uses an image generation model based on a deep convolutional generative adversarial network (GAN) structure or a diffusion model structure. The server sends the prompt as text input and its size as a control parameter to the model. Internally, the generative AI model first text-encodes the prompt. The server then uses a word embedding layer within the model to convert the prompt into a vector sequence, extracting high-level semantic features via a multi-layer attention network or convolutional network. During the image generation stage, the model uses a deconvolutional network, an upsampling module, or a diffusion decoder to map the semantic features into a two-dimensional pixel matrix.
[0451] The server accelerates tensor operations using the GPU during model inference. In one embodiment, the server groups batches of prompts into mini-inputs, thereby generating multiple images in a single forward propagation. After receiving the generated results, the server compresses and encodes the images, generating thumbnails of different sizes, and assigns a unique identifier and metadata record to each generated result.
[0452] In another embodiment, the server simultaneously uses a text generation model to generate title text or explanatory text. The server sends the following example prompt to the text generation model: "Please write an advertising headline and subheading for a new summer product that is suitable for women in their 20s, with a light and positive tone." The server associates and stores the text generation results with the image generation results for combined presentation on the terminal.
[0453] After receiving multiple visual information candidates from the server, the terminal displays these candidate options as thumbnails on the graphical interface. The terminal assigns an interactive area to each candidate in the visual layout, allowing the user to click and select. When each candidate is displayed, the terminal sends display event data back to the server, including the candidate identifier, terminal identifier, timestamp, and interface position. When the user clicks on a candidate, the terminal generates a selection event and reports the candidate identifier and session identifier to the server.
[0454] After receiving display and selection events, the server maintains a statistics table in the database. The server records the number of times each candidate is displayed and clicked in the statistics table. During periodic calculations, the server calculates click-through rate and other metrics for each candidate and writes these metrics into an evaluation information table. The server stores the correspondence between "component vectors" and "metric values" in this evaluation information table, where the component vectors include features such as color coding, layout type coding, text length, and button size.
[0455] In one implementation, the server uses a machine learning model to analyze and evaluate information. The server uses a multilayer perceptron or gradient boosting tree as the prediction model. The server takes the features of each candidate feature as an input vector and the corresponding click-through rate or whether it was clicked as a supervision signal. During training, the server uses cross-entropy loss or mean squared error as the error function and updates the model weights using stochastic gradient descent or adaptive learning rate optimization algorithms. During training, the server partitions the samples, introduces a validation set to prevent overfitting, and persists the model parameters in a model parameter library after training.
[0456] In other implementations, the server uses an attention mechanism to analyze the importance of different components. The server introduces attention weights within the model, allowing it to assign different weights to different feature dimensions when predicting click-through rates, thereby automatically identifying the relative contribution of features such as "background color," "button size," and "text length" to user behavior. During the inference phase, the server calls upon this model to score new component combinations, guiding the generation of the next round of prompts.
[0457] When generating the second prompt statement, the server reads the latest evaluation information and model scoring results. The server sorts the weights of each element, selecting elements with higher weights to add to the new prompt statement, while reducing or removing elements with lower weights or poorer performance. In one embodiment, the server generates the following second prompt statement based on the analysis results: "Please generate a 300x250 pixel advertising banner, prioritizing a light blue background and large rounded corner buttons. The button text should be concise and impactful, the overall layout should be simple, and the product image should be centered." In an implementation related to emotion inference, the server processes facial expression images and speech data collected from the terminal. The terminal sends the collected raw or compressed images to the server, which performs face detection and expression classification on the image frames, either locally or through an external emotion analysis service. In one implementation, the server uses a convolutional neural network to extract features from the facial region and uses fully connected layers to output a multidimensional emotion probability distribution (such as joy, calmness, tension, etc.). For speech data, the server can perform acoustic feature extraction (e.g., Mel-frequency cepstral coefficients) and input it into the emotion classification network. The server then fuses the facial expression classification results with the speech emotion results to obtain the final emotion state label and its confidence level.
[0458] When generating prompts based on emotional states, the server embeds an emotion-dependent description into the text template. For example, when the server detects that a user is in an excited state, it generates the following prompt: "The user is currently excited. Please generate a 300x250 pixel ad banner with the content of new product launch and limited-time sale. Use vibrant colors and high contrast to encourage the user to click immediately." When the server detects that a user is in a relaxed state, it generates the following message: "The user is currently in a relaxed state. Please generate a 300x250 pixel advertising banner with content related to vacation travel or relaxation services. The colors should be soft, the layout simple, and the overall atmosphere comfortable." In another embodiment, the server inputs emotional state and evaluation information as features into the machine learning model. The server assigns each candidate record a triple of "emotional label—component vector—click result". During model training, the server encodes the emotional label as an additional feature dimension, allowing the model to learn the optimal combination of elements under different emotional conditions. During inference, given emotional states such as "joy" or "tension" and abstract attribute information, the server predicts the component combinations with higher click-through rates and automatically generates corresponding prompts accordingly.
[0459] In multiple embodiments, the server employs specific data structures and algorithm sequences to achieve technical improvements within the computer. The server uses an index structure in its evaluation information storage to accelerate the querying of statistical data by dimensions such as color and layout combination. During model training, the server prioritizes incremental learning, updating only newly added data intervals to reduce retraining costs and lower processor load. The server uses a template-weighted combination algorithm in prompt generation. Based on the element weights output by the model, the server automatically selects different sub-templates and concatenates them to generate the final text. This weight-based template concatenation method, compared to simple rule replacement, reduces the enumeration of invalid combinations, improving the efficiency and consistency of prompt construction.
[0460] The server achieves several improvements in technical performance through the aforementioned structure and algorithms. It systematically models the relationship between visual elements and behavioral data using machine learning models, significantly improving element selection accuracy compared to manual experience-based rules, thereby enhancing click-through rate prediction and visual content matching. The server shortens the model update cycle through batch inference and incremental training, enabling rapid responses to behavioral feedback and achieving higher throughput with the same hardware resources. By incorporating emotional states into the feature space, the server allows for fine-grained adaptation of the model along the psychological state dimension, reducing invalid displays and lowering network bandwidth and terminal rendering resource consumption when many users access the platform concurrently.
[0461] In several alternative implementations, the server can employ different generative artificial intelligence model architectures. For example, in one implementation, the server uses a diffusion model, which generates images internally through a multi-step noise removal process, and the server can control intermediate features at each step based on prompts. In another implementation, the server uses an autoregressive image generation network, outputting images by generating them pixel-by-pixel or block-by-block. During the text encoding stage, the server can employ recurrent neural networks or encoder structures based on self-attention mechanisms, selecting a suitable architecture to balance performance and resource consumption in different application environments.
[0462] In some implementations, the terminal caches some visual information candidates and their statistical data locally to reduce the frequency of interaction with the server. When a user frequently browses within a short period, the terminal can pre-arrange the display order locally based on the most recent sorting result sent by the server. Only when the statistical data exceeds a set threshold or the time interval meets a certain condition will the terminal request the server for a new round of optimization. This design technically reduces the number of network round trips, lowers the communication load, and improves the overall response speed.
[0463] The user's role in the system is primarily limited to providing high-level abstract attribute input and behavioral feedback. Internally, the server automatically constructs prompts and generates visual information through specific data structures, machine learning algorithms, and generation processes. Therefore, the server does not simply automate manual work; rather, it establishes a closed loop within the computer—"abstract attributes—prompts—output generation—behavioral feedback—element evaluation—prompt update"—forming a novel adaptive generation control method. In this method, the computer system autonomously learns and updates the prompt text itself used to control the generative artificial intelligence model, thereby bringing substantial technical improvements in processing speed, prediction accuracy, data management efficiency, and communication load.
[0464] use Figure 14 The processing flow is explained.
[0465] Step 1: The user inputs abstract attribute information on the terminal. Users open the application or webpage interface on their device and enter the dimensions, constituent elements, and abstract attributes of visual information in a form. Input examples include: width, height, background color, theme type, target audience, and desired text and image elements.
[0466] The terminal receives the fields entered by the user in the interface controls, combines these fields into a structured data object, and performs basic validation locally (such as checking whether the size is a positive integer and whether the color field is not empty).
[0467] The terminal sends the structured data as input in the form of a request message to the server via the network interface.
[0468] Input: User-inputted size data (e.g., 300×250), component data (e.g., background color, image, text, button), and abstract attribute data (e.g., theme, target audience, style preference).
[0469] Output: A request data packet containing the above fields sent to the server.
[0470] Step 2: The server parses the request and constructs internal data records. The server receives request data packets from the terminal, parses the packet body, and extracts the size field, constituent element field, and abstract attribute field.
[0471] The server maps these fields to internal data structures according to a predefined data pattern, creates a request object in memory, persists the object to the database, and assigns a request identifier to it.
[0472] During this process, the server can complete the default fields. For example, if the user does not specify a color, the server can select a color from the preset default values and fill it in.
[0473] Input: The request data packet sent by the terminal in step 1.
[0474] Output: Request records stored in the database and request objects residing in memory, including request identifiers and complete attribute fields.
[0475] Step 3: The server generates an initial prompt statement based on requests and history. The server reads the size, constituent elements, and abstract attribute information of the requested object from memory, and reads historical prompts and performance data related to the user or topic from the database.
[0476] The server selects a prompt template that matches the current business type, inserts the size, theme, and key components into the template through string replacement or placeholder filling, and generates an initial prompt in natural language.
[0477] The server packages the generated prompts and structured parameters into a "generate request unit" for later transmission to the generative artificial intelligence model.
[0478] Input: Size and attribute fields in the request object, and historical prompt records.
[0479] Output: The initial prompt text corresponding to the current request, and the generated request unit containing the prompt text and parameters.
[0480] Step 4: The server invokes a generative artificial intelligence model to generate visual information candidates. The server takes the generated request unit as input and sends the prompt and related parameters to the generative artificial intelligence model deployed on the inference service via network protocol.
[0481] When calling the interface, the server specifies parameters such as the output image size and the number of images generated, and performs necessary encoding or transcoding on the prompt text.
[0482] After receiving input, the generative artificial intelligence model encodes the prompts, feeds the text vectors into the internal neural network structure for semantic feature extraction, and performs an image decoding process to generate multiple image data.
[0483] The server receives image data from generative artificial intelligence models, performs format conversion, size correction, and compression on the images, and assigns a unique candidate identifier to each image.
[0484] Input: A generation request unit containing prompts and size parameters.
[0485] Output: Multiple visual information candidate images and their candidate identifier list.
[0486] Step 5: The server stores candidate data and generates terminal response data. The server writes the metadata (candidate identifier, associated request identifier, image storage path, and prompt statement used) of each visual information candidate into the database.
[0487] The server constructs a response data structure based on the candidate list, packaging the image access address, size information, candidate identifier, and brief description into a response message.
[0488] The server attaches a statistical tracking tag to each candidate in the response message for subsequent identification of display and click events.
[0489] Input: Visual information candidate images and candidate identifiers generated in step 4.
[0490] Output: A response message containing multiple candidate visual information, their identifiers, and tracking markers.
[0491] Step 6: The terminal receives and displays visual information candidates. The terminal receives a response message from the server, parses the candidate list, and reads the image address, size, and candidate identifier from it.
[0492] The terminal loads image data from remote storage, renders each candidate image onto the corresponding image control in the interface, and arranges them according to a predetermined layout rule.
[0493] The terminal configures click event handling logic for each candidate, binding the candidate identifier with the user's click operation.
[0494] Input: The candidate visual information response message returned by the server.
[0495] Output: Multiple visual information candidates that are displayed on the terminal interface, as well as user interaction events registered for each candidate.
[0496] Step 7: Terminal reports and displays event data When each candidate image completes rendering or enters the visible area, the terminal sends a display event message to the server.
[0497] The terminal writes fields such as candidate identifier, request identifier, display time, and terminal environment information (such as device type and screen resolution) into the event message.
[0498] The terminal transmits a set of display events back to the server in batches or in real time via the network interface.
[0499] Input: A list of candidates whose local rendering is complete and their identifiers.
[0500] Output: A report data packet containing multiple event logs.
[0501] Step 8: Users select candidates on the terminal. Users browse various visual information candidates on the terminal interface and select one or more candidates by clicking or touching.
[0502] The terminal captures the user's click action and searches for the corresponding candidate identifier and request identifier based on the clicked area.
[0503] The terminal constructs a selection event message based on this information, which includes the user operation time, candidate identifier, request identifier, and session information.
[0504] Input: The actual location and time of the user's click on the interface.
[0505] Output: Local selection event logs and selection event messages ready to be sent to the server.
[0506] Step 9: Terminal reports selection event data The terminal sends the constructed selection event message to the server via the network interface.
[0507] In some implementations, the terminal can buffer multiple selection events and then package them for reporting to reduce the number of requests.
[0508] Input: The selection event message generated in step 8.
[0509] Output: Selection event data packets sent to the server.
[0510] Step 10: The server records, displays, and selects data and calculates metric values. The server receives the display event data packet and the selection event data packet, and locates the corresponding record in the statistics table based on the candidate identifier and the request identifier.
[0511] The server accumulates the number of times each candidate is displayed and clicked, and updates the time of the most recent event.
[0512] In a periodic or triggered calculation program, the server takes the number of impressions and clicks as input, calculates the click-through rate, selection rate, and other metrics for each candidate, and writes these metrics back to the evaluation information table.
[0513] Input: Display event logs and selection event logs transmitted from the terminal.
[0514] Output: Updated statistical table records and evaluation information records containing indicator values.
[0515] Step 11: Server construction components and characteristics, forming evaluation information. The server reads the configuration of each candidate's constituent elements from the database, including fields such as color code, layout type, text length, image type, and button size.
[0516] The server encodes these fields into feature vectors according to predetermined rules, such as mapping colors to discrete integer codes, mapping layout types to one-hot vectors, and normalizing text lengths to floating-point numbers.
[0517] The server associates feature vectors with their corresponding index values to form "feature vector-index value" pairs, and writes them into the evaluation information table, thus forming a training data structure that can be directly used by machine learning models.
[0518] Input: Configuration of candidate components and record of indicator values.
[0519] Output: A set of evaluation information entries containing feature vectors and index values.
[0520] Step 12: Server training or updating machine learning models The server extracts a batch of samples from the evaluation information table, uses the feature vector as the model input, and uses the click rate or whether it was clicked as the supervision label.
[0521] The server sets the network structure during model initialization, for example, the input layer dimension is equal to the feature dimension, the hidden layer is one or more fully connected layers, the activation function is a rectified linear unit, and the output layer is a one-dimensional click probability.
[0522] During training, the server uses the cross-entropy loss function to calculate the error between the predicted output and the true label, and calculates the gradient of each weight parameter through the backpropagation algorithm.
[0523] The server uses stochastic gradient descent or adaptive optimization algorithms to update network parameters until the error on the validation set converges or the preset number of iterations is reached.
[0524] Input: Feature vectors and corresponding label values from the evaluation information table.
[0525] Output: The parameters of the trained and stored machine learning model.
[0526] Step 13: The server generates feature weights based on the model output. During the model inference phase, the server inputs the feature vectors of different combinations of constituent elements into the trained model to obtain the predicted click probability for each combination.
[0527] The server sorts the combinations based on the predicted probabilities, selects the combinations with higher predicted values, and counts the frequency of each component in the combinations with high predicted values.
[0528] The server calculates weight scores for elements such as color, layout, button size, and text length based on these statistical results, forming an element weight table.
[0529] Input: A set of feature vectors combining model parameters and constituent elements.
[0530] Output: A table of element weights containing the weights of each component.
[0531] Step 14: The server generates a second prompt statement and adjusts the generation strategy. When generating the second prompt statement, the server takes the element weight table as input, automatically selects the constituent elements with higher weights, and incorporates them into the prompt statement content.
[0532] The server replaces or adds descriptive statements in the text template, such as explicitly specifying high-weight elements like "light blue background," "large button," and "short copy."
[0533] The server generates a new second prompt text and packages it together with the original request attributes into a new generated request unit, which is then used to invoke the generative AI model again.
[0534] Input: Feature weight table and original request attributes.
[0535] Output: A second prompt statement that prioritizes efficient components and a new generation request unit.
[0536] Step 15: The server adjusts prompts and sorting strategies based on emotional state. The server receives the user's current emotional state label and its confidence level from the sentiment analysis module, and encodes the emotional label as feature input.
[0537] The server inserts emotion-related descriptions into the prompt templates based on different emotional states, such as "emphasizing bright colors and urgent text when excited" and "emphasizing soft colors and soothing themes when relaxed".
[0538] When sorting the ad list, the server feeds the emotional state as an additional dimension into the sorting model, calculates the expected click probability of different ads under the current emotional state, and adjusts the output order accordingly.
[0539] Input: User's current sentiment tag, element weight table, and ad candidate list.
[0540] Output: The prompts after emotion adaptation and the order of ad candidates after emotion optimization.
[0541] Step 16: The server calls a generative artificial intelligence model to regenerate. The server uses a second prompt statement and an emotion-enhancing description to form the final prompt statement, which is then sent to the generative AI model along with size parameters.
[0542] The server triggers a new round of image generation on the model side, re-executes text encoding, feature extraction, and image decoding operations within the model, and generates a new batch of visual information.
[0543] The server checks and preprocesses the newly generated image, and stores the results in association with the prompts in this round, as candidates for new visual information.
[0544] Input: The final prompt and size parameters after element optimization and emotion adaptation.
[0545] Output: Visual information candidate images and their candidate identifiers for the new batch.
[0546] Step 17: The terminal receives the regenerated result and displays it to the user. The terminal receives a new batch of visual information candidate response data from the server and parses the image address and associated information.
[0547] The terminal replaces or supplements old candidates with new ones on the interface, and then displays them again in the updated sorting order for users to browse and select.
[0548] The terminal continues to perform display event reporting and selection event reporting for these new candidates, thus forming a closed loop in the entire optimization process.
[0549] Input: Candidate data for regenerated visual information returned by the server.
[0550] Output: Optimized visual information presented in the terminal interface, and subsequent user behavior data reporting.
[0551] The specific processing unit 290 sends the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires sound representing user input regarding the result of the specific processing. The control unit 46A sends the sound data representing user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the sound data.
[0552] Data generation model 58 is a so-called generative AI (Artificial Intelligence). Examples of data generation models 58 include ChatGPT (registered trademark) (accessible via the internet (URL: https: / / openai.com / blog / chatgpt)). Data generation model 58 is obtained through deep learning on a neural network. Input to data generation model 58 are prompt words containing instructions, and inference data such as sound data representing sound, text data representing text, and image data representing images (e.g., still image data or animation data). Data generation model 58 infers from the input inference data based on the instructions represented by the prompt words and outputs the inference result in one or more data forms, such as sound data, text data, and image data. Data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or induction. The specific processing unit 290 performs the aforementioned specific processing while using data generation model 58. The data generation model 58 can also be a model finely tuned to output inference results from prompts that do not contain instructions. In this case, the data generation model 58 can output inference results based on prompts that do not contain instructions. The data processing apparatus 12, etc., includes various data generation models 58, including AI other than the generation AI. AI other than the generation AI can be, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or Naive Bayes, and can perform various processes, but is not limited to this example. Furthermore, the AI can also be an AI agent. Furthermore, when the processing of the above-mentioned parts is performed by AI, the processing can be performed partially or entirely by AI, but is not limited to this example. Furthermore, the processing performed by the AI including the generation AI can be replaced by processing in the rule base, and the processing in the rule base can also be replaced by processing performed by the AI including the generation AI.
[0553] Furthermore, the processing of the aforementioned data processing system 10 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the smart device 14, but it can also be performed by both the specific processing unit 290 of the data processing device 12 and the control unit 46A of the smart device 14. Additionally, the specific processing unit 290 of the data processing device 12 acquires or collects information required for processing from the smart device 14 or external devices, and the smart device 14 acquires or collects information required for processing from the data processing device 12 or external devices.
[0554] For example, the collection unit is implemented by the control unit 46A of the smart device 14 or the specific processing unit 290 of the data processing device 12. For example, the acquisition unit uses the camera 42 or communication I / F 44 of the smart device 14 to acquire step data, which is then processed by the specific processing unit 290 of the data processing device 12. For example, the analysis unit is implemented by the specific processing unit 290 of the data processing device 12, which analyzes the data from the collection unit and the acquisition unit. For example, the generation unit is implemented by the specific processing unit 290 of the data processing device 12, which uses a generation AI to generate a menu. For example, the serving unit is implemented by the output device 40 of the smart device 14 or the specific processing unit 290 of the data processing device 12, which provides the generated menu to the user. The correspondence between each unit and the device or control unit is not limited to the above examples and various changes can be made.
[0555] In the above embodiments, examples of specific processing by the data processing device 12 are given, but the technology disclosed herein is not limited to this, and specific processing may also be performed by the smart device 14.
[0556] Second Implementation Method Figure 3 An example of the configuration of the data processing system 210 according to the second embodiment is shown.
[0557] like Figure 3 As shown, the data processing system 210 includes a data processing device 12 and smart glasses 214. A server can be cited as an example of the data processing device 12.
[0558] The data processing apparatus 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" as understood in this disclosure. The computer 22 includes a processor 28, RAM 30, and memory 32. The processor 28, RAM 30, and memory 32 are connected to a bus 34. Furthermore, the database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0559] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, RAM 48, and memory 50. The processor 46, RAM 48, and memory 50 are connected to a bus 52. Furthermore, the microphone 238, speaker 240, camera 42, and communication I / F 44 are also connected to the bus 52.
[0560] Microphone 238 receives instructions from user 20 by receiving sounds emitted by user 20. Microphone 238 captures sounds emitted by user 20 and converts the captured sounds into sound data, which is then output to processor 46. Speaker 240 outputs sound according to instructions from processor 46.
[0561] Camera 42 is a small digital camera equipped with an optical system such as a lens, aperture and shutter, and imaging elements such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, which captures images of the user 20's surroundings (e.g., the field of view defined by an angle equivalent to the field of vision of an average healthy person).
[0562] Communication I / F44 is connected to network 54. Communication I / F44 and 26 are responsible for the transmission and reception of various information between processor 46 and processor 28 via network 54. The transmission and reception of various information between processor 46 and processor 28 using communication I / F44 and 26 is performed in a secure state.
[0563] Figure 4 This illustrates an example of the main functions of the data processing device 12 and the smart glasses 214. For example... Figure 4 As shown, in the data processing device 12, specific processing is performed by the processor 28. The specific processing program 56 is stored in the memory 32.
[0564] The specific processing program 56 is an example of a "program" involved in the technology of this disclosure. The processor 28 reads the specific processing program 56 from the memory 32 and executes the read specific processing program 56 on the RAM 30. Specific processing is implemented by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0565] The memory 32 stores a data generation model 58 and an emotion-specific model 59. The data generation model 58 and the emotion-specific model 59 are used by the specific processing unit 290. The specific processing unit 290 can use the emotion-specific model 59 to infer the user's emotions and perform specific processing based on the user's emotions. In the emotion inference function (emotion-specific function) using the emotion-specific model 59, various inferences and predictions related to the user's emotions are performed, including inferences and predictions of the user's emotions, but this is not limited to this example. Furthermore, emotion inference and prediction may also include, for example, emotion analysis (parsing).
[0566] In the smart glasses 214, the processor 46 performs reception and output processing. The memory 50 stores the reception and output program 60. The processor 46 reads the reception and output program 60 from the memory 50 and executes the read reception and output program 60 on the RAM 48. The reception and output processing is implemented by the processor 46 operating as a control unit 46A according to the reception and output program 60 executed on the RAM 48. Furthermore, the smart glasses 214 has the same data generation model and emotion-specific model as the data generation model 58 and the emotion-specific model 59, and these models can also be used to perform the same processing as the specific processing unit 290.
[0567] Next, the specific processing of the specific processing unit 290 of the data processing device 12 will be described. Each part of the system described below is implemented by the data processing device 12 and the smart glasses 214. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0568] Example 1 The process is the same as that of the specific process described in Embodiment 1 in the first embodiment above, so the description is omitted.
[0569] Application Example 1 The process is the same as that in the specific processing described in Application Example 1 of the first embodiment above, so the description is omitted.
[0570] Example 2 The process is the same as that of the specific process in Embodiment 2 described in the first embodiment above, so the description is omitted.
[0571] Application Example 2 The process is the same as that in the specific processing described in Application Example 2 of the first embodiment above, so the description is omitted.
[0572] The specific processing unit 290 sends the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A outputs the result of the specific processing to the speaker 240. The microphone 238 acquires sound input representing the user's input regarding the result of the specific processing. The control unit 46A sends the sound data representing the user's input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the sound data.
[0573] Data generation model 58 is a so-called generative AI (Artificial Intelligence). Examples of data generation models 58 include ChatGPT (registered trademark) (accessible via the internet (URL: https: / / openai.com / blog / chatgpt)). Data generation model 58 is obtained through deep learning on a neural network. Input to data generation model 58 are prompt words containing instructions, and inference data such as sound data representing sound, text data representing text, and image data representing images (e.g., still image data or animation data). Data generation model 58 infers from the input inference data based on the instructions represented by the prompt words and outputs the inference result in one or more data forms, such as sound data, text data, and image data. Data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or induction. The specific processing unit 290 performs the aforementioned specific processing while using data generation model 58. The data generation model 58 can also be a model finely tuned to output inference results from prompts that do not contain instructions. In this case, the data generation model 58 can output inference results based on prompts that do not contain instructions. The data processing apparatus 12, etc., includes various data generation models 58, including AI other than the generation AI. AI other than the generation AI can be, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or Naive Bayes, and can perform various processes, but is not limited to this example. Furthermore, the AI can also be an AI agent. Furthermore, when the processing of the above-mentioned parts is performed by AI, the processing can be performed partially or entirely by AI, but is not limited to this example. Furthermore, the processing performed by the AI including the generation AI can be replaced by processing in the rule base, and the processing in the rule base can also be replaced by processing performed by the AI including the generation AI.
[0574] Furthermore, the processing of the aforementioned data processing system 10 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the smart glasses 214, but it can also be performed by both the specific processing unit 290 of the data processing device 12 and the control unit 46A of the smart glasses 214. Additionally, the specific processing unit 290 of the data processing device 12 acquires or collects information required for processing from the smart glasses 214 or external devices, and the smart glasses 214 acquires or collects information required for processing from the data processing device 12 or external devices.
[0575] For example, the collection unit is implemented by the control unit 46A of the smart glasses 214 or the specific processing unit 290 of the data processing device 12. For example, the acquisition unit uses the camera 42 or communication I / F 44 of the smart glasses 214 to acquire step data, which is then processed by the specific processing unit 290 of the data processing device 12. For example, the analysis unit is implemented by the specific processing unit 290 of the data processing device 12, which analyzes the data from the collection unit and the acquisition unit. For example, the generation unit is implemented by the specific processing unit 290 of the data processing device 12, which uses a generation AI to generate a menu. For example, the serving unit is implemented by the speaker 240 of the smart glasses 214 or the specific processing unit 290 of the data processing device 12, which provides the generated menu to the user. The correspondence between each unit and the device or control unit is not limited to the above examples and various changes can be made.
[0576] In the above embodiments, examples of specific processing by the data processing device 12 are given, but the technology disclosed herein is not limited to this, and specific processing may also be performed by the smart glasses 214.
[0577] Third Implementation Method Figure 5 An example of the configuration of the data processing system 310 according to the third embodiment is shown.
[0578] like Figure 5 As shown, the data processing system 310 includes a data processing device 12 and a head-mounted terminal 314. A server can be cited as an example of the data processing device 12.
[0579] The data processing apparatus 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" as understood in this disclosure. The computer 22 includes a processor 28, RAM 30, and memory 32. The processor 28, RAM 30, and memory 32 are connected to a bus 34. Furthermore, the database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0580] The head-mounted terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and memory 50. The processor 46, RAM 48, and memory 50 are connected to a bus 52. Furthermore, the microphone 238, speaker 240, camera 42, display 343, and communication I / F 44 are also connected to the bus 52.
[0581] Microphone 238 receives instructions from user 20 by receiving sounds emitted by user 20. Microphone 238 captures sounds emitted by user 20 and converts the captured sounds into sound data, which is then output to processor 46. Speaker 240 outputs sound according to instructions from processor 46.
[0582] Camera 42 is a small digital camera equipped with an optical system such as a lens, aperture and shutter, and imaging elements such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, which captures images of the user 20's surroundings (e.g., the field of view defined by an angle equivalent to the field of vision of an average healthy person).
[0583] Communication I / F44 is connected to network 54. Communication I / F44 and 26 are responsible for the transmission and reception of various information between processor 46 and processor 28 via network 54. The transmission and reception of various information between processor 46 and processor 28 using communication I / F44 and 26 is performed in a secure state.
[0584] Figure 6 This illustrates an example of the main functions of the data processing device 12 and the head-mounted terminal 314. For example... Figure 6 As shown, in the data processing device 12, specific processing is performed by the processor 28. The specific processing program 56 is stored in the memory 32.
[0585] The specific processing program 56 is an example of a "program" involved in the technology of this disclosure. The processor 28 reads the specific processing program 56 from the memory 32 and executes the read specific processing program 56 on the RAM 30. Specific processing is implemented by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0586] The memory 32 stores the data generation model 58 and the emotion-specific model 59. The data generation model 58 and the emotion-specific model 59 are used by the specific processing unit 290.
[0587] In the head-mounted terminal 314, the processor 46 performs the acceptance / output processing. The memory 50 stores the acceptance / output program 60. The processor 46 reads the acceptance / output program 60 from the memory 50 and executes the read acceptance / output program 60 on the RAM 48. The acceptance / output processing is implemented by the processor 46 operating as a control unit 46A according to the acceptance / output program 60 executed on the RAM 48.
[0588] Next, the specific processing of the specific processing unit 290 of the data processing device 12 will be described. Each part of the system described below is implemented by the data processing device 12 and the head-mounted terminal 314. In the following description, the data processing device 12 will be referred to as the "server" and the head-mounted terminal 314 will be referred to as the "terminal".
[0589] Example 1 The process is the same as that of the specific process described in Embodiment 1 in the first embodiment above, so the description is omitted.
[0590] Application Example 1 The process is the same as that in the specific processing described in Application Example 1 of the first embodiment above, so the description is omitted.
[0591] Example 2 The process is the same as that of the specific process in Embodiment 2 described in the first embodiment above, so the description is omitted.
[0592] Application Example 2 The process is the same as that in the specific processing described in Application Example 2 of the first embodiment above, so the description is omitted.
[0593] The specific processing unit 290 sends the result of the specific processing to the head-mounted terminal 314. In the head-mounted terminal 314, the control unit 46A outputs the result of the specific processing to the speaker 240 and the display 343. The microphone 238 acquires sound input representing the user's input regarding the result of the specific processing. The control unit 46A sends the sound data representing the user's input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the sound data.
[0594] Data generation model 58 is a so-called generative AI (Artificial Intelligence). Examples of data generation models 58 include ChatGPT (registered trademark) (accessible via the internet (URL: https: / / openai.com / blog / chatgpt)). Data generation model 58 is obtained through deep learning on a neural network. Input to data generation model 58 includes prompt words containing instructions, and inference data such as sound data representing sound, text data representing text, and image data representing images (e.g., still image data or animation data). Data generation model 58 infers the input inference data based on the instructions represented by the prompt words and outputs the inference result in one or more data forms such as sound data, text data, and image data. Data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or induction. The specific processing unit 290 performs the aforementioned specific processing while using data generation model 58. The data generation model 58 can also be a model finely tuned to output inference results from prompts that do not contain instructions. In this case, the data generation model 58 can output inference results based on prompts that do not contain instructions. The data processing apparatus 12, etc., includes various data generation models 58, including AI other than the generation AI. AI other than the generation AI can be, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or Naive Bayes, and can perform various processes, but is not limited to this example. Furthermore, the AI can also be an AI agent. Furthermore, when the processing of the above-mentioned parts is performed by AI, the processing can be performed partially or entirely by AI, but is not limited to this example. Furthermore, the processing performed by the AI including the generation AI can be replaced by processing in the rule base, and the processing in the rule base can also be replaced by processing performed by the AI including the generation AI.
[0595] Furthermore, the processing of the aforementioned data processing system 10 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the head-mounted terminal 314, but it can also be performed by both the specific processing unit 290 of the data processing device 12 and the control unit 46A of the head-mounted terminal 314. Additionally, the specific processing unit 290 of the data processing device 12 acquires or collects information required for processing from the head-mounted terminal 314 or external devices, and the head-mounted terminal 314 acquires or collects information required for processing from the data processing device 12 or external devices.
[0596] For example, the collection unit is implemented by the control unit 46A of the head-mounted terminal 314 or the specific processing unit 290 of the data processing device 12. For example, the acquisition unit uses the camera 42 or communication I / F 44 of the head-mounted terminal 314 to acquire step data, which is then processed by the specific processing unit 290 of the data processing device 12. For example, the analysis unit is implemented by the specific processing unit 290 of the data processing device 12, which analyzes the data from the collection unit and the acquisition unit. For example, the generation unit is implemented by the specific processing unit 290 of the data processing device 12, which uses a generation AI to generate a menu. For example, the serving unit is implemented by the speaker 240 and display 343 of the head-mounted terminal 314 or the specific processing unit 290 of the data processing device 12, which provides the generated menu to the user. The correspondence between each unit and the device or control unit is not limited to the above examples and various changes can be made.
[0597] In the above embodiments, examples of specific processing by the data processing device 12 are given, but the technology disclosed herein is not limited to this, and specific processing may also be performed by the head-mounted terminal 314.
[0598] Fourth Implementation Method Figure 7 An example of the configuration of the data processing system 410 according to the fourth embodiment is shown.
[0599] like Figure 7 As shown, the data processing system 410 includes a data processing device 12 and a robot 414. A server can be cited as an example of the data processing device 12.
[0600] The data processing apparatus 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" as understood in this disclosure. The computer 22 includes a processor 28, RAM 30, and memory 32. The processor 28, RAM 30, and memory 32 are connected to a bus 34. Furthermore, the database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0601] Robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and memory 50. The processor 46, RAM 48, and memory 50 are connected to a bus 52. Furthermore, the microphone 238, speaker 240, camera 42, controlled object 443, and communication I / F 44 are also connected to the bus 52.
[0602] Microphone 238 receives instructions from user 20 by receiving sounds emitted by user 20. Microphone 238 captures sounds emitted by user 20 and converts the captured sounds into sound data, which is then output to processor 46. Speaker 240 outputs sound according to instructions from processor 46.
[0603] Camera 42 is a small digital camera equipped with an optical system such as a lens, aperture and shutter, and imaging elements such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, to photograph the area around robot 414 (e.g., the field of view defined by a perspective equivalent to the field of vision of an average healthy person).
[0604] Communication I / F44 is connected to network 54. Communication I / F44 and 26 are responsible for the transmission and reception of various information between processor 46 and processor 28 via network 54. The transmission and reception of various information between processor 46 and processor 28 using communication I / F44 and 26 is performed in a secure state.
[0605] The controlled object 443 includes a display device, LEDs (light-emitting diodes) for the eyes, and motors for driving the arms, hands, and feet. The posture or movement of the robot 414 is controlled by controlling the motors in the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. In addition, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.
[0606] Figure 8 This illustrates an example of the main functions of the data processing device 12 and the robot 414. For example... Figure 8 As shown, in the data processing device 12, specific processing is performed by the processor 28. The specific processing program 56 is stored in the memory 32.
[0607] The specific processing program 56 is an example of a "program" involved in the technology of this disclosure. The processor 28 reads the specific processing program 56 from the memory 32 and executes the read specific processing program 56 on the RAM 30. Specific processing is implemented by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0608] The memory 32 stores the data generation model 58 and the emotion-specific model 59. The data generation model 58 and the emotion-specific model 59 are used by the specific processing unit 290.
[0609] In robot 414, the processor 46 performs the acceptance and output processing. The memory 50 stores the acceptance and output program 60. The processor 46 reads the acceptance and output program 60 from the memory 50 and executes the read acceptance and output program 60 on RAM 48. The acceptance and output processing is implemented by the processor 46 acting as the control unit 46A according to the acceptance and output program 60 executed on RAM 48.
[0610] Next, the specific processing of the specific processing unit 290 of the data processing device 12 will be described. Each part of the system described below is implemented by the data processing device 12 and the robot 414. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 will be referred to as the "terminal".
[0611] Example 1 The process is the same as that of the specific process described in Embodiment 1 in the first embodiment above, so the description is omitted.
[0612] Application Example 1 The process is the same as that in the specific processing described in Application Example 1 of the first embodiment above, so the description is omitted.
[0613] Example 2 The process is the same as that of the specific process in Embodiment 2 described in the first embodiment above, so the description is omitted.
[0614] Application Example 2 The process is the same as that in the specific processing described in Application Example 2 of the first embodiment above, so the description is omitted.
[0615] The specific processing unit 290 sends the result of the specific processing to the robot 414. In the robot 414, the control unit 46A outputs the result of the specific processing to the speaker 240 and the controlled object 443. The microphone 238 acquires sound input representing the result of the specific processing. The control unit 46A sends the sound data representing the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the sound data.
[0616] Data generation model 58 is a so-called generative AI (Artificial Intelligence). Examples of data generation models 58 include ChatGPT (registered trademark) (accessible via the internet (URL: https: / / openai.com / blog / chatgpt)). Data generation model 58 is obtained through deep learning on a neural network. Input to data generation model 58 are prompt words containing instructions, and inference data such as sound data representing sound, text data representing text, and image data representing images (e.g., still image data or animation data). Data generation model 58 infers from the input inference data based on the instructions represented by the prompt words and outputs the inference result in one or more data forms, such as sound data, text data, and image data. Data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or induction. The specific processing unit 290 performs the aforementioned specific processing while using data generation model 58. The data generation model 58 can also be a model finely tuned to output inference results from prompts that do not contain instructions. In this case, the data generation model 58 can output inference results based on prompts that do not contain instructions. The data processing apparatus 12, etc., includes various data generation models 58, including AI other than the generation AI. AI other than the generation AI can be, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or Naive Bayes, and can perform various processes, but is not limited to this example. Furthermore, the AI can also be an AI agent. Furthermore, when the processing of the above-mentioned parts is performed by AI, the processing can be performed partially or entirely by AI, but is not limited to this example. Furthermore, the processing performed by the AI including the generation AI can be replaced by processing in the rule base, and the processing in the rule base can also be replaced by processing performed by the AI including the generation AI.
[0617] Furthermore, the processing of the aforementioned data processing system 10 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the robot 414, but it can also be performed by both the specific processing unit 290 of the data processing device 12 and the control unit 46A of the robot 414. Additionally, the specific processing unit 290 of the data processing device 12 acquires or collects information required for processing from the robot 414 or external devices, and the robot 414 acquires or collects information required for processing from the data processing device 12 or external devices.
[0618] For example, the collection unit is implemented by the control unit 46A of the robot 414 or the specific processing unit 290 of the data processing device 12. For example, the acquisition unit uses the camera 42 or communication I / F 44 of the robot 414 to acquire step data, which is then processed by the specific processing unit 290 of the data processing device 12. For example, the analysis unit is implemented by the specific processing unit 290 of the data processing device 12, which analyzes the data from the collection unit and the acquisition unit. For example, the generation unit is implemented by the specific processing unit 290 of the data processing device 12, which uses a generation AI to generate a menu. For example, the serving unit is implemented by the speaker 240 of the robot 414 and the control object 443 or the specific processing unit 290 of the data processing device 12, which provides the generated menu to the user. The correspondence between each unit and the device or control unit is not limited to the above examples and various changes can be made.
[0619] In the above embodiments, examples of specific processing by the data processing device 12 are given, but the technology disclosed herein is not limited to this, and specific processing may also be performed by the robot 414.
[0620] Furthermore, the emotion-specific model 59, acting as an emotion engine, can determine a user's emotion based on a specific mapping. Specifically, the emotion-specific model 59 can determine a user's emotion based on an emotion graph that serves as a specific mapping (see [reference]). Figure 9 The emotion-specific model 59 can also determine the robot's emotion, and the specific processing unit 290 performs specific processing based on the robot's emotions.
[0621] Figure 9 This is a diagram representing an emotion map 400 that maps multiple emotions. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotion is. On the outer side of the concentric circles, emotions representing states or behaviors arising from mood are arranged. Emotions are concepts that include feelings and mental states. Emotions generated by reactions occurring in the brain are arranged roughly to the left of the concentric circles. Emotions derived from situational judgments are arranged roughly to the right of the concentric circles. Emotions generated by reactions occurring in the brain and derived from situational judgments are arranged roughly above and below the concentric circles. Furthermore, "pleasant" emotions are arranged above the concentric circles, and "unpleasant" emotions are arranged below them. Thus, in the emotion map 400, multiple emotions are mapped based on the structure that generates emotions, and emotions that are likely to occur simultaneously are mapped close to each other.
[0622] These emotions are distributed at the three o'clock position of the emotion map 400, typically fluctuating between peace and anxiety. In the right half of the emotion map 400, situational awareness dominates over internal sensation, thus resulting in an impression of calm.
[0623] The inner side of the emotion map 400 represents the inner state, while the outer side represents behavior. Therefore, the further outward you are from the emotion map 400, the more visible the emotion becomes (manifested in behavior).
[0624] Here, human emotions are based on various balances such as posture and blood sugar levels. When these balances deviate from an ideal state, it indicates an unpleasant state; when they approach the ideal state, it indicates a pleasant state. Emotions in robots, cars, motorcycles, etc., can also be created in the following way: based on various balances such as posture and remaining battery power, when these balances deviate from an ideal state, it indicates an unpleasant state; when they approach the ideal state, it indicates a pleasant state. Emotion maps can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a Brain Physiological Signal Analysis System for Voice Emotion Recognition and Emotion, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). In the left half of the emotion map, emotions belonging to the sensory-dominated region, called "response," are arranged. Furthermore, in the right half of the emotion map, emotions belonging to the situational cognition-dominated region, called "situation," are arranged.
[0625] In the emotion map, two types of emotions that promote learning are defined. One is a negative emotion on the situational side, in the middle or peripheral region of "repentance" or "reflection." This occurs when the robot experiences negative emotions such as "I don't want to experience this feeling again" or "I don't want to be blamed again." The other is a positive emotion on the response side, near the "desire" region. This occurs when there are positive feelings such as "wanting more" or "wanting to know more."
[0626] The emotion-specific model 59 inputs user input into a pre-trained neural network to obtain emotion values representing each emotion shown in the emotion map 400, thereby determining the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values representing each emotion shown in the emotion map 400. Furthermore, this neural network... Figure 10 As shown in the sentiment graph 900, it was trained in a way that sentiments that are configured close to each other have similar values. Figure 10 The text shows examples of emotions such as "peace of mind", "stability", and "reassurance" that have similar emotion values.
[0627] The above description focuses on the functions of the data processing device 12, but the system of this disclosure is not necessarily installed on a server. The system of this disclosure can also be installed as a general information processing system. This disclosure can also be installed, for example, as a software program running on a personal computer, an application running on a smartphone, etc. The method of this disclosure can also be provided to users in the form of SaaS (Software as a Service).
[0628] In the above embodiments, an example of a specific process being performed by a single computer 22 is given. However, the technology disclosed herein is not limited to this, and the specific process can also be distributed among multiple computers, including computer 22. For example, the data generation model 58 can be located on an external device of the data processing apparatus 12, where data is generated based on the input data.
[0629] In the above embodiments, examples of storing a specific processing program 56 in the memory 32 have been described, but the technology disclosed herein is not limited thereto. For example, the specific processing program 56 may also be stored in a portable computer-readable non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed into the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.
[0630] Alternatively, a specific processing program 56 may be pre-stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 according to the requirements of the data processing device 12.
[0631] In addition, it is not necessary to store all the specific processing program 56 in the storage device such as the server connected to the data processing device 12 via the network 54 or in the memory 32; a portion of the specific processing program 56 may be stored in advance.
[0632] As hardware resources for performing specific processes, various processors, as shown below, can be used. For example, a CPU can be listed as a processor, which functions as a general-purpose processor that performs specific processes by executing software, i.e., a program. Furthermore, processors can be listed as special-purpose circuits such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application-Specific Integrated Circuits), which are processors with circuitry specifically designed to perform specific processes. Each processor has built-in or connected memory, and each processor executes specific processes using that memory.
[0633] The hardware resources for performing a specific process can consist of one of these various processors, or a combination of two or more processors of the same or different types (e.g., a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resources for performing a specific process can be a single processor.
[0634] As an example of a single processor, there are two approaches: First, a processor is composed of a combination of one or more CPUs and software, which functions as a hardware resource to perform a specific process; second, as represented by a SoC (System-on-a-chip), a processor is used to implement the functionality of the entire system, which includes multiple hardware resources for performing a specific process, using a single IC (Integrated Circuit) chip. In this way, the specific process is implemented by using one or more of the aforementioned processors as hardware resources.
[0635] Furthermore, the hardware architecture of these various processors, more specifically, can utilize circuits that combine semiconductor elements and other circuit components. Moreover, the specific process described above is just one example. Therefore, without departing from the main point, unnecessary steps can certainly be deleted, new steps added, or the processing order changed.
[0636] The descriptions and illustrations above are detailed explanations of a portion of the technology disclosed herein, and are merely one example of the technology disclosed herein. For example, the above descriptions of the structure, function, effect, and results are just one example of the structure, function, effect, and results of a portion of the technology disclosed herein. Therefore, without departing from the spirit of the technology disclosed herein, unnecessary parts may be deleted, new elements added, or replacements may be made to the descriptions and illustrations above. Furthermore, to avoid confusion and facilitate understanding of a portion of the technology disclosed herein, explanations of common technical knowledge that do not require special explanation under the premise of being able to implement the technology disclosed herein have been omitted from the descriptions and illustrations above.
[0637] All documents, patent applications and technical specifications set forth in this specification are incorporated herein by reference to the same extent that each document, patent application and technical specification is specifically and individually described therein and referenced by reference.
[0638] In addition, the following notes are provided in response to the above explanation.
[0639] Example 1 (Note 1) An information processing system, characterized in that it comprises: A device for obtaining size and composition information from a user via an input device; An apparatus for generating structured data containing the size information and the composition information, and for automatically generating prompts for generative artificial intelligence models based on the structured data; A means for sending generation instruction information containing the prompt statement and the size information to the generative artificial intelligence model via a communication device, and for obtaining visual information generated in response to the prompt statement from the generative artificial intelligence model; A means for storing the visual information into a storage device, or for generating reference information about the visual information; A device for transmitting the visual information or the reference information to a terminal device via a communication device, and for outputting display control information for displaying the visual information on a display device of the terminal device; A means for sending the prompt statement to the terminal device and outputting information for displaying the prompt statement on the terminal device.
[0640] (Note 2) The information processing system according to Appendix 1 is characterized in that, Includes: an apparatus for acquiring user selection operation information for multiple visual information displayed on a terminal device via a communication device, and calculating the selection rate of each visual information based on the selection operation information; An apparatus for generating change information based on the selection rate to modify the composition information, and for automatically generating new prompt statements based on the modified composition information and the size information; A device for issuing visual information regeneration instructions to the generative artificial intelligence model using the new prompting statements.
[0641] (Note 3) The information processing system according to Appendix 1 is characterized in that, Includes: compositional information contained in visual information for a specific selection rate to meet predetermined conditions, and means for adjusting the content of the prompt statement in order to give priority to the use of the specific compositional information; An apparatus for automatically optimizing the composition information of the visual information by inputting the adjusted prompt statement and the size information into the generative artificial intelligence model and repeatedly performing the regeneration process.
[0642] Application Example 1 (Note 1) An information processing system, characterized in that it comprises: A unit for receiving input information, including size information and composition information, from a user terminal by the processing unit of an information processing device; A unit for generating, by the processing unit, a prompt statement in natural language, which is used to issue instructions to a generative artificial intelligence model to generate a visual display based on the size information and the composition information; A unit for decomposing the prompt statement by the processing unit, extracting text information, color information, image information and numerical information representing size information from the prompt statement, and encoding the information into a conditional vector of a generative artificial intelligence model; A unit for generating visual display data in the form of a pixel array corresponding to the size information by taking the condition vector as input by the processing unit and performing inference processing including matrix operations, convolution operations and nonlinear transformations on the computing device. A unit for generating editable visual display data by having the processing unit perform drawing processing based on text information contained in the constituent information and compositing processing based on image information contained in the constituent information on the visual display data using image processing software. A unit for sending the editable visual display data to the user terminal and displaying it on the user terminal by the processing unit, updating the prompt statement and / or the condition vector based on the edited content received from the user terminal, and issuing a regeneration instruction to the generative artificial intelligence model through the updated prompt statement.
[0643] (Note 2) The information processing system according to Appendix 1 is characterized in that, The processing unit is configured to acquire selection and editing operations performed on the visual display on the user terminal, calculate evaluation indicators of the composition information based on the operations, automatically correct the prompt statement so that composition information with higher evaluation indicators is reflected first, and use the corrected prompt statement to regenerate the visual display through the generative artificial intelligence model.
[0644] (Note 3) The information processing system according to Appendix 1 is characterized in that, The processing unit is configured to parse the edited content on the user terminal, extract the arrangement parameters of text elements, color elements and image elements from the visual display, and add the arrangement parameters as constraints when generating the prompt statement, so as to continuously generate visual displays reflecting the editing results through the generative artificial intelligence model.
[0645] Example 2 (Note 1) An information processing system, characterized in that it comprises: A means for acquiring input from a user in an information processing device, the input including attribute information for generating conditions for generating a visual representation, the attribute information including size information and component information, and natural language instruction statements related to the generation conditions, and a means for generating instruction information by constructing prompt statements for a generative artificial intelligence model to generate a visual representation based on the attribute information and the instruction statements. An apparatus for running the generative artificial intelligence model on computing resources, inputting the instruction information containing the prompt statement into the generative artificial intelligence model, and performing matrix operations and image generation processing through the generative artificial intelligence model to generate multiple visual representation data corresponding to the prompt statement; An apparatus for storing the plurality of visual representations in a storage device, and for transmitting the plurality of visual representations to the user terminal in a form that can be presented on the user terminal via a communication device and displaying them on the user terminal; An apparatus for acquiring behavior logs representing display events and selection events of the visual presentation from the user terminal, and summarizing the display counts and selection counts of each visual presentation based on the behavior logs to calculate the selection rate; An apparatus for performing statistical processing on constituent elements, including color elements, text elements, layout elements and operation component elements, based on constituent element information corresponding to the visual presentation and the selection rate, evaluating the correlation between the constituent elements and the selection rate, thereby extracting constituent elements with high selection rates. An apparatus for automatically adjusting the statements in the prompt statement and the generation conditions to reflect the extracted high-selectivity components, thereby generating an updated prompt statement for the next generation of visual representations. An apparatus for repeatedly regenerating the visual presentation by inputting the updated prompt statement back into the generative artificial intelligence model, and optimizing the constituent elements of the visual presentation based on the selection rate.
[0646] (Note 2) According to the information processing system described in Appendix 1, the information processing device is configured to: acquire visual presentation evaluation information and information indicating a regeneration request from the behavior log; adjust the expression intensity, emphasized object, and display position-related conditions contained in the prompt statement according to the evaluation information; and branch the prompt statement into multiple variants according to the regeneration request and input them into the generative artificial intelligence model to generate multiple visual presentation candidates.
[0647] (Note 3) According to the information processing system described in Appendix 1, the information processing device is configured to: manage metadata corresponding to the visual presentation as constituent element attributes, wherein the constituent element attributes include at least a primary color attribute, a button color attribute, a text size attribute, a layout type attribute, and a price information presence / absence attribute; summarize the selection rate for each constituent element attribute to calculate a grouped statistical value; determine constituent element attributes that meet predetermined evaluation criteria as preferred constituent elements based on the grouped statistical value; and perform self-learning updates on the instruction content for the generative artificial intelligence model by generating or updating a prompt statement template that includes the preferred constituent elements as mandatory selection conditions.
[0648] Application Example 2 (Note 1) An information processing system, characterized in that it comprises: A device for acquiring abstract attribute information about the size and constituent elements of visual information from a user in an information processing apparatus; A device for constructing prompt statements to instruct a generative artificial intelligence model to generate visual information based on the abstract attribute information and previous prompt history information; A device for inputting the prompt statement into the generative artificial intelligence model to obtain multiple visual information candidates; A device for presenting the visual information candidates to the user's terminal via a display device, and for acquiring display event information and selection event information related to the presentation; An apparatus for calculating the index values corresponding to each visual information candidate based on the display event information and the selection event information, and storing the evaluation information characterizing the correspondence between the index values and the constituent elements into a recording medium. A device for automatically generating a second prompt statement based on the evaluation information to adjust at least a portion of the adoption order, color scheme, layout, and text structure of the constituent elements, and instructing the generative artificial intelligence model to regenerate using the second prompt statement; A device for presenting the visual information obtained through the regeneration back to the user terminal.
[0649] (Note 2) The information processing system according to Appendix 1 is characterized in that, The information processing device is configured to: infer the emotional state based on the facial expression or voice information of the user obtained from the user terminal, and generate a second prompt statement based on the emotional state and the evaluation information, such that the presentation order of the visual information candidates and the content recorded about the constituent elements in the prompt statement are changed according to different emotional states.
[0650] (Note 3) The information processing system according to Appendix 1 is characterized in that, The information processing device is configured to: use the evaluation information and the inference results related to the emotional state as learning data for a machine learning model; predict favorable combinations of visual information components given a predetermined emotional state and abstract attribute information through the machine learning model; and automatically generate or modify the prompt statement based on the prediction results.
Claims
1. An information processing system, characterized in that, include: processor; The processor is configured to: receive input from the user regarding the size and constituent elements; generate prompt text based on the received input to instruct the generating artificial intelligence model to generate a visual display; and obtain the visual display from the generating artificial intelligence model using the generated prompt text, and display the visual display on the user's terminal.
2. The information processing system according to claim 1, characterized in that, The processor is also configured to: monitor the user's selection actions; measure the selection rate based on the monitored selection actions; and generate new prompt text for automatically replacing the constituent elements of the visual display.
3. The information processing system according to claim 1, characterized in that, The processor is further configured to: adjust the prompt text based on the constituent elements of the visual display with a high selection rate, so as to give priority to the constituent elements with a high selection rate; and instruct the generative artificial intelligence model to regenerate the visual display based on the adjusted prompt text.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A