Scene layout optimization method and system based on multiple visual elements and emotional perception

By acquiring image data of the target scene, identifying the scene type, extracting visual element features, dynamically determining weights and calculating emotional intensity values, generating layout optimization suggestions, and utilizing image generation models to optimize scene layouts, the problems of rigid visual elements and insufficient emotional intensity in existing technologies are solved, achieving personalized and comfortable scene layout optimization effects.

CN120339465BActive Publication Date: 2025-09-12BEIJING UNIV OF POSTS & TELECOMM +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510828540.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-20
Publication Date
2025-09-12
Estimated Expiration
2045-06-20

AI Technical Summary

Technical Problem

The existing technology in scene layout optimization has the problems of rigid static weight allocation mechanism, insufficient consideration of the comprehensive influence of multiple visual elements, and insufficient ability to dynamically quantify emotional intensity, resulting in poor scene layout optimization effect.

Method used

By acquiring image data of the target scene, identifying the scene type, extracting the feature information of visual elements, dynamically determining the weight of visual elements, and calculating the emotional intensity value, layout optimization suggestions are generated and the scene layout is optimized using the image generation model.

Benefits of technology

It achieves more personalized and comfortable scene layout optimization, improves the pertinence and flexibility of optimization results, and enhances users' sense of participation and identification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120339465B_ABST
    Figure CN120339465B_ABST
Patent Text Reader

Abstract

The present application discloses a scene layout optimization method and system based on multiple visual elements and emotional perception. It relates to the field of management optimization technology. The method includes: obtaining image data and scene type of the target scene; extracting feature information of each category of visual elements in the image data, and determining weight information of each category of visual elements according to the scene type and user feedback information; determining the emotional intensity value of the target scene according to the feature information and weight information of all categories of visual elements, and determining layout optimization suggestions for the target scene according to the emotional intensity value; and generating image data of the target scene after optimized layout using an image generation model according to the layout optimization suggestions. It solves the technical problems that the scene layout optimization method in the prior art has the technical defects of rigid static weight allocation mechanism, insufficient consideration of the comprehensive influence between multiple visual elements, and insufficient dynamic quantification capability of emotional intensity, which lead to poor scene layout optimization effect.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of management optimization technology, and in particular to a scene layout optimization method and system based on multiple visual elements and emotional perception. Background Art

[0002] Linguistic landscape refers to the system of visible linguistic symbols in public spaces. It encompasses various material media used to display language and text, such as signs, advertisements, road signs, and logos. It carries the social significance of urban space and expresses power relations. Its core functions encompass multiple dimensions, including information transmission, cultural identity formation, and ideological encoding. As the scope of the concept of "linguistic landscape" continues to expand, scholars have come to view space as a semiotic resource, further proposing the concept of "semiotic landscape," which uses text as a medium and draws on semiotic resources to construct spatial discourse. This approach is no longer limited to a single written text, but rather focuses on multi-visual semiotic resources, focusing on the interplay between textual elements, image elements, layout elements, and cultural dimensions.

[0003] The language landscape, through the combination of visual elements such as text, graphics, color, and layout, forms a lexical unit comprised of multiple visual elements. For example, the "No Parking" text (a textual element symbol), the red circle with a slash (an image element symbol), the high-contrast color scheme (a color symbol, which can be classified as part of the image element category), and the appropriate layout position (a layout element symbol) in a traffic sign together form a complete semantic-emotional expression system. Thus, the language landscape is composed of both linguistic and non-linguistic symbols, presenting a symbolic aggregate through single or multiple visual elements such as text, graphics, color, and layout. When studying this landscape, we should not only focus on static, fixed linguistic modalities but also interpret the meaning of the various visual elements in the language landscape as a whole. For example, in subway signage, the warning slogan (a textual element), the triangular warning icon (an image element), and the ceiling hanging position (a layout element) jointly influence passengers' safety perceptions and emotional responses. As the complexity of scene design increases, how to quantify and analyze the synergistic effects of multiple visual elements has become a key challenge in optimizing the user's emotional experience.

[0004] Based on in-depth research on language landscape, existing technologies have also explored scene planning and scene graph generation. However, these technologies still have some problems.

[0005] The publication number is CN118840648A, and the name is "Target Area Viewing Scene Planning Method, Device, Computer Equipment and Readable Storage Medium Based on Multimodality (Multimodality refers to the integration of multiple information modes, such as text, images, audio, etc.). The disclosed method includes: first, obtaining viewing video data of the target area, covering multiple preset viewing scene pictures. Then, by obtaining and analyzing the multimodal feedback results of users on the viewing video data, the visual landscape quality of each preset viewing scene picture is evaluated. At the same time, considering the pre-set target area landscape visual impact factor, the visual landscape quality evaluation result and the landscape visual impact factor are correlated and analyzed, and finally the landscape scene planning result of the target area is obtained. However, this solution has the following defects:

[0006] (1) Insufficient capture and expression of emotional information: The visual quality of the preset scene images is mainly analyzed by viewing video data, and there is a lack of accurate capture and effective expression mechanism of the emotional information contained in each visual element in the language landscape. It is difficult to accurately reflect the impact of different visual element combinations on user emotions, resulting in poor optimization effect of the final scene layout.

[0007] (2) Lack of dynamic adjustment capabilities: The differences in requirements of different scenarios are not fully considered, and the combination of multiple visual elements cannot be dynamically adjusted according to the actual scenario, making it difficult to achieve the best emotional expression effect, resulting in insufficient flexibility and adaptability in scene planning.

[0008] The publication number is CN118334414A, and its title is "Scene Graph Generation Method and System Supporting Historical and Cultural Block Scenes." The disclosed method includes: analyzing historical and cultural geographic elements and related relationships to construct a dataset for generating scene graphs of historical and cultural blocks; training a landscape semantic recognition model to extract the visual, semantic, and spatial features of target objects; using the co-occurrence frequency of target relationships (co-occurrence frequency refers to the frequency of two or more targets appearing simultaneously in the same scene) as prior knowledge, and predicting target relationships based on an attention mechanism and feature fusion methods; using spatiotemporally continuous street imagery as input, generating a scene graph corresponding to each street image based on a landscape semantic recognition model and a target relationship prediction model, and using this as input to achieve scene graph fusion using a scene graph fusion model. Although this solution integrates geographic elements and spatial features, it has the following shortcomings:

[0009] (1) Weak emotional information processing: It relies on fixed co-occurrence frequencies in the prediction of element relationships, lacks in-depth exploration and analysis of the emotional information of each visual element in the language landscape, cannot accurately capture the emotional connotations contained in different combinations of visual elements, and is difficult to achieve effective expression of emotional information, resulting in poor scene layout optimization effects.

[0010] (2) Limited dynamic adjustment capability: The weight information of multiple visual elements cannot be dynamically adjusted according to the requirements of different scenarios and user feedback information, making it difficult to achieve the best emotional expression effect, resulting in a disconnect between emotional expression and functional requirements, and poor scene layout optimization effect.

[0011] The existing scene layout optimization methods have technical defects such as rigid static weight allocation mechanism, insufficient consideration of the comprehensive impact of multiple visual elements, and insufficient dynamic quantification ability of emotional intensity, which lead to poor scene layout optimization effects. No effective solution has been proposed so far. Summary of the Invention

[0012] The disclosed embodiments provide a method and system for scene layout optimization based on multiple visual elements and emotional perception. This method addresses at least the technical deficiencies in existing scene layout optimization methods, which suffer from a rigid static weight allocation mechanism, a failure to fully consider the combined impact of multiple visual elements, and insufficient dynamic quantification of emotional intensity, leading to poor scene layout optimization results.

[0013] According to one aspect of an embodiment of the present disclosure, a scene layout optimization method based on multiple visual elements and emotion perception is provided, comprising: acquiring image data of a target scene, and determining a scene type of the target scene; wherein the image data includes visual elements of multiple categories; extracting feature information of the visual elements of each category in the image data, and determining weight information of the visual elements of each category in the image data based on the scene type and pre-acquired user feedback information on the layout of the target scene; determining an emotion intensity value of the target scene based on the feature information and corresponding weight information of all categories of visual elements in the image data, and determining a layout optimization suggestion for the target scene based on the emotion intensity value; wherein the emotion intensity value is used to indicate the comprehensive impact of all categories of visual elements in the target scene on the user's emotional experience; and generating image data of the target scene after optimized layout using an image generation model based on the layout optimization suggestion.

[0014] According to another aspect of the embodiments of the present disclosure, a storage medium is further provided. The storage medium includes a stored program, wherein the above method is executed by a processor when the program is running.

[0015] According to another aspect of an embodiment of the present disclosure, a scene layout optimization system based on multiple visual elements and emotional perception is provided, including: a data acquisition module for acquiring image data of a target scene and determining the scene type of the target scene; wherein the image data includes visual elements of multiple categories; an extraction and parsing module for extracting feature information of visual elements of each category in the image data, and determining weight information of visual elements of each category in the image data based on the scene type and pre-acquired user feedback information on the layout of the target scene; an emotion calculation and optimization decision module for determining the emotion intensity value of the target scene based on the feature information and corresponding weight information of visual elements of all categories in the image data, and determining a layout optimization suggestion for the target scene based on the emotion intensity value; wherein the emotion intensity value is used to indicate the comprehensive impact of visual elements of all categories in the target scene on the user's emotional experience; and a layout optimization module for generating image data of the target scene after optimized layout using an image generation model based on the layout optimization suggestion.

[0016] According to another aspect of an embodiment of the present disclosure, a scene layout optimization system based on multiple visual elements and emotional perception is also provided, including a processor; and a memory connected to the processor, for providing the processor with instructions for processing the following processing steps: obtaining image data of a target scene, and determining the scene type of the target scene; wherein the image data includes multiple categories of visual elements; extracting feature information of each category of visual elements in the image data, and determining weight information of each category of visual elements in the image data based on the scene type and pre-acquired user feedback information on the layout of the target scene; determining an emotional intensity value of the target scene based on the feature information and corresponding weight information of all categories of visual elements in the image data, and determining a layout optimization suggestion for the target scene based on the emotional intensity value; wherein the emotional intensity value is used to indicate the comprehensive impact of all categories of visual elements in the target scene on the user's emotional experience; and generating image data of the target scene after optimized layout using an image generation model based on the layout optimization suggestion.

[0017] This application first obtains image data of the target scene and determines the scene type, allowing for customized processing based on the characteristics of different scenarios (such as transportation, landscape, and medical care). This provides data support for subsequent visual element feature extraction, weight assignment, and layout optimization, ensuring that layout optimization measures are more closely aligned with actual scenario requirements. Next, feature information for each category of visual elements (such as text elements, image elements, and layout elements) in the image data is extracted. Based on the scene type and user feedback on the target scene layout, the weight information for each category of visual elements is dynamically determined. This allows for flexible adjustment of the importance of each visual element in the scene layout based on different scenarios and user preferences, thereby achieving more personalized scene layout optimization. Secondly, by integrating the feature information of all categories of visual elements and their corresponding weight information, the emotional intensity value of the target scene is calculated. This intuitively reflects the combined impact of all visual elements in the target scene on the user's emotional experience. By analyzing the characteristics of different visual elements, studying their interrelationships, and considering their comprehensive weights, this ensures that each visual element plays its optimal role in the scene layout, providing an objective and measurable basis for layout optimization. Then, based on the sentiment intensity values, specific layout optimization suggestions are generated to guide the adjustment of the scene layout, making the optimization process more directional and operational. Finally, based on the layout optimization suggestions, an image generation model is used to generate image data of the target scene after the optimized layout, translating the optimization suggestions into actual visual representations. This allows users to intuitively see the optimized scene effects, providing data support for subsequent visual and interactive presentations of optimization results. This enhances user participation and recognition in the optimization process, while also facilitating further evaluation and adjustment of the optimization results.

[0018] This application achieves a multi-dimensional improvement in scene layout optimization through the organic combination of multiple links such as accurate scene type identification, dynamic weight allocation, full consideration of the comprehensive influence between multiple visual elements, quantitative evaluation of emotional intensity values, generation of layout optimization suggestions, and application of image generation models. This method not only improves the pertinence and effectiveness of the optimization results, but also enhances the flexibility and adaptability of the optimization process, greatly improves the scene layout optimization effect, and provides users with a more personalized and comfortable scene experience. This solves the technical problems that the scene layout optimization methods existing in the prior art have rigid static weight allocation mechanisms, do not consider the comprehensive influence between multiple visual elements, and lack the ability to dynamically quantify emotional intensity, resulting in poor scene layout optimization effects. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] The drawings described herein are used to provide a further understanding of the present disclosure and constitute a part of this application. The illustrative embodiments of the present disclosure and their descriptions are used to explain the present disclosure and do not constitute an improper limitation of the present disclosure. In the drawings:

[0020] Figure 1 is a hardware structure block diagram of a computing device for implementing the method according to embodiment 1 of the present application;

[0021] Figure 2 1 is a flow chart of the scene layout optimization method based on multiple visual elements and emotional perception according to Example 1 of the present application;

[0022] Figure 3 is a schematic diagram of a scene layout optimization system based on multiple visual elements and emotional perception according to Example 2 of the present application; and

[0023] Figure 4 This is a schematic diagram of the scene layout optimization system based on multiple visual elements and emotional perception described in Example 3 of the present application. DETAILED DESCRIPTION

[0024] In order to enable those skilled in the art to better understand the technical solutions of the present disclosure, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below in conjunction with the drawings in the embodiments of the present disclosure. Obviously, the described embodiments are only part of the embodiments of the present disclosure, not all of the embodiments. Based on the embodiments of the present disclosure, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of the present disclosure.

[0025] It should be noted that the terms "first", "second", etc. in the specification and claims of the present disclosure and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that the numbers used in this way can be interchanged where appropriate, so that the embodiments of the present disclosure described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0026] Example 1

[0027] According to this embodiment, a method embodiment of a scene layout optimization method based on multiple visual elements and emotional perception is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.

[0028] The method embodiment provided in this embodiment can be executed in a mobile terminal, a computer terminal, a server or a similar computing device. Figure 1 The hardware structure block diagram of a computing device for implementing a scene layout optimization method based on multiple visual elements and emotion perception is shown. Figure 1 As shown, a computing device may include one or more processors (the processor may include, but is not limited to, a microprocessor (MCU) or a programmable logic device (FPGA) or other processing device), a memory for storing data, a transmission device for communication functions, and an input / output interface. The memory, transmission device, and input / output interface are connected to the processor via a bus. In addition, it may also include: a display, a keyboard, and a cursor control device connected to the input / output interface. Those skilled in the art will understand that Figure 1 The structure shown is only for illustration and does not limit the structure of the above electronic device. Figure 1 More or fewer components than shown, or with Figure 1 Different configurations shown.

[0029] It should be noted that the one or more processors and / or other data processing circuits described above may generally be referred to herein as "data processing circuitry." The data processing circuitry may be embodied in whole or in part as software, hardware, firmware, or any other combination thereof. Furthermore, the data processing circuitry may be a single, independent processing module, or may be fully or partially integrated into any of the other components of the computing device. As discussed in the embodiments of the present disclosure, the data processing circuitry functions as a processor control (e.g., selection of a variable resistor terminal path connected to an interface).

[0030] The memory can be used to store software programs and modules of application software, such as the program instructions / data storage device corresponding to the scene layout optimization method based on multiple visual elements and emotional perception in the embodiment of the present disclosure. The processor executes various functional applications and data processing by running the software programs and modules stored in the memory, that is, realizing the scene layout optimization method based on multiple visual elements and emotional perception of the above-mentioned application. The memory may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory may further include a memory remotely located relative to the processor, and these remote memories may be connected to the computing device via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.

[0031] The transmission device is used to receive or send data via a network. Specific examples of the aforementioned network may include a wireless network provided by a communications provider of the computing device. In one embodiment, the transmission device includes a network interface controller (NIC), which can be connected to other network devices via a base station to enable communication with the Internet. In another embodiment, the transmission device may be a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.

[0032] The display may be, for example, a touch screen liquid crystal display (LCD) that enables a user to interact with a user interface of the computing device.

[0033] It should be noted that, in some optional embodiments, the above Figure 1 The computing device shown may include hardware elements (including circuits), software elements (including computer code stored on a computer-readable medium), or a combination of both hardware and software elements. Figure 1 This is merely one example of a particular embodiment and is intended to illustrate the types of components that may be present in the computing device described above.

[0034] In the above operating environment, according to the first aspect of this embodiment, a scene layout optimization method based on multiple visual elements and emotion perception is provided. Figure 2 A schematic diagram showing the process of the method is shown in FIG. Figure 2 As shown, the method includes:

[0035] S102: Acquire image data of a target scene and determine the scene type of the target scene; wherein the image data includes visual elements of multiple categories;

[0036] S104: extracting feature information of each category of visual elements in the image data, and determining weight information of each category of visual elements in the image data according to the scene type and pre-acquired user feedback information on the target scene layout;

[0037] S106: Determining an emotional intensity value of the target scene based on feature information and corresponding weight information of all categories of visual elements in the image data, and determining a layout optimization suggestion for the target scene based on the emotional intensity value; wherein the emotional intensity value is used to indicate the comprehensive impact of all categories of visual elements in the target scene on the user's emotional experience; and

[0038] S108: Based on the layout optimization suggestion, generate image data of the target scene after the layout is optimized using an image generation model.

[0039] Specifically, in this embodiment, the target scene is a scene that needs to be optimized in layout. The language landscape of these scenes is composed of language symbols and non-language symbols, and a symbol aggregate is presented in the form of single or multiple symbols such as text, images, graphics, and colors to form a scene identifier. Reasonable scene layout can enable users in the target scene to have a good sense of interaction and experience. This technical solution aims to optimize the scene layout by processing and analyzing the image data of the target scene, so as to provide users with a better sense of interaction and experience. Scene layout optimization is mainly reflected in the optimization of the scene identifier of a certain scene, such as but not limited to optimizing the number of scene identifiers, setting positions, text content, graphic symbols, and color matching.

[0040] Therefore, when a scene requires layout optimization, it is identified as a target scene. Image data of the target scene is then collected using a camera, drone, or public map API, and the scene type of the target scene is determined (corresponding to step S102). Scene types include, but are not limited to, medical scenes, street scenes, cultural and tourism scenes, and transportation scenes. The image data of the target scene includes multiple categories of visual elements, such as text elements, image elements, and spatial layout elements.

[0041] Taking into account that the degree of influence of various visual elements on scene layout optimization and user emotional experience varies in different scene types, and the user's feedback information on the scene layout also reflects their attention and expectations on different visual elements, it is necessary to extract the feature information of each category of visual elements in the image data, and determine the weight information of each category of visual elements in the image data based on the scene type and the pre-acquired user feedback information on the target scene layout (corresponding to step S104).

[0042] Specifically, feature information is extracted for each category of visual elements in the image data. Taking text elements as an example, features such as the font, font size, color, and semantic content of the text can be extracted. Font and font size affect the readability and visual appeal of the text, color can convey an emotional atmosphere, and semantic content is directly related to the information conveyed by the scene. For image elements, features such as graphic symbols and color matching are extracted from the image. Graphic symbols reflect the information conveyed by the scene, and color matching creates a specific emotional tone. In terms of spatial layout elements, features such as the size, number, and location of scene signs in the space are extracted. These features determine the practicality and comfort of the scene.

[0043] Next, based on the scene type and pre-collected user feedback on the target scene layout, the weighting of each visual element category is determined. For example, in medical scenarios, the emphasis is generally on the accuracy, acceptability, and legibility of signage and warning information, as well as the density of signage. Therefore, the weights of text elements and spatial layout elements are relatively high. If user feedback emphasizes the acceptability of text content in medical signage (for example, avoiding the use of "strictly prohibited" in favor of "please do not"), the weight of text elements should be increased accordingly. In cultural and tourism scenarios, the emphasis is generally on attractive graphics and aesthetically pleasing layouts. Therefore, image elements and spatial layout elements are more critical to attracting tourists and creating a positive experience, so the weights of these two elements are relatively high. If user feedback emphasizes the need for novel and vivid graphic symbols, the weight of image elements should be increased accordingly. By comprehensively considering scene type and user feedback, the importance of each visual element in scene layout optimization and user emotional experience can be more accurately determined, providing a basis for subsequent emotional intensity value calculation and layout optimization recommendations.

[0044] After completing the visual element feature extraction and weight determination, the emotional intensity value of the target scene is determined based on the feature information of all categories of visual elements in the image data and the corresponding weight information, and the layout optimization suggestion of the target scene is determined based on the emotional intensity value (corresponding to step S106).

[0045] Specifically, the target scene's emotional intensity value is first determined based on the feature information and corresponding weights of all visual elements in the image data. A weighted summation method can be used to multiply the feature value of each visual element by its corresponding weight, and then add all the results to obtain the emotional intensity value. For example, for text elements, if their feature values ​​are scores for accuracy, acceptability, and readability, and their weights are values ​​determined based on the scene type and user feedback, the two are multiplied together to obtain the text element's contribution to the emotional intensity value. A similar calculation is performed for image elements and spatial layout elements, and finally, the contributions of all elements are summed to obtain the target scene's emotional intensity value. The contribution of text elements to the emotional intensity value is then calculated in synergy with the contribution of image elements. For example, the red warning text (with an emotional intensity contribution of +0.8) and the exclamation point icon (with an emotional intensity contribution of +0.6) synergistically reinforce each other, resulting in a total emotional intensity contribution of 0.8 + 0.6 = 1.4. For example, the blue guide text (with a contribution value of -0.5) and the upward arrow icon (with a contribution value of +0.5) cancel each other out, resulting in a total contribution value of 0. Regarding the contribution of layout elements to emotional intensity: if the layout density is ≤ 5 elements / ㎡, the contribution value is 0. If the layout density is greater than 5 elements / ㎡, the contribution value decreases by -0.2 for each additional element / ㎡.

[0046] It should be noted that in this technical solution, a higher emotional intensity value indicates that the target scene layout evokes a higher level of resistance in the user, that is, a stronger negative emotional experience in the user. Conversely, a lower emotional intensity value indicates that the target scene evokes a stronger positive emotional experience in the user.

[0047] Next, layout optimization recommendations for the target scene are determined based on the emotional intensity value. A high emotional intensity value indicates that the current scene's visual element combination fails to effectively meet user needs and enhance the user experience. Optimization recommendations are then made for each visual element. For text elements, if their expression carries a negative emotional bias (e.g., phrases like "strictly prohibited"), consider adjusting them to more positive and gentle expressions (e.g., "please do not" can be further optimized into a more welcoming reminder, or negative text can be replaced with positive and encouraging text). If visual presentation, such as font and size, impacts the emotional experience, adjustments can be made to more positive emotional cues. For image elements, if the theme, composition, or color scheme creates a negative atmosphere, the theme can be reselected, the composition adjusted, or the color scheme adjusted to create a more positive and pleasing visual effect. For spatial layout elements, if functional zoning is unreasonable or the relative positioning of elements is inappropriate, the spatial layout can be replanned to improve practicality and comfort. These layout optimization recommendations aim to enhance the positive emotional experience value of the target scene for users, providing them with a better sense of interaction and experience.

[0048] After calculating the emotional intensity value and determining the layout optimization recommendations, the image generation model is used to generate image data of the target scene after the optimized layout, based on the layout optimization recommendations (corresponding to step S108). This step aims to convert the optimization recommendations into input instructions that can be recognized by generative AI (such as Stable Diffusion), and use the image generation model (such as the diffusion model) to generate image data that meets the emotional intensity target, so that it can be subsequently presented to the user in a visual manner to intuitively demonstrate the optimized scene effect.

[0049] The image generation model adjusts the image data according to the input instructions, for example but not limited to: for color control, when the input instruction is "reduce the red saturation to 30%", the image generation model adjusts the S value in the HSV of the image to ≤ 30%. For symbol adaptation, when the input instruction is "add guide arrows", the image generation model calls the preset icon library (↑, ) and match the scene style (e.g., flattening / stereoscopic processing). For text optimization, when the input command is "Replace with friendly terms," ​​the image generation model automatically replaces sensitive terms (e.g., "strictly prohibited" with "please do not").

[0050] Furthermore, during the display process, users can evaluate the generated image. They can observe whether the image meets the emotional intensity target, whether the combination of text, graphics, and color is harmonious, and whether the overall effect meets expectations. If the user is dissatisfied with the generated image, they can return to step S108, triggering a new round of weight adjustment and design optimization for each visual element. Alternatively, by deploying the old and new design schemes (i.e., the scene layout scheme before optimization and the scene layout scheme after optimization) in parallel in a real scene, user behavior data (such as user dwell time, violation rate) and subjective ratings can be collected, and then returning to step S106. After updating the user feedback information, a new round of weight adjustment and design optimization for each visual element can be triggered. This closed-loop mechanism ensures that the optimization process remains closely aligned with user needs, thereby converting abstract optimization suggestions into concrete image data. Through visual display, users can intuitively experience the optimized scene effect, providing a basis for subsequent decision-making and improvement.

[0051] As described in the background technology, existing scene layout optimization methods have technical defects such as rigid static weight allocation mechanism, insufficient consideration of the comprehensive impact of multiple visual elements, and insufficient dynamic quantification ability of emotional intensity, which leads to poor scene layout optimization effect.

[0052] In light of this, the technical solution of this application first acquires image data of the target scene and determines the scene type, enabling customized processing based on the characteristics of different scenarios (such as transportation, landscape, and medical care). This provides data support for subsequent visual element feature extraction, weight assignment, and layout optimization, ensuring that layout optimization measures are more closely aligned with actual scenario requirements. Next, feature information is extracted for each category of visual elements (such as text elements, image elements, and layout elements) in the image data. Based on the scene type and user feedback on the target scene layout, the weight information for each category of visual elements is dynamically determined. This allows for flexible adjustment of the importance of each visual element in the scene layout based on different scenarios and user preferences, thereby achieving more personalized scene layout optimization. Secondly, by integrating the feature information of all categories of visual elements and their corresponding weight information, the emotional intensity value of the target scene is calculated. This intuitively reflects the combined impact of all visual elements in the target scene on the user's emotional experience. By analyzing the features of different visual elements, studying their interrelationships, and considering their comprehensive weights, this ensures that each visual element plays its optimal role in the scene layout, providing an objective and measurable basis for layout optimization. Then, based on the sentiment intensity values, specific layout optimization suggestions are generated to guide the adjustment of the scene layout, making the optimization process more directional and operational. Finally, based on the layout optimization suggestions, an image generation model is used to generate image data of the target scene after the optimized layout, translating the optimization suggestions into actual visual representations. This allows users to intuitively see the optimized scene effects, providing data support for subsequent visual and interactive presentations of optimization results. This enhances user participation and recognition in the optimization process, while also facilitating further evaluation and adjustment of the optimization results.

[0053] This application achieves a multi-dimensional improvement in scene layout optimization through the organic combination of multiple links such as accurate scene type identification, dynamic weight allocation, full consideration of the comprehensive influence among multiple visual elements, quantitative evaluation of emotional intensity values, generation of layout optimization suggestions, and application of image generation models. This method not only improves the pertinence and effectiveness of the optimization results, but also achieves dynamic adaptability through a feedback loop, greatly improving the scene layout optimization effect and providing users with a more personalized and comfortable scene experience. This solves the technical problems that the scene layout optimization methods existing in the prior art have technical defects such as the rigid static weight allocation mechanism, the failure to consider the comprehensive influence among multiple visual elements, and the insufficient dynamic quantification capability of emotional intensity, which lead to poor scene layout optimization effects.

[0054] Optionally, the multiple categories of visual elements include text elements, image elements and layout elements; and the operation of extracting feature information of each category of visual elements in the image data includes: extracting text content in the image data through optical character recognition technology, calculating the text emotion contribution value corresponding to the text content through a pre-trained language model, and using it as the feature information of the text elements in the image data; identifying graphic symbols in the image data based on a target detection algorithm, determining the semantic type of the graphic symbols, extracting the HSV value of the main color tone of the image data, and determining the image emotion contribution value of the image data according to the semantic type and the HSV value, as the feature information of the image elements in the image data; and based on the image data, counting the number of logos of the target scene within a unit area, determining the visual focus of the image data through a saliency detection model, and determining the layout emotion contribution value of the image data according to the number of logos and the visual focus, as the feature information of the layout elements in the image data.

[0055] Specifically, optical character recognition (OCR) technology is used to extract text from the image data. OCR technology converts text from images into an editable text format, providing foundational data for subsequent analysis. A pre-trained language model (such as the ERNIE-3.0 model) is then used to calculate the text sentiment contribution corresponding to the text. This pre-trained language model has been trained on a large amount of text data and is able to understand the semantics and sentiment of text. Using this model to calculate the text sentiment contribution value converts text into a numerical representation with emotional significance, thereby quantifying the impact of the text element on the overall emotional expression of the image. For example, if the text extracted from the image data is "strictly prohibited," the pre-trained language model calculates a text sentiment contribution value of 4.5. If the text extracted from the image data is "please do not," the pre-trained language model calculates a text sentiment contribution value of 3.2. This text sentiment contribution value is then used as feature information for the text element in the image data.

[0056] Next, an object detection algorithm (such as the YOLOv5 model) is used to identify the graphical symbols in the image data and determine their semantic types. Object detection algorithms can locate specific objects in the image and identify the graphical symbols within them. The graphical symbols are then classified to determine their semantic types. Examples of semantic types include, but are not limited to, prohibition, guidance, and warning. Determining the semantic type of a graphical symbol helps understand its meaning and provides a semantic foundation for subsequent analysis.

[0057] Next, the HSV values ​​(hue, saturation, and brightness) of the dominant hues in the image data are extracted. Dominant hues vary across different scene types. For example, in traffic scenarios, hazard warning signs often use high-contrast warm hues (red and orange) to enhance warnings, while traffic signs often use cool hues (blue and green) to provide clear guidance. Another example is that in medical settings, emergency rescue signs often use vibrant warm hues (red and yellow) to emphasize their urgency, while medical area signs often use softer hues (light blue and light green) to create a comfortable atmosphere. Because the HSV color space better aligns with human color perception, extracting the HSV values ​​of the dominant hues accurately describes the primary color features of the image, providing critical color information for the subsequent calculation of the image's emotional contribution value. The image's emotional contribution value is then determined based on the semantic type and HSV values. By combining the semantic type of the graphic symbol and the HSV values ​​of the dominant hue of the image, the image's emotional contribution value is calculated, taking into account both the semantic and color characteristics of the image elements. For example, a hue of H = 0° in the HSV value of the dominant hue in image data indicates red, a saturation of S = 80%, and a lightness of V of 60%. In this case, combining the semantic type of the graphic symbol and the HSV value of the dominant hue of the image, the calculated image emotion contribution value is 0.8. This value quantifies the influence of the image element on the overall emotional expression of the image, making the characteristic information of the image element more emotionally representative. Therefore, the image emotion contribution value can be used as the characteristic information of the image element in the image data.

[0058] Then, based on the image data, the number of signs per unit area (per square meter) in the target scene is counted. Sign count is a key indicator of image layout characteristics. By counting the number of signs per unit area, we can understand the density of signs in the target scene and thus infer the image's layout style and information density. The visual focus of the image data is determined using a saliency detection model (such as the DeepGaze II model). This saliency detection model simulates the human visual attention mechanism and identifies the most attention-grabbing area in an image, known as the visual focus. The visual focus is a key element in image layout, influencing the viewer's focus and browsing order. The layout emotional contribution value of the image data is then determined based on the number of signs and the visual focus. The layout emotional contribution value is calculated by comprehensively considering the characteristics of these two layout elements: the number of signs and the visual focus. This value quantifies the impact of these layout elements on the overall emotional expression of the image, making the feature information of these layout elements more emotionally representative. Therefore, the layout emotional contribution value can be used as characteristic information of the layout elements in the image data.

[0059] By extracting feature information from text, image, and layout elements, we provide rich feature data for tasks such as sentiment analysis and content understanding of image data. This feature information effectively characterizes the characteristics of the image and the emotional semantics conveyed by the scene layout from multiple dimensions, laying a critical data foundation for subsequent in-depth processing tasks such as image scene analysis, layout reconstruction, and sentiment calculation.

[0060] In addition, before extracting the feature information of each category of visual elements in the image data, the image data is preprocessed, such as standardizing the image resolution (such as 1920×1080) and removing noise (such as reflections and blurred areas).

[0061] Optionally, the operation of determining the image emotion contribution value of the image data based on the semantic type and the HSV value includes: determining the symbol emotion contribution value corresponding to the graphic symbol based on the semantic type; calculating the color emotion contribution value of the image data based on the HSV value; and determining the image emotion contribution value of the image data based on the symbol emotion contribution value and the color emotion contribution value.

[0062] Specifically, different graphic symbols have different semantic connotations, which are linked to human emotional cognition. For example, in traffic scenes, a "No Entry" sign (a red circle with a white horizontal bar) is often associated with negative emotions such as danger and prohibition, while a green arrow-shaped sign is associated with positive emotions such as guidance and safety. By analyzing and classifying the semantics of a large number of traffic scene graphic symbols and establishing a mapping relationship between semantics and emotions, it is possible to determine the corresponding symbolic emotional contribution value based on the semantic type of the graphic symbol. This symbolic emotional contribution value reflects the contribution of the emotional information contained in the graphic symbol itself to the overall emotional expression of the image. By determining the symbolic emotional contribution value, the semantic emotion of the graphic symbol can be quantified, providing basic data for the subsequent calculation of the image emotional contribution value. For example, by collecting a variety of common traffic scene graphic symbols and inviting professionals in the traffic field or conducting a large-scale questionnaire survey, the semantics represented by each graphic symbol, as well as the corresponding emotional tendency (positive, negative, or neutral) and emotional intensity, can be determined. This information is compiled into a semantic emotional dictionary, and each graphic symbol is assigned an initial symbolic emotional contribution value. In practical applications, the semantics of traffic scene graphic symbols may be influenced by other elements in the image or the overall context. Therefore, it's necessary to dynamically adjust the symbol's emotional contribution based on the specific context of the image. For example, in an image of a road construction scene, even if a "No Entry" sign appears (which might originally be associated with danger), the surrounding construction fences, warning lights, and busy construction workers may give the sign a special, non-purely negative emotional meaning. It also serves as a reminder of construction safety and, in some ways, a positive guidance tool, guiding people to safely detour.

[0063] HSV values ​​(hue, saturation, and value) are important parameters for describing color, and different colors have different emotional characteristics. For example, in traffic scenes, red is often associated with emotions such as warning and danger. For example, the red light in a traffic signal provides a strong visual reminder to drivers and pedestrians to stop. Blue is associated with emotions such as calmness and reliability. For example, some traffic signs use a blue background to convey a professional and stable message. Saturation and value also influence the emotional expression of color. Highly saturated colors are generally more vivid and eye-catching. For example, a highly saturated yellow warning sign in a traffic scene can evoke a stronger emotional response and quickly capture attention. High-value colors often convey a sense of brightness and openness, such as the bright white road markings in traffic scenes, which clearly provide directions in daylight. Low-value colors, on the other hand, can create a heavy and solemn atmosphere, such as the dark gray road barriers in traffic scenes, which convey a sense of calmness and inviolability. In other words, the color emotion contribution value reflects the contribution of the color information in the image to the overall emotional expression of the image. By calculating the color emotion contribution value, the color characteristics of the image can be quantified, providing another important dimension of data for the subsequent calculation of the image emotion contribution value.

[0064] Furthermore, before calculating the image's emotional contribution, a color emotion model can be established by conducting emotional testing and analysis on a large number of colors with different HSV values ​​in traffic scenes. This model can predict the corresponding emotional tendency and intensity based on the color's HSV value and calculate the color's emotional contribution. For example, a function can be defined that takes a color's HSV value as input and outputs the corresponding color's emotional contribution. Specifically, the red color in traffic lights (assuming its HSV value is within a specific range) is analyzed to have a strong emotional tendency to warn of danger, so a higher emotional contribution value (e.g., +0.8, with positive values ​​indicating negative emotion) can be assigned. Conversely, the blue color in traffic signs (also within a specific HSV value range) is assigned a negative emotional contribution value (e.g., -0.3, indicating positive emotion) because it conveys a sense of calmness and reliability. This quantifies the emotional characteristics of colors in traffic scenes and facilitates more accurate calculation of image emotional contribution values.

[0065] Considering that the emotional expression of an image is the result of the combined effects of multiple elements, with graphic symbols and color being two key factors, the symbolic and color emotional contribution values ​​reflect the contribution of graphic symbols and color to the overall emotional expression of the image, respectively. Therefore, the image emotional contribution value of image data can be determined by comprehensively considering these two factors. In other words, the image emotional contribution value is a comprehensive quantitative indicator of the overall emotional expression of an image, which can be used in applications such as image sentiment analysis, classification, and retrieval. By calculating the image emotional contribution value, the emotional characteristics of an image can be more objectively assessed, providing richer information for image processing and understanding. The image emotional contribution value is obtained by performing a weighted summation of the symbolic and color emotional contribution values. For example, two weight coefficients can be defined to represent the importance of the symbolic and color emotional contribution values ​​in the emotional expression of the image, respectively. The image emotional contribution value is then obtained by multiplying the two contribution values ​​by the corresponding weight coefficients and adding them together. The weight coefficients can be adjusted based on specific application scenarios and requirements.

[0066] Alternatively, machine learning algorithms can be used to determine image emotion contributions. For example, a large amount of image data can be collected and labeled with emotion labels for each image. Machine learning algorithms such as support vector machines (SVMs) and neural networks can then be used to train this image data and establish a mapping relationship between symbolic emotion contributions, color emotion contributions, and image emotion labels. In practical applications, the symbolic emotion contribution and color emotion contribution values ​​of a new image can be input into the trained model to predict the image emotion contribution value.

[0067] Therefore, by comprehensively considering the semantic emotion contribution value and color emotion contribution value of graphic symbols, the image emotion contribution value of image data in specific scenarios is accurately quantified.

[0068] Optionally, the operation of determining the layout emotion contribution value of the image data based on the number of identifications and the visual focus includes: determining the basic density score of each sub-area of ​​the target scene based on the number of identifications; wherein, the target scene is divided into multiple sub-areas according to the unit area; according to the visual focus, the identification significance weight of each sub-area is determined; and according to the basic density score and the identification significance weight, the layout emotion contribution value of each sub-area is determined, and the sum of the layout emotion contribution values ​​of each sub-area is used as the layout emotion contribution value of the image data.

[0069] Specifically, taking the target scene as an example of a traffic scene, the number of signs in a traffic scene reflects the density of traffic information in a specific area. For example, at a busy intersection in the city center, due to the large traffic volume and complex road conditions, a large number of traffic signs are set up, such as traffic lights, traffic signs (including direction instructions, speed limit information, etc.), zebra crossing signs, etc. The target scene (such as the entire intersection area) is divided into multiple sub-areas according to the unit area. The number of signs in each sub-area can reflect the distribution density of traffic information in the area. By counting the number of signs in each sub-area and determining a basic density score based on this, the differences in traffic information distribution between different sub-areas can be quantified, which helps to identify areas with dense and sparse traffic information, laying the foundation for further analysis of the emotional impact of signs in image layout.

[0070] Visual focus is the key area in an image that attracts the user's attention. Different scene types have different visual focus areas. In traffic scene images, visual focus varies depending on the specific situation. For example, when a vehicle is waiting at a red light at an intersection, the driver and pedestrians typically focus on the traffic light, making it the key visual element of the scene. Meanwhile, while driving, the driver may pay more attention to the traffic sign ahead, making it the visual focal point. In medical scene images, visual focus also varies depending on the context. For example, in an image depicting a hospital registration hall, signs like the registration window and the call display become the visual focal point, as they provide key information guiding patients to register and seek medical treatment. In an image depicting a hospital ward, bedside signs (indicating patient information, care level, etc.) and emergency call button signs attract attention and become the visual focal point.

[0071] Determining the salience weights of signs in each sub-region based on visual focus means taking into account the importance and appeal of the signs within the image layout. In traffic scenarios, signs near the visual focus, such as those near traffic lights where drivers and pedestrians focus their attention while waiting at a red light, will have relatively high salience weights. In medical scenarios, signs near the registration window signs in the registration hall or near bedside signs in the ward will also have high salience weights. This makes the calculation of the layout's emotional contribution value more consistent with human visual perception and attention allocation, and can more accurately reflect the emotional impact of signs within the image layout.

[0072] Next, the layout sentiment contribution value of each sub-region is calculated by combining the basic density score and the sign salience weight. The basic density score reflects the distribution of signs, while the sign salience weight reflects the visual importance of the signs. Combining these two allows for a more comprehensive assessment of the emotional impact of signs in the image layout. For example, if a sub-region has a large number of signs (a high basic density score) and is located near the visual focal point (a high sign salience weight), the layout sentiment contribution value of this sub-region will be relatively high. Finally, the layout sentiment contribution values ​​of each sub-region are summed to obtain the layout sentiment contribution value of the entire image data, achieving a quantitative assessment of the emotional impact of the image layout.

[0073] Optionally, the operation of determining the weight information of each category of visual elements in the image data based on the scene type and pre-acquired user feedback information on the target scene layout includes: determining the initial weight information of each category of visual elements in the image data based on a preset rule library and the scene type; wherein the preset rule library stores the initial weight information of each category of visual elements in image data of different scene types; and based on the feedback information, using a reinforcement learning mechanism to optimize the initial weight information of each category of visual elements in the image data to obtain the weight information of each category of visual elements in the image data; wherein the reward signal of the reinforcement learning mechanism is obtained based on the feedback information.

[0074] Specifically, the preset rule base is a database that stores a large amount of information about the initial weights of various visual elements in image data of different scene types. Because different scene types (such as traffic scenes, medical scenes, natural scenery scenes, etc.) have different characteristics and requirements, the importance of visual elements in these scenes also varies. The preset rule base assigns corresponding initial weights to each category of visual elements for different scene types. When processing new image data, a reasonable initial weight distribution scheme can be quickly obtained, providing a starting point for subsequent weight optimization, making the entire weight determination process more orderly and efficient.

[0075] Therefore, based on the input scene type, matching rules can be searched in the preset rule library to obtain the initial weight information for each category of elements (including text elements, image elements, and layout elements) in the image data for that scene type. Taking the traffic scene as an example, the initial weight information α, β, and γ corresponding to text elements, image elements, and layout elements stored in the preset rule library are 0.6, 0.3, and 0.1, respectively. Taking the medical scene as an example, the initial weight information α, β, and γ corresponding to text elements, image elements, and layout elements stored in the preset rule library are 0.5, 0.3, and 0.2, respectively. Taking the cultural and tourism scene as an example, the initial weight information α, β, and γ corresponding to text elements, image elements, and layout elements stored in the preset rule library are 0.5, 0.2, and 0.3, respectively.

[0076] Next, considering that feedback information reflects the user's actual perception and needs regarding image layout, it serves as an important basis for optimizing visual element weights. By analyzing this feedback, we can understand the user's attention, preferences, and dissatisfaction with different visual elements, allowing for targeted adjustments to the initial weights. This technical solution utilizes a reinforcement learning mechanism to automatically adjust the weights of each visual element based on feedback. Reinforcement learning is a machine learning method that uses an intelligent agent to interact with its environment and learn optimal behavior strategies based on environmental feedback (rewards or penalties). During the dynamic weight adjustment process, the intelligent agent can be viewed as an algorithm that adjusts the visual element weights, while the environment is a collection of image data and user feedback. The reinforcement learning mechanism automatically adjusts the visual element weights based on feedback, ensuring that the weight distribution better meets the user's actual needs. Compared to traditional rule-based weight adjustment methods, reinforcement learning offers greater adaptability and learning capabilities, enabling it to gradually optimize the weight distribution scheme through continuous learning.

[0077] In this embodiment, the spatial state of the reinforcement learning strategy is: the current weight parameters (α, β, γ), the scenario type, and historical effect data (i.e., user feedback information). The action space of the reinforcement learning strategy is the incremental adjustment of the weight parameters (e.g., Δα = +0.1, Δβ = +0.1, Δγ = -0.1). The reward function is composed of user satisfaction (weight 70%) and the reduction in violation rate (weight 30%).

[0078] The reward signal is the core driving force of the reinforcement learning mechanism, which guides the intelligent agent to continuously explore and optimize the weight distribution scheme to obtain the maximum reward. By reasonably designing the calculation method of the reward signal, the reinforcement learning mechanism can pay more attention to the key needs of the user and improve the effect of weight optimization. The reward signal of this technical solution is obtained based on feedback information and is used to guide the reinforcement learning mechanism to adjust the weights of visual elements. Among them, the reward signal can be positive or negative. A positive reward signal indicates that the current weight distribution scheme has been recognized by the user and should continue to be maintained or strengthened; a negative reward signal indicates that the current weight distribution scheme does not meet the needs of the user and needs to be adjusted.

[0079] Therefore, the reinforcement learning mechanism can continuously adjust the weights of visual elements based on reward signals. After multiple iterations and learning, it ultimately obtains weight information that better meets user needs. Using user satisfaction (rating 1-5) and scene effect (such as the reduction in violation rate) as reward signals, the weight parameters (α, β, γ) are dynamically optimized through a policy gradient algorithm (such as PPO). The specific adjustment logic is as follows:

[0080] Taking a traffic scenario as an example, let's assume that visual elements include traffic sign text warnings (weight parameter α), graphical guidance (weight parameter β), and sign layout (weight parameter γ). If feedback indicates that the text warnings on traffic signs are too strong, causing user resistance (such as drivers), resulting in lower user satisfaction scores, and potentially leading to potential violations due to driver distraction (the violation rate does not decrease as expected), the reinforcement learning mechanism will reduce α and increase β. Reducing the strength of text warnings can reduce user resistance, while increasing the weight of graphical guidance can more intuitively guide driver behavior, thereby optimizing the visual presentation of the traffic scene and improving the user (driver) experience.

[0081] As a result, the weight information optimized by the reinforcement learning mechanism better meets the actual needs of users, improving the visual effects and user experience of image data. At the same time, this weight determination method is also universal and scalable and can be applied to different types of scene image data.

[0082] Optionally, the operation of determining the emotional intensity value of the target scene based on the feature information of all categories of visual elements in the image data and the corresponding weight information includes: calculating the emotional intensity value of the target scene according to the feature information of all categories of visual elements in the image data and the corresponding weight information using the following formula:

[0083]

[0084] Wherein, S is the emotional intensity value of the target scene, is the feature information of the text elements in the image data, is the weight information of the text elements in the image data, is the feature information of the image elements in the image data, is the weight information of the image elements in the image data, is the characteristic information of the layout elements in the image data, is the weight information of the layout elements in the image data.

[0085] Optionally, the operation of generating image data of the target scene after optimized layout using an image generation model according to the layout optimization suggestion includes: generating prompt words according to the layout optimization suggestion; and inputting the image data and the prompt words into a pre-trained image generation model, and outputting the image data of the target scene after optimized layout.

[0086] Specifically, optimization suggestions can be converted into specific prompts based on the input requirements of generative AI (such as Stable Diffusion). These prompts are used to guide the model in adjusting the corresponding visual elements in the image data according to specific requirements, such as background color, element shape, icon symbol style, text content, and setting location. For example, the background color can be explicitly specified as "blue background"; the element shape can be required to have a "rounded border"; and the icon symbol style can be described as "simple arrow icon."

[0087] The converted prompt words and the initially acquired image data (such as the original scene's image features and layout information) are then fed into an image generation model (such as a diffusion model). Upon receiving this input, the image generation model uses its internal learning algorithms and parameters to process and analyze the original image data and prompt words. The image data provides the model with foundational information, while the prompt words guide the model in generating images according to specific requirements. Through a series of calculations and inferences, the model generates new image data, representing the target scene image with an optimized layout, which the model returns as output.

[0088] The following uses a city park as an example to illustrate the various steps of this technical solution:

[0089] There are many problems with the layout and design of the guide signs in this city park. For example, from a textual perspective, expressions such as "Tourists Stop" have a high emotional intensity (an emotional intensity value of 4.5, which is a high-resistance emotional trigger word), which can easily make tourists feel disgusted. In terms of images, the red prohibition symbol has a high intensity (an emotional intensity value of 0.9), giving people a strong sense of prohibition, and the layout of the signs is too dense (6 / ㎡), resulting in visual confusion in the space and a layout penalty (-0.3). These problems combined have resulted in an 18% probability of tourists mistakenly entering the ecological protection area. At the same time, tourists' satisfaction rating for the park's guide signs is only 3.5 points out of a full 7.0 points, reflecting that the current guide sign system fails to effectively guide tourists and brings them a bad visual and psychological experience. Therefore, it is necessary to optimize the layout of the guide signs in this target scene. The optimization steps are as follows:

[0090] Step 1: Image data acquisition

[0091] The image data of the city park is collected through a camera, a drone or a public map API interface, and the scene type of the city park is determined to be a cultural and tourism scene.

[0092] Step 2: Feature extraction and analysis

[0093] (1) Text feature extraction and analysis: We used optical character recognition (OCR) technology to extract the text content of guide signs from image data. We then used the ERNIE-3.0 natural language processing model to analyze the text in the guide signs, finding that the sentiment contribution value for text such as "Tourists Stop" was 4.5. This model allows us to deeply understand the emotional tendencies and influence conveyed by the text, providing basic data for subsequent optimization.

[0094] (2) Image feature extraction and analysis: The YOLOv5 target detection algorithm combined with color analysis methods were used to process the image elements in the guide signs. YOLOv5 can quickly and accurately detect image elements such as the red prohibition sign. Color analysis further evaluates the visual impact of image elements, and the emotional contribution value of the red prohibition sign is 0.9.

[0095] (3) Layout feature extraction and analysis: By statistically analyzing the actual distribution of guide signs in the park, we found that the sign density is 6 per square meter. According to the preset rules, when the sign density exceeds a certain threshold, a layout penalty will be generated. Here, the layout penalty is calculated to be -0.3, that is, the layout sentiment contribution value is -0.3.

[0096] Step 3: Emotional computing and optimized decision making

[0097] (1) Initial weight setting: By calling the preset rule library, we can find that the initial weight of text in the cultural tourism scene is α = 0.5, the initial weight of image is β = 0.2, and the initial weight of layout is γ = 0.3. This weight distribution reflects that in the cultural tourism scene, image elements usually have a more significant visual impact on tourists, so they are given a higher weight; text and layout elements also have certain weights to ensure that all factors are taken into consideration.

[0098] (2) User feedback collection and initial weight optimization: Set up multiple feedback points in the park, such as the visitor service center and the entrance to major attractions, and collect tourists' opinions and suggestions on guide signs through paper questionnaires, electronic questionnaires, and on-site interviews. For example, ask tourists whether the text content of the existing guide signs is easy to understand, whether the images are clear, and whether the layout is reasonable, and ask tourists to score the importance of different aspects. Organize and analyze the collected user feedback data, and use data analysis tools (such as Excel, SPSS, etc.) to statistically analyze the distribution of opinions on different aspects and find out the issues and needs that tourists generally care about. For example, if it is found that most tourists think that the text content is more important than the image, and they also have more opinions on the rationality of the layout, then the initial weight needs to be adjusted. Then, according to the analysis results of the feedback data, the initial weight is optimized and adjusted. Assume that after analysis, it is found that the text weight should be appropriately increased, and the image weight and layout weight also need to be fine-tuned according to tourist feedback. After optimization, the new weights are set to α = 0.3, β = 0.5, and γ = 0.2. Through optimization, the weight distribution is made more in line with the actual needs and concerns of tourists.

[0099] (3) Comprehensive Strength Calculation: Based on the optimized weights and feature strengths, the comprehensive strength is calculated as 0.3×4.5 + 0.5×0.9 + 0.2×(-0.3) = 1.35 + 0.45 - 0.06 = 1.74. Since the comprehensive strength is too high (needs to be reduced to below 1.0), it indicates that the current guide sign design has a significant negative impact on tourists and needs to be optimized.

[0100] Step 4: Generate optimization suggestions

[0101] Based on the previous data analysis and comprehensive strength calculation results, optimization suggestions were determined. For example, the text could be replaced with "Ecological Protection Area, Thank You for Your Cooperation" (strength 2.8), which is gentler and friendlier, potentially alleviating visitors' resistance. A green background (strength -0.6) was used, as green often creates a natural and comfortable feeling, helping to alleviate visitors' anxiety. Furthermore, the sign density could be reduced to 4 signs per square meter to create a more rational sign layout and avoid visual clutter.

[0102] Step 5: Image data generation

[0103] Based on the optimization suggestions, corresponding prompt words were generated. These prompt words and the original image data were input into a diffusion model to generate scene image data featuring a new sign with a green shading, leaf icons, and a dispersed layout, along with a density of 4 signs per square meter. This was then presented to users through a visual representation. The green shading blends in with the park's natural environment, the leaf icons further reinforce the ecological conservation theme, and the dispersed layout makes the sign clearer and easier to read, enhancing the visual appeal and guiding function of the guide sign.

[0104] Step 6: Effect Verification

[0105] The effectiveness of the optimized guide signs has been verified in practice. After a period of observation and statistical analysis, it was found that the probability of tourists mistakenly entering the ecological protection zone has dropped from 18% to 6%. This shows that the optimized guide signs can more effectively guide tourists and reduce the risk of tourists mistakenly entering the restricted area.

[0106] At the same time, visitor satisfaction scores for the park's guide signs increased from 3.5 / 7.0 to 6.2 / 7.0. This indicates that visitors highly evaluate the optimized guide signs in terms of visual effects, information transmission, and guidance functions, further proving the effectiveness of the technical solution.

[0107] By optimizing the layout of guide signs in a city park, this paper applied data collection and feature extraction techniques to accurately analyze existing problems with the signage. A dynamic weight allocation method was employed, and initial weights were optimized based on user feedback, comprehensively considering the impact of factors such as text, images, and layout on visitors. Generative AI visualization technology was used to propose reasonable optimization suggestions and generate new guide signs. Finally, through performance verification, it was demonstrated that the optimized guide signs significantly reduced the rate of visitors straying from the park and improved visitor satisfaction. This provides a scientific and effective approach for optimizing guide signs in city parks, contributing to improved park management and service quality.

[0108] In addition, according to this embodiment, a storage medium is also provided, which includes a stored program, wherein when the program is run, a processor executes any one of the above methods.

[0109] The beneficial effects of this application include:

[0110] (1) Breaking through the limitations of traditional fixed weight models, a dynamic weight adjustment method based on scenario types (covering different fields such as transportation, culture and tourism) and real-time feedback (including key indicators such as user ratings and violation rates) is proposed. During the dynamic weight adjustment process, a reinforcement learning algorithm (such as the PPO algorithm) is introduced to optimize the weights (α, β, γ) of multiple visual elements, thereby achieving precise adaptation of the emotional strategy. This allows the model to flexibly adjust according to different scenarios and real-time feedback, thereby improving the accuracy and adaptability of decision-making.

[0111] (2) Comprehensively consider multiple visual elements such as text, images, and layout, deeply explore the interactive effects between them, including synergy and conflict, and conduct quantitative analysis. Through this integration and quantification, we can more comprehensively and accurately grasp the impact of each element on the overall effect, providing a scientific basis for subsequent optimization.

[0112] (3) Combining generative AI technology with reinforcement learning to create an automated design closed loop of “analysis → optimization → generation → feedback → iteration”. This closed loop can achieve automated monitoring and continuous optimization of the design process, improve design efficiency and quality, reduce manual intervention, and lower design costs.

[0113] (4) Detailed classification of emotional intensity, for example, by subdividing it into values ​​between 1 and 5, and accurately matching it according to the needs of different scenarios. For example, in a hospital scenario, the combination of "moderate warning + soothing color" can not only effectively convey necessary information, but also relieve patients' tension and improve user experience.

[0114] It should be noted that for the aforementioned method embodiments, for simplicity of description, they are all expressed as a series of action combinations. However, those skilled in the art should be aware that the present invention is not limited by the order of the actions described, because according to the present invention, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in this specification are all preferred embodiments, and the actions and modules involved are not necessarily required by the present invention.

[0115] Example 2

[0116] Figure 3The schematic diagram of the structure of the scene layout optimization system based on multiple visual elements and emotion perception according to this embodiment is shown. The scene layout optimization system includes: a data acquisition module 310, which is used to acquire image data of a target scene and determine the scene type of the target scene; wherein the image data includes multiple categories of visual elements; an extraction and analysis module 320, which is used to extract feature information of each category of visual elements in the image data, and determine the weight information of each category of visual elements in the image data based on the scene type and pre-acquired user feedback information on the layout of the target scene; an emotion calculation and optimization decision module 330, which is used to determine the emotion intensity value of the target scene based on the feature information and corresponding weight information of all categories of visual elements in the image data, and determine the layout optimization suggestion of the target scene based on the emotion intensity value; wherein the emotion intensity value is used to indicate the comprehensive impact of all categories of visual elements in the target scene on the user's emotional experience; and a layout optimization module 340, which is used to generate the image data of the target scene after the layout optimization suggestion using an image generation model.

[0117] Optionally, the multiple categories of visual elements include text elements, image elements and layout elements; and the extraction and parsing module 320 is specifically used to: extract the text content in the image data through optical character recognition technology, calculate the text emotion contribution value corresponding to the text content through a pre-trained language model, and use it as the feature information of the text element in the image data; identify the graphic symbols in the image data based on the target detection algorithm, determine the semantic type of the graphic symbols, extract the HSV value of the main color tone of the image data, determine the image emotion contribution value of the image data according to the semantic type and the HSV value, and use it as the feature information of the image element in the image data; and based on the image data, count the number of signs of the target scene within a unit area, determine the visual focus of the image data through a saliency detection model, and determine the layout emotion contribution value of the image data according to the number of signs and the visual focus, and use it as the feature information of the layout element in the image data.

[0118] Optionally, the operation of determining the image emotion contribution value of the image data based on the semantic type and the HSV value includes: determining the symbol emotion contribution value corresponding to the graphic symbol based on the semantic type; calculating the color emotion contribution value of the image data based on the HSV value; and determining the image emotion contribution value of the image data based on the symbol emotion contribution value and the color emotion contribution value.

[0119] Optionally, the operation of determining the layout emotion contribution value of the image data based on the number of identifications and the visual focus includes: determining the basic density score of each sub-area of ​​the target scene based on the number of identifications; wherein, the target scene is divided into multiple sub-areas according to the unit area; according to the visual focus, the identification significance weight of each sub-area is determined; and according to the basic density score and the identification significance weight, the layout emotion contribution value of each sub-area is determined, and the sum of the layout emotion contribution values ​​of each sub-area is used as the layout emotion contribution value of the image data.

[0120] Optionally, the extraction and parsing module 320 is specifically used to: determine the initial weight information of the visual elements of each category in the image data according to a preset rule base and the scene type; wherein the preset rule base stores the initial weight information of the visual elements of each category in the image data of different scene types; and according to the feedback information, use a reinforcement learning mechanism to optimize the initial weight information of the visual elements of each category in the image data to obtain the weight information of the visual elements of each category in the image data; wherein the reward signal of the reinforcement learning mechanism is obtained based on the feedback information.

[0121] Optionally, the emotion calculation and optimization decision module 330 is specifically configured to calculate the emotion intensity value of the target scene according to the feature information and corresponding weight information of all categories of visual elements in the image data using the following formula:

[0122]

[0123] Wherein, S is the emotional intensity value of the target scene, is the feature information of the text elements in the image data, is the weight information of the text elements in the image data, is the feature information of the image elements in the image data, is the weight information of the image elements in the image data, is the characteristic information of the layout elements in the image data, is the weight information of the layout elements in the image data.

[0124] Optionally, the layout optimization module 340 is specifically configured to: generate prompt words according to the layout optimization suggestion; and input the image data and the prompt words into a pre-trained image generation model to output image data of the target scene after the layout is optimized.

[0125] Therefore, according to this embodiment, through the organic combination of multiple links such as accurate scene type identification, dynamic weight allocation, full consideration of the comprehensive influence between multiple visual elements, quantitative evaluation of emotional intensity values, generation of layout optimization suggestions, and application of image generation models, a multi-dimensional improvement of scene layout optimization is achieved. This method not only improves the pertinence and effectiveness of the optimization results, but also enhances the flexibility and adaptability of the optimization process, greatly improves the scene layout optimization effect, and provides users with a more personalized and comfortable scene experience. This solves the technical defects of the scene layout optimization method in the existing technology, such as the rigid static weight allocation mechanism, the failure to consider the comprehensive influence between multiple visual elements, and the insufficient dynamic quantification ability of emotional intensity, which leads to the technical problem of poor scene layout optimization effect.

[0126] Example 3

[0127] Figure 4 A scene layout optimization system based on multiple visual elements and emotional perception according to the present embodiment is shown, comprising: a processor 410; and a memory 420, connected to the processor 410, for providing instructions for the processor 410 to process the following processing steps: obtaining image data of a target scene, and determining a scene type of the target scene; wherein the image data includes visual elements of multiple categories; extracting feature information of each category of visual elements in the image data, and determining weight information of each category of visual elements in the image data based on the scene type and pre-acquired user feedback information on the layout of the target scene; determining an emotional intensity value of the target scene based on the feature information and corresponding weight information of all categories of visual elements in the image data, and determining a layout optimization suggestion for the target scene based on the emotional intensity value; wherein the emotional intensity value is used to indicate the comprehensive impact of all categories of visual elements in the target scene on the user's emotional experience; and generating image data of the target scene after optimized layout using an image generation model based on the layout optimization suggestion.

[0128] Optionally, the multiple categories of visual elements include text elements, image elements and layout elements; and the operation of extracting feature information of each category of visual elements in the image data includes: extracting text content in the image data through optical character recognition technology, calculating the text emotion contribution value corresponding to the text content through a pre-trained language model, and using it as the feature information of the text elements in the image data; identifying graphic symbols in the image data based on a target detection algorithm, determining the semantic type of the graphic symbols, extracting the HSV value of the main color tone of the image data, and determining the image emotion contribution value of the image data according to the semantic type and the HSV value, as the feature information of the image elements in the image data; and based on the image data, counting the number of logos of the target scene within a unit area, determining the visual focus of the image data through a saliency detection model, and determining the layout emotion contribution value of the image data according to the number of logos and the visual focus, as the feature information of the layout elements in the image data.

[0129] Optionally, the operation of determining the image emotion contribution value of the image data based on the semantic type and the HSV value includes: determining the symbol emotion contribution value corresponding to the graphic symbol based on the semantic type; calculating the color emotion contribution value of the image data based on the HSV value; and determining the image emotion contribution value of the image data based on the symbol emotion contribution value and the color emotion contribution value.

[0130] Optionally, the operation of determining the layout emotion contribution value of the image data based on the number of identifications and the visual focus includes: determining the basic density score of each sub-area of ​​the target scene based on the number of identifications; wherein, the target scene is divided into multiple sub-areas according to the unit area; according to the visual focus, the identification significance weight of each sub-area is determined; and according to the basic density score and the identification significance weight, the layout emotion contribution value of each sub-area is determined, and the sum of the layout emotion contribution values ​​of each sub-area is used as the layout emotion contribution value of the image data.

[0131] Optionally, the operation of determining the weight information of each category of visual elements in the image data based on the scene type and pre-acquired user feedback information on the target scene layout includes: determining the initial weight information of each category of visual elements in the image data based on a preset rule library and the scene type; wherein the preset rule library stores the initial weight information of each category of visual elements in image data of different scene types; and based on the feedback information, using a reinforcement learning mechanism to optimize the initial weight information of each category of visual elements in the image data to obtain the weight information of each category of visual elements in the image data; wherein the reward signal of the reinforcement learning mechanism is obtained based on the feedback information.

[0132] Optionally, the operation of determining the emotional intensity value of the target scene based on the feature information of all categories of visual elements in the image data and the corresponding weight information includes: calculating the emotional intensity value of the target scene according to the feature information of all categories of visual elements in the image data and the corresponding weight information using the following formula:

[0133]

[0134] Wherein, S is the emotional intensity value of the target scene, is the feature information of the text elements in the image data, is the weight information of the text elements in the image data, is the feature information of the image elements in the image data, is the weight information of the image elements in the image data, is the characteristic information of the layout elements in the image data, is the weight information of the layout elements in the image data.

[0135] Optionally, the operation of generating image data of the target scene after optimized layout using an image generation model according to the layout optimization suggestion includes: generating prompt words according to the layout optimization suggestion; and inputting the image data and the prompt words into a pre-trained image generation model, and outputting the image data of the target scene after optimized layout.

[0136] Therefore, according to this embodiment, through the organic combination of multiple links such as accurate scene type identification, dynamic weight allocation, full consideration of the comprehensive influence between multiple visual elements, quantitative evaluation of emotional intensity values, generation of layout optimization suggestions, and application of image generation models, a multi-dimensional improvement of scene layout optimization is achieved. This method not only improves the pertinence and effectiveness of the optimization results, but also enhances the flexibility and adaptability of the optimization process, greatly improves the scene layout optimization effect, and provides users with a more personalized and comfortable scene experience. This solves the technical defects of the scene layout optimization method in the existing technology, such as the rigid static weight allocation mechanism, the failure to consider the comprehensive influence between multiple visual elements, and the insufficient dynamic quantification ability of emotional intensity, which leads to the technical problem of poor scene layout optimization effect.

[0137] The serial numbers of the above embodiments of the present invention are for description only and do not represent the advantages or disadvantages of the embodiments.

[0138] In the above embodiments of the present invention, the description of each embodiment has its own focus. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0139] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are only schematic. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of units or modules, which can be electrical or other forms.

[0140] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.

[0141] In addition, the functional units in the various embodiments of the present invention may be integrated into a single processing unit, each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.

[0142] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the portion that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes various media that can store program code, such as a USB flash drive, read-only memory (ROM), random access memory (RAM), a mobile hard drive, a magnetic disk, or an optical disk.

[0143] The above is only a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present invention. These improvements and modifications should also be regarded as within the scope of protection of the present invention.

Claims

1. A scene layout optimization method based on multiple visual elements and emotional perception, characterized in that: include: Acquire image data of a target scene with layout optimization requirements, and determine the scene type of the target scene; wherein the image data includes visual elements of multiple categories; Extracting feature information of each category of visual elements in the image data, and determining weight information of each category of visual elements in the image data based on the scene type and pre-acquired user feedback information on the target scene layout; wherein the feedback information is used to represent the user's attention level, preferences, and / or dissatisfaction with different visual elements; determining an emotional intensity value of the target scene based on feature information and corresponding weight information of all categories of visual elements in the image data, and determining a layout optimization suggestion for the target scene based on the emotional intensity value; wherein the emotional intensity value is used to indicate the comprehensive impact of all categories of visual elements in the target scene on the user's emotional experience, a higher emotional intensity value indicates a higher level of resistance to the layout of the target scene for the user, and a lower emotional intensity value indicates a stronger positive emotional experience for the user caused by the target scene; the layout optimization suggestion includes a layout optimization suggestion for at least one category of visual elements; and Generate prompt words based on the layout optimization suggestion; input the image data and the prompt words into a pre-trained image generation model to generate pixel-level image data of the target scene after the layout is optimized.

2. The method according to claim 1, characterized in that The plurality of categories of visual elements include text elements, image elements, and layout elements; and the operation of extracting feature information of each category of visual elements in the image data includes: Extracting text content from the image data using optical character recognition technology, and calculating text sentiment contribution values ​​corresponding to the text content using a pre-trained language model as feature information of text elements in the image data; Identifying graphic symbols in the image data based on an object detection algorithm, determining a semantic type of the graphic symbol, extracting an HSV value of a dominant color tone of the image data, and determining an image emotion contribution value of the image data based on the semantic type and the HSV value as feature information of an image element in the image data; and Based on the image data, the number of logos in the target scene per unit area is counted, the visual focus of the image data is determined through a saliency detection model, and the layout emotion contribution value of the image data is determined based on the number of logos and the visual focus as feature information of the layout elements in the image data.

3. The method according to claim 2, characterized in that The operation of determining the image emotion contribution value of the image data according to the semantic type and the HSV value includes: Determining a symbolic emotion contribution value corresponding to the graphic symbol according to the semantic type; Calculating a color emotion contribution value of the image data according to the HSV value; and An image emotion contribution value of the image data is determined according to the symbol emotion contribution value and the color emotion contribution value.

4. The method according to claim 2, characterized in that The operation of determining the layout emotion contribution value of the image data according to the number of identifiers and the visual focus includes: Determining a basic density score of each sub-area of ​​the target scene according to the number of identifiers; wherein the basic density score is determined by dividing the target scene into a plurality of sub-areas according to unit area; Determining the identification significance weight of each sub-region according to the visual focus; and The layout emotion contribution value of each sub-region is determined according to the basic density score and the identification significance weight, and the sum of the layout emotion contribution values ​​of each sub-region is used as the layout emotion contribution value of the image data.

5. The method according to claim 1, wherein The operation of determining weight information of visual elements of each category in the image data according to the scene type and pre-acquired user feedback information on the target scene layout includes: Determining initial weight information of visual elements of each category in the image data according to a preset rule library and the scene type; wherein the preset rule library stores initial weight information of visual elements of each category in image data of different scene types; and Based on the feedback information, a reinforcement learning mechanism is used to optimize the initial weight information of the visual elements of each category in the image data to obtain the weight information of the visual elements of each category in the image data; wherein the reward signal of the reinforcement learning mechanism is obtained based on the feedback information.

6. The method according to claim 1, characterized in that The operation of determining the emotional intensity value of the target scene according to the feature information of all categories of visual elements in the image data and the corresponding weight information includes: According to the feature information of all categories of visual elements in the image data and the corresponding weight information, the emotional intensity value of the target scene is calculated by the following formula: Wherein, S is the emotional intensity value of the target scene, is the feature information of the text elements in the image data, is the weight information of the text elements in the image data, is the feature information of the image elements in the image data, is the weight information of the image elements in the image data, is the characteristic information of the layout elements in the image data, is the weight information of the layout elements in the image data.

7. A storage medium, characterized in that: The storage medium includes a stored program, wherein when the program is run, the processor executes the method according to any one of claims 1 to 6.

8. A scene layout optimization system based on multiple visual elements and emotional perception, comprising: A data acquisition module, configured to acquire image data of a target scene requiring layout optimization and determine the scene type of the target scene; wherein the image data includes visual elements of multiple categories; an extraction and parsing module, configured to extract feature information of each category of visual elements in the image data, and determine weight information of each category of visual elements in the image data based on the scene type and pre-acquired user feedback information on the target scene layout; wherein the feedback information is used to represent the user's attention level, preferences, and / or dissatisfaction with different visual elements; an emotion calculation and optimization decision module, configured to determine an emotion intensity value of the target scene based on feature information and corresponding weight information of all categories of visual elements in the image data, and determine a layout optimization suggestion for the target scene based on the emotion intensity value; wherein the emotion intensity value is used to indicate the combined impact of all categories of visual elements in the target scene on the user's emotional experience, a higher emotion intensity value indicates a higher degree of resistance to the target scene's layout in the user, and a lower emotion intensity value indicates a stronger positive emotional experience brought to the user by the target scene; the layout optimization suggestion includes a layout optimization suggestion for at least one category of visual elements; and The layout optimization module is used to generate prompt words according to the layout optimization suggestion; input the image data and the prompt words into a pre-trained image generation model to generate pixel-level image data of the target scene after the layout is optimized.

9. A scene layout optimization system based on multiple visual elements and emotional perception, characterized by: include: processor; as well as A memory, connected to the processor, configured to provide the processor with instructions for processing the following processing steps: Acquire image data of a target scene with layout optimization requirements, and determine the scene type of the target scene; wherein the image data includes visual elements of multiple categories; Extracting feature information of each category of visual elements in the image data, and determining weight information of each category of visual elements in the image data based on the scene type and pre-acquired user feedback information on the target scene layout; wherein the feedback information is used to represent the user's attention level, preferences, and / or dissatisfaction with different visual elements; determining an emotional intensity value of the target scene based on feature information and corresponding weight information of all categories of visual elements in the image data, and determining a layout optimization suggestion for the target scene based on the emotional intensity value; wherein the emotional intensity value is used to indicate the comprehensive impact of all categories of visual elements in the target scene on the user's emotional experience, a higher emotional intensity value indicates a higher level of resistance to the layout of the target scene for the user, and a lower emotional intensity value indicates a stronger positive emotional experience for the user caused by the target scene; the layout optimization suggestion includes a layout optimization suggestion for at least one category of visual elements; and Generate prompt words based on the layout optimization suggestion; input the image data and the prompt words into a pre-trained image generation model to generate pixel-level image data of the target scene after the layout is optimized.

Citation Information

Patent Citations

  • Multi-mode-based target area viewing scene planning method and device, computer equipment and readable storage medium

    CN118840648A

  • Scene graph generation method and system supporting historical and cultural block scene

    CN118334414A