Scene layout optimization method and system based on multiple visual elements and emotion perception

By obtaining scene image data, identifying visual element characteristics and dynamically adjusting weights, calculating emotional intensity values, and generating optimization suggestions, the problem of rigid scene layout optimization in the existing technology is solved, and a personalized and comfortable scene experience is achieved.

CN120339465AActive Publication Date: 2025-07-18BEIJING UNIV OF POSTS & TELECOMM +1

Patent Information

Application Number
CN202510828540.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-20
Publication Date
2025-07-18
Estimated Expiration
2045-06-20

AI Technical Summary

Technical Problem

In the prior art, the scene layout optimization method has a rigid static weight allocation mechanism, insufficient comprehensive impact between multiple visual elements, and insufficient dynamic quantization ability of emotional intensity, resulting in poor scene layout optimization effect.

Method used

By obtaining the image data of the target scene, identifying the scene type, extracting the feature information of visual elements, dynamically adjusting the weight of visual elements based on user feedback, calculating emotional intensity values, generating layout optimization suggestions, and optimizing scene layout using image generation model.

Benefits of technology

It realizes a more personalized and comfortable scene experience, improves the pertinence and flexibility of optimization results, enhances the adaptability of the optimization process, and improves the optimization effect of scene layout.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120339465A_ABST
    Figure CN120339465A_ABST
Patent Text Reader

Abstract

The invention discloses a scene layout optimization method and system based on multiple visual elements and emotion perception. Relates to the technical field of management optimization. The method comprises the steps of obtaining image data and a scene type of a target scene; feature information of visual elements of each category in the image data is extracted, and weight information of the visual elements of each category is determined according to the scene type and user feedback information; determining an emotion intensity value of the target scene according to the feature information and the weight information of all the types of visual elements, and determining a layout optimization suggestion of the target scene according to the emotion intensity value; and according to the layout optimization suggestion, generating image data of the target scene after layout optimization by using an image generation model. The technical problem that in the prior art, a scene layout optimization mode has the technical defects that a static weight distribution mechanism is rigid, comprehensive influence among a plurality of visual elements is not fully considered, and dynamic quantification capacity of emotional intensity is insufficient, so that the scene layout optimization effect is poor is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of management optimization technology, and in particular to a scene layout optimization method and system based on multiple visual elements and emotional perception. Background Art

[0002] Linguistic Landscape refers to the visible language symbol system in public space, including various material carriers for displaying language and text, such as signs, advertisements, road signs, and logos. It carries the construction of social significance and expression of power relations in urban space. Its core functions include multiple dimensions such as information transmission, cultural identity shaping, and ideological coding. With the continuous expansion of the cognitive scope of "linguistic landscape", scholars regard space as a symbolic resource and put forward the concept of "semiotic landscape", which is to construct spatial discourse with text as the medium and with the help of symbolic resources. It is no longer limited to a single written text. The research focus has shifted to multi-visual element symbolic resources, focusing on the interaction between text elements, image elements, layout elements, and cultural dimensions.

[0003] The language landscape forms a semantic unit composed of multiple visual element symbols through the combination of visual elements such as text, graphics, color, and layout. For example, the "No Parking" text (text element symbol), red circle plus slash (image element symbol), high contrast color matching (color symbol, which can be classified as part of the image element category) and reasonable layout position (layout element symbol) in the traffic sign together constitute a complete semantic-emotional expression system. It can be seen that the language landscape is composed of language symbols and non-language symbols, and presents a symbol aggregate in the form of single or multiple visual elements such as text, graphics, color, and layout. When studying it, we should not only focus on static and fixed language modes, but also interpret the arrangement and combination meaning of various visual elements in the language landscape as a whole. For example, in the subway guide scene, the warning slogans (text elements), triangular warning icons (image elements), and ceiling hanging positions (layout elements) jointly affect passengers' safety cognition and emotional response. With the increase in the complexity of scene design, how to quantify and analyze the synergistic effects of multiple categories of visual elements has become a key challenge in optimizing the user's emotional experience.

[0004] Based on the in-depth study of language landscape, existing technologies have also explored scene planning and scene graph generation. However, these technologies still have some problems.

[0005] The publication number is CN118840648A, and the title is a method, device, computer device, and readable storage medium for planning a viewing scene of a target area based on multimodality (multimodality refers to the integration of multiple information modalities, such as text, images, audio, etc.). The disclosed method includes: First, obtain the viewing video data of the target area, covering multiple preset viewing scene pictures. Then, by obtaining and analyzing the multimodal feedback results of the user for the viewing video data, evaluate the visual landscape quality of each preset viewing scene picture. At the same time, considering the preset visual impact factors of the target area landscape, perform a correlation analysis on the visual landscape quality evaluation results and the landscape visual impact factors, and finally obtain the landscape scene planning result of the target area. However, this solution has the following defects: (1) Insufficient capture and expression of emotional information: It mainly relies on analyzing the visual quality of preset scene pictures from viewing video data, lacking an accurate capture and effective expression mechanism for the emotional information contained in each visual element in the language landscape, making it difficult to accurately reflect the impact of different combinations of visual elements on the user's emotions, resulting in poor optimization effects of the final scene layout.

[0006] (2) Lack of dynamic adjustment ability: It does not fully consider the differences in different scene requirements, cannot dynamically adjust the combination of multiple visual elements according to the actual scene, making it difficult to achieve the best emotional expression effect, resulting in insufficient flexibility and adaptability in scene planning.

[0007] The publication number is CN118334414A, and the title is a method and system for generating a scene graph to support the scene of a historical and cultural block. The disclosed method includes: Analyze the historical and cultural geographical elements and their relationships, and construct a dataset for generating a scene graph of the historical and cultural block; Train a landscape semantic recognition model to extract the visual, semantic, and spatial features of the target object; Use the co-occurrence frequency of target relationships (co-occurrence frequency refers to the frequency of two or more targets appearing simultaneously in the same scene) as prior knowledge, and based on the attention mechanism and feature fusion method, realize the prediction of target relationships; Use the temporally and spatially continuous street view images as input, and based on the landscape semantic recognition model and the target relationship prediction model, generate the scene graph corresponding to each street view image, and use it as input to utilize the scene graph fusion model to realize the fusion of the scene graphs. Although this solution integrates geographical elements and spatial features, it has the following deficiencies: (1) Weak processing of emotional information: It relies on a fixed co-occurrence frequency in the prediction of element relationships, lacking in-depth exploration and analysis of the emotional information of each visual element in the language landscape, unable to accurately capture the emotional connotations contained in different combinations of visual elements, making it difficult to effectively express emotional information, and resulting in poor optimization effects of the scene layout.

[0008] (2) Limited dynamic adjustment ability: It fails to dynamically adjust the weight information of multiple visual elements according to different scenario requirements and user feedback information, making it difficult to achieve the best emotional expression effect, resulting in a disconnection between emotional expression and functional requirements and poor optimization effect of the scenario layout.

[0009] Regarding the technical defects of the existing scenario layout optimization methods, such as the rigid static weight allocation mechanism, the insufficient consideration of the comprehensive influence among multiple visual elements, and the lack of dynamic quantification ability of emotional intensity, which lead to the problem of poor scenario layout optimization effect, no effective solution has been proposed yet. Summary of the Invention

[0010] Embodiments of the present disclosure provide a method and system for optimizing a scenario layout based on multiple visual elements and emotional perception, which can at least solve the technical problems in the prior art that the existing scenario layout optimization methods have technical defects such as a rigid static weight allocation mechanism, insufficient consideration of the comprehensive influence among multiple visual elements, and lack of dynamic quantification ability of emotional intensity, resulting in poor scenario layout optimization effect.

[0011] According to one aspect of the embodiments of the present disclosure, a method for optimizing a scenario layout based on multiple visual elements and emotional perception is provided, including: obtaining image data of a target scenario and determining the scenario type of the target scenario; wherein the image data includes multiple categories of visual elements; extracting feature information of each category of visual elements in the image data, and determining weight information of each category of visual elements in the image data according to the scenario type and pre-obtained user feedback information on the layout of the target scenario; determining an emotional intensity value of the target scenario according to the feature information and corresponding weight information of all categories of visual elements in the image data, and determining a layout optimization suggestion for the target scenario according to the emotional intensity value; wherein the emotional intensity value is used to indicate the comprehensive influence of all categories of visual elements in the target scenario on the user's emotional experience; and generating image data of the optimized-layout target scenario by using an image generation model according to the layout optimization suggestion.

[0012] According to another aspect of the embodiments of the present disclosure, a storage medium is further provided, and the storage medium includes a stored program, wherein the above-mentioned method is executed by a processor when the program runs.

[0013] According to another aspect of the embodiments of the present disclosure, there is also provided a scene layout optimization system based on multi-visual elements and emotion perception, including: a data acquisition module, configured to acquire image data of a target scene and determine the scene type of the target scene; wherein the image data includes visual elements of multiple categories; an extraction and parsing module, configured to extract the feature information of the visual elements of each category in the image data, and determine the weight information of the visual elements of each category in the image data according to the scene type and the feedback information of the user on the layout of the target scene acquired in advance; an emotion calculation and optimization decision module, configured to determine the emotion intensity value of the target scene according to the feature information and the corresponding weight information of all categories of visual elements in the image data, and determine the layout optimization suggestion of the target scene according to the emotion intensity value; wherein the emotion intensity value is used to indicate the comprehensive influence of all categories of visual elements in the target scene on the user's emotional experience; and a layout optimization module, configured to generate the image data of the target scene with an optimized layout by using an image generation model according to the layout optimization suggestion.

[0014] According to another aspect of the embodiments of the present disclosure, there is also provided a scene layout optimization system based on multi-visual elements and emotion perception, including a processor; and a memory, connected to the processor, for providing instructions for the processor to perform the following processing steps: acquiring image data of a target scene and determining the scene type of the target scene; wherein the image data includes visual elements of multiple categories; extracting the feature information of the visual elements of each category in the image data, and determining the weight information of the visual elements of each category in the image data according to the scene type and the feedback information of the user on the layout of the target scene acquired in advance; determining the emotion intensity value of the target scene according to the feature information and the corresponding weight information of all categories of visual elements in the image data, and determining the layout optimization suggestion of the target scene according to the emotion intensity value; wherein the emotion intensity value is used to indicate the comprehensive influence of all categories of visual elements in the target scene on the user's emotional experience; and generating the image data of the target scene with an optimized layout by using an image generation model according to the layout optimization suggestion.

[0015] This application first obtains the image data of the target scene and determines the scene type, so as to perform customized processing according to the characteristics of different scenes (such as traffic, landscape, medical, etc.), providing data support for subsequent visual element feature extraction, weight assignment, and layout optimization, and ensuring that the layout optimization measures are more in line with the actual scene requirements. Then, it extracts the feature information of the visual elements of each category (such as text elements, image elements, and layout elements) in the image data, and dynamically determines the weight information of the visual elements of each category according to the scene type and the user's feedback information on the target scene layout. It can flexibly adjust the importance of each visual element in the scene layout according to the changes in different scenes and user preferences, so as to achieve more personalized scene layout optimization. Secondly, by comprehensively considering the feature information of all categories of visual elements and their corresponding weight information, it calculates the emotional intensity value of the target scene, which can intuitively reflect the comprehensive impact of all visual elements in the target scene on the user's emotional experience. Thus, through the feature analysis, relationship study, and comprehensive weight consideration of different visual elements, it ensures that each visual element plays the best role in the scene layout, providing an objective and measurable basis for layout optimization. After that, according to the emotional intensity value, it generates specific layout optimization suggestions, which can guide the adjustment of the scene layout, making the optimization process more directional and operable. Finally, according to the layout optimization suggestions, it uses an image generation model to generate the image data of the target scene after layout optimization, so as to transform the optimization suggestions into actual visual presentations, enabling users to intuitively see the optimized scene effects, providing data support for subsequent visual and interactive optimization result displays, enhancing the user's sense of participation and recognition in the optimization process, and also facilitating further evaluation and adjustment of the optimization effects.

[0016] Through the organic combination of multiple links such as accurate scene type recognition, dynamic weight assignment, full consideration of the comprehensive impact among multiple visual elements, quantitative evaluation of emotional intensity values, generation of layout optimization suggestions, and application of image generation models, this application realizes the multi-dimensional improvement of scene layout optimization. This method not only improves the pertinence and effectiveness of the optimization results, but also enhances the flexibility and adaptability of the optimization process, greatly improving the scene layout optimization effect, and providing a more personalized and comfortable scene experience for users. Thus, it solves the technical defects in the existing technology that the static weight assignment mechanism of the scene layout optimization method is rigid, the comprehensive impact among multiple visual elements is not considered, and the dynamic quantification ability of emotional intensity is insufficient, resulting in poor scene layout optimization effects. Brief Description of the Drawings

[0017] The drawings described herein are used to provide a further understanding of the present disclosure, and constitute a part of this application. The illustrative embodiments of the present disclosure and their descriptions are used to explain the present disclosure and do not constitute an improper limitation to the present disclosure. In the drawings: Figure 1 It is a hardware structure block diagram of a computing device for implementing the method described in Embodiment 1 of the present application; Figure 2 It is a schematic flowchart of a method for optimizing scene layout based on multiple visual elements and emotion perception according to Embodiment 1 of the present application; Figure 3 It is a schematic diagram of a system for optimizing scene layout based on multiple visual elements and emotion perception according to Embodiment 2 of the present application; and Figure 4 It is a schematic diagram of a system for optimizing scene layout based on multiple visual elements and emotion perception according to Embodiment 3 of the present application. Detailed implementation manners

[0018] In order to enable those skilled in the art to better understand the technical solutions of the present disclosure, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present disclosure. Obviously, the described embodiments are only a part of the embodiments of the present disclosure, rather than all of the embodiments. Based on the embodiments in the present disclosure, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of the present disclosure.

[0019] It should be noted that the terms "first", "second", etc. in the specification and claims of the present disclosure and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily need to be used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present disclosure described herein can be implemented in an order different from those illustrated or described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device comprising a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0020] Embodiment 1 According to this embodiment, a method embodiment of a method for optimizing scene layout based on multiple visual elements and emotion perception is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that here.

[0021] The method embodiment provided in this embodiment can be executed on a mobile terminal, a computer terminal, a server or a similar computing device. Figure 1The hardware structural block diagram of a computing device for implementing a method for optimizing scene layout based on multiple visual elements and emotion perception is shown. As Figure 1 shown, the computing device may include one or more processors (the processor may include, but is not limited to, processing devices such as a microprocessor MCU or a field programmable gate array FPGA), a memory for storing data, a transmission device for communication functions, and an input / output interface. The memory, the transmission device, and the input / output interface are connected to the processor through a bus. In addition, it may further include: a display, a keyboard, and a cursor control device connected to the input / output interface. Those of ordinary skill in the art can understand that Figure 1 the structure shown is only schematic and does not limit the structure of the above electronic device. For example, the computing device may further include more or fewer components than Figure 1 shown, or have a different configuration from Figure 1 shown.

[0022] It should be noted that the above one or more processors and / or other data processing circuits are generally referred to as "data processing circuits" in this article. The data processing circuit may be embodied in software, hardware, firmware, or any combination thereof, in whole or in part. In addition, the data processing circuit may be a single independent processing module, or be incorporated in whole or in part into any one of other elements in the computing device. As involved in the embodiments of the present disclosure, the data processing circuit is a kind of processor control (such as the selection of a variable resistance terminal path connected to an interface).

[0023] The memory can be used to store software programs and modules of application software, such as the program instructions / data storage device corresponding to the method for optimizing scene layout based on multiple visual elements and emotion perception in the embodiments of the present disclosure. The processor executes various functional applications and data processing by running the software programs and modules stored in the memory, that is, implements the method for optimizing scene layout based on multiple visual elements and emotion perception of the above application program. The memory may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memories, or other non-volatile solid-state memories. In some instances, the memory may further include a memory remotely disposed relative to the processor, and these remote memories may be connected to the computing device through a network. Examples of the above network include, but are not limited to, the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.

[0024] The transmission device is used to receive or send data via a network. Specific examples of the above-mentioned network may include a wireless network provided by a communication provider of a computing device. In one example, the transmission device includes a network adapter (Network Interface Controller, NIC), which can be connected to other network devices through a base station to communicate with the Internet. In one example, the transmission device can be a Radio Frequency (RF) module, which is used to communicate with the Internet wirelessly.

[0025] The display can be, for example, a touch-screen liquid crystal display (LCD), which enables users to interact with the user interface of the computing device.

[0026] It should be noted here that in some alternative embodiments, the above Figure 1 illustrated computing device may include hardware elements (including circuits), software elements (including computer code stored on a computer-readable medium), or a combination of both hardware elements and software elements. It should be pointed out that Figure 1 is only an example of a specific specific instance and is intended to illustrate the types of components that may exist in the above-mentioned computing device.

[0027] Under the above operating environment, according to the first aspect of this embodiment, a method for optimizing scene layout based on multiple visual elements and emotional perception is provided. Figure 2 The flowchart of the method is shown. Referring to Figure 2 shown, the method includes: S102: Obtain the image data of the target scene and determine the scene type of the target scene; wherein, the image data includes multiple categories of visual elements; S104: Extract the feature information of the visual elements of each category in the image data, and determine the weight information of the visual elements of each category in the image data according to the scene type and the feedback information of the user on the layout of the target scene obtained in advance; S106: Determine the emotional intensity value of the target scene according to the feature information and the corresponding weight information of all categories of visual elements in the image data, and determine the layout optimization suggestion for the target scene according to the emotional intensity value; wherein the emotional intensity value is used to indicate the comprehensive impact of all categories of visual elements in the target scene on the user's emotional experience; and S108: According to the layout optimization suggestion, use an image generation model to generate the image data of the target scene after optimizing the layout.

[0028] Specifically, in this embodiment, the target scenarios are those that require layout optimization. The linguistic landscapes of these scenarios are composed of linguistic and non-linguistic symbols, presenting a symbol aggregate through single or multiple symbol forms such as text, images, graphics, colors, etc., forming a scene identifier. A reasonable scene layout can enable users in the target scenario to have a good sense of interaction and experience. This technical solution aims to optimize the scene layout by processing and analyzing the image data of the target scenario, providing users with a better sense of interaction and experience. The scene layout optimization is mainly reflected in optimizing the scene identifier of a certain scene, such as but not limited to optimizing the number, setting position, text content, graphic symbols, and color matching of the scene identifier.

[0029] Therefore, when a certain scene has a layout optimization requirement, it is determined as the target scene. Then, the image data of the target scene is collected through a camera, drone, or public map API interface, and the scene type of the target scene is determined (corresponding to step S102). Among them, the scene type is, for example but not limited to, medical scenes, street scenes, cultural and tourism scenes, traffic scenes, etc. The image data of the target scene includes multiple categories of visual elements, such as text elements, image elements, and spatial layout elements.

[0030] Considering that under different scene types, the influence degrees of various visual elements on scene layout optimization and user emotional experience are different, and the feedback information of users on the scene layout also reflects their attention and expectations for different visual elements. Therefore, it is necessary to extract the feature information of each category of visual elements in the image data, and determine the weight information of each category of visual elements in the image data according to the scene type and the pre-obtained feedback information of users on the target scene layout (corresponding to step S104).

[0031] Specifically, the feature information is extracted for each category of visual elements in the image data. Taking text elements as an example, features such as the font, font size, color, and semantic content of the text can be extracted. The font and font size affect the readability and visual attractiveness of the text, the color can convey the emotional atmosphere, and the semantic content is directly related to the information conveyed by the scene. For image elements, features such as graphic symbols and color matching in the image are extracted. The graphic symbols reflect the information conveyed by the scene, and the color matching creates a specific emotional tone. In terms of spatial layout elements, features such as the size, number, and setting position of the scene identifier in the space are extracted, and these features determine the practicality and comfort of the scene.

[0032] Next, based on the scene type and the feedback information of the user on the target scene layout obtained in advance, determine the weight information of each category of visual elements. For example, in a medical scene, usually the accuracy, acceptability, readability of indication signs and warning information, as well as the density of indication signs are emphasized. In this case, the weight of text elements and spatial layout elements is relatively high. If the user feedback emphasizes the requirement for the acceptability of the text content in medical guidance signs (for example, avoiding using "strictly prohibited" and suggesting using "please do not"), then the weight of text elements needs to be increased accordingly. In a cultural and tourism scene, usually the graphic attractiveness and layout aesthetics are emphasized. In this case, image elements and spatial layout elements are more crucial for attracting tourists and creating a good experience. Therefore, the weights of these two types of elements are relatively high. And if the user feedback emphasizes that the graphic symbols need to be novel and vivid, then the weight of image elements needs to be increased accordingly. By comprehensively considering the scene type and user feedback, the importance of each visual element in scene layout optimization and user emotional experience can be determined more accurately, providing a basis for subsequent emotional intensity value calculation and layout optimization suggestions.

[0033] After completing the extraction of visual element features and determination of weights, based on the feature information and corresponding weight information of all categories of visual elements in the image data, determine the emotional intensity value of the target scene, and determine the layout optimization suggestions for the target scene according to the emotional intensity value (corresponding to step S106).

[0034] Specifically, first, based on the feature information and corresponding weight information of all categories of visual elements in the image data, determine the emotional intensity value of the target scene. Among them, the method of weighted summation can be used. Multiply the feature value of each visual element by its corresponding weight, and then add all the results to obtain the emotional intensity value. For example, for text elements, if its feature value is the score of accuracy, acceptability, and readability, and the weight is the value determined according to the scene type and user feedback, multiplying the two gives the contribution of this text element to the emotional intensity value. Similarly, similar calculations are also performed on image elements and spatial layout elements. Finally, add the contributions of all elements to obtain the emotional intensity value of the target scene. Then, perform collaborative calculation on the contribution of text elements to the emotional intensity value and the contribution of image elements to the emotional intensity value. For example, the red warning text (the contribution value of its emotional intensity is +0.8) and the exclamation mark icon (the contribution value of its emotional intensity is +0.6) are collaboratively strengthened, and the total contribution value of emotional intensity = 0.8 + 0.6 = 1.4. Another example is that the blue guiding text (the contribution value of its emotional intensity is -0.5) and the upward arrow icon (the contribution value of its emotional intensity is +0.5) cancel each other out, and the total contribution value of emotional intensity = 0. For the contribution of layout elements to emotional intensity: if the layout density ≤ 5 pieces / ㎡, then the contribution value of its emotional intensity is 0; if the layout density > 5 pieces / ㎡, then for each additional 1 piece / ㎡, the contribution value of its emotional intensity is -0.2.

[0035] It should be noted that in this technical solution, the higher the emotional intensity value, the higher the resistance emotion brought by the layout of the target scene to the user, that is, the stronger the negative emotional experience brought to the user. On the contrary, the lower the emotional intensity value, the stronger the positive emotional experience brought by the target scene to the user.

[0036] Then, based on the emotional intensity value, optimization suggestions for the layout of the target scene are determined. If the emotional intensity value is high, it indicates that the combination of visual elements in the current scene fails to effectively meet the user's needs and improve the user experience. Therefore, corresponding optimization suggestions need to be put forward for different visual elements. For text elements, if their expression style has a negative emotional tendency (such as expressions like "strictly prohibited"), it can be considered to be adjusted to a more positive and gentle expression (for example, "please do not" can be further optimized to a more warm reminder expression, or the negative text can be replaced with positive guiding text); if the visual presentation forms such as font and font size affect the emotional experience, it can be adjusted to a style more in line with the positive emotional atmosphere. For image elements, if the theme, composition, or color scheme creates a negative atmosphere, the theme can be reselected, the composition method can be adjusted, or the color scheme can be adjusted to create a more positive and pleasant visual effect. For spatial layout elements, if the functional partition is unreasonable or the relative position of the elements is inappropriate, the spatial layout can be re-planned to improve the practicality and comfort of the space. Through these layout optimization suggestions, the aim is to enhance the positive emotional experience value brought by the target scene to the user and provide a better sense of interaction and experience for the user.

[0037] After calculating the emotional intensity value and determining the layout optimization suggestions, according to the layout optimization suggestions, image data of the target scene after the optimized layout is generated using an image generation model (corresponding to step S108). This step aims to convert the optimization suggestions into input instructions recognizable by generative AI (such as Stable Diffusion), and generate image data that meets the emotional intensity target with the help of an image generation model (such as a diffusion model), so that it can be presented to the user in a visual way later to intuitively display the optimized scene effect.

[0038] Among them, the adjustment rules of the image generation model for the image data according to the input instructions are for example but not limited to: for color control, when the input instruction is "reduce the red saturation to 30%", the image generation model adjusts the S value in the image HSV ≤ 30%. For symbol adaptation, when the input instruction is "add a guiding arrow", the image generation model calls the preset icon library (↑, ) and matches the scene style (such as flat / stereoscopic processing). For text optimization, when the input instruction is "replace with friendly terms", the image generation model automatically replaces the sensitive word library (such as replacing "strictly prohibited" with "please do not").

[0039] In addition, during the display process, the user can evaluate the generated image. Observe whether the image meets the emotional intensity target, whether the combination of text, graphics, and colors is coordinated, and whether the overall effect meets the expectations. If the user is not satisfied with the generated image, they can return to step S108 to trigger a new round of adjustment of the weights of each visual element and design optimization. Alternatively, by deploying the old and new design schemes (i.e., the scene layout scheme before optimization and the optimized scene layout scheme) in parallel in the real scenario, then collecting user behavior data (such as user stay time, violation rate) and subjective scores, and then returning to step S106, after updating the user feedback information, trigger a new round of adjustment of the weights of each visual element and design optimization. Through this closed-loop mechanism, it is ensured that the optimization process continuously adheres to the user's needs. Thus, the abstract optimization suggestions are transformed into specific image data, and the optimized scene effect is visually presented to the user through visualization, providing a basis for subsequent decision-making and improvement.

[0040] As described in the background art, the existing scene layout optimization methods have technical defects such as a rigid static weight allocation mechanism, insufficient consideration of the comprehensive influence among multiple visual elements, and insufficient dynamic quantification ability of emotional intensity, resulting in the problem of poor scene layout optimization effect.

[0041] In view of this, the technical solution of this application first obtains the image data of the target scene and determines the scene type, so as to perform customized processing according to the characteristics of different scenes (such as traffic, landscape, medical, etc.), providing data support for subsequent visual element feature extraction, weight assignment, and layout optimization, and ensuring that the layout optimization measures are more in line with the actual scene requirements. Then, it extracts the feature information of each category of visual elements (such as text elements, image elements, and layout elements) in the image data, and dynamically determines the weight information of each category of visual elements according to the scene type and the feedback information of the user on the target scene layout. It can flexibly adjust the importance of each visual element in the scene layout according to the changes of different scenes and user preferences, so as to achieve more personalized scene layout optimization. Secondly, by comprehensively considering the feature information of all categories of visual elements and their corresponding weight information, it calculates the emotional intensity value of the target scene, which can intuitively reflect the comprehensive impact of all visual elements in the target scene on the user's emotional experience. Thus, through the feature analysis, mutual relationship research, and comprehensive weight consideration of different visual elements, it ensures that each visual element plays the best role in the scene layout, providing an objective and measurable basis for layout optimization. After that, according to the emotional intensity value, it generates specific layout optimization suggestions, which can guide the adjustment of the scene layout, making the optimization process more directional and operable. Finally, according to the layout optimization suggestions, it uses an image generation model to generate the image data of the target scene after layout optimization, so as to transform the optimization suggestions into actual visual presentations, enabling users to intuitively see the optimized scene effects, providing data support for subsequent visual and interactive optimization result displays, enhancing the user's sense of participation and recognition in the optimization process, and at the same time facilitating further evaluation and adjustment of the optimization effects.

[0042] This application realizes the multi-dimensional improvement of scene layout optimization through the organic combination of multiple links such as accurate scene type recognition, dynamic weight assignment, full consideration of the comprehensive influence among multiple visual elements, emotional intensity value quantitative evaluation, layout optimization suggestion generation, and image generation model application. This method not only improves the pertinence and effectiveness of the optimization results, but also realizes dynamic adaptability through a feedback loop, greatly enhancing the scene layout optimization effect and providing users with a more personalized and comfortable scene experience. Thus, it solves the technical defects in the existing technology that the scene layout optimization method has a rigid static weight assignment mechanism, does not consider the comprehensive influence among multiple visual elements, and has insufficient emotional intensity dynamic quantification ability, resulting in poor scene layout optimization effects.

[0043] Optionally, the multiple categories of visual elements include text elements, image elements, and layout elements; and the operation of extracting the feature information of each category of visual elements in the image data includes: extracting the text content in the image data through optical character recognition technology, calculating the text sentiment contribution value corresponding to the text content through a pre-trained language model, and using it as the feature information of the text elements in the image data; identifying graphic symbols in the image data based on an object detection algorithm, determining the semantic type of the graphic symbols, extracting the HSV value of the main color tone of the image data, and determining the image sentiment contribution value of the image data according to the semantic type and the HSV value, and using it as the feature information of the image elements in the image data; and based on the image data, counting the number of identifiers per unit area in the target scene, determining the visual focus of the image data through a saliency detection model, and determining the layout sentiment contribution value of the image data according to the number of identifiers and the visual focus, and using it as the feature information of the layout elements in the image data.

[0044] Specifically, the text content in the image data is extracted through optical character recognition (OCR) technology. Among them, OCR technology can convert the text in the image into an editable text format, providing basic data for subsequent analysis. Then, a pre-trained language model (such as the ERNIE-3.0 model) is used to calculate the text sentiment contribution value corresponding to the text content. Among them, the pre-trained language model has been pre-trained with a large amount of text data and can understand the semantics and sentiment tendency of the text. By calculating the text sentiment contribution value through this model, the text content can be converted into a numerical representation with emotional meaning, thereby quantifying the influence degree of the text elements on the overall emotional expression of the image. For example, the text content extracted from the image data is "strictly prohibited", and the text sentiment contribution value corresponding to "strictly prohibited" calculated using the pre-trained language model is 4.5. The text content extracted from the image data is "please do not", and the text sentiment contribution value corresponding to "please do not" calculated using the pre-trained language model is 3.2. After that, this text sentiment contribution value is used as the feature information of the text elements in the image data.

[0045] Then, graphic symbols in the image data are identified based on an object detection algorithm (such as the YOLOv5 model), and their semantic types are determined. The object detection algorithm can locate specific objects in the image and identify the graphic symbols therein. After that, the graphic symbols are classified to determine their semantic types. Among them, the semantic types include, for example but not limited to, types such as prohibition, guidance, and warning. Determining the semantic type of the graphic symbols helps to understand the meaning represented by the graphic symbols and provides a semantic basis for subsequent analysis.

[0046] Next, extract the HSV values (hue, saturation, value) of the dominant color of the image data. Among them, the dominant colors in different types of scenes have their own characteristics. For example, in traffic scenes, danger warning signs often use high-contrast warm colors (red, orange) to enhance the warning effect, while traffic indication signs often use cool colors (blue, green) to provide clear guidance. Another example is that in medical scenes, emergency rescue signs often use bright warm colors (red, yellow) to highlight the emergency characteristics, and medical area reminder signs often use soft colors (light blue, light green) to create a comfortable atmosphere. Since the HSV color space is more in line with the human perception of colors, by extracting the HSV values of the dominant color, the main color characteristics of the image can be accurately described, providing key color information for calculating the image emotional contribution value. After that, determine the image emotional contribution value of the image data according to the semantic type and HSV values. By combining the semantic type of the graphic symbol and the HSV value of the dominant color of the image, comprehensively considering the semantic and color characteristics of the image elements, the image emotional contribution value is calculated. For example, the hue H in the HSV value of the dominant color in the image data is 0°, indicating that the dominant color is red, the saturation S is 80%, and the value V is 60%. In this case, by combining the semantic type of the graphic symbol and the HSV value of the dominant color of the image, the calculated image emotional contribution value is 0.8, which can quantify the influence degree of the image elements on the overall emotional expression of the image, making the characteristic information of the image elements more emotionally representative. Thus, the image emotional contribution value can be used as the characteristic information of the image elements in the image data.

[0047] Then, based on the image data, count the number of signs in the target scene per unit area (per square meter). The number of signs is one of the important indicators reflecting the layout characteristics of the image. By counting the number of signs per unit area, the density of signs in the target scene can be understood, thereby inferring the layout style and information density of the image. Determine the visual focus of the image data through a saliency detection model (such as the DeepGaze II model). Among them, the saliency detection model can simulate the human visual attention mechanism and find the area that attracts the most attention in the image, that is, the visual focus. The visual focus is a key element in the image layout, which affects the viewer's attention points and browsing order of the image. After that, determine the layout emotional contribution value of the image data according to the number of signs and the visual focus. By comprehensively considering the characteristics of these two layout elements, the number of signs and the visual focus, the layout emotional contribution value is calculated. This value can quantify the influence degree of the layout elements on the overall emotional expression of the image, making the characteristic information of the layout elements more emotionally representative. Thus, the layout emotional contribution value can be used as the characteristic information of the layout elements in the image data.

[0048] Thus, by separately extracting the feature information of text elements, image elements, and layout elements, rich feature data is provided for tasks such as sentiment analysis and content understanding of image data. These feature information can effectively represent the characteristics of the image and the emotional semantics conveyed by the scene layout from multiple dimensions, laying a key data foundation for subsequent in-depth processing work such as image scene parsing, layout reconstruction, and emotion calculation.

[0049] In addition, before extracting the feature information of visual elements of each category in the image data, preprocessing of the image data is also included. For example, operations such as standardizing the image resolution (such as 1920×1080) and removing noise (such as reflective and blurred areas).

[0050] Optionally, the operation of determining the image emotion contribution value of the image data according to the semantic type and the HSV value includes: determining the symbol emotion contribution value corresponding to the graphic symbol according to the semantic type; calculating the color emotion contribution value of the image data according to the HSV value; and determining the image emotion contribution value of the image data according to the symbol emotion contribution value and the color emotion contribution value.

[0051] Specifically, different graphical symbols have different semantic connotations, which are associated with human emotional cognition. For example, in traffic scenarios, the "No Entry" sign with a white horizontal bar in the middle of a red circle is usually associated with negative emotions such as danger and prohibition; while the green arrow-shaped guiding sign is associated with positive emotions such as guidance and safety. By analyzing and classifying the semantics of a large number of traffic-scenario graphical symbols and establishing a mapping relationship between semantics and emotions, the corresponding symbolic emotional contribution value can be determined according to the semantic type of the graphical symbol. This symbolic emotional contribution value reflects the degree of contribution of the emotional information contained in the graphical symbol itself to the overall emotional expression of the image. By determining the symbolic emotional contribution value, the semantic emotions of the graphical symbol can be quantified, providing basic data for subsequent calculation of the image emotional contribution value. For example, by collecting various common traffic-scenario graphical symbols and inviting traffic-domain professionals or through large-scale questionnaire surveys, determine the semantics represented by each graphical symbol, as well as the corresponding emotional tendency (positive, negative or neutral) and emotional intensity. Organize this information into a semantic-emotion dictionary and assign an initial symbolic emotional contribution value to each graphical symbol. In practical applications, the semantics of traffic-scenario graphical symbols may be affected by other elements in the image or the overall context. Therefore, it is necessary to dynamically adjust the symbolic emotional contribution value according to the specific situation of the image and in combination with the context information. For example, in an image of a road construction scenario, even if the "No Entry" sign appears (which may originally be associated with danger), but considering the overall scene with construction fences, warning lights and construction workers busy working around, this sign may be given a special, non-simply negative emotional meaning. It also reminds people to pay attention to construction safety and, from a certain perspective, has a positive guiding role in guiding people to detour safely.

[0052] HSV values (hue, saturation, value) are important parameters for describing colors, and different colors have different emotional characteristics. For example, in traffic scenarios, red is usually associated with emotions such as warning and danger. For instance, the red light in traffic lights reminds drivers and pedestrians to stop moving forward through strong visual stimuli; blue is associated with emotions such as calmness and reliability. For example, some traffic signs use a blue background to convey professional and stable information. Saturation and value also affect the emotional expression of colors. High-saturation colors are usually more vivid and eye-catching. For example, high-saturation yellow warning signs in traffic scenarios can cause stronger emotional reactions in people and quickly attract attention; high-value colors often give people a feeling of brightness and openness. For example, bright white road markings in traffic scenarios can clearly guide the way during the day; while low-value colors may create a heavy and serious atmosphere. For example, dark gray road dividers in traffic scenarios give people an impression of being steady and not to be crossed casually. That is to say, the color emotional contribution value reflects the contribution degree of color information in the image to the overall emotional expression of the image. By calculating the color emotional contribution value, the color characteristics of the image can be quantified, providing another important dimension of data for subsequent calculation of the image emotional contribution value.

[0053] In addition, before calculating the image emotional contribution value, an emotional test and analysis can be carried out on the application of a large number of colors with different HSV values in traffic scenarios to establish a color emotion model. This model can predict the corresponding emotional tendency and emotional intensity according to the HSV values of colors and calculate the color emotional contribution value. For example, a function can be defined that takes the HSV values of a color as input and outputs the corresponding color emotional contribution value. Specifically, for the red color in traffic lights (assuming its HSV values are within a certain specific range), according to the model analysis, its emotional tendency of warning and danger is relatively strong, and a relatively high color emotional contribution value can be assigned (such as +0.8, a positive value indicates a negative emotion); while for the blue color in traffic signs (also setting a specific HSV value range), because it conveys calm and reliable information, a negative color emotional contribution value can be assigned (such as -0.3, a negative value indicates a positive emotion), so as to quantify the emotional characteristics of colors in traffic scenarios and help calculate the image emotional contribution value more accurately.

[0054] Considering that the emotional expression of an image is the result of the combined action of multiple elements in the image, where graphic symbols and colors are two important elements. The symbolic emotional contribution value and the color emotional contribution value respectively reflect the contribution degrees of graphic symbols and colors to the overall emotional expression of the image. Therefore, the image emotional contribution value of the image data can be determined by comprehensively considering these two factors. That is to say, the image emotional contribution value is a comprehensive quantitative index for the overall emotional expression of the image, and it can be used in applications such as image emotional analysis, classification, and retrieval. By calculating the image emotional contribution value, the emotional characteristics of the image can be evaluated more objectively, providing richer information for image processing and understanding. Among them, the symbolic emotional contribution value and the color emotional contribution value are weighted and summed to obtain the image emotional contribution value. For example, two weight coefficients can be defined to represent the importance of the symbolic emotional contribution value and the color emotional contribution value in the image emotional expression respectively, and then the two contribution values are multiplied by the corresponding weight coefficients and added to obtain the image emotional contribution value. The weight coefficients can be adjusted according to specific application scenarios and requirements.

[0055] Alternatively, machine learning algorithms can be used to determine the image emotional contribution value. For example, a large amount of image data can be collected and the emotional labels of each image are annotated, and then machine learning algorithms such as support vector machines (SVM) and neural networks are used to train the image data to establish the mapping relationship between the symbolic emotional contribution value, the color emotional contribution value and the image emotional label. In practical applications, the symbolic emotional contribution value and the color emotional contribution value of a new image are input into the trained model, and the image emotional contribution value of the image can be predicted.

[0056] Thus, by comprehensively considering the semantic emotional contribution value of graphic symbols and the color emotional contribution value, the accurate quantification of the image emotional contribution value of image data in a specific scenario is achieved.

[0057] Optionally, the operation of determining the layout emotional contribution value of the image data according to the number of identifiers and the visual focus includes: determining the basic density score of each sub-region of the target scenario according to the number of identifiers; where the target scenario is divided into multiple sub-regions by unit area; determining the identifier significance weight of each sub-region according to the visual focus; and determining the layout emotional contribution value of each sub-region according to the basic density score and the identifier significance weight, and taking the sum of the layout emotional contribution values of each sub-region as the layout emotional contribution value of the image data.

[0058] Specifically, taking the traffic scene as an example of the target scene, in the traffic scene, the number of signs reflects the density of traffic information in a specific area. For example, at the crossroads in the bustling city center, due to the large traffic flow and complex road conditions, a large number of traffic signs will be set up, such as traffic lights, traffic signs (including direction signs, speed limit information, etc.), zebra crossing signs, etc. Divide the target scene (such as the entire crossroads area) into multiple sub-areas according to the unit area, and the number of signs in each sub-area can reflect the distribution density of traffic information in this area. By counting the number of signs in each sub-area and determining the basic density score accordingly, to quantify the differences in traffic information distribution in different sub-areas, it helps to identify areas with dense and sparse traffic information, laying a foundation for further analyzing the emotional impact of signs in the image layout.

[0059] The visual focus is the key area in the image that attracts the user's attention, and there are differences in the visual focus of different types of scene images. In traffic scene images, the visual focus will vary according to the specific situation. For example, when the vehicle is waiting for the red light at the intersection, the visual focus of the driver and pedestrians usually focuses on the traffic light, and at this time, the traffic light sign becomes the key visual element of the scene image; during driving, the driver may pay more attention to the traffic sign in front, and the traffic sign becomes the visual focus sign at this time. In medical scene images, the visual focus also changes according to the situation. For example, in the image showing the hospital registration hall, the indication signs of the registration window, the call display screen and other signs will become the visual focus because they are the key information for guiding patients to register and seek medical treatment; while in the image presenting the hospital ward scene, the bedside signs (marking patient information, nursing level, etc.), the emergency call button signs, etc. will attract people's attention and become the visual focus signs.

[0060] Determining the sign saliency weight of each sub-area according to the visual focus means considering the importance and attractiveness of the sign in the image layout. In the traffic scene, the signs near the visual focus, such as the signs near the traffic light that the driver and pedestrians focus on when waiting for the red light, will have a relatively high saliency weight; in the medical scene, the signs near the indication sign of the registration window in the registration hall or near the bedside sign in the ward will also have a high saliency weight. This makes the calculation of the layout emotional contribution value more in line with the laws of human visual perception and attention distribution, and can more accurately reflect the impact of the sign on the emotion in the image layout.

[0061] Next, by comprehensively considering the basic density score and the logo salience weight, calculate the layout sentiment contribution value of each sub-region. The basic density score reflects the distribution of logos, while the logo salience weight represents the visual importance of the logos. Combining the two can more comprehensively evaluate the impact of logos on sentiment in the image layout. For example, if a sub-region has a large number of logos (high basic density score) and is near the visual focus (high logo salience weight), then the layout sentiment contribution value of this sub-region will be relatively high. Finally, add up the layout sentiment contribution values of each sub-region to obtain the layout sentiment contribution value of the entire image data, realizing a quantitative evaluation of the emotional impact of the image layout.

[0062] Optionally, the operation of determining the weight information of visual elements of each category in the image data according to the scene type and the pre-obtained feedback information of the user on the target scene layout includes: determining the initial weight information of visual elements of each category in the image data according to a preset rule library and the scene type; wherein, the preset rule library stores the initial weight information of visual elements of each category in the image data of different scene types; and optimizing the initial weight information of visual elements of each category in the image data by using a reinforcement learning mechanism according to the feedback information to obtain the weight information of visual elements of each category in the image data; wherein, the reward signal of the reinforcement learning mechanism is obtained based on the feedback information.

[0063] Specifically, the preset rule library is a database that stores a large amount of initial weight information of various visual elements in the image data of different scene types. Since different scene types (such as traffic scenes, medical scenes, natural scenery scenes, etc.) have different characteristics and requirements, the importance of visual elements in these scenes also varies. The preset rule library assigns corresponding initial weights to visual elements of each category for different scene types, enabling a reasonable initial weight allocation scheme to be quickly obtained when processing new image data, providing a starting point for subsequent weight optimization, and making the entire weight determination process more orderly and efficient.

[0064] Therefore, according to the input scene type, the corresponding rules can be searched in the preset rule library, so as to obtain the initial weight information of various types of elements (including text elements, image elements, and layout elements) in the image data under this scene type. Taking the traffic scene as an example, the initial weight information α, β, γ corresponding to the text element, image element, and layout element stored in the preset rule library are 0.6, 0.3, and 0.1 respectively. Taking the medical scene as an example, the initial weight information α, β, γ corresponding to the text element, image element, and layout element stored in the preset rule library are 0.5, 0.3, and 0.2 respectively. Taking the cultural and tourism scene as an example, the initial weight information α, β, γ corresponding to the text element, image element, and layout element stored in the preset rule library are 0.5, 0.2, and 0.3 respectively.

[0065] Then, considering that the feedback information reflects the actual feelings and needs of users for the image layout and is an important basis for optimizing the weights of visual elements. By analyzing the feedback information, it is possible to understand the degree of attention, preferences, and dissatisfaction of users with different visual elements, so as to adjust the initial weights targeted. This technical solution uses a reinforcement learning mechanism to automatically adjust the weights of each visual element according to the feedback information. Reinforcement learning is a machine learning method that enables an agent to interact with the environment and learn the optimal behavior strategy based on the feedback (reward or punishment) of the environment. In the process of dynamically adjusting the weights, the agent can be regarded as an algorithm for adjusting the weights of visual elements, and the environment is the set of image data and user feedback information. The reinforcement learning mechanism can automatically adjust the weights of visual elements according to the feedback information, making the weight allocation more in line with the actual needs of users. Compared with the traditional rule-based weight adjustment method, reinforcement learning has stronger adaptability and learning ability, and can gradually optimize the weight allocation scheme in the continuous learning process.

[0066] In this embodiment, the state space of the reinforcement learning strategy is: the current weight parameters (α, β, γ), the scene type, and the historical effect data (i.e., user feedback information). The action space of the reinforcement learning strategy is the incremental adjustment of the weight parameters (such as Δα = +0.1, Δβ = +0.1, Δγ = -0.1). The reward function is jointly composed of user satisfaction (weight 70%) and the decline rate of the violation rate (weight 30%).

[0067] The reward signal is the core driving force of the reinforcement learning mechanism. It guides the agent to continuously explore and optimize the weight allocation scheme to obtain the maximum reward. By reasonably designing the calculation method of the reward signal, the reinforcement learning mechanism can pay more attention to the key needs of users and improve the effect of weight optimization. The reward signal of this technical solution is obtained based on feedback information and is used to guide the reinforcement learning mechanism to adjust the weights of visual elements. Among them, the reward signal can be positive or negative. A positive reward signal indicates that the current weight allocation scheme has been recognized by the user and should be maintained or strengthened; a negative reward signal indicates that the current weight allocation scheme does not meet the user's needs and needs to be adjusted.

[0068] Thus, the reinforcement learning mechanism can continuously adjust the weights of visual elements according to the reward signal. After multiple iterations and learning, it can finally obtain a weight information that can better meet the user's needs. Taking user satisfaction (rating from 1 to 5) and scene effect (such as the decline rate of violation rate) as the reward signal, the weight parameters (α, β, γ) are dynamically optimized through the policy gradient algorithm (such as PPO). The specific adjustment logic is as follows: Taking the traffic scene as an example, in the traffic scene, assume that the visual elements include the text warning of traffic signs (weight parameter is α), graphic guidance (weight parameter is β), and sign layout display (weight parameter is γ). If the feedback information shows that the text warning intensity of the traffic sign is too high, causing the user (such as a driver) to have a resistant emotion and the user satisfaction rating to decrease, and at the same time, some potential violation behaviors may be caused due to the driver's distracted attention (the violation rate does not reach the expected decline rate), the reinforcement learning mechanism will then decrease α and increase β. Reducing the text warning intensity can reduce user resistance, and increasing the weight of graphic guidance can more intuitively guide the driver's behavior, thus optimizing the visual display effect of the traffic scene and improving the user's (driver's) experience.

[0069] Thus, the weight information optimized by the reinforcement learning mechanism is more in line with the actual needs of users, which can improve the visual effect of image data and the user experience. At the same time, this method of determining weights also has a certain degree of generality and scalability and can be applied to different types of scene image data.

[0070] Optionally, the operation of determining the emotional intensity value of the target scene according to the feature information and corresponding weight information of all categories of visual elements in the image data includes: according to the feature information and corresponding weight information of all categories of visual elements in the image data, calculating the emotional intensity value of the target scene through the following formula:

[0071] where S is the emotional intensity value of the target scene, is the feature information of the text element in the image data, is the weight information of the text elements in the image data, is the feature information of the image elements in the image data, is the weight information of the image elements in the image data, is the feature information of the layout elements in the image data, is the weight information of the layout elements in the image data.

[0072] Optionally, according to the layout optimization suggestion, the operation of generating the image data of the target scene with the optimized layout by using the image generation model includes: generating a prompt according to the layout optimization suggestion; and inputting the image data and the prompt into a pre-trained image generation model to output the image data of the target scene with the optimized layout.

[0073] Specifically, according to the input requirements of generative AI (such as Stable Diffusion), the optimization suggestion can be converted into a specific prompt. The prompt is used to guide the model to adjust the corresponding visual elements in the image data according to the optimization suggestion according to specific requirements, such as background color, element shape, icon symbol style, text content and setting position, etc. For example, for the background color, it can be specified as "blue background"; for the element shape, it can be required as "rounded border"; for the style of the icon symbol, it can be described as "simple arrow icon", etc.

[0074] Then, the converted prompt and the initially obtained image data (such as the image features and layout information of the original scene) are used as inputs and input into an image generation model (such as a diffusion model). After receiving the inputs, the image generation model will use its internal learning algorithms and parameters to process and analyze the original image data and the prompt. Among them, the image data provides the basic information for the model, and the prompt guides the model to generate images according to specific requirements. The model generates new image data through a series of calculations and inferences. This new image data is the image of the target scene with the optimized layout, and the model returns it as the output result.

[0075] The following will take a certain city park as an example of the target scene to illustrate each step of the technical solution: There are many problems with the layout and design of the guide signs in this city park. For example, from the textual point of view, expressions such as "Tourists Stop" have a high emotional intensity (the emotional intensity value is 4.5, which is a high-resistance emotional trigger word), which can easily make tourists feel disgusted. In terms of the image, the red prohibition symbol has a high intensity (the emotional intensity value is 0.9), giving people a strong sense of prohibition, and the layout of the signs is too dense (6 / ㎡), resulting in visual confusion in the space and a layout penalty (-0.3). These problems combined make the probability of tourists mistakenly entering the ecological protection area reach 18%. At the same time, the tourists' satisfaction score for the park's guide signs is only 3.5 points, with a full score of 7.0 points, reflecting that the current guide sign system fails to effectively guide tourists and brings tourists a bad visual and psychological experience. Therefore, it is necessary to optimize the layout of the guide signs in this target scene. The optimization steps are as follows: Step 1: Image data acquisition The image data of the city park is collected through cameras, drones or public map API interfaces, and the scene type of the city park is determined to be a cultural and tourism scene.

[0076] Step 2: Feature extraction and analysis (1) Text feature extraction and analysis: The text content of the guide signs in the image data is extracted through optical character recognition (OCR) technology, and then the ERNIE-3.0 natural language processing model is used to analyze the text in the guide signs, and the text sentiment contribution value of the text such as "Tourists stop" is 4.5. Through this model, we can deeply understand the emotional tendency and influence conveyed by the text, providing basic data for subsequent optimization.

[0077] (2) Image feature extraction and analysis: The YOLOv5 target detection algorithm combined with the color analysis method was used to process the image elements in the guide signs. YOLOv5 can quickly and accurately detect image elements such as the red prohibition sign, while the color analysis further evaluates the visual impact of the image elements and finds that the image emotion contribution value of the red prohibition sign is 0.9.

[0078] (3) Layout feature extraction and analysis: By statistically analyzing the actual distribution of guide signs in the park, the density of signs is 6 / ㎡. According to the preset rules, when the sign density exceeds a certain threshold, a layout penalty will be generated. The layout penalty calculated here is -0.3, that is, the layout emotional contribution value is -0.3.

[0079] Step 3: Emotional computing and optimized decision making (1)Initial weight setting: By calling the preset rule library, the initial weight of the text in the cultural and tourism scenario can be found as α = 0.5, the initial weight of the image as β = 0.2, and the initial weight of the layout as γ = 0.3. This weight allocation reflects that in the cultural and tourism scenario, image elements usually have a more significant visual impact on tourists, so a higher weight is given; text and layout elements also have certain weights respectively to ensure that all factors are considered comprehensively.

[0080] (2)User feedback collection and initial weight optimization: Set up multiple feedback points in the park, such as the tourist service center, main scenic spot entrances, etc., and collect tourists' opinions and suggestions on the guiding signs through paper questionnaires, electronic questionnaires, and on-site interviews. For example, ask tourists about their views on whether the text content of the existing guiding signs is easy to understand, whether the images are clear, and whether the layout is reasonable, and let tourists score the importance of different aspects. Organize and analyze the collected user feedback data, and use data analysis tools (such as Excel, SPSS, etc.) to count the distribution of opinions in different aspects, and find out the problems and needs that tourists generally pay attention to. For example, if it is found that most tourists think the importance of text content is higher than that of images, and many opinions are also put forward on the rationality of the layout, then the initial weights need to be adjusted. Then, according to the analysis results of the feedback data, optimize and adjust the initial weights. Suppose after analysis, it is found that the text weight should be appropriately increased, and the image weight and layout weight also need to be fine-tuned according to tourists' feedback. After optimization, the new weight settings are α = 0.3, β = 0.5, γ = 0.2. Through optimization, the weight allocation is made more in line with the actual needs and concerns of tourists.

[0081] (3)Comprehensive intensity calculation: According to the optimized weights and the intensity of each feature above, calculate the comprehensive intensity as 0.3×4.5 + 0.5×0.9 + 0.2×(-0.3) = 1.35 + 0.45 - 0.06 = 1.74. Since the comprehensive intensity is too high (it needs to be reduced to below 1.0), it indicates that the current guiding sign design has a greater negative impact on tourists and needs to be optimized and adjusted.

[0082] Step Four: Generation of optimization suggestions Based on the previous data analysis and comprehensive intensity calculation results, determine the optimization suggestions. For example, the optimization suggestions include replacing the text with "Ecological protection area, thank you for your cooperation" (intensity 2.8), which is a milder and friendlier expression and can reduce tourists' resistance; using a green background (intensity -0.6), as green usually gives people a natural and comfortable feeling and helps to relieve tourists' nervousness; at the same time, reducing the sign density to 4 per square meter to make the sign layout more reasonable and avoid visual chaos.

[0083] Step Five: Image data generation According to the optimization suggestions, corresponding prompting words are generated. The prompting words and the original image data are input into the diffusion model to generate new logo images with green shading, leaf icons, and a scattered layout, and the logo density of the new logo is 4 per square meter. Then, it is presented to the user in a visual way. Among them, the green shading blends with the natural environment of the park, the leaf icons further strengthen the theme of ecological protection, the scattered layout makes the logos clearer and easier to read, and improves the visual effect and guiding function of the navigation logos.

[0084] Step Six: Effect Verification In actual applications, the effects of the optimized navigation logos are verified. After a period of observation and statistics, it is found that the probability of tourists straying into the ecological protection area has dropped from the original 18% to 6%. This indicates that the optimized navigation logos can guide tourists more effectively and reduce the risk of tourists straying into restricted areas.

[0085] At the same time, the satisfaction score of tourists for the park navigation logos has increased from 3.5 / 7.0 to 6.2 / 7.0. This shows that tourists have given higher evaluations to the optimized navigation logos in terms of visual effects, information transmission, and guiding functions, further proving the effectiveness of this technical solution.

[0086] Through the layout optimization of the navigation logos in a certain urban park, data collection and feature extraction technologies are used to accurately analyze the problems existing in the navigation logos; a dynamic weight allocation method is adopted, and the initial weights are optimized according to user feedback, comprehensively considering the impacts of factors such as text, images, and layout on tourists; generative AI visualization technology is used to put forward reasonable optimization suggestions and generate new navigation logos; finally, through effect verification, it is proved that the optimized navigation logos can significantly reduce the rate of tourists straying and improve tourist satisfaction. Thus, it provides a scientific and effective method for the optimization of urban park navigation logos, which helps to improve the management level and service quality of the park.

[0087] In addition, according to this embodiment, a storage medium is also provided. The storage medium includes a stored program, wherein, when the program runs, the method of any one of the above is executed by a processor.

[0088] The beneficial effects of this application include: (1) Break through the limitations of the traditional fixed-weight model, and propose a dynamic weight adjustment method based on scene types (covering different fields such as transportation and culture and tourism) and real-time feedback (including key indicators such as user ratings and violation rates). In the process of dynamic weight adjustment, a reinforcement learning algorithm (such as the PPO algorithm) is introduced to optimize the weights (α, β, γ) of multiple visual elements, so as to achieve the precise adaptation of the emotion strategy, enabling the model to flexibly adjust according to different scenes and real-time feedback, and improving the accuracy and adaptability of decision-making.

[0089] (2)Comprehensively consider various visual elements such as text, images, and layout, deeply explore the interaction effects among them, including synergistic effects and conflict relationships, and conduct quantitative analysis. Through this integration and quantification, the impact of each element on the overall effect can be grasped more comprehensively and accurately, providing a scientific basis for subsequent optimization.

[0090] (3)Combine generative AI technology with reinforcement learning to create an automated design closed-loop of "analysis → optimization → generation → feedback → iteration". This closed-loop can achieve automated monitoring and continuous optimization of the design process, improve design efficiency and quality, reduce manual intervention, and lower design costs.

[0091] (4)Carefully divide the emotional intensity, for example, divide it into numerical values between 1 - 5, and make precise matches according to the requirements of different scenarios. For example, in a hospital scenario, the combination of "medium warning + soothing color" can not only effectively convey necessary information but also relieve the patient's nervousness and enhance the user experience.

[0092] It should be noted that for the foregoing method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that the present invention is not limited by the described action sequence, because according to the present invention, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to the present invention.

[0093] Embodiment 2 Figure 3 The structural schematic diagram of the scenario layout optimization system based on multi-visual elements and emotion perception according to this embodiment is shown. The scenario layout optimization system includes: a data acquisition module 310, configured to acquire image data of a target scenario and determine the scenario type of the target scenario; wherein the image data includes multiple categories of visual elements; an extraction and parsing module 320, configured to extract the feature information of each category of visual elements in the image data and determine the weight information of each category of visual elements in the image data according to the scenario type and the feedback information of the user on the layout of the target scenario obtained in advance; an emotion calculation and optimization decision module 330, configured to determine the emotional intensity value of the target scenario according to the feature information and the corresponding weight information of all categories of visual elements in the image data, and determine the layout optimization suggestion for the target scenario according to the emotional intensity value; wherein the emotional intensity value is used to indicate the comprehensive impact of all categories of visual elements in the target scenario on the user's emotional experience; and a layout optimization module 340, configured to generate the image data of the target scenario after optimized layout by using an image generation model according to the layout optimization suggestion.

[0094] Optionally, the visual elements of the multiple categories include text elements, image elements, and layout elements; and the extraction and parsing module 320 is specifically configured to: extract the text content in the image data through optical character recognition technology, calculate the text sentiment contribution value corresponding to the text content through a pre-trained language model, and use it as the feature information of the text elements in the image data; identify the graphic symbols in the image data based on the object detection algorithm, determine the semantic type of the graphic symbols, extract the HSV value of the main color tone of the image data, and determine the image sentiment contribution value of the image data according to the semantic type and the HSV value, and use it as the feature information of the image elements in the image data; and based on the image data, count the number of identifiers per unit area in the target scene, determine the visual focus of the image data through a saliency detection model, and determine the layout sentiment contribution value of the image data according to the number of identifiers and the visual focus, and use it as the feature information of the layout elements in the image data.

[0095] Optionally, the operation of determining the image sentiment contribution value of the image data according to the semantic type and the HSV value includes: determining the symbol sentiment contribution value corresponding to the graphic symbol according to the semantic type; calculating the color sentiment contribution value of the image data according to the HSV value; and determining the image sentiment contribution value of the image data according to the symbol sentiment contribution value and the color sentiment contribution value.

[0096] Optionally, the operation of determining the layout sentiment contribution value of the image data according to the number of identifiers and the visual focus includes: determining the basic density score of each sub-region of the target scene according to the number of identifiers; where the target scene is divided into multiple sub-regions per unit area; determining the identifier saliency weight of each sub-region according to the visual focus; and determining the layout sentiment contribution value of each sub-region according to the basic density score and the identifier saliency weight, and using the sum of the layout sentiment contribution values of each sub-region as the layout sentiment contribution value of the image data.

[0097] Optionally, the extraction and parsing module 320 is specifically configured to: determine the initial weight information of the visual elements of each category in the image data according to a preset rule library and the scene type; where the preset rule library stores the initial weight information of the visual elements of each category in the image data of different scene types; and optimize the initial weight information of the visual elements of each category in the image data by adopting a reinforcement learning mechanism according to the feedback information to obtain the weight information of the visual elements of each category in the image data; where the reward signal of the reinforcement learning mechanism is obtained based on the feedback information.

[0098] Optionally, the emotion calculation and optimization decision module 330 is specifically configured to: calculate the emotion intensity value of the target scene according to the feature information and corresponding weight information of all categories of visual elements in the image data through the following formula:

[0099] where S is the emotion intensity value of the target scene, is the feature information of the text element in the image data, is the weight information of the text element in the image data, is the feature information of the image element in the image data, is the weight information of the image element in the image data, is the feature information of the layout element in the image data, is the weight information of the layout element in the image data.

[0100] Optionally, the layout optimization module 340 is specifically configured to: generate a prompt word according to the layout optimization suggestion; and input the image data and the prompt word into a pre-trained image generation model to output the image data of the target scene with an optimized layout.

[0101] Therefore, according to this embodiment, through the organic combination of multiple links such as accurate scene type recognition, dynamic weight allocation, full consideration of the comprehensive influence between multiple visual elements, quantitative evaluation of emotion intensity value, generation of layout optimization suggestions, and application of image generation models, a multi-dimensional improvement of scene layout optimization is achieved. This method not only improves the pertinence and effectiveness of the optimization result, but also enhances the flexibility and adaptability of the optimization process, greatly improving the scene layout optimization effect and providing users with a more personalized and comfortable scene experience. Thus, it solves the technical problems in the prior art that the scene layout optimization method has a rigid static weight allocation mechanism, does not consider the comprehensive influence between multiple visual elements, and has insufficient dynamic quantification ability of emotion intensity, resulting in poor scene layout optimization effect.

[0102] Embodiment 3 Figure 4The scene layout optimization system based on multi-visual elements and emotion perception according to the present embodiment is shown, including: a processor 410; and a memory 420, connected to the processor 410, for providing instructions for the processor 410 to process the following processing steps: obtaining image data of a target scene and determining the scene type of the target scene; wherein the image data includes multiple categories of visual elements; extracting the feature information of each category of visual elements in the image data, and determining the weight information of each category of visual elements in the image data according to the scene type and the feedback information of the user on the layout of the target scene obtained in advance; determining the emotion intensity value of the target scene according to the feature information and the corresponding weight information of all categories of visual elements in the image data, and determining the layout optimization suggestion of the target scene according to the emotion intensity value; wherein the emotion intensity value is used to indicate the comprehensive influence of all categories of visual elements in the target scene on the user's emotional experience; and generating the image data of the target scene after optimized layout by using an image generation model according to the layout optimization suggestion.

[0103] Optionally, the multiple categories of visual elements include text elements, image elements, and layout elements; and the operation of extracting the feature information of each category of visual elements in the image data includes: extracting the text content in the image data by optical character recognition technology, calculating the text emotion contribution value corresponding to the text content by a pre-trained language model, and using it as the feature information of the text elements in the image data; identifying the graphic symbols in the image data based on a target detection algorithm, determining the semantic type of the graphic symbols, extracting the HSV value of the main color of the image data, and determining the image emotion contribution value of the image data according to the semantic type and the HSV value, and using it as the feature information of the image elements in the image data; and based on the image data, counting the number of identifiers in the target scene per unit area, determining the visual focus of the image data through a saliency detection model, and determining the layout emotion contribution value of the image data according to the number of identifiers and the visual focus, and using it as the feature information of the layout elements in the image data.

[0104] Optionally, the operation of determining the image emotion contribution value of the image data according to the semantic type and the HSV value includes: determining the symbol emotion contribution value corresponding to the graphic symbol according to the semantic type; calculating the color emotion contribution value of the image data according to the HSV value; and determining the image emotion contribution value of the image data according to the symbol emotion contribution value and the color emotion contribution value.

[0105] Optionally, the operation of determining the layout emotional contribution value of the image data according to the number of identifiers and the visual focus includes: determining the basic density scores of the sub-regions of the target scene according to the number of identifiers; wherein the target scene is divided into multiple sub-regions by unit area; determining the identifier salience weights of the sub-regions according to the visual focus; and determining the layout emotional contribution values of the sub-regions according to the basic density scores and the identifier salience weights, and taking the sum of the layout emotional contribution values of the sub-regions as the layout emotional contribution value of the image data.

[0106] Optionally, the operation of determining the weight information of the visual elements of each category in the image data according to the scene type and the pre-obtained feedback information of the user on the layout of the target scene includes: determining the initial weight information of the visual elements of each category in the image data according to the preset rule library and the scene type; wherein the preset rule library stores the initial weight information of the visual elements of each category in the image data of different scene types; and optimizing the initial weight information of the visual elements of each category in the image data by using a reinforcement learning mechanism according to the feedback information to obtain the weight information of the visual elements of each category in the image data; wherein the reward signal of the reinforcement learning mechanism is obtained based on the feedback information.

[0107] Optionally, the operation of determining the emotional intensity value of the target scene according to the feature information and the corresponding weight information of the visual elements of all categories in the image data includes: calculating the emotional intensity value of the target scene according to the feature information and the corresponding weight information of the visual elements of all categories in the image data by the following formula:

[0108] where S is the emotional intensity value of the target scene, is the feature information of the text elements in the image data, is the weight information of the text elements in the image data, is the feature information of the image elements in the image data, is the weight information of the image elements in the image data, is the feature information of the layout elements in the image data, is the weight information of the layout elements in the image data.

[0109] Optionally, the operation of generating the image data of the target scene with the optimized layout according to the layout optimization suggestion includes: generating a prompt word according to the layout optimization suggestion; and inputting the image data and the prompt word into a pre-trained image generation model to output the image data of the target scene with the optimized layout.

[0110] Therefore, according to this embodiment, through the organic combination of multiple links such as accurate scene type recognition, dynamic weight allocation, fully considering the comprehensive influence among multiple visual elements, quantification and evaluation of emotional intensity values, generation of layout optimization suggestions, and application of image generation models, multi-dimensional improvement of scene layout optimization is achieved. This method not only improves the pertinence and effectiveness of the optimization result, but also enhances the flexibility and adaptability of the optimization process, greatly improving the scene layout optimization effect and providing users with a more personalized and comfortable scene experience. Thus, it solves the technical problems in the prior art that the scene layout optimization method has a rigid static weight allocation mechanism, does not consider the comprehensive influence among multiple visual elements, and has insufficient dynamic quantification ability of emotional intensity, resulting in poor scene layout optimization effect.

[0111] The serial numbers of the above embodiments of the present invention are only for description and do not represent the advantages or disadvantages of the embodiments.

[0112] In the above embodiments of the present invention, the descriptions of each embodiment have their own emphases. For the parts not detailed in a certain embodiment, reference may be made to the relevant descriptions of other embodiments.

[0113] In the several embodiments provided in the present application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are only illustrative. For example, the division of the units is only a logical function division, and there may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed mutual coupling or direct coupling or communication connection can be through some interfaces. The indirect coupling or communication connection of units or modules can be in an electrical or other form.

[0114] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0115] In addition, in each embodiment of the present invention, each functional unit can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of a software functional unit.

[0116] If the above-mentioned integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of the present invention. The foregoing storage medium includes: various media such as USB flash drives, read-only memories (ROMs), random access memories (RAMs), mobile hard disks, magnetic disks, or optical discs that can store program codes.

[0117] The above are only the preferred embodiments of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.

Claims

1. A method for optimizing scene layout based on multi-visual elements and emotion perception, characterized in that Including: Obtain the image data of the target scene and determine the scene type of the target scene; wherein, the image data includes visual elements of multiple categories; Extract the feature information of the visual elements of each category in the image data, and determine the weight information of the visual elements of each category in the image data according to the scene type and the feedback information of the user on the layout of the target scene obtained in advance; Determine the emotional intensity value of the target scene according to the feature information and the corresponding weight information of all categories of visual elements in the image data, and determine the layout optimization suggestion of the target scene according to the emotional intensity value; wherein the emotional intensity value is used to indicate the comprehensive impact of all categories of visual elements in the target scene on the user's emotional experience; and According to the layout optimization suggestion, use an image generation model to generate the image data of the target scene with the optimized layout.

2. The method according to claim 1, characterized in that, The visual elements of the multiple categories include text elements, image elements, and layout elements; and the operation of extracting the feature information of the visual elements of each category in the image data includes: Extract the text content in the image data through optical character recognition technology, and calculate the text emotion contribution value corresponding to the text content through a pre-trained language model as the feature information of the text elements in the image data; Based on an object detection algorithm, identify the graphic symbols in the image data, determine the semantic type of the graphic symbols, extract the HSV value of the main color of the image data, and determine the image emotion contribution value of the image data according to the semantic type and the HSV value as the feature information of the image elements in the image data; and Based on the image data, count the number of identifiers in the target scene per unit area, determine the visual focus of the image data through a saliency detection model, and determine the layout emotion contribution value of the image data according to the number of identifiers and the visual focus as the feature information of the layout elements in the image data.

3. The method according to claim 2, characterized in that, The operation of determining the image emotion contribution value of the image data according to the semantic type and the HSV value includes: Determine the symbol emotion contribution value corresponding to the graphic symbol according to the semantic type; Calculate the color emotion contribution value of the image data according to the HSV value; and Determine the image emotion contribution value of the image data according to the symbol emotion contribution value and the color emotion contribution value.

4. The method according to claim 2, characterized in that, The operation of determining the layout emotion contribution value of the image data according to the number of identifiers and the visual focus includes: Determine the basic density score of each sub-region of the target scene according to the number of identifiers; wherein, the target scene is divided into multiple sub-regions per unit area; Determine the identifier saliency weight of each sub-region according to the visual focus; and Determine the layout emotion contribution value of each sub-region according to the basic density score and the identifier saliency weight, and use the sum of the layout emotion contribution values of each sub-region as the layout emotion contribution value of the image data.

5. The method according to claim 1, wherein The operation of determining the weight information of visual elements of each category in the image data according to the scene type and the feedback information of the user on the target scene layout obtained in advance includes: Determining the initial weight information of visual elements of each category in the image data according to a preset rule library and the scene type; wherein, the preset rule library stores the initial weight information of visual elements of each category in the image data of different scene types; and Optimizing the initial weight information of visual elements of each category in the image data by using a reinforcement learning mechanism according to the feedback information to obtain the weight information of visual elements of each category in the image data; wherein, the reward signal of the reinforcement learning mechanism is obtained based on the feedback information.

6. The method according to claim 1, wherein The operation of determining the emotional intensity value of the target scene according to the feature information and the corresponding weight information of visual elements of all categories in the image data includes: Calculating the emotional intensity value of the target scene according to the feature information and the corresponding weight information of visual elements of all categories in the image data through the following formula: ; Among them, S is the emotional intensity value of the target scenario, is the feature information of the text element in the image data, is the weight information of the text element in the image data, is the feature information of the image element in the image data, is the weight information of the image element in the image data, is the feature information of the layout element in the image data, is the weight information of the layout element in the image data.

7. The method according to claim 1, characterized in that, The operation of generating the image data of the target scene with an optimized layout by using an image generation model according to the layout optimization suggestion includes: Generating a prompt according to the layout optimization suggestion; and Inputting the image data and the prompt into a pre-trained image generation model to output the image data of the target scene with an optimized layout.

8. A storage medium, characterized in that, The storage medium includes a stored program, wherein, when the program runs, the method according to any one of claims 1 to 7 is executed by a processor.

9. A scene layout optimization system based on multi-visual elements and emotion perception, comprising: A data acquisition module, configured to acquire the image data of the target scene and determine the scene type of the target scene; wherein, the image data includes visual elements of multiple categories; An extraction and parsing module, configured to extract the feature information of visual elements of each category in the image data, and determine the weight information of visual elements of each category in the image data according to the scene type and the feedback information of the user on the target scene layout obtained in advance; An emotion calculation and optimization decision module, configured to determine the emotional intensity value of the target scene according to the feature information and the corresponding weight information of visual elements of all categories in the image data, and determine the layout optimization suggestion of the target scene according to the emotional intensity value; wherein the emotional intensity value is used to indicate the comprehensive impact of visual elements of all categories in the target scene on the user's emotional experience; and A layout optimization module, configured to generate the image data of the target scene with an optimized layout by using an image generation model according to the layout optimization suggestion.

10. A scene layout optimization system based on multi-visual elements and emotion perception, characterized in that, Comprising: A processor; And A memory, connected to the processor, configured to provide instructions for the processor to perform the following processing steps: Acquire the image data of the target scene and determine the scene type of the target scene; wherein, the image data includes visual elements of multiple categories; Extract the feature information of the visual elements of each category in the image data, and determine the weight information of the visual elements of each category in the image data according to the scene type and the feedback information of the user on the target scene layout obtained in advance; Determine the emotional intensity value of the target scene according to the feature information and the corresponding weight information of the visual elements of all categories in the image data, and determine the layout optimization suggestion for the target scene according to the emotional intensity value; wherein the emotional intensity value is used to indicate the comprehensive influence of the visual elements of all categories in the target scene on the user's emotional experience; and According to the layout optimization suggestion, use an image generation model to generate the image data of the target scene with the optimized layout.

Citation Information

Patent Citations

  • Public subjective emotion quantification method and system and electronic equipment

    CN114565300A

  • Scene graph generation method and system supporting historical and cultural block scene

    CN118334414A

  • Image generation method, and method and device for generating target text graph generative model

    CN118840447A

  • Multi-mode-based target area viewing scene planning method and device, computer equipment and readable storage medium

    CN118840648A

  • Image generation method and device, electronic equipment and storage medium

    CN119359850A

Cited By

  • Suspension control method, device and equipment based on emotional interaction and storage medium

    CN121133327A

  • Training-free video editing method based on frequency enhancement and conflict adaptive model

    CN121937570A