Intelligent optimization method and system for picture-text layout combined with digital multimedia
By using spatiotemporal collaborative modeling of text and images and a dynamic adaptive typesetting engine, the problem of text and image typesetting methods being unable to be adjusted in real time has been solved, enabling personalized optimization of text and image content and improving the display quality and efficiency of digital multimedia.
Patent Information
- Application Number
- CN202610030751.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-01-12
- Publication Date
- 2026-03-20
- Estimated Expiration
- 2046-01-12
AI Technical Summary
Existing graphic layout methods cannot adjust in real time according to device screen characteristics, ambient light intensity, and audience behavior intentions, resulting in blurry images and misaligned text, failing to meet the needs of different devices and audiences, and reducing user experience.
By acquiring the set of digital multimedia graphics and text to be optimized, real-time display environment parameters, and audience behavior and intent data, spatiotemporal collaborative modeling of graphics and text is performed to build a dynamic self-adaptive typesetting engine and generate a target graphic and text typesetting optimization scheme that conforms to the real-time environment and audience intent.
It enables the coordinated display of text and images across time, space, and emotion, enhancing the content's appeal and attractiveness, improving the rationality and stability of layout, and providing a personalized digital multimedia content display experience.
Smart Images

Figure CN121527249B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of digital multimedia, in particular to a graphic-text layout intelligent optimization method and system combined with digital multimedia. BACKGROUND
[0002] In the current booming development of digital multimedia, graphic-text layout plays a crucial role in various types of digital media content display. Whether it is web design, mobile application interface display, or electronic publications, high-quality graphic-text layout is indispensable. However, existing graphic-text layout methods have many limitations.
[0003] Traditional graphic-text layout methods are often based on fixed templates and preset rules, lacking consideration for real-time display environments. Different device screen characteristics, such as screen size and resolution, vary greatly, and fixed layout schemes may not present the best results on different devices, leading to issues such as blurred images and misaligned text. Meanwhile, environmental light intensity also affects user viewing experience of graphic-text content. In strong light environments, color contrast may need to be adjusted, while in weak light environments, brightness may need to be increased. However, traditional methods cannot automatically adjust according to changes in light.
[0004] In addition, existing methods do not fully consider the behavior intentions of the audience. Interactive behavior sequences, emotional tendencies, and content demand expressions are all important factors affecting layout effectiveness. Different audiences have different focuses on graphic-text content, emotional tendencies affect their acceptance of layout styles, and content demand expressions directly determine the presentation focus of graphic-text content. However, traditional layout methods cannot adjust in real-time according to these dynamic changes in audience needs, resulting in layout schemes that do not meet audience expectations and reducing user experience. SUMMARY
[0005] Therefore, the purpose of the present application is to provide a graphic-text layout intelligent optimization method and system combined with digital multimedia.
[0006] According to a first aspect of the present application, a graphic-text layout intelligent optimization method combined with digital multimedia is provided, which includes:
[0007] Obtaining a set of digital multimedia graphic-texts to be optimized, real-time display environment parameters, and audience behavior intention data. The set of digital multimedia graphic-texts to be optimized includes multiple image units and multiple text units. The real-time display environment parameters include device screen characteristics, environmental light intensity, and network transmission rate. The audience behavior intention data includes interactive behavior sequences, emotional tendency labels, and content demand expressions.
[0008] image units and text units in the digital multimedia collection to be optimized are subjected to image-text spatiotemporal collaborative modeling, and a spatiotemporal association relationship and an emotional collaborative relationship between the image units and the text units are established in combination with the emotional tendency labels in the audience behavior intention data, to obtain an image-text collaborative model;
[0009] Based on the real-time display environment parameters, the interactive behavior sequences in the audience behavior intention data, and the image-text collaborative model, a dynamic self-adaptive layout engine is constructed, which includes an environment parameter mapping module, a behavior intention decoding module, a collaborative relationship application module, and a layout parameter evolution module.
[0010] The dynamic self-adaptive layout engine is applied to dynamically calculate layout parameters of the image units and the text units in the digital multimedia collection to be optimized, to generate an initial layout scheme, and the layout parameter evolution module of the dynamic self-adaptive layout engine is applied to perform multi-dimensional conflict pre-performance and adaptive adjustment on the initial layout scheme, to obtain an intermediate layout scheme.
[0011] The intermediate layout scheme is deployed to a real-time display environment, and the dynamic self-adaptive layout engine is driven to continuously evolve the layout scheme in combination with the content demand expressions in the audience behavior intention data and the real-time updated display environment parameters, to generate a target image-text layout optimization scheme that conforms to the real-time environment and the audience intention, and the target image-text layout optimization scheme is output for digital multimedia content display.
[0012] According to a second aspect of the present application, an image-text layout intelligent optimization system combined with digital multimedia is provided, which includes a machine-readable storage medium and a processor, the machine-readable storage medium stores machine-executable instructions, and the processor, when executing the machine-executable instructions, implements the aforementioned image-text layout intelligent optimization method combined with digital multimedia.
[0013] According to a third aspect of the present application, a computer-readable storage medium is provided, which stores computer-executable instructions, and when the computer-executable instructions are executed, the aforementioned image-text layout intelligent optimization method combined with digital multimedia is implemented.
[0014] According to any one of the above aspects, the technical effect of the present application is that:
[0015] The image-text collaborative model is constructed by acquiring a to-be-optimized digital multimedia image-text set, real-time display environment parameters and audience behavior intention data, then performing image-text space-time collaborative modeling on image units and text units, and combining audience emotional tendency labels to establish space-time correlation and emotional collaborative relationship. The image-text collaborative model effectively solves the problem of collaborative display of image-text content in the space-time and emotional aspects, makes the collocation between images and texts more natural and harmonious, and enhances the appeal and attraction of the content. The dynamic self-adaptive layout engine based on real-time display environment parameters, audience interactive behavior sequences and the image-text collaborative model has strong environmental adaptability and audience demand interpretation ability. The multiple modules contained therein cooperate with each other, can dynamically calculate layout parameters according to different environmental conditions and audience behaviors, and generate an initial layout scheme. Through the multi-dimensional conflict pre-performance and adaptive adjustment of the layout parameter evolution module, an intermediate layout scheme is obtained, which effectively avoids various conflict problems that may occur in layout, and improves the rationality and stability of the layout. After the intermediate layout scheme is deployed to the real-time display environment, combined with the audience content demand expression and the real-time updated display environment parameters, the layout engine is continuously evolved to generate a target image-text layout optimization scheme that meets the real-time environment and audience intention, which can always keep the layout scheme highly matched with the actual display environment and audience demand, provides users with more high-quality and personalized digital multimedia content display experience, and significantly improves the quality and efficiency of digital multimedia image-text layout. BRIEF DESCRIPTION OF DRAWINGS
[0016] Figure 1 A flowchart of an image-text layout intelligent optimization method combined with digital multimedia provided by an embodiment of the application is shown.
[0017] Figure 2 A component structure diagram of an image-text layout intelligent optimization system combined with digital multimedia provided by an embodiment of the application is shown. DETAILED DESCRIPTION
[0018] Figure 1 A flowchart of an image-text layout intelligent optimization method combined with digital multimedia provided by an embodiment of the application is shown. The detailed steps include:
[0019] Step S110: acquiring a to-be-optimized digital multimedia image-text set, real-time display environment parameters and audience behavior intention data, the to-be-optimized digital multimedia image-text set containing multiple image units and multiple text units, the real-time display environment parameters containing device screen characteristics, environmental light intensity and network transmission rate, and the audience behavior intention data containing interactive behavior sequences, emotional tendency labels and content demand expressions.
[0020] This embodiment takes the online menu text and image layout optimization scenario of the catering industry as an example for illustration. The online menu, as a specific form of digital multimedia text and image collection, contains text units for dish introduction and image units for dish pictures. In this scenario, the digital multimedia text and image collection to be optimized is a certain restaurant's online electronic menu, which contains multiple dish categories, each category having multiple dish corresponding text units and image units. For example, under the "Signature Hot Dishes" category, the dish "Braised Ribs" has a corresponding text unit containing dish name, taste description, ingredient composition, price information, etc. At the same time, there is a high-definition real object image of the dish as an image unit.
[0021] The acquisition of real-time display environment parameters is completed through the mobile terminal device currently used by the user. The device screen characteristics include the physical size of the screen, the resolution parameter, the current screen display direction (portrait or landscape), and the screen pixel density. For example, the physical size of the screen of the smartphone used by the user is commonly rectangular, the resolution parameter reflects the number of pixels in the horizontal and vertical directions of the screen, the screen direction is detected in real time by the built-in gravity sensor of the device, and the pixel density affects the level of detail of the screen display content. The intensity of the ambient light is collected by the light sensor of the device, and its value reflects the brightness of the environment in which the user is currently located, for example, in bright outdoor environments and dim indoor environments, the collected ambient light intensity values differ significantly. The network transmission rate is monitored in real time by the network module of the device, including the current network connection type (such as Wi-Fi, 4G, 5G) and the corresponding download and upload rates, which affect the loading speed of large-capacity data such as image units in the online menu.
[0022] The acquisition of audience behavior intention data needs to be carried out in accordance with relevant laws and regulations, and privacy protection measures need to be taken when dealing with privacy sensitive data. For the interaction behavior sequence, through the user behavior collection module integrated in the catering APP, the user's operations in browsing the online menu are recorded under the condition of user authorization, such as the sequence and properties of clicking dish category buttons, sliding dish lists, long pressing dish pictures to view details, zooming dish pictures, etc. The sentiment tendency label is obtained by analyzing the user's past dish evaluation content, for example, the use of positive words such as "delicious" and "satisfied" in the user's evaluation corresponds to a positive sentiment tendency label, and the use of negative words such as "tasteless" and "disappointed" corresponds to a negative sentiment tendency label. The content demand expression is obtained by comprehensive analysis of user search records, favorite dish types, and historical ordering records in the APP, etc. For example, if the user searches for "vegetariandishes" multiple times, it indicates that the content demand expression contains a preference for vegetariandishes.
[0023] In acquiring the above data, for privacy-sensitive data (such as user's geographic location information, detailed consumption records, etc.), data desensitization technology is used for processing, the real identity information of the user is anonymized, and at the same time, encrypted transmission is used to ensure that the data is not leaked during transmission. When storing, use encryption storage technology, limit data access permissions, and only authorize relevant modules to access the processed non-sensitive data when necessary.
[0024] Step S120: Image units and text units in the to-be-optimized digital multimedia text collection are subjected to image-text spatio-temporal collaborative modeling, and the image-text spatio-temporal association relationship and emotional collaborative relationship of the image units and the text units are established in combination with the emotional tendency labels in the audience behavior intention data, to obtain an image-text collaborative model.
[0025] Step S121: Each text unit is subjected to reading rhythm analysis processing, and sentence pause labels, vocabulary density changes, and semantic turning point positions in the text unit are extracted to determine reading time length segmentation and key semantic intervals of the text unit, and form text spatio-temporal features.
[0026] Step S1211: Each text unit is subjected to sentence structure analysis processing, and comma, period, exclamation mark, and question mark sentence pause labels in the text unit are identified, and the occurrence frequency and interval distance of different pause labels are counted.
[0027] Taking the text unit of the dish "fish-flavored shredded pork" in the online menu as an example, the content of the text unit is: "fish-flavored shredded pork, classic Sichuan cuisine, salty and slightly spicy, tender and juicy shredded pork, rich in ingredients, including green peppers, carrots, and mushrooms, etc., and the price is affordable." First, the sentence structure of the text unit is analyzed to identify the pause labels, which are comma and period. The comma appears after "fish-flavored shredded pork", "classic Sichuan cuisine", "salty and slightly spicy", "tender and juicy shredded pork", "rich in ingredients", and "including green peppers, carrots, and mushrooms, etc."; the period appears at the end of the text unit. The occurrence frequency of different pause labels is counted, with 6 commas and 1 period. The interval distance is obtained by counting the number of characters between two adjacent pause labels, for example, the comma interval between "fish-flavored shredded pork" and "classic Sichuan cuisine" is 4 characters ("fish-flavored shredded pork" is 4 characters), the comma interval between "classic Sichuan cuisine" and "salty and slightly spicy" is 5 characters ("classic Sichuan cuisine" is 5 characters), and so on.
[0028] Step S1212: According to the interval character number of the pause label, the text unit is divided into multiple sentence segments, and each sentence segment is bounded by two adjacent pause labels.
[0029] According to the identified pause marks and the number of interval characters, the text unit of "Yuxiang Rou Si" is divided into sentence segments. The first sentence segment is "Yuxiang Rou Si" (bounded by the beginning and the first comma); the second sentence segment is "Jingdian Chuancai" (bounded by the first comma and the second comma); the third sentence segment is "Kougan Xianxian Weila" (bounded by the second comma and the third comma); the fourth sentence segment is "Rou Si Xianren" (bounded by the third comma and the fourth comma); the fifth sentence segment is "Peiliang Fengfu" (bounded by the fourth comma and the fifth comma); the sixth sentence segment is "Bianhuang Qingjiao, Huoluobo, Muer Etc." (bounded by the fifth comma and the sixth comma); and the seventh sentence segment is "Jiaqian Shihui" (bounded by the sixth comma and the period).
[0030] Step S1213: Perform vocabulary density calculation processing on each sentence segment, count the number of words and the number of characters in each sentence segment, and calculate the vocabulary density value per unit character length.
[0031] For each sentence segment divided in step S1212, vocabulary density calculation is performed. The number of words refers to the number of independent words in the sentence segment, and the number of characters refers to the total number of all characters (including Chinese characters, punctuation marks, etc.) contained in the sentence segment. For example, in the sentence segment "Kougan Xianxian Weila", the number of words is 4 ("Kougan", "Xianxian", "Weila"), and the number of characters is 6 ("Kougan Xianxian Weila" has a total of 6 Chinese characters), so the vocabulary density value per unit character length is the number of words divided by the number of characters, i.e. 4 divided by 6, resulting in a corresponding vocabulary density value. The above calculation is performed on all sentence segments to obtain their respective vocabulary density values.
[0032] Step S1214: Analyze the change trend of the vocabulary density values, and identify the sentence segments whose vocabulary density values exceed the preset density threshold, which correspond to the areas that need to be paid attention to when reading.
[0033] The vocabulary density values of each sentence segment calculated in step S1213 are arranged in order to form the change trend of the vocabulary density values. A density threshold is preset, which is set according to the general reading habits and information importance of the catering menu text. After analyzing the change trend, the sentence segments whose vocabulary density values exceed the preset density threshold are identified. For example, in the text unit of "Yuxiang Rou Si", the sentence segment "Peiliang Fengfu, Bianhuang Qingjiao, Huoluobo, Muer Etc." may have a higher vocabulary density value because it contains information about multiple ingredients. If it exceeds the preset density threshold, this sentence segment is identified as an area that needs to be paid attention to when reading, because users usually pay attention to the composition of ingredients when selecting dishes.
[0034] Step S1215: Perform semantic analysis on the text unit to identify semantic turning connecting words and semantic emphasis words, and determine semantic turning positions and semantic emphasis positions.
[0035] The text unit is subjected to semantic analysis using natural language processing technology. The semantic turning connecting words include "but", "however", "but", etc., and the semantic emphasis words include "most", "especially", "classic", "signature", etc. In the text unit of "Yuxiang Rou Si", "classic" in "classic Sichuan cuisine" is a semantic emphasis word, indicating the tradition and popularity of the dish; if the text unit contains a statement such as "although the price is slightly high, the taste is excellent", "although" and "but" are semantic turning connecting words, and the positions of the two words are semantic turning positions, and "excellent taste" is a semantic emphasis position. By identifying these words, the semantic turning positions and semantic emphasis positions in the text unit are determined.
[0036] Step S1216: Combine the sentence segment division result, the lexical density trend, and the semantic analysis result to divide the text unit into multiple reading time segments. The sentence segments with lexical density values exceeding the preset density threshold and the sentence segments containing semantic emphasis positions correspond to areas that require attention during reading.
[0037] The text unit is divided into reading time segments based on the sentence segment division result (step S1212), the sentence segments with lexical density values exceeding the preset density threshold (step S1214), and the semantic emphasis positions obtained by semantic analysis (step S1215). The division of reading time segments is based on the importance and reading complexity of different sentence segments. Sentence segments with high importance and high reading complexity are assigned longer reading time. For example, in the text unit of "Yuxiang Rou Si", the sentence segments containing "classic Sichuan cuisine" (containing semantic emphasis words) and "rich ingredients, including green peppers, carrots, and wood ear fungus" (with lexical density values exceeding the threshold) are divided into segments that require longer reading time, while other sentence segments such as "affordable price" have relatively short reading time. At the same time, the intervals containing these important sentence segments are the key semantic intervals.
[0038] Step S1217: Integrate the reading time segments, key semantic intervals, sentence pause markers distribution, and lexical density change data to form text spatiotemporal features.
[0039] The reading time segment and the key semantic interval obtained in step S1216, the sentence pause mark distribution in step S1211, and the lexical density change data in step S1213 are integrated. The text space-time feature is a multi-dimensional feature vector, which includes the reading time allocation information (reading time segment) of the text unit in the time dimension and the key information distribution (key semantic interval) in the space dimension, and also includes the related data of the sentence structure and the lexical distribution, which are used for subsequent matching with the image space-time feature.
[0040] Step S122: performing visual presentation timing analysis processing on each image unit, extracting the visual focus switching path, color transition rhythm, and detail information level in the image unit, determining the optimal display duration and visual attention order of the image unit, and forming an image space-time feature.
[0041] Step S1221: performing visual focus detection processing on each image unit, identifying multiple visual focuses in the image unit based on color contrast, brightness difference, edge complexity, and object saliency.
[0042] Taking the image unit of the dish "Yu Xiang Rou Si" as an example, the image is a real object photograph of the dish, including the dish body, the plate edge, the background, and other elements. The visual focus detection analyzes the color contrast of the image, and there is a significant color difference between the dish body (Yu Xiang Rou Si) and the plate and background, so the area with high color contrast is easy to become a visual focus. In terms of brightness difference, the dish body part is brighter after light processing, and forms a contrast with the relatively dark background. The edge complexity is detected by analyzing the complexity of the object contour in the image, and the edges of the meat and vegetables in the dish are relatively complex, so the corresponding area has high edge complexity. The object saliency is based on the characteristics of the catering image, and the dish itself is the salient object of the image. By comprehensively analyzing the above factors, multiple visual focuses in the image unit are identified, such as the center area of the dish (where the meat and main ingredients are concentrated), the color bright part of the dish (such as the carrot block), etc.
[0043] Step S1222: analyzing the position relationship, size difference, and attraction strength between the visual focuses, simulating the visual line movement trajectory when the audience watches the image, and determining the visual focus switching path.
[0044] After identifying multiple visual focal points, analyze their positional relationships, such as which visual focal points are on the left, right, top, or bottom of the image, and their relative distances. Size differences refer to the area proportions occupied by different visual focal points in the image. Attraction strengths are evaluated by integrating factors such as color contrast, brightness differences, and edge complexity. The stronger the attraction, the higher the color contrast, the greater the brightness difference, and the higher the edge complexity. Based on these analyses, simulate the viewer's eye movement trajectory when viewing the image, starting with the strongest visual focal point and moving to other visual focal points, forming a visual focal point switching path. For example, in the "fish-flavored shredded pork" image, the shredded pork and main ingredients in the center area are the strongest visual focal points, which are first focused on, and then the eyes may move to the brightly colored carrot pieces, then to other ingredients, and finally may sweep across the edge of the plate and the background, forming a specific visual focal point switching path.
[0045] Step S1223: Perform color analysis processing on the image unit to extract the color distribution and color transition mode of different regions in the image unit, identify regions with significant color transitions, and calculate their areas to determine the color transition rhythm feature.
[0046] Color space conversion is performed on the image unit to convert the image from RGB color space to HSV color space for more accurate color property analysis. The hue, saturation, and lightness values of different regions in the image are extracted, and the color distribution is calculated, such as the main colors of the dish area and the colors of the background area. Color transition modes include gradual transition and abrupt transition, with gradual transition referring to a slow change in color between adjacent regions and abrupt transition referring to a significant jump in color between adjacent regions. Regions with significant color transitions are identified, i.e., regions with abrupt transition and significant color differences. The areas of these significant color transition regions are calculated as a proportion of the total image area. The color transition rhythm feature is determined by the size, distribution, and transition mode of the significant color transition regions, such as a larger area of significant color transition regions resulting in a more intense color transition rhythm.
[0047] Step S1224: Perform detail information extraction processing on the image unit based on image resolution, texture complexity, and object detail richness to divide the image unit into multiple detail information levels, with levels with rich detail elements corresponding to different visual attention depths.
[0048] Image resolution determines the sharpness of the image, and a high-resolution image contains more detailed information. Texture complexity is determined by analyzing the texture features of different regions in the image, such as the texture of the dish surface, the texture of the plate, the texture of the background, etc. The more complex the texture, the more detailed information. The richness of object details refers to the number of details contained in the image, such as the texture of the shredded pork in the "shredded pork with fish sauce" image, the texture of the vegetables, and the gloss of the soup. All of these belong to object details. According to these factors, the image unit is divided into multiple levels of detailed information, for example, the highest level is the core detail area of the dish (such as the texture and gloss of the shredded pork), which contains the most detailed elements; the middle level is the secondary detail area of the dish (such as the shape and color of the vegetables); the lowest level is the background and the edge area of the plate, which has relatively simple detailed elements. Different levels of detailed information correspond to different depths of visual focus, and the audience will first focus on the high-level area with rich detailed elements when watching the image, and then gradually focus on the low-level area with simple detailed elements.
[0049] Step S1225: According to the physical length of the visual focus switching path and the number of visual focuses, calculate the basic time required for the audience to completely browse all visual focuses.
[0050] The physical length of the visual focus switching path is obtained by quantifying the simulated visual line movement trajectory in the image coordinate system, that is, the sum of the distances between the coordinates of each point on the trajectory. The number of visual focuses is the total number of visual focuses identified in step S1221. The calculation of the basic time is based on the average time required for a general human eye to move from one visual focus to another when watching an image, and the average time spent at each visual focus. Divide the physical length of the visual focus switching path by the average movement speed of the human eye to get the total time of visual line movement; multiply the number of visual focuses by the average time spent at each visual focus to get the total time of visual line stay; the sum of the two is the basic time required for the audience to completely browse all visual focuses.
[0051] Step S1226: Adjust the basic time by combining the proportion of color transition area in the total image area, and increase the browsing time when the proportion of color transition area exceeds the preset proportion threshold.
[0052] A color transition area proportion threshold is preset, which is set according to the visual characteristics of the catering image. Calculate the proportion of the significant color transition area obtained in step S1223 in the total image area. If the proportion exceeds the preset proportion threshold, it means that the color change of the image is relatively rich, and the audience needs more time to perceive and understand these color transitions, so a certain browsing time is added to the basic time. The amount of time added is related to the degree to which the proportion of color transition area exceeds the threshold, the higher the proportion, the more time added.
[0053] Step S1227: Further adjust the browsing time according to the number and complexity of the detail information levels. When the proportion of the levels with rich detail elements exceeds the preset proportion threshold, the browsing time is increased accordingly, and the optimal display duration of the image unit is finally determined.
[0054] The more the number of detail information levels, the richer the detail levels contained in the image, and the more time the audience needs to pay attention to these details layer by layer. The complexity of the detail information level is evaluated by the number and complexity of the detail elements in each level. A proportion threshold of detail element rich level is preset, and the area proportion of the levels with rich detail elements (such as the highest level and the intermediate level in step S1224) in the entire image is calculated. If the proportion exceeds the preset proportion threshold, it means that the image has rich detail information, and the browsing time needs to be increased to allow the audience to fully observe the details. Add the time adjusted in step S1226 to the time adjusted according to the detail information level to obtain the optimal display duration of the image unit.
[0055] Step S1228: Based on the visual focus switching path and the detail information level, determine the visual attention order when the audience watches the image, first pay attention to the areas with visual focus attraction greater than the preset attraction threshold and the areas with rich detail information levels, and then pay attention to other areas.
[0056] A threshold of attraction is preset, and the areas with visual focus attraction greater than the threshold are screened out. These areas are the focus of the audience's first attention. At the same time, the areas with rich detail information levels are also the focus of visual attention. Combined with the visual focus switching path, the visual attention order is determined, starting from the area with the strongest visual focus attraction and belonging to the level with rich detail elements, and then following the visual focus switching path to pay attention to other areas that meet the conditions, and then paying attention to the areas with weaker visual focus attraction and the levels with simple detail elements. For example, in the "fish-flavored shredded pork" image, the shredded pork in the center area (strong attraction and rich detail) is first paid attention to, then the carrot pieces (stronger attraction and some details), then other ingredients, and finally the edge of the plate and the background.
[0057] Step S1229: Integrate the visual focus switching path, color transition rhythm, detail information level, optimal display duration, and visual attention order to form the spatiotemporal characteristics of the image.
[0058] The visual focus switching path determined in step S1222, the color transition rhythm feature determined in step S1223, the detail information level divided in step S1224, the optimal display duration determined in step S1227, and the visual attention order determined in step S1228 are integrated to form an image space-time feature. The image space-time feature is also a multi-dimensional feature vector, which contains the optimal display duration in the time dimension and the visual attention order, the focus switching path, the color transition rhythm, and the detail information level in the space dimension.
[0059] Step S123: The text space-time feature and the image space-time feature are time sequence matched, and the space-time association relationship between the image unit and the text unit is established according to the corresponding relationship between the reading duration segmentation and the optimal display duration and the association relationship between the key semantic interval and the visual attention order.
[0060] Taking the text unit and the image unit of “Yuxiang Rou Si” as an example, the reading duration segmentation in the text space-time feature represents the time required to read each part of the dish text, and the optimal display duration in the image space-time feature is the time during which the image of the dish should be displayed. The total duration of the reading duration segmentation is matched with the optimal display duration. If the total duration of the reading is close to the optimal display duration, it is considered that the two are well matched in time; if the difference is large, the reading duration segmentation or the image display duration needs to be adjusted so that the two can be coordinated in time, for example, the reading duration of the key semantic interval in the text should correspond to the display duration of the corresponding visual attention order in the image. The key semantic interval corresponds to the content that the user needs to read, such as the ingredients, and the visual attention order in the image displays the area of the ingredients. When the user reads the ingredients part of the text, the image should display the corresponding visual area of the ingredients, thereby establishing the space-time association relationship between the image unit and the text unit.
[0061] Step S124: The sentiment tendency of each text unit is extracted, and the sentiment feature value of the text unit is determined by analyzing the sentiment word distribution, the tone expression manner, and the semantic sentiment intensity in the text unit.
[0062] Step S1241: The text unit is segmented, the stop words are removed, and the words with actual semantics are retained.
[0063] The text unit is segmented using a Chinese segmentation tool, and the continuous text sequence is divided into independent words. For example, “Yuxiang Rou Si, classic Sichuan cuisine, salty and spicy, tender and fresh, rich ingredients” is segmented to obtain “Yuxiang Rou Si”, “classic”, “Sichuan cuisine”, “taste”, “salty”, “spicy”, “tender”, “rich ingredients”, and the like. Then, the stop words, including “of”, “is”, “in”, and the like, which do not contribute to the actual emotion and semantics, are removed, and the above-mentioned words with actual semantics are retained after processing.
[0064] Step S1242: Construct an emotional vocabulary dictionary, which contains positive emotional vocabulary, negative emotional vocabulary and neutral vocabulary, and assigns corresponding emotional intensity values to different emotional vocabulary.
[0065] The construction of the emotional vocabulary dictionary is based on existing Chinese emotional dictionaries and is extended and adjusted in combination with the characteristics of the catering field. Positive emotional vocabulary such as "delicious", "tender", "classic", "rich", "affordable", etc. is assigned a positive emotional intensity value; negative emotional vocabulary such as "inedible", "greasy", "stiff", etc. is assigned a negative emotional intensity value; neutral vocabulary such as "price", "ingredients", "contains", etc. has an emotional intensity value of zero. The size of the emotional intensity value is set according to the intensity of the emotional expression of the vocabulary in the catering context, for example, the emotional intensity value of "classic" is higher than that of "not bad".
[0066] Step S1243: Match the processed vocabulary in step S1241 with the emotional vocabulary dictionary, and count the number of positive emotional vocabulary and negative emotional vocabulary in the text unit and the corresponding emotional intensity values.
[0067] Match the segmented and stop word removed vocabulary with the constructed emotional vocabulary dictionary one by one, and identify the positive emotional vocabulary and negative emotional vocabulary in the text unit. For example, in the "fish fragrant shredded pork" text unit, "classic", "tender", "rich", "affordable" are positive emotional vocabulary. Count the number of these positive emotional vocabulary and obtain the corresponding emotional intensity value of each vocabulary; at the same time, count the number of negative emotional vocabulary (if any) and its emotional intensity value.
[0068] Step S1244: Analyze the manner of expressing the tone of the text unit, and adjust the corresponding emotional intensity value if there is an exclamation sentence, a rhetorical question, etc. to strengthen the tone.
[0069] The manner of expressing the tone will affect the intensity of emotional expression. Exclamation sentences usually express strong emotions, and rhetorical questions may also have a certain emotional tendency. For example, if there is an exclamation sentence such as "This dish is too delicious!" in the text unit, "delicious" is a positive emotional vocabulary, and its emotional intensity value should be appropriately increased due to the strengthening effect of the exclamation sentence. If there is a rhetorical question such as "Isn't this dish not delicious?", it actually expresses positive emotions, so the emotional intensity value of the corresponding emotional vocabulary also needs to be adjusted.
[0070] Step S1245: Calculate the semantic emotional intensity of the text unit, subtract the sum of the emotional intensity values of the negative emotional vocabulary from the sum of the emotional intensity values of the positive emotional vocabulary, and obtain the preliminary emotional value of the text unit.
[0071] The positive emotional intensity values of the positive emotional words in step S1243 are added to obtain a positive emotional sum; the negative emotional intensity values of the negative emotional words are added to obtain a negative emotional sum. The preliminary emotional value of the text unit is obtained by subtracting the negative emotional sum from the positive emotional sum. For example, the sum of the intensity values of the positive emotional words in the text unit of "fish fragrant shredded pork" is the intensity value of "classic" plus the intensity value of "tender" plus the intensity value of "abundant" plus the intensity value of "affordable", and the preliminary emotional value is the sum assuming that the negative emotional sum is zero.
[0072] Step S1246: The preliminary emotional value is corrected based on the adjustment result of the tone expression mode to obtain the emotional feature value of the text unit.
[0073] The preliminary emotional value is corrected based on the analysis result of the tone expression mode. If there is an expression of strengthening tone, a certain correction value is added to the preliminary emotional value; if there is an expression of weakening tone, a certain correction value is reduced. The corrected emotional value is the emotional feature value of the text unit, which is a numerical value, a positive number indicating positive emotion, a negative number indicating negative emotion, and the absolute value of the numerical value indicating the emotional intensity.
[0074] Step S125: The emotional tendency extraction processing is performed on each image unit to analyze the color emotional attributes, composition emotional expressions, and visual element emotional symbols in the image unit to determine the emotional feature value of the image unit.
[0075] Step S1251: The color emotional attributes of the image unit are analyzed to extract the dominant color, color saturation, and brightness value of the image, and the emotional tendency and intensity corresponding to each color are determined based on the color psychology theory.
[0076] Different colors have different emotional tendencies in color psychology, for example, red is usually associated with passion and appetite, green is associated with health and freshness, and yellow is associated with warmth and brightness. The color analysis of the image unit extracts the dominant color, i.e., the color with the highest proportion in the image; the color saturation reflects the brightness of the color, and the high saturation color has a stronger emotional expression; the brightness value reflects the lightness and darkness of the color. Based on these color attributes and the color psychology theory, the emotional tendency (positive, negative, or neutral) and emotional intensity value corresponding to each color are determined. For example, the dominant color of the "fish fragrant shredded pork" image may include red (shredded pork and sauce) and green (green pepper), red corresponds to positive appetite emotion with a positive intensity value, and green corresponds to positive health emotion with a positive intensity value.
[0077] Step S1252: The composition emotional expression analysis of the image unit is performed to analyze the composition mode (such as symmetrical composition, diagonal line composition, white space, etc.) of the image, and different composition modes convey different emotional atmosphere.
[0078] Symmetrical composition gives people a stable, solemn feeling; diagonal composition has dynamic and vitality; white space composition gives people a simple, comfortable feeling. Analyzing the composition of the image unit, for example, the "fish fragrant shredded pork" image may use center composition, placing the main body of the dish in the center of the image, highlighting the dish and conveying a direct, focused emotional atmosphere. According to the emotional atmosphere corresponding to different composition methods, the corresponding emotional intensity value is given, for example, the positive emotional intensity value conveyed by center composition is medium.
[0079] Step S1253: identifying visual elements in the image unit, analyzing the emotional symbolic meaning of the visual elements, such as the freshness of the ingredients, the delicacy of the dish arrangement, etc.
[0080] Identify visual elements in the image, such as ingredients, dish arrangement, tableware, etc. The freshness of the ingredients is judged by the color and shape of the ingredients, and fresh ingredients have positive emotional symbolic meaning; the delicacy of the dish arrangement reflects the care of the restaurant, and delicate dish arrangement has positive emotional symbolic meaning. For example, in the "fish fragrant shredded pork" image, the ingredients are bright in color, indicating that the ingredients are fresh, and the dish is neatly arranged, all of which have positive emotional symbolic meaning and are given corresponding emotional intensity values.
[0081] Step S1254: Weighted sum of color emotional intensity value, composition emotional intensity value and visual element emotional intensity value to get the emotional feature value of the image unit.
[0082] The color emotional intensity value, composition emotional intensity value and visual element emotional intensity value are set with weights, and the weight size is determined according to their importance in the emotional expression of the food image, for example, the color emotion and the visual element emotion (freshness of ingredients, dish arrangement) may have higher weight. After multiplying each emotional intensity value by the corresponding weight and adding them up, the emotional feature value of the image unit is obtained. The emotional feature value is also a numerical value, a positive number represents positive emotion, a negative number represents negative emotion, and the absolute value size represents the emotional intensity.
[0083] Step S126: Extract the emotional preference feature value of the audience for similar image-text from the emotional tendency label of the audience behavior intention data as a reference benchmark for emotional coordination.
[0084] The sentiment tendency label in the audience behavior intention data is based on the past behaviors of the user, such as evaluation, likes, and collections, on similar food menu texts and images. For example, when browsing other Sichuan dishes, the user gives a high evaluation and more collections to the dishes that contain descriptions such as “spicy and fresh” and “rich in ingredients”, and the pictures are bright in color and fresh in ingredients. These behaviors are labeled as positive sentiment tendency. By analyzing the above-mentioned sentiment tendency labels, the audience's emotional preferences for similar texts and images (such as Sichuan dishes) are extracted, such as preference for positive emotional expression, emotional focus on food freshness and taste description, etc. The above-mentioned preferences are converted into specific emotional preference feature values as reference benchmarks for subsequent emotional coordination.
[0085] Step S127: The sentiment feature values of the text unit and the sentiment feature values of the image unit are matched and processed in coordination with the audience's emotional preference feature values, the deviation of the sentiment feature values of the image unit and the text unit is adjusted, and the emotional coordination relationship is established.
[0086] The deviation of the sentiment feature values of the text unit and the audience's emotional preference feature values, and the deviation of the sentiment feature values of the image unit and the audience's emotional preference feature values are calculated. If the deviation is within the preset acceptable range, it is considered that the emotional expression of the text unit and the image unit conforms to the audience's preference; if the deviation exceeds the preset range, the sentiment feature values of the text unit and the image unit need to be adjusted. The adjustment method can be to modify the use of emotional words in the text unit, or to adjust the color, contrast, etc. of the image unit to change its sentiment feature value, so that the sentiment feature values of the two are close to the audience's emotional preference feature value, and finally the deviation between the sentiment feature values of the text unit and the image unit is also within the preset range, thereby establishing an emotional coordination relationship and ensuring that the text and image are consistent in emotional expression and conform to the audience's preference.
[0087] Step S128: Integrate the spatio-temporal correlation relationship and the emotional coordination relationship to construct a text-image coordination model containing time sequence mapping rules, emotional matching parameters, and coordination weights.
[0088] The spatio-temporal correlation relationship established in step S123 and the emotional coordination relationship established in step S127 are integrated. The time sequence mapping rules define the correspondence between the text reading time segmentation and the image optimal display time, as well as the association rules between the key semantic intervals and the visual attention order; the emotional matching parameters include the matching threshold and the adjustment coefficient between the text emotional feature value, the image emotional feature value, and the audience's emotional preference feature value; the coordination weight gives different weight values according to the importance of the text and image units in spatio-temporal correlation and emotional coordination, for example, for key dishes, the coordination weight of the text and image units is higher. By integrating these elements, a text-image coordination model is constructed, which can comprehensively reflect the spatio-temporal and emotional correlation between the image unit and the text unit.
[0089] Step S130: based on the real-time display environment parameters, the interactive behavior sequence in the audience behavior intention data, and the text-image collaborative model, a dynamic self-adaptive layout engine is constructed, which includes an environment parameter mapping module, a behavior intention decoding module, a collaborative relationship application module, and a layout parameter evolution module.
[0090] Step S131: the device screen characteristics in the real-time display environment parameters are analyzed and processed, the screen size specification, resolution parameter, screen direction, and pixel density are extracted, the initial layout space is calculated based on the screen size specification, resolution parameter, and pixel density, and the layout space elasticity coefficient is determined in combination with the screen direction, which is used for dynamically adjusting the size adaptation range of the text-image unit.
[0091] Step S1311: the screen size specification is converted from physical size to pixel size, and the pixel size of the screen in the horizontal and vertical directions is obtained by multiplying the screen physical size by the pixel density.
[0092] The screen size specification is usually expressed in diagonal length in inches, which needs to be converted into pixel size in the horizontal and vertical directions. First, according to the width-height ratio (such as 16:9) and the diagonal physical size of the screen, the horizontal and vertical physical sizes of the screen (in centimeters) are calculated, then the physical size is converted into inches (1 inch = 2.54 centimeters), and then multiplied by the pixel density (in pixels / inch) to obtain the pixel size of the screen in the horizontal and vertical directions. For example, the screen physical size is diagonal in inches, the width-height ratio is 16:9, the horizontal and vertical physical sizes are calculated, and then multiplied by the pixel density to obtain the horizontal and vertical pixel numbers, i.e. the pixel size.
[0093] Step S1312: according to the resolution parameter and the pixel size, the effective display area of the screen is calculated, and the pixel occupation of the non-content display area such as the status bar and the navigation bar is removed.
[0094] The resolution parameter is the total number of pixels of the screen, and the effective display area refers to the pixel area that can be used to display online menu content. The status bar and the navigation bar occupy part of the pixels, and the number of pixels in these areas needs to be subtracted from the total pixel size. For example, the status bar is located at the top of the screen with a height of several pixels, and the navigation bar is located at the bottom of the screen with a height of several pixels. The vertical pixel size of the effective display area is obtained by subtracting the height of the status bar and the navigation bar from the vertical pixel size of the screen, and the horizontal pixel size is usually the total horizontal pixel size of the screen, thereby determining the pixel range of the effective display area, i.e. the initial layout space.
[0095] Step S1313: analyze the influence of the screen orientation on the layout space, when the screen orientation is landscape, the horizontal layout space increases and the vertical layout space decreases; when the screen orientation is portrait, the vertical layout space increases and the horizontal layout space decreases.
[0096] The screen orientation is detected by the device gravity sensor, and the horizontal and vertical pixel sizes of the screen are interchanged in the landscape and portrait states (ignoring the influence of the status bar and the navigation bar). In the landscape state, the horizontal pixel size of the effective display area is larger, which is suitable for displaying multiple image-text units side by side; in the portrait state, the vertical pixel size is larger, which is suitable for arranging the image-text units in a top-down manner. According to the change of the screen orientation, the change trend of the layout space is determined.
[0097] Step S1314: based on the change trend of the screen orientation and the size of the initial layout space, the layout space elasticity coefficient is calculated, which reflects the adjustable range of the layout space in different directions.
[0098] The layout space elasticity coefficient includes a horizontal elasticity coefficient and a vertical elasticity coefficient. When the initial layout space is large, the elasticity coefficient is large, allowing the size of the image-text unit to be adjusted in a large range; when the initial layout space is small, the elasticity coefficient is small, and the adjustment range of the image-text unit size is limited. For example, in the landscape state, the horizontal layout space increases, and the horizontal elasticity coefficient increases accordingly, allowing the image unit to have a larger adjustment range in the horizontal direction; in the portrait state, the vertical elasticity coefficient increases. The specific calculation of the elasticity coefficient considers factors such as the screen orientation, the size of the initial layout space, and the average size requirement of the image-text unit.
[0099] Step S132: detecting the ambient light intensity in the real-time display environment parameter, and determining the visual comfort parameter according to the change of the light intensity, the visual comfort parameter being used to adjust the brightness contrast of the image unit and the font clarity of the text unit.
[0100] Step S1321: dividing the ambient light intensity into multiple intervals, such as a low light interval, a medium light interval, and a high light interval.
[0101] According to the ambient light intensity value collected by the light sensor, multiple intervals are divided. The low light interval corresponds to a dark environment, such as a night indoor without light; the medium light interval corresponds to a general indoor light environment; and the high light interval corresponds to a bright outdoor environment or a strong light irradiation environment. The light intensity range of different intervals is set according to the actual application scenario and the sensitivity of the device sensor.
[0102] Step S1322: presetting a corresponding visual comfort parameter reference value for each light interval, including an image brightness reference value, an image contrast reference value, and a text font clarity reference value.
[0103] The human eye perceives image brightness, contrast, and text font clarity differently under different lighting environments. In the low light interval, to avoid the screen being too bright and stimulating the eyes, the image brightness reference value and the contrast reference value are set low, and the text font clarity reference value is set high to ensure readability; in the high light interval, to improve the visibility of images and text, the image brightness reference value and the contrast reference value are set high, and the text font clarity reference value is also increased accordingly; the reference values in the medium light interval are between the two.
[0104] Step S1323: Real-time monitoring of changes in ambient light intensity, when the light intensity switches between different intervals, the visual comfort parameters are adjusted to the reference values corresponding to the interval.
[0105] The light intensity is monitored in real time by a light sensor, and when it is detected that the light intensity has changed from one interval to another, such as from the medium light interval to the high light interval, the image brightness reference value and the contrast reference value in the visual comfort parameters are immediately adjusted to the reference values corresponding to the high light interval, and the text font clarity reference value is also adjusted to the reference value of the high light interval, to adapt to the new lighting environment and ensure the comfort of the user's viewing.
[0106] Step S133: Monitoring the network transmission rate in the real-time display environment parameter, determining the content loading priority parameter according to the transmission rate fluctuation, the content loading priority parameter is used to sort the loading order of the graphic and text units.
[0107] Step S1331: Set multiple levels of network transmission rate, such as low speed level, medium speed level, and high speed level, each level corresponding to a different rate range.
[0108] According to the common network transmission rate, set the low speed level (such as download speed lower than a certain value), medium speed level (download speed within a certain value range), and high speed level (download speed higher than a certain value), the above rate range is set considering the general transmission capacity of different network types (such as 2G, 3G, 4G, 5G, Wi-Fi).
[0109] Step S1332: Analyze the loading performance of graphic and text units under different network transmission rate levels, under the low speed level, large capacity image units load slowly; under the high speed level, all graphic and text units can be loaded quickly.
[0110] In the low speed level network environment, image units have large data volume and long loading time, which may cause the user to wait for a long time; text units have small data volume and load quickly. Under the medium speed level, most graphic and text units can be loaded within an acceptable time. Under the high speed level, both image and text units can be loaded quickly with almost no delay.
[0111] Step S1333: Determine the content loading priority parameter based on the importance of the graphic-text unit, the data size, and the network transmission rate level.
[0112] The importance of the graphic-text unit is measured by the coordination weight (the coordination weight in step S128), and the graphic-text unit with a large coordination weight has high importance. In terms of data size, the data size of the image unit is usually larger than that of the text unit. In combination with the network transmission rate level, at a low speed level, the text unit with high importance and small data size is preferentially loaded, then the image unit with high importance but large data size is loaded, and finally the graphic-text unit with low importance is loaded; at a high speed level, multiple graphic-text units can be loaded simultaneously, and the priority is mainly sorted according to the importance. The content loading priority parameter is represented by a numerical value, and the higher the numerical value, the higher the loading priority.
[0113] Step S134: Input the layout space elasticity coefficient, the visual comfort parameter, and the content loading priority parameter into the environment parameter mapping module to establish a dynamic mapping relationship between the environment parameters and the layout adjustment parameters.
[0114] The environment parameter mapping module internally includes a mapping rule library, which defines the corresponding relationship between the layout space elasticity coefficient, the visual comfort parameter, the content loading priority parameter, and the layout adjustment parameters (such as the size, position, brightness, contrast, and loading order of the graphic-text unit). When the layout space elasticity coefficient is input, the module determines the adjustable range of the size of the graphic-text unit according to the coefficient size; when the visual comfort parameter is input, the adjustment target values of the image brightness, contrast, and text font clarity are determined; and when the content loading priority parameter is input, the loading order of the graphic-text unit is determined. Through these mapping relationships, the changes in the environment parameters can be reflected in real time on the layout adjustment parameters, realizing the dynamic influence of the environment parameters on the layout.
[0115] Step S135: Decode the interaction behavior sequence in the audience behavior intention data, extract the interaction features of the interaction behavior, and analyze the audience intention corresponding to the interaction features. The interaction features include frequency, time length, operation path, and operation force.
[0116] Step S1351: Perform type identification processing on each interaction operation in the interaction behavior sequence to distinguish different interaction types such as click operation, sliding operation, scaling operation, and long press operation.
[0117] The interactive behavior sequence is a continuous operation record of the user in a period of time, and each operation has a corresponding operation type identifier. By analyzing the triggering mode and device feedback of the operation, different interaction types such as click operation (user's finger touches the screen and then lifts up), sliding operation (user's finger moves a distance on the screen), zoom operation (double finger opens or pinches on the screen), and long press operation (user's finger keeps contact at the same position on the screen for a period of time) are identified.
[0118] Step S1352: Count the number of occurrences of each interaction type in a preset time window to obtain the interaction behavior frequency.
[0119] A time window is preset, such as the past 30 seconds or 1 minute, and the number of occurrences of click operation, sliding operation, zoom operation, and long press operation in the time window, i.e. the interaction behavior frequency, is counted. For example, in 30 seconds, the user performs 3 click operations on the "signature hot dishes" category, 5 sliding operations on the dish list, 1 zoom operation on a dish image, and no long press operation, so the click frequency is 3 times / 30 seconds, the sliding frequency is 5 times / 30 seconds, the zoom frequency is 1 time / 30 seconds, and the long press frequency is 0 times / 30 seconds.
[0120] Step S1353: Record the duration of each interaction operation from start to end to obtain the interaction behavior duration.
[0121] For click operation, the duration is from the time when the finger contacts the screen to the time when it leaves the screen; for sliding operation, the duration is from the time when the finger starts to move to the time when it stops moving and leaves the screen; for zoom operation, the duration is from the time when the double finger starts to contact the screen to the time when the zooming action is completed and the screen is left; for long press operation, the duration is the total time from the time when the finger contacts the screen to the time when it leaves the screen. The duration of each interaction operation is recorded to obtain the distribution of the interaction behavior duration.
[0122] Step S1354: Track the position coordinate changes of consecutive interaction operations to form an interaction operation path and analyze the straightness, tortuosity, and coverage range of the operation path.
[0123] The starting position coordinates and ending position coordinates of each interaction operation are recorded by the touch sensor of the device, and for sliding operation, the changes of intermediate position coordinates are also recorded. The position coordinates of consecutive interaction operations are connected in time sequence to form an interaction operation path. The straightness is measured by calculating the deviation degree of the actual trajectory in the path from the straight trajectory, the smaller the deviation, the higher the straightness; the tortuosity is measured by calculating the number of direction changes and the angle in the path, the more the number of direction changes and the larger the angle, the higher the tortuosity; the coverage range is determined by calculating the area of the screen region involved in the path, the larger the area, the wider the coverage range.
[0124] Step S1355: For the device supporting pressure sensing, collect the pressure data in the interactive operation process to determine the operation strength. For the device not supporting pressure sensing, indirectly infer the operation strength based on the interaction time length and operation speed.
[0125] The device supporting pressure sensing can directly collect the pressure value of the user's finger pressing the screen. The greater the pressure value, the greater the operation strength. The device not supporting pressure sensing infers the operation strength based on the interaction time length and operation speed. Generally, the operation with shorter interaction time length and faster operation speed may correspond to greater operation strength; the operation with longer interaction time length and slower operation speed may correspond to smaller operation strength. For example, the fast sliding operation may have greater strength than the slow sliding operation.
[0126] Step S1356: Analyze the combination relationship among the interactive behavior frequency, interactive behavior time length, operation path characteristics, and operation strength. When the sliding operation frequency exceeds the preset frequency threshold, the single sliding time length is lower than the preset time length threshold, the straightness of the operation path is higher than the preset straightness threshold, and the coverage range of the operation path is greater than the preset range threshold, it is determined that the corresponding fast browsing intention.
[0127] The preset sliding operation frequency threshold, single sliding time length threshold, operation path straightness threshold, and coverage range threshold. When the sliding operation frequency of the user exceeds the preset frequency threshold, it means that the user has performed multiple sliding operations in a short time; when the single sliding time length is lower than the preset time length threshold, it means that the sliding speed is fast each time; when the straightness of the operation path is higher than the preset threshold, it means that the sliding direction is relatively fixed; and when the coverage range of the operation path is greater than the preset threshold, it means that the user browses a wide range of content. The combination of the above characteristics indicates that the user is quickly browsing the menu content and does not in-depth view a certain dish, so it is determined as a fast browsing intention.
[0128] Step S1357: The click operation frequency is lower than the preset frequency standard, the single click after the stay time length exceeds the preset stay time length standard, the zoom operation magnification operation proportion exceeds the preset magnification proportion standard, and the operation path is concentrated in the area meeting the preset local range standard, corresponding to the detailed viewing intention.
[0129] The preset click operation frequency standard, single click after the stay time length standard, zoom operation magnification proportion standard, and operation path local range standard. The low click operation frequency indicates that the user does not frequently switch dishes; the long stay time length after single click indicates that the user stays for a long time after clicking a certain dish to view the details; the high magnification operation proportion in the zoom operation indicates that the user tends to magnify to view the details of the dish image; and the operation path is concentrated in a local area, indicating that the user focuses on one or a few dishes. The combination of the above characteristics corresponds to the detailed viewing intention, and the user wants to in-depth understand the information of a specific dish.
[0130] Step S1358: If the click operation frequency exceeds the preset frequency standard, the operation path dispersion degree meets the preset dispersion standard, the short distance sliding proportion in the sliding operation exceeds the preset short sliding proportion standard, and the operation force fluctuation amplitude exceeds the preset force fluctuation standard, a selection comparison intention is selected.
[0131] The preset click operation frequency standard, operation path dispersion degree standard, sliding operation short distance proportion standard, and operation force fluctuation amplitude standard. The high click operation frequency indicates that the user frequently clicks different dishes; the operation path dispersion indicates that the user operates in multiple areas of the menu; the high short distance sliding proportion indicates that the user switches dishes in a small range for comparison; and the large operation force fluctuation amplitude indicates that the user changes the operation force greatly between different dishes, which may show hesitation and comparison mentality. The combination of the above characteristics corresponds to the selection comparison intention, and the user compares between multiple dishes to make a selection.
[0132] Step S1359: Verify the correspondence between different interaction feature combinations and audience intentions, and correct the misjudged intention classification results.
[0133] A large number of user interaction behavior samples are collected, and the above samples are manually annotated to determine the actual audience intention. The intention classification results obtained through steps S1356 to S1358 are compared with the manual annotation results, and the classification accuracy is calculated. For misjudged samples, analyze the deviation of the interaction feature combination and the classification rule, adjust the preset threshold and feature weight, optimize the classification rule, and thus correct the misjudged intention classification results, and improve the accuracy of audience intention recognition.
[0134] Step S136: Input the audience intention classification results and corresponding interaction features into the behavior intention decoding module to establish the correspondence between the interaction behavior sequence and the layout strategy, and different audience intentions correspond to different graphic-text layout strategies.
[0135] The behavior intention decoding module internally stores a mapping table of audience intentions and layout strategies. For the quick browsing intention, the layout strategy should adopt a concise and clear layout, highlight the key dishes, the graphic-text unit size is moderate, reduce redundant information, and speed up the content loading speed; for the detailed viewing intention, the layout strategy should increase the graphic-text unit size of the currently viewed dish to display more detailed information, such as material close-up images, detailed taste descriptions, etc.; for the selection comparison intention, the layout strategy should display the dishes that the user may compare side by side, which facilitates user comparison, for example, arranging the images and key information of multiple dishes horizontally or vertically. When the audience intention classification results and corresponding interaction features are input, the module selects the corresponding layout strategy according to the mapping table, thereby establishing the correspondence between the interaction behavior sequence and the layout strategy.
[0136] Step S137: input the time sequence mapping rule, sentiment matching parameter and coordination weight in the graphic-text coordination model into the coordination relationship application module, establish the association logic of coordination relationship and layout parameters, and realize the effective application of space-time association and sentiment coordination in the layout process.
[0137] The coordination relationship application module receives the time sequence mapping rule, sentiment matching parameter and coordination weight in the graphic-text coordination model. The time sequence mapping rule guides how to arrange the display time sequence and duration of text units and image units during layout; the sentiment matching parameter guides how to adjust the sentiment expression of graphic-text units to meet the audience preference; and the coordination weight determines the priority processing order of different graphic-text units in the case of layout conflict or limited resources. The module establishes the association logic of these coordination relationships and layout parameters (such as position, size, display duration, brightness, contrast, etc.) inside, for example, according to the time sequence mapping rule to determine the arrangement order and display duration parameters of graphic-text units, according to the sentiment matching parameter to adjust the brightness and contrast of images and the font color of text to enhance the sentiment expression, and according to the coordination weight to prioritize the placement of graphic-text units with high coordination weight when the layout space is limited.
[0138] Step S138: construct a layout parameter evolution module, which includes a conflict pre-play algorithm, an adaptive adjustment rule and an evolution iteration mechanism, and can dynamically update the layout parameters based on the output of the environment parameter mapping module, the behavior intention decoding module and the coordination relationship application module.
[0139] Step S1381: design a conflict pre-play algorithm that can simulate the layout effect under different display environment changes (such as screen direction rotation, light intensity change, network rate fluctuation) and different audience interaction behaviors (such as clicking, sliding, zooming).
[0140] The conflict pre-play algorithm constructs a virtual layout environment based on the current layout parameters and the output of the environment parameter mapping module, the behavior intention decoding module and the coordination relationship application module. When simulating screen direction rotation, the algorithm adjusts the size and position of graphic-text units according to the layout space elasticity coefficient; when simulating light intensity change, it adjusts the image brightness and contrast and the text font clarity according to the visual comfort parameter; when simulating network rate fluctuation, it adjusts the loading order and loading method of graphic-text units according to the content loading priority parameter. At the same time, it simulates the influence of different audience interaction behaviors on layout, such as adjusting the size and position of the graphic-text unit of a dish to highlight it when the user clicks on it. Through these simulations, the layout effect and possible conflicts are predicted.
[0141] Step S1382: develop an adaptive adjustment rule. When a layout conflict is detected during the pre-play process, automatically select the corresponding adjustment strategy according to the conflict type and severity, such as adjusting the position, size, display duration, loading order, etc. of graphic-text units.
[0142] Specific adaptive adjustment rules are formulated for different types of layout conflicts (such as display area overlap, loading delay imbalance, etc.). For example, for display area overlap conflicts, the rules stipulate adjusting the positions of the graphic-text units according to the levels of the coordination weights, with the graphic-text units having higher coordination weights being given priority in retaining their positions; for loading delay imbalance conflicts, the rules stipulate reordering the loading sequence according to the content loading priority parameters. At the same time, the priority and amplitude of the adjustments are determined according to the severity of the conflicts, with severe conflicts being given priority and the adjustments having a larger amplitude.
[0143] Step S1383: An evolutionary iteration mechanism is established, through which the conflict pre-play algorithm and the adaptive adjustment rules are executed multiple times to continuously optimize the layout parameters until the layout effect meets the preset optimization target.
[0144] The evolutionary iteration mechanism sets an upper limit on the number of iterations and a threshold for the optimization target. In each iteration process, the conflict pre-play algorithm is first executed to detect layout conflicts; then the layout parameters are adjusted according to the adaptive adjustment rules; and then the conflict pre-play algorithm is executed again to check whether the adjusted layout effect still has conflicts. This process is repeated until the layout effect meets the optimization target (such as having no conflicts or the conflicts being within an acceptable range) or the upper limit on the number of iterations is reached. In each iteration, the adjustment strategy is optimized based on the pre-play results of the previous iteration, causing the layout parameters to gradually evolve towards the optimal direction.
[0145] Step S139: The environment parameter mapping module, the behavior intention decoding module, the coordination relationship application module, and the layout parameter evolution module are connected through data interaction interfaces, the data transmission timing and triggering conditions between the modules are set, and a dynamic self-adaptive layout engine is formed.
[0146] A unified data interaction interface is designed for the four modules, the data format and communication protocol are defined, and the correct transmission of data between the modules is ensured. The data transmission timing between the modules is set, for example, the environment parameter mapping module converts the environment parameters into layout adjustment parameters in real time and transmits them to the coordination relationship application module and the layout parameter evolution module; the behavior intention decoding module transmits the recognized audience intentions to the coordination relationship application module and the layout parameter evolution module at regular intervals (such as every few seconds); the coordination relationship application module generates preliminary layout parameters based on the received environment layout adjustment parameters and audience intentions and transmits them to the layout parameter evolution module. The triggering conditions include triggering the environment parameter mapping module to update data when the environment parameters change by more than a preset threshold, triggering the behavior intention decoding module to update audience intentions when the interaction behavior sequence accumulates to a certain number, triggering the layout parameter evolution module to optimize the preliminary layout parameters after they are generated, etc. Through these connections and settings, the four modules work together to form a dynamic self-adaptive layout engine.
[0147] Step S140: applying the dynamic self-adaptive layout engine to dynamically calculate the layout parameters of the image units and text units in the to-be-optimized digital multimedia text-image set, to generate an initial layout scheme, and performing multi-dimensional conflict pre-performance and self-adaptive adjustment on the initial layout scheme by a layout parameter evolution module of the dynamic self-adaptive layout engine to obtain an intermediate layout scheme.
[0148] Step S141: inputting the attribute information of the image units and text units in the to-be-optimized digital multimedia text-image set into a cooperative relationship application module of the dynamic self-adaptive layout engine, determining the basic layout parameters of each text-image unit by combining the timing mapping rules and emotional matching parameters in the text-image cooperative model, wherein the basic layout parameters include initial position, initial size, and initial display duration.
[0149] The attribute information of the image units includes resolution, size, format, and data volume of the image; and the attribute information of the text units includes text length, character number, and number of emotional words contained. The cooperative relationship application module determines the initial display duration of the text units and image units according to the attribute information and the timing mapping rules in the text-image cooperative model, so as to match the text reading duration with the image display duration; adjusts the initial size of the text-image units according to the emotional matching parameters, and the initial size of the text-image units with high emotional feature values and meeting the audience preferences is larger. The determination of the initial position is based on the dish classification and the default layout order, for example, the text-image units under the “signature hot dishes” classification are arranged from top to bottom and from left to right according to the default order of the dishes in the menu. Through these processes, the basic layout parameters of each text-image unit are obtained.
[0150] Step S142: inputting the real-time display environment parameters into an environment parameter mapping module to obtain a layout space elasticity coefficient, a visual comfort parameter, and a content loading priority parameter, and performing environment adaptation adjustment on the basic layout parameters based on the layout space elasticity coefficient, the visual comfort parameter, and the content loading priority parameter to obtain environment adaptation layout parameters.
[0151] The environment parameter mapping module calculates the layout space elasticity coefficient, the visual comfort parameter, and the content loading priority parameter according to the input real-time display environment parameters. According to the layout space elasticity coefficient, the initial size of the text-image units is adjusted, and if the elasticity coefficient is large, the size of the text-image units can be appropriately increased or decreased to adapt to the layout space; according to the visual comfort parameter, the initial brightness and contrast of the image units and the font size and definition of the text units are adjusted; and according to the content loading priority parameter, the initial loading order of the text-image units is adjusted, and the text-image units with high priority are loaded first. The above adjustments are applied to the basic layout parameters to obtain the environment adaptation layout parameters.
[0152] Step S143: input the interactive behavior sequence in the audience behavior intention data into the behavior intention decoding module, determine the audience intention classification result, perform strategy adaptation adjustment on the environment adaptation layout parameter based on the audience intention classification result, and obtain the strategy adaptation layout parameter.
[0153] The behavior intention decoding module decodes the input interactive behavior sequence to determine the audience intention classification result (such as quick browsing, detailed viewing, and selection comparison). For the quick browsing intention, the strategy adaptation adjustment reduces the size of the image-text unit, increases the layout density, and accelerates the content loading; for the detailed viewing intention, the size of the currently viewed image-text unit is increased to highlight the details; for the selection comparison intention, the image-text units that can be compared are adjusted to adjacent positions to facilitate comparison. The above strategy adjustment is applied to the environment adaptation layout parameter to obtain the strategy adaptation layout parameter.
[0154] Step S144: integrate the strategy adaptation layout parameters of all image-text units to determine the final position coordinates, display size, display duration, and loading order of each image-text unit, and generate an initial layout scheme.
[0155] The strategy adaptation layout parameters of all image-text units are summarized. For the position coordinates, it is necessary to ensure that the image-text units do not overlap (initial judgment) and meet the layout space restrictions; the display size is determined according to the adjusted parameters; the display duration is set according to the time sequence mapping rule; and the loading order is arranged according to the content loading priority parameter. The above information is organized into a structured data format, including the unique identifier of each image-text unit and the corresponding final position coordinates, display size, display duration, and loading order, to form an initial layout scheme.
[0156] Step S145: input the initial layout scheme into the layout parameter evolution module, start the conflict pre-play algorithm, and simulate the layout effect under different display environment changes and different audience interactive behaviors.
[0157] After receiving the initial layout scheme, the layout parameter evolution module starts the conflict pre-play algorithm. The algorithm simulates various possible display environment changes, such as screen orientation from portrait to landscape, environment light intensity from low light to high light, network transmission rate from high speed to low speed, etc.; and simulates different audience interactive behaviors, such as user clicking on a dish, sliding the menu, zooming the image, etc. In each simulation scenario, the algorithm calculates the display effect of the image-text unit based on the initial layout scheme, and records relevant data such as the position, size, loading time, and visual comfort parameters of each image-text unit.
[0158] Step S146: based on the pre-play results, identify possible layout conflicts, including display area overlap, loading delay imbalance, visual comfort decline, and coordination relationship rupture.
[0159] For example, step S1461: Extract the display area coordinates of all graphic-text units in different simulation scenarios from the rehearsal result. Compare the display area coordinates of any two graphic-text units. If there is a coordinate intersection, and the area of the intersection accounts for more than a preset threshold proportion of the area of either graphic-text unit, it is determined that there is a display area overlap conflict.
[0160] The display area coordinates are represented by the top-left and bottom-right pixel coordinates of the graphic-text unit, forming a rectangular area. For any two graphic-text units, calculate the intersection of their display area coordinates. The area of the intersection is calculated by multiplying the width and height of the rectangular intersection. Divide the intersection area by the area of the smaller of the two graphic-text units to obtain the intersection area proportion. If the proportion exceeds the preset threshold, it indicates that the two graphic-text units overlap to a high degree, affecting user viewing, and it is determined that there is a display area overlap conflict.
[0161] Step S1462: Extract the loading completion time of each graphic-text unit in the rehearsal result. Calculate the difference between the maximum and minimum values of all graphic-text unit loading completion times. If the difference exceeds a preset time difference threshold, it is determined that there is a loading delay imbalance conflict.
[0162] The loading completion time refers to the time required from starting to load a graphic-text unit to the unit being fully displayed on the screen. In the rehearsal result, the loading completion time of each graphic-text unit is recorded. Find the maximum and minimum values of all loading completion times and calculate their difference. If the difference exceeds the preset time difference threshold, it indicates that the loading speed of different graphic-text units differs too much, and some graphic-text units are still loading while others have completed loading, resulting in inconsistent user experience, and it is determined that there is a loading delay imbalance conflict.
[0163] Step S1463: According to the visual comfort parameter change curve in the rehearsal result, after the environmental light intensity changes, the brightness and contrast of the image unit are not adjusted in a timely manner, resulting in a visual comfort parameter below a preset comfort threshold, or the font clarity of the text unit below a preset clarity threshold, determining a visual comfort decline conflict.
[0164] The visual comfort parameter change curve records the changes in visual comfort parameters during changes in environmental light intensity. When the environmental light intensity changes, the brightness and contrast of the image unit should be adjusted accordingly to maintain the visual comfort parameter above the preset comfort threshold; the font clarity of the text unit should also be maintained above the preset clarity threshold. If, in the rehearsal result, after the light intensity changes, the visual comfort parameter is below the preset threshold, or the font clarity is below the preset threshold, it indicates that the layout scheme has failed to adapt to the light changes in a timely manner, and it is determined that there is a visual comfort decline conflict.
[0165] Step S1464: Check the spatio-temporal correlation relationship and emotional coordination relationship of the text and image units in the preview result. If the corresponding relationship deviation between the text reading time length segmentation and the image optimal display time length caused by screen rotation exceeds the preset corresponding deviation threshold, or the deviation between the text emotional feature value and the image emotional feature value caused by environmental parameter changes exceeds the preset emotional deviation threshold, it is determined that the coordination relationship is broken.
[0166] Screen rotation will change the layout space, which may cause the corresponding relationship between the text reading time length segmentation and the image optimal display time length to change. Calculate the corresponding relationship deviation after rotation. If it exceeds the preset corresponding deviation threshold, it means that the spatio-temporal correlation relationship is destroyed. Environmental parameter changes (such as light changes affecting image color) may cause the deviation between the text emotional feature value and the image emotional feature value to increase. If it exceeds the preset emotional deviation threshold, it means that the emotional coordination relationship is destroyed. Both of these two cases are determined as coordination relationship broken conflict.
[0167] Step S1465: Establish a layout conflict list. Sort all identified layout conflicts by severity. Conflicts with severity higher than the preset high priority threshold are prioritized.
[0168] According to the influence degree of layout conflict on user experience, set a severity scoring standard for each conflict type. For example, the severity of the display area overlap conflict is positively correlated with the overlap area ratio, and the severity of the loading delay imbalance conflict is positively correlated with the time difference. Score each identified layout conflict, sort the scoring results from high to low, and establish a layout conflict list. A high priority threshold is preset. Conflicts with scores higher than the threshold are marked as high priority conflicts and are prioritized in subsequent adjustment.
[0169] Step S147: According to the adaptive adjustment rule, adjust the layout parameters that have conflict risks. The adjustment operation includes adjusting the position coordinates of overlapping units, optimizing the loading order to balance the delay, adjusting the brightness contrast to improve comfort, and correcting the timing mapping rule to repair the coordination relationship.
[0170] Step S1471: For display area overlap conflicts, query the coordination weight and audience intention association degree of the text and image units involved in the conflict. Text and image units with coordination weight greater than the preset high weight threshold and audience intention association degree greater than the preset high association degree threshold are prioritized to retain the original position coordinates. Text and image units with coordination weight less than the preset low weight threshold and audience intention association degree less than the preset low association degree threshold are moved away from the conflict area. The moving distance is positively correlated with the size of the conflict intersection area, and the moved position meets the layout space elasticity coefficient constraint.
[0171] The coordination weight is obtained from the image-text coordination model, and the audience intention correlation degree is obtained by analyzing the matching degree of the image-text unit and the current audience intention. For the image-text units involved in the conflict, the coordination weight and the audience intention correlation degree are compared, and the image-text units with high weight and high correlation degree are kept in the original position. For the image-text units with low weight and low correlation degree, the moving distance is determined according to the size of the conflict intersection area, and the larger the intersection area, the farther the moving distance. The moving direction is selected to be away from the conflict area, such as moving to the left if the conflict area is on the right, or moving down if the conflict area is on the top. The position after moving cannot exceed the size adaptation range determined by the layout space elasticity coefficient.
[0172] Step S1472: For the loading delay imbalance conflict, the loading order of the image-text units is reordered according to the content loading priority parameter and the display duration of the image-text units. The image-text units with a display duration exceeding a preset display duration threshold and an audience intention correlation degree greater than a preset high correlation degree threshold are loaded preferentially. For the image-text units with a loading speed lower than a preset loading speed threshold, the size parameter is adjusted to reduce the data volume, or a progressive loading method is used, i.e., first loading low-resolution image data to meet the basic display requirements, and then loading high-resolution image data to balance the loading delay.
[0173] The content loading priority parameter and the audience intention correlation degree determine the loading priority of the image-text units. The image-text units with long display duration and high correlation degree should be loaded preferentially. For the image-text units with slow loading speed, if they are image units, their size can be appropriately reduced (within the range allowed by the layout space elasticity coefficient) to reduce the image data volume and improve the loading speed; or a progressive loading method can be used. For text units, the initial loaded content can be simplified, i.e., first loading the core information (such as dish name and price), and then loading the detailed description. Through these methods, the loading completion time of different image-text units is balanced, the time difference is reduced, and the loading delay imbalance conflict is solved.
[0174] Step S1473: For the visual comfort decrease conflict, the image unit brightness and contrast parameters are adjusted to reach the baseline values in the visual comfort parameter, i.e., the image brightness and contrast baseline values. The font size, character spacing, and line spacing of the text unit are adjusted to ensure that the font clarity is not lower than the preset clarity threshold under different lighting conditions.
[0175] According to the visual comfort parameter baseline value corresponding to the ambient light intensity, the brightness and contrast of the image unit are adjusted. For example, in a high-light environment, the image brightness and contrast are increased; in a low-light environment, the image brightness and contrast are decreased. For the text unit, the font clarity is improved by increasing the font size and adjusting the character spacing and line spacing, so that the user can clearly read the text content under various lighting conditions, and the visual comfort parameter is restored to above the preset threshold.
[0176] Step S1474: To address the conflict in the collaborative relationship, recalculate the temporal matching relationship between the spatiotemporal features of the text and the spatiotemporal features of the image, correct the temporal mapping rules based on the current display environment parameters, and adjust the deviation between the reading time segment and the optimal display time to no more than the preset corresponding deviation threshold; adjust the emotional feature values of the image unit and the text unit, adjust the emotional deviation between the two to no more than the preset emotional deviation threshold, and repair the emotional collaborative relationship.
[0177] Based on the current display environment parameters (such as screen orientation and size), the text reading time segments and optimal image display time are recalculated, and the correspondence in the time-series mapping rules is corrected to keep the deviation within a preset range. For emotional synergy, the deviation between the emotional feature values of the text and images and the audience's emotional preference feature values is analyzed, and the emotional words in the text or the color and composition of the images are adjusted to reduce the deviation of the emotional feature values to below a preset emotional deviation threshold, thereby repairing the synergy.
[0178] Step S148: Repeat the conflict pre-simulation algorithm and adaptive adjustment rules until there are no layout conflicts in the pre-simulation results, record the adjusted layout parameters, and generate an intermediate layout scheme.
[0179] After completing one round of layout parameter adjustments, the conflict pre-simulation algorithm is executed again to check for any new layout conflicts. If conflicts still exist, the adaptive adjustment rules are applied, and this process is repeated until no layout conflicts appear in the pre-simulation results, or the severity of the conflicts is below the preset acceptable threshold. At this point, all adjusted layout parameters are recorded, including the position coordinates, size, display duration, loading order, brightness, and contrast of the text and image units. These parameters are then organized into structured data to generate an intermediate layout scheme.
[0180] Step S150: Deploy the intermediate layout scheme to the real-time display environment, combine the content demand description in the audience behavior intent data and the real-time updated display environment parameters, drive the dynamic self-adaptive layout engine to continuously evolve the layout scheme, generate a target graphic and text layout optimization scheme that conforms to the real-time environment and audience intent, and output the target graphic and text layout optimization scheme for digital multimedia content display.
[0181] Step S151: Deploy the intermediate layout scheme to the digital multimedia real-time display environment and establish a real-time data connection between the layout scheme and the display environment parameter acquisition module and the audience behavior intent acquisition module.
[0182] The intermediate layout scheme is deployed to the mobile terminal device of the user through the digital multimedia content management system, so that the user can display it in the menu interface of the catering APP. At the same time, a data connection channel is established between the layout scheme and the display environment parameter acquisition module and the audience behavior intention acquisition module in the APP, so as to ensure that the layout scheme can receive data updates from the two modules in real time.
[0183] Step S152: Through the display environment parameter acquisition module, the device screen characteristic change, the environmental light intensity fluctuation, and the network transmission rate change are collected in real time to form a real-time updated display environment parameter stream.
[0184] The display environment parameter acquisition module continuously monitors the screen characteristics of the device (such as sudden rotation of the screen direction), the environmental light intensity (such as the user walking from indoors to outdoors causing the light intensity to suddenly increase), and the network transmission rate (such as switching from Wi-Fi to 4G causing the rate to change), and organizes the above change data in chronological order into a real-time updated display environment parameter stream, which is continuously sent to the dynamic self-adaptive layout engine.
[0185] Step S153: Through the audience behavior intention acquisition module, the new interactive behavior sequence of the audience in browsing the intermediate layout scheme is collected in real time, and the content demand expressions input by the audience are collected to form a real-time updated audience behavior intention data stream.
[0186] The audience behavior intention acquisition module records the new interactive operations of the user in browsing the intermediate layout scheme, such as new clicks, swipes, zooms, etc., to form a new interactive behavior sequence. At the same time, through the search box, filtering condition setting, etc. in the APP, the content demand expressions input by the user are collected, such as "light dishes" "discount activities" etc. The above new interactive behavior sequence and content demand expressions are organized into a real-time updated audience behavior intention data stream, which is sent to the dynamic self-adaptive layout engine.
[0187] Step S154: The real-time updated display environment parameter stream is input into the environment parameter mapping module of the dynamic self-adaptive layout engine, the dynamic mapping relationship between the environment parameters and the layout adjustment parameters is updated, and the new layout space elasticity coefficient, visual comfort parameter and content loading priority parameter are output.
[0188] The environment parameter mapping module receives the real-time updated display environment parameter stream, recalculates the layout space elasticity coefficient, visual comfort parameter and content loading priority parameter. For example, after the screen direction is rotated, the elasticity coefficient is recalculated; after the light intensity fluctuates, the baseline value of the visual comfort parameter is updated; after the network rate changes, the content loading priority is adjusted. The above new parameters are output to the collaborative relationship application module and the layout parameter evolution module.
[0189] Step S155: input the real-time updated audience behavior intention data stream into the behavior intention decoding module, update the correspondence between the interactive behavior sequence and the layout strategy, correct the audience intention classification result combined with the content demand expression, and output a new layout strategy.
[0190] The behavior intention decoding module decodes the new interactive behavior sequence, corrects the previous audience intention classification result combined with the content demand expression. For example, the user previously showed a comparison intention, and after inputting "light dishes", the intention may be corrected to a detailed viewing intention for light dishes. According to the corrected audience intention, the correspondence between the interactive behavior sequence and the layout strategy is updated, and a new layout strategy is output.
[0191] Step S156: The collaborative relationship application module adjusts the timing mapping rule, emotion matching parameter and collaboration weight in the image-text collaborative model based on the new layout strategy and the updated environment parameter, and outputs a new collaborative application logic.
[0192] The collaborative relationship application module adjusts the timing mapping rule in the image-text collaborative model according to the new layout strategy and the updated environment parameter to adapt to the new display environment and audience intention; updates the emotion matching parameter to make the emotion expression of the image-text unit more consistent with the user's current content demand expression; adjusts the collaboration weight to increase the collaboration weight of the image-text unit related to the content demand expression. Based on these adjustments, a new collaborative application logic is output to guide the calculation of layout parameters.
[0193] Step S157: The layout parameter evolution module updates the layout parameters of the intermediate layout scheme in real time based on the new environment parameter mapping relationship, layout strategy and collaborative application logic, adjusts the position coordinates, display size, display duration and loading order of the image-text unit.
[0194] The layout parameter evolution module updates the layout parameters of the intermediate layout scheme in real time based on the new environment parameter mapping relationship (layout space elasticity coefficient, visual comfort parameter, content loading priority parameter), new layout strategy and new collaborative application logic. Adjust the position coordinates of the image-text unit to adapt to the screen direction change, adjust the display size to meet the elasticity coefficient and layout strategy, adjust the display duration to match the new timing mapping rule, and adjust the loading order to reflect the new content loading priority.
[0195] Step S158: During the layout parameter updating process, the conflict pre-play algorithm is executed synchronously, and no new conflict is generated in the updated layout scheme.
[0196] While adjusting the layout parameters, the layout parameter evolution module synchronously executes the conflict pre-play algorithm, simulates the display effect of the updated layout scheme under the current environment and audience behavior, and checks whether a new layout conflict is generated. If there is a conflict, the adaptive adjustment rules are immediately used for re-adjustment until the updated layout scheme has no conflict.
[0197] Step S159: The real-time data collection, module parameter updating, layout parameter adjustment and conflict pre-play steps are continuously repeated until there is no new content demand expression in the audience behavior intention data stream, and the fluctuation ranges of each parameter in the real-time display environment parameter stream are continuously stable within the respective preset threshold range.
[0198] The dynamic adaptive layout engine continuously executes steps S152 to S158, continuously collects real-time data, updates module parameters, adjusts layout parameters, and performs conflict pre-play. When the audience does not input new content demand expressions for a period of time, and the fluctuation ranges of the display environment parameters (screen characteristics, light intensity, network rate) are within the preset threshold (such as the light intensity changes less than the preset value within a few minutes), it is indicated that the user demand and the display environment tend to be stable.
[0199] Step S1510: Determine the final stable layout scheme as the target graphic-text layout optimization scheme.
[0200] When the stable condition of step S159 is met, the dynamic adaptive layout engine stops adjusting the layout parameters, and determines the current layout scheme as the target graphic-text layout optimization scheme. The target graphic-text layout optimization scheme can adapt to the real-time display environment and audience intention, and achieve the best layout effect of the online menu graphic-text.
[0201] Step S1511: Output the target graphic-text layout optimization scheme for digital multimedia content display.
[0202] The target graphic-text layout optimization scheme is sent to the user's mobile terminal device through a data interface, and the catering APP displays the optimized online menu on the screen according to the scheme, including the position, size, brightness, contrast, loading order, etc. of the graphic-text unit, to provide a good menu browsing experience for the user.
[0203] Figure 2The application provides a combination of digital multimedia and intelligent optimization system for picture-text layout, which comprises a processor 1001, a memory 1003 and program codes stored in the memory 1003. The processor 1001 executes the program codes to realize the steps of the combination of digital multimedia and intelligent optimization method for picture-text layout. The processor 1001 and the memory 1003 are connected, for example, through a bus 1002. Optionally, the combination of digital multimedia and intelligent optimization system for picture-text layout can further comprise a transceiver 1004, which can be used for data interaction, such as data sending and / or data receiving, between the combination of digital multimedia and intelligent optimization system for picture-text layout and other combination of digital multimedia and intelligent optimization systems for picture-text layout. It should be noted that the transceiver 1004 is not limited to one in actual scheduling, and the structure of the combination of digital multimedia and intelligent optimization system for picture-text layout does not constitute a limitation to the embodiments of the application.
[0204] The memory 1003 is used for storing program codes for executing the embodiments of the application and is controlled by the processor 1001 to execute. The processor 1001 is used for executing the program codes stored in the memory 1003 to realize the steps shown in the foregoing method embodiments.
[0205] The embodiments of the application provide a computer readable storage medium, which stores program codes. When the program codes are executed by a processor, the steps and corresponding contents of the foregoing method embodiments can be realized.
[0206] The above description is only optional implementation of some implementation scenarios of the application. It should be noted that, for those skilled in the technical field, other similar implementation manners according to the technical concept of the application can be adopted without departing from the technical concept of the application, which also belongs to the protection scope of the embodiments of the application.
Claims
1. A method for intelligent optimization of text and image layout combining digital multimedia, characterized in that, The method includes: The system acquires a set of digital multimedia images and texts to be optimized, real-time display environment parameters, and audience behavioral intent data. The set of digital multimedia images and texts to be optimized includes multiple image units and multiple text units. The real-time display environment parameters include device screen characteristics, ambient light intensity, and network transmission rate. The audience behavioral intent data includes interactive behavior sequences, sentiment markers, and content demand expressions. The image units and text units in the digital multimedia image and text set to be optimized are modeled in a spatiotemporal manner. Combined with the sentiment tendency markers in the audience behavior intention data, the spatiotemporal correlation and sentiment synergy relationship between the image units and text units are established to obtain the image and text synergy model. Based on the real-time display environment parameters, the interaction behavior sequence in the audience behavior intent data, and the graphic-text collaboration model, a dynamic self-adaptive typesetting engine is constructed. The dynamic self-adaptive typesetting engine includes an environment parameter mapping module, a behavior intent decoding module, a collaboration relationship application module, and a typesetting parameter evolution module. The dynamic adaptive typesetting engine is used to dynamically calculate the typesetting parameters of the image units and text units in the digital multimedia graphic and text collection to be optimized, and generate an initial typesetting scheme. The typesetting parameter evolution module of the dynamic adaptive typesetting engine is used to perform multi-dimensional conflict pre-simulation and adaptive adjustment on the initial typesetting scheme to obtain an intermediate typesetting scheme. The intermediate layout scheme is deployed to the real-time display environment. Combining the content demand description in the audience behavior intent data and the real-time updated display environment parameters, the dynamic self-adaptive layout engine is driven to continuously evolve the layout scheme, generate a target graphic and text layout optimization scheme that conforms to the real-time environment and audience intent, and output the target graphic and text layout optimization scheme for digital multimedia content display.
2. The intelligent optimization method for graphic and text layout combining digital multimedia as described in claim 1, characterized in that, The step involves performing spatiotemporal collaborative modeling of image and text units in the digital multimedia image and text set to be optimized. This is combined with sentiment markers from the audience's behavioral intent data to establish spatiotemporal correlations and sentiment collaborations between image and text units, resulting in an image-text collaboration model. The model includes: Each text unit is analyzed for reading rhythm, extracting sentence pause markers, changes in word density, and semantic transition positions to determine the reading time segments and key semantic intervals of the text unit, thus forming the spatiotemporal characteristics of the text. Visual presentation timing analysis is performed on each image unit to extract the visual focus switching path, color transition rhythm and detail information hierarchy in the image unit, determine the optimal display duration and visual attention order of the image unit, and form the spatiotemporal characteristics of the image. The spatiotemporal features of text and images are matched in a temporal sequence. Based on the correspondence between reading time segments and optimal display time, and the correlation between key semantic intervals and visual attention order, the spatiotemporal relationship between image units and text units is established. Sentiment extraction is performed on each text unit to analyze the distribution of sentiment words, tone expression, and semantic sentiment intensity in the text unit, and to determine the sentiment feature value of the text unit. For each image unit, sentiment extraction processing is performed to analyze the color sentiment attributes, compositional sentiment expression, and visual element sentiment symbolism in the image unit, and to determine the sentiment feature value of the image unit. From the sentiment tendency tags of the audience's behavioral intention data, the sentiment preference feature values of the audience for similar images and texts are extracted as a reference benchmark for sentiment coordination; The emotional feature values of text units, image units, and audience emotional preference feature values are collaboratively matched to adjust the deviation between the emotional feature values of image units and text units, and to establish an emotional synergy relationship. By integrating spatiotemporal relationships and emotional collaboration, a text-image collaboration model is constructed, which includes temporal mapping rules, emotional matching parameters, and collaboration weights.
3. The intelligent optimization method for graphic and text layout combining digital multimedia as described in claim 1, characterized in that, The construction of a dynamic self-adaptive typesetting engine based on the real-time display environment parameters, the interaction behavior sequence in the audience behavior intent data, and the graphic-text collaboration model includes: The device screen characteristics in the real-time display environment parameters are analyzed and processed to extract the screen size specifications, resolution parameters, screen orientation and pixel density. The initial layout space is calculated based on the screen size specifications, resolution parameters and pixel density, and the layout space elasticity coefficient is determined in combination with the screen orientation. The layout space elasticity coefficient is used to dynamically adjust the size adaptation range of the graphic unit. The ambient light intensity in the real-time display environment parameters is detected and processed, and the visual comfort parameter is determined based on the change in light intensity. The visual comfort parameter is used to adjust the brightness and contrast of the image unit and the font clarity of the text unit. The network transmission rate in the real-time display environment parameters is monitored and processed, and the content loading priority parameter is determined based on the transmission rate fluctuation. The content loading priority parameter is used to sort the loading order of graphic units. Input the layout space flexibility coefficient, visual comfort parameter, and content loading priority parameter into the environment parameter mapping module to establish a dynamic mapping relationship between environment parameters and layout adjustment parameters; The interaction sequence in the audience behavior intent data is decoded, the interaction features of the interaction are extracted, and the audience intent corresponding to the interaction features is analyzed. The interaction features include frequency, duration, operation path and operation intensity. The audience intent classification results and corresponding interaction features are input into the behavior intent decoding module to establish the correspondence between the interaction behavior sequence and the layout strategy. Different audience intents correspond to different graphic layout strategies. The temporal mapping rules, sentiment matching parameters, and collaborative weights in the graphic-text collaboration model are input into the collaborative relationship application module to establish the association logic between collaborative relationships and layout parameters, thereby realizing the effective application of spatiotemporal association and sentiment collaboration during the layout process. A layout parameter evolution module is constructed, which includes a conflict pre-simulation algorithm, adaptive adjustment rules, and an evolution iteration mechanism. It can dynamically update layout parameters based on the outputs of the environment parameter mapping module, the behavior intention decoding module, and the collaborative relationship application module. The environmental parameter mapping module, behavioral intent decoding module, collaborative relationship application module, and typesetting parameter evolution module are connected through a data interaction interface. The data transmission timing and triggering conditions between modules are set to form a dynamic self-adaptive typesetting engine.
4. The intelligent optimization method for graphic and text layout combining digital multimedia as described in claim 1, characterized in that, The application of the dynamic adaptive typesetting engine dynamically calculates the typesetting parameters of image units and text units in the digital multimedia graphic and text set to be optimized, generating an initial typesetting scheme. The typesetting parameter evolution module of the dynamic adaptive typesetting engine then performs multi-dimensional conflict pre-simulation and adaptive adjustment on the initial typesetting scheme to obtain an intermediate typesetting scheme, including: The attribute information of the image units and text units in the digital multimedia graphic and text set to be optimized is input into the collaborative relationship application module of the dynamic self-adaptive typesetting engine. Combined with the temporal mapping rules and sentiment matching parameters in the graphic and text collaboration model, the basic typesetting parameters of each graphic and text unit are determined. The basic typesetting parameters include the initial position, initial size and initial display duration. The real-time display environment parameters are input into the environment parameter mapping module to obtain the layout space flexibility coefficient, visual comfort parameter and content loading priority parameter. Based on the layout space flexibility coefficient, visual comfort parameter and content loading priority parameter, the basic layout parameters are adjusted for environmental adaptation to obtain the environment-adapted layout parameters. The interaction behavior sequence in the audience behavior intent data is input into the behavior intent decoding module to determine the audience intent classification result. Based on the audience intent classification result, the environment adaptation layout parameters are adjusted to obtain the strategy adaptation layout parameters. Integrate the strategy adaptation and layout parameters of all graphic and text units, determine the final position coordinates, display size, display duration and loading order of each graphic and text unit, and generate an initial layout scheme; Input the initial layout scheme into the layout parameter evolution module, start the conflict pre-simulation algorithm, and simulate the layout effect under different display environment changes and different audience interaction behaviors; Based on the pre-simulation results, potential layout conflicts are identified, including overlapping display areas, unbalanced loading delays, decreased visual comfort, and broken collaborative relationships. Based on the adaptive adjustment rules, the layout parameters that have conflict risks are adjusted. The adjustment operations include adjusting the position coordinates of overlapping units, optimizing the loading order to balance the delay, adjusting the brightness and contrast to improve comfort, and correcting the timing mapping rules to repair the collaborative relationship. Repeat the conflict pre-simulation algorithm and adaptive adjustment rules until there are no layout conflicts in the pre-simulation results. Record the adjusted layout parameters and generate an intermediate layout scheme.
5. The intelligent optimization method for graphic and text layout combining digital multimedia as described in claim 2, characterized in that, The process of analyzing the reading rhythm of each text unit involves extracting sentence pause markers, changes in lexical density, and semantic transition locations within the text unit. This determines the reading time segments and key semantic intervals of the text unit, forming the spatiotemporal features of the text, including: Perform sentence structure parsing on each text unit, identify commas, periods, exclamation marks, and question marks as sentence pause markers in the text unit, and count the frequency and interval of different pause markers. The text unit is divided into multiple sentence segments based on the number of characters between the pause marks, with each sentence segment bounded by two adjacent pause marks. For each sentence segment, perform word density calculation processing, count the number of words and characters in each sentence segment, and calculate the word density value per unit character length; Analyze the changing trend of word density values, identify sentence fragments whose word density values exceed a preset density threshold, and identify the sentence fragments that correspond to areas that need to be focused on during reading; Semantic analysis is performed on text units to identify semantic transition conjunctions and semantic emphasis words, and to determine the positions of semantic transitions and semantic emphasis. Based on the results of sentence segmentation, the trend of word density change, and the results of semantic analysis, the text unit is divided into multiple reading time segments. The segments with word density values exceeding the preset density threshold and prominent semantic points correspond to segments with reading times exceeding the preset time threshold. Mark the sentence segments where the semantic focus is located to determine the key semantic range of the text unit; By integrating data on reading time segments, key semantic intervals, sentence pause marker distribution, and word density changes, spatiotemporal features of the text are formed.
6. The intelligent optimization method for graphic and text layout combining digital multimedia as described in claim 2, characterized in that, The step involves performing visual presentation timing analysis on each image unit, extracting the visual focus switching path, color transition rhythm, and detail information hierarchy within the image unit, determining the optimal display duration and visual attention sequence for each image unit, and forming the spatiotemporal characteristics of the image, including: Visual focus detection is performed on each image unit, and multiple visual focuses in the image unit are identified based on color contrast, brightness difference, edge complexity and object salience. Analyze the positional relationships, size differences, and attractiveness of visual focal points, simulate the eye movement trajectory of the audience when viewing the image, and determine the visual focal point switching path; Color analysis is performed on image units to extract the color distribution and color transition patterns in different regions of the image units, identify regions with significant color transitions, calculate their areas, and determine the color transition rhythm characteristics. The image unit is processed for detail information extraction. Based on the image resolution, texture complexity and object detail richness, the image unit is divided into multiple detail information levels. The levels with rich detail elements to the levels with simple detail elements correspond to different visual attention depths. Based on the physical length of the visual focus switching path and the number of visual focuses, calculate the basic time required for the audience to fully browse all visual focuses; The base time is adjusted based on the proportion of the color transition area to the total image area. When the proportion of the color transition area exceeds the preset threshold, the browsing time is increased accordingly. Based on the number and complexity of the detailed information levels, the browsing time is further adjusted. When the proportion of levels with rich detailed elements exceeds the preset proportion threshold, the browsing time is increased accordingly, and the optimal display time of the image unit is finally determined. Based on the visual focus switching path and the level of detail information, the order of visual attention when the audience views the image is determined. First, attention is paid to areas where the visual focus attractiveness is greater than the preset attractiveness threshold and areas with rich detail information levels, and then attention is paid to other areas. By integrating the visual focus switching path, color transition rhythm, detail information hierarchy, optimal display duration, and visual attention order, the spatiotemporal characteristics of the image are formed.
7. The intelligent optimization method for graphic and text layout combining digital multimedia as described in claim 3, characterized in that, The process of decoding the interaction sequence in the audience behavior intent data, extracting the interaction features of the interaction behaviors, and analyzing the audience intent corresponding to the interaction features includes: Perform type identification processing on each interactive operation in the interactive behavior sequence to distinguish different interaction types such as click operation, swipe operation, zoom operation, and long press operation; Count the number of times each type of interaction occurs within a preset time window to obtain the frequency of interaction behavior; Record the duration of each interaction from start to finish to obtain the duration of the interaction behavior; Track the position coordinate changes of continuous interactive operations to form interactive operation paths, and analyze the straightness, tortuosity and coverage of the operation paths; For devices that support pressure sensing, pressure data during interactive operation is collected to determine the intensity of the interactive operation. For devices that do not support pressure sensing, the intensity of the operation is indirectly inferred based on the interaction duration and operation speed. Analyze the combination relationship between interaction frequency, interaction duration, operation path characteristics and operation force. When the swipe operation frequency exceeds the preset frequency threshold and the single swipe duration is less than the preset duration threshold, and at the same time the straightness of the operation path is higher than the preset straightness threshold and the coverage of the operation path is greater than the preset range threshold, the corresponding quick browsing intent is determined. The click frequency is lower than the preset frequency standard and the dwell time after a single click exceeds the preset dwell time standard. The zoom operation ratio exceeds the preset zoom ratio standard. The operation path is concentrated in an area that meets the preset local range standard, which corresponds to a type of audience intent. The frequency of click operations exceeds the preset frequency standard and the dispersion of operation paths meets the preset dispersion standard. The proportion of short-distance swipes in swipe operations exceeds the preset short-swipe proportion standard. Combined with the fluctuation of operation force exceeding the preset force fluctuation standard, it corresponds to a type of audience intent. The correspondence between different combinations of interactive features and audience intent is verified, and misjudged intent classification results are corrected.
8. The intelligent optimization method for graphic and text layout combining digital multimedia as described in claim 4, characterized in that, The adjustment of layout parameters with potential conflict risks according to adaptive adjustment rules includes: To address overlapping conflicts in the display area, the system queries the collaborative weight and audience intent relevance of the conflicting graphic and text units. Graphic and text units with a collaborative weight greater than a preset high weight threshold and an audience intent relevance greater than a preset high relevance threshold are given priority to retain their original position coordinates. Graphic and text units with a collaborative weight less than a preset low weight threshold and an audience intent relevance less than a preset low relevance threshold are moved away from the conflict area. The moving distance is positively correlated with the size of the conflict intersection area, and the moved position conforms to the layout space elasticity coefficient constraint. To address loading delay imbalance conflicts, the loading order of text and image units is reordered based on content loading priority parameters and display duration of text and image units. Text and image units with display duration exceeding the preset display duration threshold and audience intent relevance greater than the preset high relevance threshold are loaded first. For graphic units whose loading speed is lower than the preset loading speed threshold, adjust their size parameters to reduce the amount of data, or adopt a progressive loading method, first loading low-resolution image data to meet basic display requirements, and then loading high-resolution image data to balance loading delay. Adjust the font size, character spacing, and line spacing of the text units to ensure that the font clarity is not lower than the preset clarity threshold under different lighting conditions; To address the conflict caused by the break in the collaborative relationship, the temporal matching relationship between the spatiotemporal features of the text and the spatiotemporal features of the image is recalculated. Based on the current display environment parameters, the temporal mapping rules are corrected, and the deviation between the reading time segments and the optimal display time is adjusted to not exceed the preset corresponding deviation threshold. Adjust the sentiment feature values of image units and text units to reduce the sentiment deviation between them to no more than a preset sentiment deviation threshold, thereby repairing the sentiment synergy relationship. After the adjustment is completed, the conflict pre-simulation algorithm is re-executed to verify whether the conflict has been resolved. If not, the adjustment steps are repeated until the conflict is eliminated.
9. The intelligent optimization method for graphic and text layout combining digital multimedia as described in claim 1, characterized in that, The process of deploying the intermediate layout scheme to the real-time display environment, combining the content demand descriptions in the audience behavior intent data and the real-time updated display environment parameters, drives the dynamic self-adaptive layout engine to continuously evolve the layout scheme, generating a target graphic and text layout optimization scheme that conforms to the real-time environment and audience intent, includes: Deploy the intermediate layout scheme to the digital multimedia real-time display environment and establish a real-time data connection between the layout scheme and the display environment parameter acquisition module and the audience behavior intent acquisition module; By displaying the environmental parameter acquisition module, changes in device screen characteristics, fluctuations in ambient light intensity, and changes in network transmission rate are collected in real time, forming a real-time updated display environmental parameter stream. The audience behavior intent collection module collects new interaction behavior sequences of the audience in real time during the browsing of intermediate layout schemes, and collects the content demand expressions input by the audience to form a real-time updated audience behavior intent data stream. The real-time updated display environment parameter stream is input into the environment parameter mapping module of the dynamic self-adaptive typesetting engine, which updates the dynamic mapping relationship between environment parameters and typesetting adjustment parameters, and outputs new typesetting space elasticity coefficient, visual comfort parameter and content loading priority parameter. The real-time updated audience behavior intent data stream is input into the behavior intent decoding module to update the correspondence between the interaction behavior sequence and the layout strategy. The audience intent classification results are corrected in combination with the content requirement description, and a new layout strategy is output. Based on the new layout strategy and updated environmental parameters, the collaborative relationship application module adjusts the temporal mapping rules, sentiment matching parameters and collaborative weights in the graphic-text collaborative model, and outputs new collaborative application logic. The layout parameter evolution module updates the layout parameters of the intermediate layout scheme in real time based on the new environmental parameter mapping relationship, layout strategy and collaborative application logic, and adjusts the position coordinates, display size, display duration and loading order of graphic units. During the typesetting parameter update process, a conflict pre-simulation algorithm is executed simultaneously, and no new conflicts are generated in the updated typesetting scheme; The process of continuously repeating real-time data collection, module parameter updates, layout parameter adjustments, and conflict rehearsals continues until there are no new content requirements expressed in the audience behavior intent data stream, and the fluctuation range of each parameter in the real-time display environment parameter stream remains stable within its respective preset threshold range. The final stable layout scheme was determined as the target text and image layout optimization scheme.
10. A smart optimization system for graphic and text layout combining digital multimedia, characterized in that, The method includes a processor and a computer-readable storage medium storing machine-executable instructions that, when executed by the processor, implement the intelligent optimization method for graphic and text layout combining digital multimedia as described in any one of claims 1-9.
Citation Information
Patent Citations
Exhibition hall content self-adaptive central control display method and exhibition hall content self-adaptive central control display system
CN120491810A
Page structure optimization method and system for PowerPoint
CN120493879A