Image-text typesetting intelligent optimization method and system combined with digital multimedia

By acquiring device parameters and audience intent data to perform spatiotemporal collaborative modeling of text and graphics, a dynamic self-adaptive typesetting engine is constructed, solving the problem that text and graphics typesetting methods cannot be adjusted in real time, and realizing high-quality personalized digital multimedia display.

CN121527249AActive Publication Date: 2026-02-13SHANGHAI MINGQI NETWORK TECH CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202610030751.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-01-12
Publication Date
2026-02-13
Estimated Expiration
2046-01-12

AI Technical Summary

Technical Problem

Existing graphic layout methods cannot adjust in real time according to device screen characteristics, ambient light intensity, and audience behavior intentions, resulting in blurry images and misaligned text. This fails to meet the personalized needs of different devices and audiences, thus affecting user experience.

Method used

By acquiring the set of digital multimedia graphics and text to be optimized, real-time display environment parameters, and audience behavior and intent data, spatiotemporal collaborative modeling of graphics and text is performed to build a dynamic self-adaptive typesetting engine and generate a target graphic and text typesetting optimization scheme that conforms to the real-time environment and audience intent.

Benefits of technology

It enables the coordinated display of text and images across time, space, and emotion, enhancing the content's appeal and attractiveness, improving the rationality and stability of the layout, and providing a personalized digital multimedia content display experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121527249A_ABST
    Figure CN121527249A_ABST
Patent Text Reader

Abstract

The invention provides an image-text typesetting intelligent optimization method and system combined with digital multimedia, and relates to the technical field of digital multimedia, and the method comprises the steps: firstly obtaining a to-be-optimized image-text set, and displaying environment parameters and audience behavior intention data in real time; performing space-time collaborative modeling on the image-text units, and establishing an image-text collaborative model; constructing a dynamic adaptive typesetting engine comprising a plurality of modules; performing typesetting parameter dynamic calculation on an image unit and a text unit in the digital multimedia image-text set to be optimized by applying the dynamic adaptive typesetting engine to generate an initial typesetting scheme, and adjusting to obtain an intermediate scheme; after the intermediate scheme is deployed, the engine is driven to evolve continuously in combination with real-time data, the target image-text typesetting optimization scheme is generated and output for display, typesetting can be optimized dynamically according to the environment and audience requirements, and the display effect and the user experience are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of digital multimedia technology, and more specifically, to a method and system for intelligent optimization of graphic and text layout that combines digital multimedia. Background Technology

[0002] In today's booming digital multimedia landscape, graphic layout plays a crucial role in the presentation of various digital media content. Whether it's web design, mobile application interface display, or electronic publications, high-quality graphic layout is indispensable. However, existing graphic layout methods have many limitations.

[0003] Traditional text and image layout methods often rely on fixed templates and preset rules, lacking consideration for the real-time display environment. Different devices have vastly different screen characteristics, such as screen size and resolution, meaning that a fixed layout scheme may not produce optimal results on different devices, leading to problems such as blurry images and misaligned text. Furthermore, ambient light intensity also affects the user's viewing experience. In bright light, it may be necessary to adjust the color contrast of the text and images, while in low light, brightness may need to be increased. However, traditional methods cannot automatically adjust according to changes in lighting conditions.

[0004] Furthermore, existing methods do not fully consider the audience's behavioral intentions. The audience's interaction sequences, emotional inclinations, and content needs are all crucial factors influencing layout effectiveness. Different audiences focus on different aspects of text and images, their emotional inclinations affect their acceptance of layout styles, and their content needs directly determine the emphasis of the text and images. However, traditional layout methods struggle to adjust in real-time to these dynamically changing audience needs, leading to layout schemes that do not meet audience expectations and diminishing the user experience. Summary of the Invention

[0005] In view of this, the purpose of this application is to provide a method and system for intelligent optimization of graphic and text layout that combines digital multimedia.

[0006] According to a first aspect of this application, a method for intelligent optimization of text and image layout combined with digital multimedia is provided, the method comprising: The system acquires a set of digital multimedia images and texts to be optimized, real-time display environment parameters, and audience behavioral intent data. The set of digital multimedia images and texts to be optimized includes multiple image units and multiple text units. The real-time display environment parameters include device screen characteristics, ambient light intensity, and network transmission rate. The audience behavioral intent data includes interactive behavior sequences, sentiment markers, and content demand expressions. The image units and text units in the digital multimedia image and text set to be optimized are modeled in a spatiotemporal manner. Combined with the sentiment tendency markers in the audience behavior intention data, the spatiotemporal correlation and sentiment synergy relationship between the image units and text units are established to obtain the image and text synergy model. Based on the real-time display environment parameters, the interaction behavior sequence in the audience behavior intent data, and the graphic-text collaboration model, a dynamic self-adaptive typesetting engine is constructed. The dynamic self-adaptive typesetting engine includes an environment parameter mapping module, a behavior intent decoding module, a collaboration relationship application module, and a typesetting parameter evolution module. The dynamic adaptive typesetting engine is used to dynamically calculate the typesetting parameters of the image units and text units in the digital multimedia graphic and text collection to be optimized, and generate an initial typesetting scheme. The typesetting parameter evolution module of the dynamic adaptive typesetting engine is used to perform multi-dimensional conflict pre-simulation and adaptive adjustment on the initial typesetting scheme to obtain an intermediate typesetting scheme. The intermediate layout scheme is deployed to the real-time display environment. Combining the content demand description in the audience behavior intent data and the real-time updated display environment parameters, the dynamic self-adaptive layout engine is driven to continuously evolve the layout scheme, generate a target graphic and text layout optimization scheme that conforms to the real-time environment and audience intent, and output the target graphic and text layout optimization scheme for digital multimedia content display.

[0007] According to a second aspect of this application, a digital multimedia-integrated intelligent optimization system for graphic and text layout is provided. The digital multimedia-integrated intelligent optimization system for graphic and text layout includes a machine-readable storage medium and a processor. The machine-readable storage medium stores machine-executable instructions. When the processor executes the machine-executable instructions, the digital multimedia-integrated intelligent optimization system for graphic and text layout implements the aforementioned digital multimedia-integrated intelligent optimization method for graphic and text layout.

[0008] According to a third aspect of this application, a computer-readable storage medium is provided, wherein computer-executable instructions are stored therein, and when the computer-executable instructions are executed, the aforementioned intelligent optimization method for graphic and text layout combined with digital multimedia is implemented.

[0009] Based on any of the above aspects, the technical effect of this application is as follows: By acquiring the digital multimedia image and text set to be optimized, real-time display environment parameters, and audience behavioral intent data, and then performing spatiotemporal collaborative modeling of image and text units, and combining audience sentiment markers to establish spatiotemporal and emotional relationships, a collaborative image-text model is constructed. This effectively solves the problem of collaborative display of image and text content at the spatiotemporal and emotional levels, making the combination of images and text more natural and harmonious, and enhancing the appeal and attractiveness of the content. Based on real-time display environment parameters, audience interaction behavior sequences, and the collaborative image-text model, a dynamic self-adaptive typesetting engine possesses strong environmental adaptability and audience demand interpretation capabilities. Its multiple modules collaborate to dynamically calculate typesetting parameters based on different environmental conditions and audience behavior, generating an initial typesetting scheme. Through the typesetting parameter evolution module, multi-dimensional conflict pre-simulation and adaptive adjustment are performed to obtain an intermediate typesetting scheme, effectively avoiding various conflict problems that may occur during typesetting and improving the rationality and stability of the typesetting. After deploying the intermediate layout scheme to the real-time display environment, the layout engine is continuously evolved by combining the audience's content needs and the real-time updated display environment parameters. This generates a target graphic and text layout optimization scheme that conforms to the real-time environment and the audience's intentions. It can always maintain a high degree of matching between the layout scheme and the actual display environment and audience needs, providing users with a higher quality and more personalized digital multimedia content display experience, and significantly improving the quality and efficiency of digital multimedia graphic and text layout. Attached Figure Description

[0010] Figure 1 A flowchart illustrating the intelligent optimization method for graphic and text layout combining digital multimedia provided in an embodiment of this application is shown. Figure 2 This illustration shows a schematic diagram of the component structure of the intelligent optimization system for graphic and text layout that combines digital multimedia, provided in an embodiment of this application. Detailed Implementation

[0011] Figure 1 This paper illustrates a flowchart of a method and system for intelligent optimization of text and graphic layout combining digital multimedia, as provided in an embodiment of this application. The detailed steps include: Step S110: Obtain the set of digital multimedia images and text to be optimized, real-time display environment parameters, and audience behavior intent data. The set of digital multimedia images and text to be optimized includes multiple image units and multiple text units. The real-time display environment parameters include device screen characteristics, ambient light intensity, and network transmission rate. The audience behavior intent data includes interactive behavior sequences, sentiment markers, and content demand statements.

[0012] This embodiment uses the optimization of online menu layout in the catering industry as an example. An online menu, as a specific form of digital multimedia graphic collection, includes text units describing the dishes and image units depicting the dishes. In this scenario, the digital multimedia graphic collection to be optimized is specifically the online electronic menu of a restaurant, which contains multiple dish categories. Each category has multiple text units and image units corresponding to various dishes. For example, under the "Signature Hot Dishes" category, the dish "Braised Pork Ribs" has a corresponding text unit containing the dish name, flavor description, ingredient composition, and price information; it also includes a high-resolution image of the dish as an image unit.

[0013] Real-time environmental parameters are acquired through the user's current mobile device. Device screen characteristics include physical screen size, resolution, current screen orientation (portrait or landscape), and pixel density. For example, a user's smartphone screen is typically rectangular; resolution reflects the number of pixels horizontally and vertically; screen orientation is detected in real-time by the device's built-in gravity sensor; and pixel density affects the sharpness of the displayed content. Ambient light intensity is collected by the device's light sensor, reflecting the brightness of the user's environment. For instance, the ambient light intensity values ​​differ significantly between bright outdoor and dim indoor environments. Network transmission speed is monitored in real-time by the device's network module, including the current network connection type (Wi-Fi, 4G, 5G) and corresponding download and upload speeds. These parameters affect the loading speed of large amounts of data, such as images, in online menus.

[0014] The acquisition of audience behavioral intent data must comply with relevant laws and regulations, and privacy protection measures must be taken when sensitive data is involved. For interaction behavior sequences, the user behavior collection module integrated into the restaurant app records user actions while browsing the online menu, such as clicking on dish category buttons, swiping through the dish list, long-pressing on dish images to view details, and zooming in and out of dish images, along with the order and attributes of these actions, with user authorization. Sentiment markers are obtained by analyzing users' past reviews of dishes. For example, the use of positive words like "delicious" and "satisfied" in reviews corresponds to a positive sentiment marker, while the use of negative words like "unpalatable" and "disappointing" corresponds to a negative sentiment marker. Content demand expressions are obtained through comprehensive analysis of user search history, saved dish types, and historical ordering records within the app. For example, multiple searches for "vegetarian food" indicate a preference for vegetarian food in the user's content demand expressions.

[0015] When obtaining the above data, for privacy-sensitive data (such as users' geographical location information, detailed consumption records, etc.), data desensitization technology is adopted for processing, the real identity information of users is anonymized, and at the same time, encrypted transmission is used to ensure that the data is not leaked during the transmission process. When storing, encrypted storage technology is adopted to restrict data access rights, and only relevant modules are authorized to access the processed non-sensitive data when necessary.

[0016] Step S120: Perform graphic-text spatio-temporal collaborative modeling on the image units and text units in the digital multimedia graphic-text set to be optimized. Combining the emotional tendency markers in the audience behavior intention data, establish the spatio-temporal association relationship and emotional collaboration relationship between the image units and text units, and obtain the graphic-text collaboration model.

[0017] Step S121: Perform reading rhythm analysis and processing on each text unit, extract the sentence pause markers, vocabulary density changes, and semantic turning points in the text unit, determine the reading duration segmentation and key semantic intervals of the text unit, and form text spatio-temporal features.

[0018] Step S1211: Perform sentence structure parsing and processing on each text unit, identify the comma, period, exclamation mark, and question mark sentence pause markers in the text unit, and count the occurrence frequencies and interval distances of different pause markers.

[0019] Taking the text unit of the "Yuxiang Shredded Pork" dish in the online menu as an example, the content of this text unit is: "Yuxiang Shredded Pork, a classic Sichuan dish, has a salty, fresh and slightly spicy taste, the shredded pork is tender, the ingredients are rich, including green peppers, carrots, black fungus, etc., and the price is affordable." First, perform sentence structure parsing on this text unit to identify the pause markers in it. Here, there are commas and periods. The comma appears after "Yuxiang Shredded Pork", after "a classic Sichuan dish", after "has a salty, fresh and slightly spicy taste", after "the shredded pork is tender", after "the ingredients are rich", and after "including green peppers, carrots, black fungus, etc."; the period appears at the end of the text unit. Count the occurrence frequencies of different pause markers. The comma appears 6 times and the period appears 1 time. The interval distance is obtained by calculating the number of characters between adjacent pause markers. For example, the comma interval between "Yuxiang Shredded Pork" and "a classic Sichuan dish" is 4 characters ("Yuxiang Shredded Pork" is 4 characters), and the comma interval between "a classic Sichuan dish" and "has a salty, fresh and slightly spicy taste" is 5 characters ("a classic Sichuan dish" is 5 characters), and so on.

[0020] Step S1212: Divide the text unit into multiple sentence segments according to the interval characters of the pause markers, and each sentence segment is bounded by two adjacent pause markers.

[0021] Continuing with the example in step S1211, according to the identified pause markers and the number of intervening characters, the text unit of "Yuxiang Shredded Pork" is divided into statement segments. The first statement segment is "Yuxiang Shredded Pork" (bounded by the start and the first comma); the second statement segment is "a classic Sichuan dish" (bounded by the first comma and the second comma); the third statement segment is "tasting salty, fresh and slightly spicy" (bounded by the second comma and the third comma); the fourth statement segment is "the shredded pork is tender" (bounded by the third comma and the fourth comma); the fifth statement segment is "rich in ingredients" (bounded by the fourth comma and the fifth comma); the sixth statement segment is "including green peppers, carrots, black fungus, etc." (bounded by the fifth comma and the sixth comma); the seventh statement segment is "affordable price" (bounded by the sixth comma and the full stop).

[0022] Step S1213: Perform a lexical density calculation process on each statement segment, count the number of words and characters in each statement segment, and calculate the lexical density value per unit character length.

[0023] For each statement segment divided in step S1212, perform a lexical density calculation. The number of words refers to the number of independent words in the statement segment, and the number of characters is the total number of all characters included in the statement segment (including Chinese characters, punctuation marks, etc.). For example, in the statement segment "tasting salty, fresh and slightly spicy", the number of words is 4 ("tasting", "salty, fresh", "slightly spicy"), and the number of characters is 6 ("tasting salty, fresh and slightly spicy" has a total of 6 Chinese characters). Then the lexical density value per unit character length is the number of words divided by the number of characters, that is, 4 divided by 6, to obtain a corresponding lexical density value. Perform the above calculation on all statement segments to obtain their respective lexical density values.

[0024] Step S1214: Analyze the changing trend of the lexical density values, identify the statement segments whose lexical density values exceed the preset density threshold, and the said statement segments correspond to the areas that need to be focused on during reading.

[0025] Arrange the lexical density values of each statement segment calculated in step S1213 in order to form the changing trend of the lexical density values. Preset a density threshold, which is set according to the general reading habits and information importance of the catering menu text. After analyzing the changing trend, identify the statement segments whose lexical density values exceed the preset density threshold. For example, in the text unit of "Yuxiang Shredded Pork", for the statement segment "rich in ingredients, including green peppers, carrots, black fungus, etc.", due to containing various ingredient information, the lexical density value may be relatively high. If it exceeds the preset density threshold, then this statement segment is identified as the area that needs to be focused on during reading because users usually pay attention to the ingredient composition when choosing dishes.

[0026] Step S1215: Perform semantic analysis on the text units, identify semantic transition conjunctions and semantic emphasis words, and determine the semantic transition position and semantic emphasis position.

[0027] Natural Language Processing (NLP) techniques were used to perform semantic analysis on text units. Semantic transition conjunctions included "but," "however," and "yet," while semantic emphasis words included "most," "especially," "classic," and "signature dish." In the text unit "Fish-flavored Shredded Pork," the word "classic" in "classic Sichuan cuisine" is a semantic emphasis word, indicating the dish's tradition and popularity. If the text unit contained expressions like "although the price is slightly high, the taste is excellent," then "although" and "but" are semantic transition conjunctions, and their positions indicate semantic transition points, while "the taste is excellent" indicates a semantic emphasis point. By identifying these words, the semantic transition and emphasis points within the text units were determined.

[0028] Step S1216: Combining the results of sentence segment division, the trend of word density change, and the results of semantic analysis, the text unit is divided into multiple reading time segments. Sentence segments with word density values ​​exceeding the preset density threshold and sentence segments containing semantic key positions correspond to areas that need to be focused on during reading.

[0029] Based on the combined results of sentence segmentation (step S1212), sentence segments exceeding the preset density threshold in the lexical density trend (step S1214), and semantic focus locations obtained from semantic analysis (step S1215), the text unit is segmented by reading time. The segmentation is based on the importance and reading complexity of different sentence segments; sentence segments with higher importance and greater reading complexity are allocated longer reading times. For example, in the "Fish-flavored Shredded Pork" text unit, the sentence segments containing "classic Sichuan cuisine" (containing semantically emphasized words) and "rich in ingredients, including green peppers, carrots, and wood ear mushrooms" (lexical density exceeding the threshold) are segmented to require longer reading times, while other sentence segments, such as "affordable," have relatively shorter reading times. Furthermore, the intervals containing these important sentence segments are designated as key semantic intervals.

[0030] Step S1217: Integrate reading time segmentation, key semantic intervals, sentence pause marker distribution, and word density change data to form text spatiotemporal features.

[0031] The reading time segments and key semantic intervals obtained in step S1216, along with the sentence pause marker distribution from step S1211 and the vocabulary density change data from step S1213, are integrated. The text spatiotemporal feature is a multi-dimensional feature vector that includes information on the distribution of reading time (reading time segments) in the time dimension and the distribution of key information (key semantic intervals) in the spatial dimension. It also includes relevant data on sentence structure and vocabulary distribution, which are used for subsequent matching with image spatiotemporal features.

[0032] Step S122: Perform visual presentation timing analysis on each image unit, extract the visual focus switching path, color transition rhythm and detail information level in the image unit, determine the optimal display duration and visual attention order of the image unit, and form the spatiotemporal characteristics of the image.

[0033] Step S1221: Perform visual focus detection processing on each image unit, and identify multiple visual focuses in the image unit based on color contrast, brightness difference, edge complexity and object salience.

[0034] Taking the image unit of "Fish-flavored Shredded Pork" as an example, this image is a photograph of the dish, including elements such as the main body of the dish, the edge of the plate, and the background. Visual focus detection analyzes the color contrast of the image. There is a significant color difference between the main body of the dish (Fish-flavored Shredded Pork) and the plate and background; areas with high color contrast are more likely to become visual focal points. Regarding brightness difference, the main body of the dish, which has been illuminated, is brighter, contrasting with the relatively dark background. Edge complexity is determined by detecting the complexity of the object contours in the image; the edges of the shredded pork, vegetables, etc., are relatively complex, corresponding to areas with high edge complexity. Object salience is based on the characteristics of food images; the dish itself is a salient object in the image. By comprehensively analyzing these factors, multiple visual focal points in the image unit are identified, such as the central area of ​​the dish (where the shredded pork and main ingredients are concentrated) and brightly colored parts of the dish (such as carrot chunks).

[0035] Step S1222: Analyze the positional relationship, size difference, and attractiveness of visual focal points, simulate the eye movement trajectory of the audience when viewing the image, and determine the visual focal point switching path.

[0036] After identifying multiple visual focal points, their positional relationships are analyzed, such as which focal points are on the left, right, top, and bottom of the image, and their relative distances. Size difference refers to the proportion of area occupied by different visual focal points in the image. Attractiveness is evaluated by considering factors such as color contrast, brightness difference, and edge complexity; visual focal points with higher color contrast, greater brightness difference, and higher edge complexity are more attractive. Based on these analyses, the eye movement trajectory of the viewer is simulated, starting from the most attractive visual focal point and moving sequentially towards other visual focal points, forming a visual focal point switching path. For example, in an image of "Fish-flavored Shredded Pork," the shredded pork and main ingredients in the central area are the most attractive visual focal points and are first noticed. Then, the gaze may move to the brightly colored carrot pieces, then to other ingredients, and finally sweep across the edge of the plate and the background, forming a specific visual focal point switching path.

[0037] Step S1223: Perform color analysis processing on the image unit, extract the color distribution and color transition mode of different regions in the image unit, identify the regions with significant color transitions, calculate their area, and determine the color transition rhythm characteristics.

[0038] Image units undergo color space conversion, transforming the image from RGB to HSV color space for more accurate color attribute analysis. Hue, saturation, and brightness values ​​are extracted from different regions of the image, and color distribution is statistically analyzed, such as identifying the main colors of the food area and the background area. Color transition modes include gradual transitions and abrupt transitions. Gradual transitions refer to slow color changes between adjacent areas, while abrupt transitions refer to significant color jumps between adjacent areas. Regions with significant color transitions are identified, i.e., areas with abrupt transitions and large color differences in the transition area. The area of ​​these significant color transition regions is calculated as a percentage of the total image area. Color transition rhythm characteristics are determined by the size, distribution, and changes in the transition mode of significant color transition regions; for example, larger significant color transition regions result in a stronger color transition rhythm.

[0039] Step S1224: Perform detail information extraction processing on the image unit. Based on image resolution, texture complexity and object detail richness, the image unit is divided into multiple detail information levels. The levels with rich detail elements to the levels with simple detail elements correspond to different visual attention depths.

[0040] Image resolution determines image clarity; high-resolution images contain more detailed information. Texture complexity is determined by analyzing the texture features of different regions in an image, such as the texture of the food surface, the plate, and the background. The more complex the texture, the richer the detail. Object detail richness refers to the number of details contained in an object in an image. For example, in an image of "fish-flavored shredded pork," the texture of the shredded pork, the texture of the vegetables, and the sheen of the broth all belong to object details. Based on these factors, image units are divided into multiple levels of detail information. For example, the highest level is the core detail area of ​​the dish (such as the texture and sheen of the shredded pork), containing the richest detail elements; the middle level is the secondary detail area of ​​the dish (such as the shape and color of the vegetables); and the lowest level is the background and the edge area of ​​the plate, where the detail elements are relatively simple. Different levels of detail information correspond to different visual attention depths. When viewing an image, the audience will first focus on the high-level areas with rich detail elements, and then gradually focus on the low-level areas with simpler detail elements.

[0041] Step S1225: Calculate the basic time required for the audience to fully browse all visual focuses based on the physical length of the visual focus switching path and the number of visual focuses.

[0042] The physical length of the visual focus switching path is obtained by quantizing the simulated gaze movement trajectory in the image coordinate system, i.e., the sum of the distances between the coordinates of all points on the trajectory. The number of visual focuses is the total number of visual focuses identified in step S1221. The calculation of the base time is based on the average time required for the human eye to move from one visual focus to another when viewing an image, and the average time spent at each visual focus. Dividing the physical length of the visual focus switching path by the average speed of human eye movement yields the total gaze movement time; multiplying the number of visual focuses by the average time spent at each visual focus yields the total gaze dwell time; the sum of the two is the base time required for the audience to fully view all visual focuses.

[0043] Step S1226: Adjust the base time based on the proportion of the color transition area to the total image area. When the proportion of the color transition area exceeds the preset threshold, increase the browsing time accordingly.

[0044] A preset threshold for the proportion of color transition areas is established, based on the visual characteristics of food and beverage images. The proportion of the area of ​​significant color transition areas obtained in step S1223 to the total image area is calculated. If this proportion exceeds the preset threshold, it indicates that the image has rich color variations, and the audience needs more time to perceive and understand these color transitions. Therefore, a certain amount of viewing time is added to the base time. The amount of added time is related to the degree to which the proportion of the color transition area exceeds the threshold; the higher the proportion, the more time is added.

[0045] Step S1227: Based on the number and complexity of the detailed information levels, further adjust the browsing time. When the proportion of levels with rich detailed elements exceeds the preset proportion threshold, increase the browsing time accordingly, and finally determine the optimal display time for the image unit.

[0046] The more layers of detail information there are, the richer the image's details, requiring more time for the audience to focus on each layer. The complexity of the detail information layers is comprehensively evaluated by the number and complexity of detail elements in each layer. A preset threshold for the proportion of detail-rich layers is established. The area ratio of detail-rich layers (such as the highest and middle layers in step S1224) in the entire image is calculated. If this ratio exceeds the preset threshold, it indicates that the image is rich in detail information, requiring more viewing time for the audience to fully observe the details. The optimal display duration for the image unit is obtained by adding the time adjusted in step S1226 to the time adjusted based on the detail information layers.

[0047] Step S1228: Based on the visual focus switching path and detail information level, determine the visual attention order of the audience when viewing the image. First, focus on the area where the visual focus attraction is greater than the preset attraction threshold and the detail information level area with rich detail elements, and then focus on other areas.

[0048] An attractiveness threshold is preset, and areas with visual appeal exceeding this threshold are selected as the primary focus of the audience. Simultaneously, areas rich in detail are also key areas of visual attention. Based on the visual focus switching path, the order of visual attention is determined, starting with the most attractive visual focus area that belongs to the rich detail level. Following the visual focus switching path, other areas meeting the criteria are then addressed, followed by areas with weaker visual appeal and areas with simpler details. For example, in an image of "Fish-flavored Shredded Pork," the shredded pork in the center (highly attractive and rich in detail) is the first focus, followed by the carrot chunks (highly attractive and with some detail), then the other ingredients, and finally the edge of the plate and the background.

[0049] Step S1229: Integrate the visual focus switching path, color transition rhythm, detail information hierarchy, optimal display duration and visual attention order to form the spatiotemporal characteristics of the image.

[0050] Integrate the visual focus switching path determined in step S1222, the color transition rhythm feature determined in step S1223, the detailed information levels divided in step S1224, the optimal display duration determined in step S1227, and the visual attention order determined in step S1228 to form the spatio-temporal feature of the image. The spatio-temporal feature of the image is also a multi-dimensional feature vector, which includes the optimal display duration of the image in the time dimension and the visual attention order, focus switching path, color transition rhythm, and detailed information levels in the space dimension, etc.

[0051] Step S123: Perform a temporal matching process on the text spatio-temporal feature and the image spatio-temporal feature, and establish a spatio-temporal association relationship between the image unit and the text unit according to the corresponding relationship between the reading duration segments and the optimal display duration, and the association relationship between the key semantic intervals and the visual attention order.

[0052] Taking the text unit and image unit of "Yuxiang shredded pork" as an example, the reading duration segments in the text spatio-temporal feature represent the time required to read each part of the text of this dish, and the optimal display duration in the image spatio-temporal feature is the time for which the image of this dish should be displayed. Match the total duration of the reading duration segments with the optimal display duration. If the total reading duration is close to the optimal display duration, it is considered that the two match well in time; if the difference is large, the reading duration segments or the image display duration need to be adjusted so that the two can cooperate in time. For example, the reading duration of the key semantic intervals in the text should correspond to the display duration of the corresponding visual attention order in the image. The key semantic intervals correspond to the content that the user needs to focus on reading in the text, such as the ingredient composition, and associate it with the area in the visual attention order of the image that displays the ingredients. When the user reads the ingredient composition part of the text, the image should display to the corresponding ingredient visual area, thus establishing the spatio-temporal association relationship between the image unit and the text unit.

[0053] Step S124: Perform an emotional tendency extraction process on each text unit, analyze the distribution of emotional words, the tone expression method, and the semantic emotional intensity in the text unit, and determine the emotional feature value of the text unit.

[0054] Step S1241: Perform word segmentation on the text unit, remove stop words, and retain the words with actual semantics.

[0055] Use a Chinese word segmentation tool to perform word segmentation on the text unit, and split the continuous text sequence into independent words. For example, after segmenting "Yuxiang shredded pork, a classic Sichuan dish, with a salty, fresh and slightly spicy taste, the shredded pork is tender, and the ingredients are rich", we get words such as "Yuxiang shredded pork", "classic", "Sichuan dish", "taste", "salty", "fresh", "slightly spicy", "shredded pork", "tender", "ingredients", "rich", etc. Then remove the stop words, which include words such as "of", "is", "in" that have no actual emotional and semantic contributions, and retain the above words with actual semantics after processing.

[0056] Step S1242: Construct an emotional lexicon that includes positive emotional words, negative emotional words, and neutral words, and assign corresponding emotional intensity values ​​to different emotional words.

[0057] The emotional lexicon is constructed based on existing Chinese emotional lexicons, with expansions and adjustments made to suit the characteristics of the catering industry. Positive emotional words such as "delicious," "fresh and tender," "classic," "abundant," and "affordable" are assigned positive emotional intensity values; negative emotional words such as "unpalatable," "greasy," and "stiff" are assigned negative emotional intensity values; neutral words such as "price," "ingredients," and "include" have an emotional intensity value of zero. The magnitude of the emotional intensity value is set according to the strength of the emotion expressed by the word in the catering context; for example, "classic" has a higher emotional intensity value than "good."

[0058] Step S1243: Match the vocabulary processed in step S1241 with the sentiment lexicon, and count the number of positive and negative sentiment words and their corresponding sentiment intensity values ​​in the text unit.

[0059] The words after word segmentation and stop word removal are matched one by one with the constructed sentiment lexicon to identify positive and negative sentiment words in the text unit. For example, in the text unit "fish-flavored shredded pork", "classic", "fresh and tender", "abundant", and "affordable" are positive sentiment words. The number of these positive sentiment words is counted, and the sentiment intensity value corresponding to each word is obtained; at the same time, the number (if any) of negative sentiment words and their sentiment intensity values ​​are also counted.

[0060] Step S1244: Analyze the tone of the text unit. If there are exclamatory sentences, rhetorical questions, or other expressions that strengthen the tone, adjust the corresponding emotional intensity value.

[0061] The tone and manner of expression affect the intensity of emotional expression. Exclamatory sentences usually express strong emotions, while rhetorical questions may also carry a certain emotional bias. For example, if a text unit contains an exclamatory sentence like "This dish is so delicious!", where "delicious" is a positive emotional word, its emotional intensity value should be appropriately increased due to the intensifying effect of the exclamation. If a rhetorical question exists, such as "Isn't this dish delicious?", it actually expresses a positive affirmation, and the corresponding emotional word intensity value also needs to be adjusted.

[0062] Step S1245: Calculate the semantic sentiment intensity of the text unit by subtracting the sum of the sentiment intensity values ​​of negative sentiment words from the sum of the sentiment intensity values ​​of positive sentiment words to obtain the preliminary sentiment value of the text unit.

[0063] Add the sentiment intensity values ​​of the positive sentiment words counted in step S1243 to obtain the total positive sentiment; add the sentiment intensity values ​​of the negative sentiment words to obtain the total negative sentiment. Subtract the total negative sentiment from the total positive sentiment to obtain the preliminary sentiment value of the text unit. For example, in the text unit "Fish-flavored Shredded Pork," the total sentiment intensity value of the positive sentiment words is the sum of the intensity values ​​of "classic," "fresh and tender," "rich," and "affordable." Assuming the total negative sentiment is zero, the preliminary sentiment value is this sum.

[0064] Step S1246: Based on the adjustment results of the tone expression, the preliminary sentiment value is corrected to obtain the sentiment feature value of the text unit.

[0065] Based on the analysis results of the tone expression in step S1244, the initial sentiment value is corrected. If there is an expression that strengthens the tone, a certain correction value is added to the initial sentiment value; if there is an expression that weakens the tone, a certain correction value is reduced. The corrected sentiment value is the sentiment feature value of the text unit. This sentiment feature value is a numerical value, where a positive number represents positive sentiment and a negative number represents negative sentiment, and the absolute value of the value represents the intensity of the sentiment.

[0066] Step S125: Perform sentiment extraction processing on each image unit, analyze the color sentiment attributes, compositional sentiment expression and visual element sentiment symbolism in the image unit, and determine the sentiment feature value of the image unit.

[0067] Step S1251: Perform color emotional attribute analysis on the image unit, extract the main color tone, color saturation and brightness value of the image, and determine the emotional tendency and intensity corresponding to each color according to the theory of color psychology.

[0068] Different colors have different emotional connotations in color psychology. For example, red is usually associated with enthusiasm and appetite, green with health and freshness, and yellow with warmth and brightness. Color analysis of image units extracts the dominant hue, which is the color with the highest proportion in the image. Color saturation reflects the vividness of a color; high-saturation colors express stronger emotions. Brightness reflects the lightness or darkness of a color. Based on these color attributes and color psychology theory, the corresponding emotional connotation (positive, negative, or neutral) and emotional intensity value for each color are determined. For example, the dominant hue of an image of "fish-flavored shredded pork" might include red (shredded pork and sauce) and green (green peppers). Red corresponds to positive emotions related to appetite, with a positive intensity value; green corresponds to positive emotions related to health, with a positive intensity value.

[0069] Step S1252: Perform compositional emotion expression analysis on the image units, analyze the composition methods of the images (such as symmetrical composition, diagonal composition, white space, etc.), and how different composition methods convey different emotional atmospheres.

[0070] Symmetrical composition conveys stability and solemnity; diagonal composition exudes dynamism and vitality; and negative space composition offers a sense of simplicity and comfort. Analyzing the composition of image units, for example, an image of "fish-flavored shredded pork" might employ a central composition, placing the main dish in the center of the image to highlight it and convey a direct and focused emotional atmosphere. Based on the emotional atmosphere corresponding to different composition methods, appropriate emotional intensity values ​​are assigned; for example, a central composition conveys a moderate level of positive emotional intensity.

[0071] Step S1253: Identify the visual elements in the image unit and analyze the emotional symbolic meaning of the visual elements, such as the freshness of the ingredients and the exquisiteness of the plating of the dishes.

[0072] Identify visual elements in images, such as the ingredients, plating, and tableware. The freshness of ingredients is judged by their color and shape; fresh ingredients carry positive emotional connotations. The refinement of the plating reflects the restaurant's attention to detail; elegant plating also carries positive emotional connotations. For example, in an image of "Fish-Fragrant Shredded Pork," the bright colors of the ingredients indicate freshness, and the neat plating further reinforces these positive emotional connotations, each assigned a corresponding emotional intensity value.

[0073] Step S1254: The color emotional intensity value, composition emotional intensity value, and visual element emotional intensity value are weighted and summed to obtain the emotional feature value of the image unit.

[0074] Weights are assigned to the emotional intensity values ​​of color, composition, and visual elements. The weights are determined based on their importance in the emotional expression of food images; for example, color emotion and visual element emotion (freshness of ingredients, plating) may have higher weights. Each emotional intensity value is multiplied by its corresponding weight and then summed to obtain the emotional feature value of the image unit. This emotional feature value is also a single numerical value; a positive number represents positive emotion, a negative number represents negative emotion, and the absolute value represents the emotional intensity.

[0075] Step S126: Extract the audience's emotional preference feature values ​​for similar images and texts from the emotional tendency tags of the audience's behavioral intention data, and use them as a reference benchmark for emotional coordination.

[0076] The sentiment markers in audience behavioral intent data are based on users' past evaluations, likes, and favorites of similar food and beverage images and text. For example, when browsing images and text of other Sichuan dishes, users give higher ratings and more favorites to dishes with descriptions such as "spicy and fragrant" and "rich in ingredients," and with brightly colored pictures and fresh ingredients. These behaviors are marked as positive sentiment. By analyzing these sentiment markers, the audience's emotional preferences for similar images and text (such as images and text of Sichuan dishes) are extracted. These preferences include a preference for positive emotional expressions and a high level of emotional attention to descriptions of ingredient freshness and taste. These preferences are then converted into specific sentiment preference feature values, serving as a reference benchmark for subsequent sentiment coordination.

[0077] Step S127: Perform collaborative matching processing on the emotional feature values ​​of text units, image units, and audience emotional preference feature values ​​to adjust the deviation between the emotional feature values ​​of image units and text units and establish an emotional collaborative relationship.

[0078] The deviations between the emotional feature values ​​of text units and audience emotional preference feature values, as well as the deviations between the emotional feature values ​​of image units and audience emotional preference feature values, are calculated. If the deviations are within a preset acceptable range, the emotional expression of the text and image units is considered to conform to audience preferences; if the deviations exceed the preset range, the emotional feature values ​​of the text and image units need to be adjusted. Adjustments can be made by modifying the use of emotional vocabulary in the text units, or by adjusting the color, contrast, etc., of the image units to change their emotional feature values, so that the emotional feature values ​​of both are closer to the audience emotional preference feature values. Ultimately, this ensures that the deviation between the emotional feature values ​​of the text and image units is also within the preset range, thereby establishing an emotional synergy and ensuring that the text and images are consistent in emotional expression and conform to audience preferences.

[0079] Step S128: Integrate spatiotemporal relationships and emotional collaboration relationships to construct a text-image collaboration model that includes temporal mapping rules, emotional matching parameters, and collaboration weights.

[0080] The spatiotemporal correlation established in step S123 and the emotional synergy established in step S127 are integrated. The temporal mapping rules clarify the correspondence between text reading time segments and optimal image display time, as well as the correlation rules between key semantic intervals and visual attention order. Emotional matching parameters include matching thresholds and adjustment coefficients between text emotional feature values, image emotional feature values, and audience emotional preference feature values. Synergy weights are assigned different weight values ​​based on the importance of the image and text units in spatiotemporal correlation and emotional synergy; for example, for key dishes, the synergy weight of the image and text units is higher. By integrating these elements, an image-text synergy model is constructed, which can comprehensively reflect the spatiotemporal and emotional correlation between image units and text units.

[0081] Step S130: Based on the real-time display environment parameters, the interaction behavior sequence in the audience behavior intent data, and the graphic-text collaboration model, a dynamic self-adaptive typesetting engine is constructed. The dynamic self-adaptive typesetting engine includes an environment parameter mapping module, a behavior intent decoding module, a collaboration relationship application module, and a typesetting parameter evolution module.

[0082] Step S131: Analyze the device screen characteristics in the real-time display environment parameters, extract the screen size specifications, resolution parameters, screen orientation and pixel density, calculate the initial layout space based on the screen size specifications, resolution parameters and pixel density, and determine the layout space elasticity coefficient in combination with the screen orientation. The layout space elasticity coefficient is used to dynamically adjust the size adaptation range of the graphic unit.

[0083] Step S1311: Convert the screen size specifications from physical size to pixel size, and obtain the horizontal and vertical pixel sizes of the screen by multiplying the physical screen size by the pixel density.

[0084] Screen dimensions are typically expressed in inches for diagonal length, which needs to be converted to horizontal and vertical pixel dimensions. First, based on the screen's aspect ratio (e.g., 16:9) and diagonal physical dimensions, calculate the screen's horizontal and vertical physical dimensions (in centimeters). Then, convert these physical dimensions to inches (1 inch = 2.54 centimeters), and multiply by the pixel density (in pixels per inch) to obtain the screen's horizontal and vertical pixel dimensions. For example, if the screen's physical dimensions are a certain number of inches diagonally and the aspect ratio is 16:9, calculate the horizontal and vertical physical dimensions, and then multiply each by the pixel density to obtain the number of horizontal and vertical pixels, i.e., the pixel dimensions.

[0085] Step S1312: Calculate the effective display area of ​​the screen based on the resolution parameters and pixel size, and remove the pixel occupancy of non-content display areas such as the status bar and navigation bar.

[0086] The resolution parameter is the total number of pixels on the screen, while the effective display area refers to the pixel area that can be used to display online menu content. The status bar and navigation bar occupy some pixels, and their pixel count needs to be subtracted from the total pixel size. For example, the status bar is located at the top of the screen with a height of several pixels, and the navigation bar is located at the bottom of the screen with a height of several pixels. Subtracting the height pixels of the status bar and navigation bar from the vertical pixel size of the screen yields the vertical pixel size of the effective display area. The horizontal pixel size is usually the total horizontal pixel size of the screen, thus determining the pixel range of the effective display area, i.e., the initial layout space.

[0087] Step S1313: Analyze the impact of screen orientation on layout space. When the screen orientation is landscape, the horizontal layout space increases and the vertical layout space decreases; when the screen orientation is portrait, the vertical layout space increases and the horizontal layout space decreases.

[0088] Screen orientation is detected by the device's gravity sensor. In landscape and portrait modes, the horizontal and vertical pixel sizes of the screen are interchanged (ignoring the influence of the status bar and navigation bar). In landscape mode, the horizontal pixel size of the effective display area is larger, suitable for displaying multiple text and graphic units side-by-side; in portrait mode, the vertical pixel size is larger, suitable for displaying text and graphic units vertically. The layout space is determined based on the change in screen orientation.

[0089] Step S1314: Based on the changing trend of screen orientation and the size of the initial layout space, calculate the layout space elasticity coefficient, which reflects the adjustable range of the layout space in different directions.

[0090] The layout space flexibility coefficient includes a horizontal flexibility coefficient and a vertical flexibility coefficient. When the initial layout space is large, the flexibility coefficient is large, allowing for a wider range of adjustments to the size of text and image units; when the initial layout space is small, the flexibility coefficient is small, limiting the adjustment range of text and image unit sizes. For example, in landscape mode, the horizontal layout space increases, and the horizontal flexibility coefficient increases accordingly, allowing for greater adjustment flexibility in the horizontal size of image units; in portrait mode, the vertical flexibility coefficient increases. The specific calculation of the flexibility coefficient comprehensively considers factors such as screen orientation, initial layout space size, and the average size requirement of text and image units.

[0091] Step S132: Detect and process the ambient light intensity in the real-time display environment parameters, and determine the visual comfort parameter based on the change in light intensity. The visual comfort parameter is used to adjust the brightness and contrast of the image unit and the font clarity of the text unit.

[0092] Step S1321: Divide the ambient light intensity into multiple intervals, such as low light interval, medium light interval, and high light interval.

[0093] Based on the ambient light intensity values ​​collected by the light sensor, the environment is divided into multiple zones. The low light zone corresponds to dim environments, such as indoor environments without lights at night; the medium light zone corresponds to normal indoor lighting environments; and the high light zone corresponds to bright outdoor environments or environments with strong light. The light intensity range of each zone is set according to the actual application scenario and the sensitivity of the device's sensors.

[0094] Step S1322: Preset corresponding visual comfort parameter benchmark values ​​for each illumination range, including image brightness benchmark value, image contrast benchmark value, and text font clarity benchmark value.

[0095] Under different lighting conditions, the human eye perceives image brightness, contrast, and text sharpness differently. In low-light conditions, to avoid excessive screen brightness that could irritate the eyes, the baseline values ​​for image brightness and contrast are set lower, while the baseline value for text sharpness is set higher to ensure readability. In high-light conditions, to improve the visibility of images and text, the baseline values ​​for image brightness and contrast are set higher, and the baseline value for text sharpness is also increased accordingly. The baseline values ​​in medium-light conditions fall between these two levels.

[0096] Step S1323: Monitor changes in ambient light intensity in real time. When the light intensity switches between different ranges, adjust the visual comfort parameters accordingly to the baseline value of the corresponding range.

[0097] By monitoring ambient light intensity in real time through a light sensor, when the light intensity is detected to change from one range to another, such as from a medium light range to a high light range, the image brightness and contrast reference values ​​in the visual comfort parameters are immediately adjusted to the reference values ​​corresponding to the high light range, and the text font clarity reference value is also adjusted to the reference value of the high light range, in order to adapt to the new lighting environment and ensure the user's viewing comfort.

[0098] Step S133: Monitor and process the network transmission rate in the real-time display environment parameters, and determine the content loading priority parameter based on the transmission rate fluctuation. The content loading priority parameter is used to sort the loading order of the graphic units.

[0099] Step S1331: Set multiple levels of network transmission rate, such as low speed, medium speed and high speed, with each level corresponding to a different rate range.

[0100] Based on common network transmission speeds, we set low-speed levels (e.g., download speeds below a certain value), medium-speed levels (download speeds within a certain range), and high-speed levels (download speeds above a certain value). These speed ranges take into account the general transmission capabilities of different network types (e.g., 2G, 3G, 4G, 5G, Wi-Fi).

[0101] Step S1332: Analyze the loading performance of image and text units under different network transmission rate levels. At low speed levels, large-capacity image units load slowly; at high speed levels, all image and text units can be loaded quickly.

[0102] In low-speed network environments, image units, due to their large data volume, take a long time to load, potentially causing excessive user wait times; text units, with their smaller data volume, load relatively quickly. At medium speeds, most image and text units can be loaded within an acceptable timeframe. At high speeds, both image and text units load rapidly with virtually no delay.

[0103] Step S1333: Determine the content loading priority parameters based on the importance of the graphic unit, the data volume, and the network transmission rate level.

[0104] The importance of image and text units is measured by a collaborative weight (the collaborative weight in step S128), with image and text units having higher collaborative weights indicating greater importance. Regarding data size, image units typically have a larger data volume than text units. Considering the network transmission rate level, at low speeds, text units with high importance and small data volume are loaded first, followed by image units with high importance but large data volume, and finally image and text units with low importance. At high speeds, multiple image and text units can be loaded simultaneously, with priority primarily based on importance. The content loading priority parameter is represented by a numerical value; a higher value indicates a higher loading priority.

[0105] Step S134: Input the layout space flexibility coefficient, visual comfort parameter, and content loading priority parameter into the environment parameter mapping module to establish a dynamic mapping relationship between environment parameters and layout adjustment parameters.

[0106] The environmental parameter mapping module contains a mapping rule library that defines the correspondence between layout space flexibility coefficients, visual comfort parameters, content loading priority parameters, and layout adjustment parameters (such as text and image unit size, position, brightness, contrast, and loading order). When the layout space flexibility coefficient is input, the module determines the adjustable range of the text and image unit size based on the coefficient value; when the visual comfort parameter is input, it determines the target adjustment values ​​for image brightness and contrast and text font clarity; when the content loading priority parameter is input, it determines the loading order of the text and image units. Through these mapping relationships, changes in environmental parameters can be reflected in the layout adjustment parameters in real time, realizing the dynamic influence of environmental parameters on layout.

[0107] Step S135: Decode the interaction sequence in the audience behavior intent data, extract the interaction features of the interaction behavior, and analyze the audience intent corresponding to the interaction features. The interaction features include frequency, duration, operation path, and operation intensity.

[0108] Step S1351: Perform type identification processing on each interactive operation in the interactive behavior sequence to distinguish different interaction types such as click operation, swipe operation, zoom operation, and long press operation.

[0109] Interaction sequence is a record of continuous user actions over a period of time, with each action having a corresponding action type identifier. By analyzing the triggering method of the action and device feedback, different interaction types are identified, such as click (user's finger lightly touches the screen and then lifts it), swipe (user's finger moves a certain distance on the screen), zoom (two fingers make an opening or pinching motion on the screen), and long press (user's finger remains in contact with the same position on the screen for a period of time).

[0110] Step S1352: Count the number of times each type of interaction occurs within a preset time window to obtain the frequency of interaction behavior.

[0111] Define a time window, such as the past 30 seconds or 1 minute, and count the number of times each of the following actions—click, swipe, zoom, and long press—occurs within that time window; this is the interaction frequency. For example, if a user clicked the "Signature Hot Dishes" category 3 times, swiped the menu 5 times, zoomed in on a dish image once, and did not long press within 30 seconds, then the click frequency is 3 times / 30 seconds, the swipe frequency is 5 times / 30 seconds, the zoom frequency is 1 time / 30 seconds, and the long press frequency is 0 times / 30 seconds.

[0112] Step S1353: Record the duration of each interactive operation from start to end to obtain the duration of the interactive behavior.

[0113] For click operations, the duration is the time from when the finger touches the screen to when it leaves the screen; for swipe operations, the duration is the time from when the finger begins to move and stops moving and leaves the screen; for zoom operations, the duration is the time from when two fingers first touch the screen to when the zoom action is completed and the finger leaves the screen; for long press operations, the duration is the total time from when the finger touches the screen to when it leaves the screen. Record the duration of each interaction operation to obtain the distribution of interaction duration.

[0114] Step S1354: Track the position coordinate changes of continuous interactive operations, form interactive operation paths, and analyze the straightness, tortuosity, and coverage of the operation paths.

[0115] The device's touch sensors record the starting and ending coordinates of each interaction. For swipe operations, changes in intermediate coordinates are also recorded. The coordinates of consecutive interactions are connected chronologically to form the interaction path. Straightness is measured by the deviation between the actual trajectory and a straight line in the calculated path; the smaller the deviation, the higher the straightness. Twists and turns are measured by the number of directional changes and the angles within the calculated path; more directional changes and larger angles result in higher twist and turns. Coverage is determined by the screen area involved in the calculated path; a larger area indicates wider coverage.

[0116] Step S1355: For devices that support pressure sensing, collect pressure data during the interactive operation to determine the intensity of the interactive operation. For devices that do not support pressure sensing, indirectly infer the intensity of the operation based on the interaction duration and operation speed.

[0117] Devices that support pressure sensing can directly collect the pressure value of a user's finger pressing the screen; the greater the pressure value, the stronger the operation. Devices that do not support pressure sensing infer the operation force based on the interaction duration and operation speed. Generally, shorter interaction times and faster operation speeds may correspond to greater operation force, while longer interaction times and slower operation speeds may correspond to less operation force. For example, a fast swipe may require more force than a slow swipe.

[0118] Step S1356: Analyze the combination relationship between interaction frequency, interaction duration, operation path characteristics and operation force. When the swipe operation frequency exceeds the preset frequency threshold and the single swipe duration is lower than the preset duration threshold, and the straightness of the operation path is higher than the preset straightness threshold and the coverage of the operation path is greater than the preset range threshold, determine the corresponding quick browsing intent.

[0119] The system presets four thresholds: swipe frequency threshold, single swipe duration threshold, operation path straightness threshold, and coverage threshold. When a user's swipe frequency exceeds the preset frequency threshold, it indicates that the user has swiped multiple times in a short period. A single swipe duration shorter than the preset duration threshold indicates that each swipe is performed quickly. An operation path straightness higher than the preset threshold indicates that the swipe direction is relatively fixed. An operation path coverage greater than the preset threshold indicates that the user is browsing a wide range of content. These characteristics combined suggest that the user is quickly browsing the menu without examining any particular dish in detail, thus indicating a rapid browsing intent.

[0120] Step S1357: The click operation frequency is lower than the preset frequency standard and the dwell time after a single click exceeds the preset dwell time standard; the proportion of zoom operations in zoom operations exceeds the preset zoom proportion standard; the operation path is concentrated in an area that meets the preset local range standard; see the corresponding detailed intent.

[0121] The system includes preset standards for click frequency, single-click dwell time, zoom ratio, and localized operation path. A low click frequency indicates that users do not frequently switch between dishes; a long single-click dwell time indicates that users spend a considerable amount of time viewing a dish after clicking on it, likely to view details; a high proportion of zoom operations indicates that users tend to zoom in to see the details of the dish image; and operation paths concentrated in a localized area indicate that users are focused on one or a few dishes. These combined characteristics correspond to a detailed viewing intent, suggesting that users want to learn more about specific dishes.

[0122] Step S1358: The click operation frequency exceeds the preset frequency standard and the operation path dispersion meets the preset dispersion standard. The proportion of short-distance sliding in the sliding operation exceeds the preset short-sliding proportion standard. Combined with the operation force fluctuation amplitude exceeding the preset force fluctuation standard, the corresponding comparison intention is selected.

[0123] The system presets standards for click frequency, operation path dispersion, short-distance swipe percentage, and operation force fluctuation. A high click frequency indicates that the user frequently clicks on different dishes; a dispersed operation path indicates that the user operates across multiple areas of the menu; a high percentage of short-distance swipes indicates that the user switches between dishes within a small area for comparison; and a large fluctuation in operation force indicates that the user's operation force varies greatly between different dishes, potentially reflecting hesitation and comparison. These characteristics combined correspond to the user's intention to compare multiple dishes in order to make a choice.

[0124] Step S1359: Verify the correspondence between different combinations of interactive features and audience intent, and correct any misjudged intent classification results.

[0125] A large number of user interaction behavior samples are collected and manually labeled to determine the actual audience intent. The intent classification results obtained through steps S1356 to S1358 are compared with the manually labeled results to calculate the classification accuracy. For misclassified samples, the deviation between their interaction feature combinations and classification rules is analyzed, and the preset thresholds and feature weights are adjusted to optimize the classification rules, thereby correcting the misclassified intent classification results and improving the accuracy of audience intent recognition.

[0126] Step S136: Input the audience intent classification results and corresponding interaction features into the behavior intent decoding module to establish the correspondence between the interaction behavior sequence and the layout strategy. Different audience intents correspond to different graphic layout strategies.

[0127] The behavioral intent decoding module internally stores a mapping table between audience intent and layout strategies. For quick browsing intents, the layout strategy should adopt a concise and clear layout, highlighting key dishes, with moderately sized image and text units, reducing redundant information, and speeding up content loading. For detailed viewing intents, the layout strategy should increase the size of the image and text units of the currently viewed dish, displaying more detailed information, such as close-up images of ingredients and detailed flavor descriptions. For selection and comparison intents, the layout strategy should display image and text units of dishes that the user may compare side by side for easy comparison, for example, arranging images and key information of multiple dishes horizontally or vertically. When the audience intent classification result and corresponding interaction characteristics are input, the module selects the appropriate layout strategy according to the mapping table, thereby establishing a correspondence between the interaction behavior sequence and the layout strategy.

[0128] Step S137: Input the temporal mapping rules, sentiment matching parameters and collaborative weights in the text-image collaboration model into the collaborative relationship application module to establish the association logic between collaborative relationships and layout parameters, and realize the effective application of spatiotemporal association and sentiment collaboration during the layout process.

[0129] The collaborative relationship application module receives the temporal mapping rules, sentiment matching parameters, and collaborative weights from the text-image collaboration model. The temporal mapping rules guide the arrangement of the display time and duration of text and image units during layout; the sentiment matching parameters guide the adjustment of the emotional expression of text and image units to match audience preferences; and the collaborative weights determine the priority order of different text and image units when there are layout conflicts or limited resources. Internally, the module establishes the association logic between these collaborative relationships and layout parameters (such as position, size, display duration, brightness, contrast, etc.). For example, the temporal mapping rules determine the arrangement order and display duration parameters of text and image units; the sentiment matching parameters adjust the brightness and contrast of images and the font color of text to enhance emotional expression; and the collaborative weights prioritize text and image units with higher collaborative weights when layout space is limited.

[0130] Step S138: Construct a typesetting parameter evolution module. The typesetting parameter evolution module includes a conflict pre-simulation algorithm, adaptive adjustment rules, and evolution iteration mechanism. It can dynamically update typesetting parameters based on the output of the environment parameter mapping module, behavior intention decoding module, and collaborative relationship application module.

[0131] Step S1381: Design a conflict simulation algorithm that can simulate the layout effect under different display environment changes (such as screen orientation rotation, light intensity changes, network speed fluctuations) and different audience interaction behaviors (such as clicking, swiping, zooming).

[0132] The conflict simulation algorithm constructs a virtual layout environment based on the outputs of the current layout parameters and environmental parameter mapping module, the behavioral intent decoding module, and the collaborative relationship application module. When simulating screen rotation, the algorithm adjusts the size and position of text and image units according to the layout space elasticity coefficient; when simulating changes in lighting intensity, it adjusts image brightness and contrast and text font clarity according to visual comfort parameters; when simulating network speed fluctuations, it adjusts the loading order and loading method of text and image units according to content loading priority parameters. Simultaneously, it simulates the impact of different audience interaction behaviors on the layout; for example, when a user clicks on a dish, the size and position of the dish's text and image unit are adjusted to highlight it. Through these simulations, the algorithm predicts the layout effect and potential conflicts.

[0133] Step S1382: Formulate adaptive adjustment rules. When layout conflicts are detected during the rehearsal, the corresponding adjustment strategy is automatically selected according to the type and severity of the conflict, such as adjusting the position, size, display duration, and loading order of graphic units.

[0134] For different types of layout conflicts (such as overlapping display areas and unbalanced loading delays), specific adaptive adjustment rules are formulated. For example, for overlapping display area conflicts, the rules stipulate that the positions of text and image units should be adjusted according to their coordination weights, with units having higher coordination weights being given priority in retaining their positions. For unbalanced loading delay conflicts, the rules stipulate that the loading order should be reordered based on the content loading priority parameter. Furthermore, the priority and magnitude of adjustments are determined based on the severity of the conflict, with severe conflicts being handled first and adjusted more significantly.

[0135] Step S1383: Establish an evolutionary iteration mechanism. By repeatedly executing the conflict pre-simulation algorithm and adaptive adjustment rules, continuously optimize the typesetting parameters until the typesetting effect meets the preset optimization target.

[0136] The evolutionary iteration mechanism sets an upper limit on the number of iterations and an optimization target threshold. In each iteration, a conflict pre-detection algorithm is first executed to detect layout conflicts; then, the layout parameters are adjusted according to adaptive adjustment rules; next, the conflict pre-detection algorithm is executed again to check if conflicts still exist in the adjusted layout. This process is repeated until the layout effect meets the optimization target (e.g., no conflicts or conflicts within an acceptable range) or the upper limit on the number of iterations is reached. In each iteration, the adjustment strategy is optimized based on the previous pre-detection results, gradually evolving the layout parameters towards the optimal direction.

[0137] Step S139: Connect the environmental parameter mapping module, behavior intent decoding module, collaborative relationship application module and typesetting parameter evolution module through the data interaction interface, set the data transmission sequence and triggering conditions between modules, and form a dynamic self-adaptive typesetting engine.

[0138] A unified data interaction interface was designed for the four modules, defining data formats and communication protocols to ensure correct data transmission between modules. The timing of data transmission between modules was set. For example, the environment parameter mapping module converts environment parameters into layout adjustment parameters in real time and passes them to the collaboration relationship application module and the layout parameter evolution module; the behavior intent decoding module periodically (e.g., every few seconds) passes the identified audience intent to the collaboration relationship application module and the layout parameter evolution module; the collaboration relationship application module generates preliminary layout parameters based on the received environment layout adjustment parameters and audience intent, and passes them to the layout parameter evolution module. Triggering conditions include the environment parameter mapping module updating data when environmental parameter changes exceed a preset threshold, the behavior intent decoding module updating audience intent when the interaction behavior sequence accumulates to a certain number, and the layout parameter evolution module optimizing after the preliminary layout parameters are generated. Through these connections and settings, the four modules work collaboratively to form a dynamic, self-adaptive layout engine.

[0139] Step S140: Apply the dynamic adaptive typesetting engine to dynamically calculate the typesetting parameters of the image units and text units in the digital multimedia graphic and text collection to be optimized, generate an initial typesetting scheme, and use the typesetting parameter evolution module of the dynamic adaptive typesetting engine to perform multi-dimensional conflict pre-simulation and adaptive adjustment of the initial typesetting scheme to obtain an intermediate typesetting scheme.

[0140] Step S141: Input the attribute information of the image units and text units in the digital multimedia graphic and text set to be optimized into the collaborative relationship application module of the dynamic self-adaptive typesetting engine. Combine the temporal mapping rules and sentiment matching parameters in the graphic and text collaboration model to determine the basic typesetting parameters of each graphic and text unit. The basic typesetting parameters include the initial position, initial size and initial display duration.

[0141] The attribute information of image units includes image resolution, size, format, and data volume; the attribute information of text units includes text length, number of characters, and number of sentiment words contained. Based on this attribute information and the temporal mapping rules in the image-text collaboration model, the collaborative relationship application module determines the initial display duration of text and image units, matching the text reading time with the image display time. According to sentiment matching parameters, the initial size of the image and text units is adjusted; image and text units with high sentiment feature values ​​and those that align with audience preferences have larger initial sizes. The initial position is determined based on the dish category and default layout order. For example, image and text units under the "Signature Hot Dishes" category are arranged from top to bottom and left to right according to the default order of dishes in the menu. Through these processes, the basic layout parameters for each image and text unit are obtained.

[0142] Step S142: Input the real-time display environment parameters into the environment parameter mapping module to obtain the layout space flexibility coefficient, visual comfort parameter and content loading priority parameter. Based on the layout space flexibility coefficient, visual comfort parameter and content loading priority parameter, adjust the basic layout parameters to adapt to the environment and obtain the environment-adapted layout parameters.

[0143] The environment parameter mapping module calculates the layout space flexibility coefficient, visual comfort parameter, and content loading priority parameter based on the input real-time display environment parameters. Based on the layout space flexibility coefficient, the initial size of the text and image units is adjusted; if the flexibility coefficient is large, the size of the text and image units can be appropriately increased or decreased to adapt to the layout space. Based on the visual comfort parameter, the initial brightness and contrast of the image units, as well as the font size and clarity of the text units, are adjusted. Based on the content loading priority parameter, the initial loading order of the text and image units is adjusted, with higher-priority units loaded first. These adjustments are applied to the basic layout parameters to obtain the environment-adaptive layout parameters.

[0144] Step S143: Input the interaction behavior sequence in the audience behavior intent data into the behavior intent decoding module, determine the audience intent classification result, and adjust the environment adaptation layout parameters based on the audience intent classification result to obtain the strategy adaptation layout parameters.

[0145] The behavioral intent decoding module decodes the input sequence of interactive behaviors to determine the audience intent classification result (e.g., quick browsing, detailed viewing, selection and comparison). For the quick browsing intent, the strategy adaptation adjustment reduces the size of the text and image units, increases the layout density, and speeds up content loading; for the detailed viewing intent, it increases the size of the currently viewed text and image unit to highlight details; for the selection and comparison intent, it adjusts potentially comparable text and image units to adjacent positions for easier comparison. Applying these strategy adjustments to the environment adaptation layout parameters yields the strategy adaptation layout parameters.

[0146] Step S144: Integrate the strategy adaptation and layout parameters of all graphic and text units, determine the final position coordinates, display size, display duration and loading order of each graphic and text unit, and generate an initial layout scheme.

[0147] Summarize the strategy adaptation and layout parameters for all text and image units. For position coordinates, ensure that the text and image units do not overlap (preliminary assessment) and comply with the layout space limitations. The display size is determined based on the adjusted parameters. The display duration is set according to the time sequence mapping rules. The loading order is arranged according to the content loading priority parameters. Organize the above information into a structured data format, including the unique identifier of each text and image unit and its corresponding final position coordinates, display size, display duration, and loading order, to form the initial layout scheme.

[0148] Step S145: Input the initial layout scheme into the layout parameter evolution module, start the conflict pre-simulation algorithm, and simulate the layout effect under different display environment changes and different audience interaction behaviors.

[0149] After receiving the initial layout scheme, the layout parameter evolution module initiates the conflict pre-simulation algorithm. The algorithm simulates various possible changes in the display environment, such as screen orientation changing from portrait to landscape, ambient light intensity changing from low to high, and network transmission speed changing from high to low. It also simulates different audience interaction behaviors, such as users clicking on a dish, swiping a menu, or zooming in and out of an image. In each simulated scenario, the algorithm calculates the display effect of the text and image units based on the initial layout scheme and records relevant data, such as the position, size, loading time, and visual comfort parameters of each text and image unit.

[0150] Step S146: Based on the pre-playback results, identify possible layout conflicts, including overlapping display areas, unbalanced loading delays, decreased visual comfort, and broken collaborative relationships.

[0151] For example, step S1461: Extract the display area coordinates of all graphic units in different simulated scenarios from the pre-play results, compare the display area coordinates of any two graphic units, if there is a coordinate intersection, and the area of ​​the intersection exceeds the proportion of the area of ​​any graphic unit itself, it is determined to be a display area overlap conflict.

[0152] The display area coordinates are represented by the pixel coordinates of the top-left and bottom-right corners of the text / image unit, forming a rectangular area. For any two text / image units, the intersection of their display area coordinates is calculated, and the area of ​​the intersection is calculated by multiplying the width and height of the rectangular intersection. The intersection area is divided by the area of ​​the smaller text / image unit to obtain the percentage of the intersection area. If this percentage exceeds a preset threshold, it indicates that the two text / image units overlap significantly, affecting user viewing, and is determined to be a display area overlap conflict.

[0153] Step S1462: Extract the loading completion time of each graphic unit in the pre-simulation results, calculate the difference between the maximum and minimum loading completion times of all graphic units, and if the difference exceeds the preset time difference threshold, it is determined that there is a loading delay imbalance conflict.

[0154] Loading completion time refers to the time required from the start of loading a text / image unit to its full display on the screen. In the pre-visualization results, the loading completion time for each text / image unit is recorded. The maximum and minimum loading completion times are identified, and their difference is calculated. If this difference exceeds a preset time difference threshold, it indicates that the loading speed of different text / image units varies too much; some units have finished loading while others are still loading, resulting in an inconsistent user experience, and this is classified as a loading latency imbalance conflict.

[0155] Step S1463: According to the visual comfort parameter change curve in the pre-simulation results, if the brightness and contrast of the image unit are not adjusted in time after the ambient light intensity changes, resulting in the visual comfort parameter being lower than the preset comfort threshold, or the font clarity of the text unit being lower than the preset clarity threshold, it is determined to be a visual comfort decrease conflict.

[0156] The visual comfort parameter change curve records the changes in visual comfort parameters during changes in ambient light intensity. When ambient light intensity changes, the brightness and contrast of image units should be adjusted accordingly to maintain visual comfort parameters above a preset comfort threshold; the font clarity of text units should also remain above a preset clarity threshold. If, in the preview results, the visual comfort parameters or font clarity fall below the preset threshold after a change in light intensity, it indicates that the layout scheme has failed to adapt to the light change in a timely manner, and this is judged as a visual comfort decrease conflict.

[0157] Step S1464: Check the spatiotemporal correlation and emotional coordination relationship of the text and image units in the pre-run results. If the screen rotation causes the deviation between the text reading time segment and the optimal display time of the image to exceed the preset corresponding deviation threshold, or if the change of environmental parameters causes the deviation between the text emotional feature value and the image emotional feature value to exceed the preset emotional deviation threshold, it is determined to be a break or conflict in the coordination relationship.

[0158] Screen rotation alters the layout space, potentially changing the correspondence between text reading time segments and optimal image display time. The deviation in this correspondence after rotation is calculated; if it exceeds a preset deviation threshold, the spatiotemporal relationship is disrupted. Changes in environmental parameters (such as lighting affecting image color) may increase the deviation between text sentiment features and image sentiment features; if this deviation exceeds a preset sentiment deviation threshold, the sentiment coordination relationship is disrupted. Both of these situations are considered disruptions or conflicts in the coordination relationship.

[0159] Step S1465: Create a list of typesetting conflicts, sort all identified typesetting conflicts by severity, and prioritize conflicts with a severity level higher than the preset high priority threshold.

[0160] Based on the impact of layout conflicts on user experience, severity scoring criteria are set for each conflict type. For example, the severity of overlapping display area conflicts is positively correlated with the percentage of overlapping area, and the severity of loading delay imbalance conflicts is positively correlated with the time difference. Each identified layout conflict is scored, and the scores are sorted from high to low to create a layout conflict list. A high-priority threshold is preset; conflicts with scores higher than this threshold are marked as high-priority conflicts and are handled first in subsequent adjustments.

[0161] Step S147: Adjust the layout parameters that have conflict risks according to the adaptive adjustment rules. The adjustment operations include adjusting the position coordinates of overlapping units, optimizing the loading order to balance the delay, adjusting the brightness and contrast to improve comfort, and correcting the timing mapping rules to repair the collaborative relationship.

[0162] Step S1471: For overlapping conflicts in the display area, query the collaborative weight and audience intent relevance of the conflicting graphic units. Graphic units with a collaborative weight greater than a preset high weight threshold and an audience intent relevance greater than a preset high relevance threshold are given priority to retain their original position coordinates. Graphic units with a collaborative weight less than a preset low weight threshold and an audience intent relevance less than a preset low relevance threshold are moved away from the conflict area. The moving distance is positively correlated with the size of the conflict intersection area, and the moved position conforms to the layout space elasticity coefficient constraint.

[0163] The collaboration weights are obtained from the text-image collaboration model, and the audience intent relevance is obtained by analyzing the matching degree between the text-image unit and the current audience intent. For text-image units involved in a conflict, their collaboration weights and audience intent relevance are compared. Text-image units with high weights and high relevance retain their original positions. For text-image units with low weights and low relevance, the movement distance is determined based on the size of the conflict intersection area; the larger the intersection area, the farther the movement distance. The movement direction is chosen to be away from the conflict area; for example, if the conflict area is on the right, it moves to the left; if the conflict area is on the top, it moves downwards. The moved position cannot exceed the size adaptation range determined by the layout space flexibility coefficient.

[0164] Step S1472: To address the loading delay imbalance conflict, based on the content loading priority parameter and the display duration of the graphic unit, the loading order of the graphic units is reordered. Graphic units whose display duration exceeds the preset display duration threshold and whose audience intent relevance is greater than the preset high relevance threshold are loaded first. For graphic units whose loading speed is lower than the preset loading speed threshold, their size parameters are adjusted to reduce the amount of data, or a progressive loading method is adopted, loading low-resolution image data first to meet basic display requirements, and then loading high-resolution image data to balance the loading delay.

[0165] The loading priority parameter and the relevance to audience intent determine the loading priority of text and image units. Units with long display durations and high relevance should be loaded first. For slow-loading text and image units, if they are image units, their size can be appropriately reduced (within the allowable range of the layout space flexibility coefficient) to reduce the amount of image data and improve loading speed; or a progressive loading method can be used. For text units, the initial loading content can be simplified, loading core information (such as dish names and prices) first, followed by detailed descriptions. These methods balance the loading completion time of different text and image units, reduce time differences, and resolve loading delay imbalances.

[0166] Step S1473: To address the conflict of decreased visual comfort, adjust the brightness and contrast parameters of the image unit according to the image brightness contrast benchmark value and text font clarity benchmark value in the visual comfort parameters to make them reach the benchmark value; adjust the font size, character spacing and line spacing of the text unit to ensure that the font clarity is not lower than the preset clarity threshold under different lighting conditions.

[0167] Based on the baseline values ​​of visual comfort parameters corresponding to ambient light intensity, the brightness and contrast of image units are adjusted. For example, in high-light environments, image brightness and contrast are increased; in low-light environments, image brightness and contrast are decreased. For text units, font size is increased, and character spacing and line spacing are adjusted to improve font clarity, ensuring that users can clearly read text content under various lighting conditions, thus restoring visual comfort parameters to above the preset threshold.

[0168] Step S1474: To address the conflict in the collaborative relationship, recalculate the temporal matching relationship between the spatiotemporal features of the text and the spatiotemporal features of the image, correct the temporal mapping rules based on the current display environment parameters, and adjust the deviation between the reading time segment and the optimal display time to no more than the preset corresponding deviation threshold; adjust the emotional feature values ​​of the image unit and the text unit, adjust the emotional deviation between the two to no more than the preset emotional deviation threshold, and repair the emotional collaborative relationship.

[0169] Based on the current display environment parameters (such as screen orientation and size), the text reading time segments and optimal image display time are recalculated, and the correspondence in the time-series mapping rules is corrected to keep the deviation within a preset range. For emotional synergy, the deviation between the emotional feature values ​​of the text and images and the audience's emotional preference feature values ​​is analyzed, and the emotional words in the text or the color and composition of the images are adjusted to reduce the deviation of the emotional feature values ​​to below a preset emotional deviation threshold, thereby repairing the synergy.

[0170] Step S148: Repeat the conflict pre-simulation algorithm and adaptive adjustment rules until there are no layout conflicts in the pre-simulation results, record the adjusted layout parameters, and generate an intermediate layout scheme.

[0171] After completing one round of layout parameter adjustments, the conflict pre-simulation algorithm is executed again to check for any new layout conflicts. If conflicts still exist, the adaptive adjustment rules are applied, and this process is repeated until no layout conflicts appear in the pre-simulation results, or the severity of the conflicts is below the preset acceptable threshold. At this point, all adjusted layout parameters are recorded, including the position coordinates, size, display duration, loading order, brightness, and contrast of the text and image units. These parameters are then organized into structured data to generate an intermediate layout scheme.

[0172] Step S150: Deploy the intermediate layout scheme to the real-time display environment, combine the content demand description in the audience behavior intent data and the real-time updated display environment parameters, drive the dynamic self-adaptive layout engine to continuously evolve the layout scheme, generate a target graphic and text layout optimization scheme that conforms to the real-time environment and audience intent, and output the target graphic and text layout optimization scheme for digital multimedia content display.

[0173] Step S151: Deploy the intermediate layout scheme to the digital multimedia real-time display environment and establish a real-time data connection between the layout scheme and the display environment parameter acquisition module and the audience behavior intent acquisition module.

[0174] Through a digital multimedia content management system, the intermediate layout scheme is deployed to users' mobile devices, enabling it to be displayed in the menu interface of the restaurant app. Simultaneously, a data connection channel is established within the app between the layout scheme and the modules for collecting display environment parameters and audience behavior intent, ensuring that the layout scheme can receive real-time data updates from these two modules.

[0175] Step S152: By using the display environment parameter acquisition module, real-time data is collected on changes in device screen characteristics, fluctuations in ambient light intensity, and changes in network transmission rate, forming a real-time updated display environment parameter stream.

[0176] The display environment parameter acquisition module continuously monitors the device's screen characteristics (such as sudden screen orientation rotation), ambient light intensity (such as a sudden increase in light intensity caused by a user moving from indoors to outdoors), and network transmission rate (such as a change in rate caused by switching from Wi-Fi to 4G). It organizes the above-mentioned change data into a real-time updated display environment parameter stream in chronological order and continuously sends it to the dynamic self-adaptive typesetting engine.

[0177] Step S153: Through the audience behavior intent collection module, new interactive behavior sequences of the audience are collected in real time during the process of browsing the intermediate layout scheme. At the same time, the content demand expressions input by the audience are collected to form a real-time updated audience behavior intent data stream.

[0178] The audience behavior intent collection module records new user interactions while browsing intermediate layout schemes, such as new clicks, swipes, and zooms, forming new interaction behavior sequences. Simultaneously, it collects user-inputted content requests, such as "light dishes" and "special offers," through the app's search box and filter settings. These new interaction behavior sequences and content request statements are organized into a real-time updated audience behavior intent data stream and sent to the dynamic adaptive layout engine.

[0179] Step S154: Input the real-time updated display environment parameter stream into the environment parameter mapping module of the dynamic self-adaptive typesetting engine, update the dynamic mapping relationship between environment parameters and typesetting adjustment parameters, and output new typesetting space elasticity coefficient, visual comfort parameter and content loading priority parameter.

[0180] The environment parameter mapping module receives a real-time updated stream of display environment parameters and recalculates the layout space flexibility coefficient, visual comfort parameter, and content loading priority parameter. For example, after screen orientation rotation, the flexibility coefficient is recalculated; after fluctuations in lighting intensity, the baseline value of the visual comfort parameter is updated; and after changes in network speed, the content loading priority is adjusted. These new parameters are then output to the collaboration application module and the layout parameter evolution module.

[0181] Step S155: Input the real-time updated audience behavior intent data stream into the behavior intent decoding module, update the correspondence between the interaction behavior sequence and the layout strategy, correct the audience intent classification results in combination with the content requirement description, and output a new layout strategy.

[0182] The behavioral intent decoding module decodes new interaction sequences and, combined with content requirement descriptions, corrects previous audience intent classification results. For example, if a user previously exhibited a comparison intent, entering "light dishes" might correct their intent to view detailed information about light dishes. Based on the corrected audience intent, the module updates the correspondence between the interaction sequence and the layout strategy, outputting a new layout strategy.

[0183] Step S156: Based on the new layout strategy and updated environmental parameters, the collaborative relationship application module adjusts the temporal mapping rules, sentiment matching parameters, and collaborative weights in the graphic-text collaborative model, and outputs new collaborative application logic.

[0184] The collaborative relationship application module adjusts the temporal mapping rules in the text-image collaboration model based on the new layout strategy and updated environmental parameters to adapt to the new display environment and audience intent; it updates the sentiment matching parameters to make the emotional expression of text-image units better match the user's current content needs; and it adjusts the collaboration weights, increasing the collaboration weights of text-image units related to content needs. Based on these adjustments, it outputs new collaborative application logic to guide the calculation of layout parameters.

[0185] Step S157: Based on the new environmental parameter mapping relationship, layout strategy and collaborative application logic, the layout parameter evolution module updates the layout parameters of the intermediate layout scheme in real time, and adjusts the position coordinates, display size, display duration and loading order of graphic units.

[0186] The layout parameter evolution module integrates new environmental parameter mapping relationships (layout space flexibility coefficient, visual comfort parameter, content loading priority parameter), new layout strategies, and new collaborative application logic to update the layout parameters of the intermediate layout scheme in real time. It adjusts the position coordinates of text and image units to adapt to screen orientation changes, adjusts display sizes to conform to the flexibility coefficient and layout strategy, adjusts display duration to match new timing mapping rules, and adjusts loading order to reflect new content loading priorities.

[0187] Step S158: During the typesetting parameter update process, the conflict pre-simulation algorithm is executed synchronously, and no new conflicts are generated in the updated typesetting scheme.

[0188] While adjusting the layout parameters, the layout parameter evolution module simultaneously executes a conflict pre-simulation algorithm to simulate the display effect of the updated layout scheme under the current environment and audience behavior, and to check for any new layout conflicts. If a conflict is found, it is immediately readjusted according to the adaptive adjustment rules until the updated layout scheme is conflict-free.

[0189] Step S159: Continuously repeat the steps of real-time data collection, module parameter update, layout parameter adjustment and conflict rehearsal until there are no new content requirements in the audience behavior intent data stream, and the fluctuation range of each parameter in the real-time display environment parameter stream remains stable within their respective preset threshold range.

[0190] The dynamic self-adaptive typesetting engine continuously loops through steps S152 to S158, constantly collecting real-time data, updating module parameters, adjusting typesetting parameters, and performing conflict rehearsals. When the audience does not input any new content requirements within a certain period of time, and the fluctuation range of display environment parameters (screen characteristics, light intensity, network speed) is within the preset threshold (e.g., the change in light intensity within a few minutes is less than the preset value), it indicates that user needs and the display environment are stabilizing.

[0191] Step S1510: Determine the final stable layout scheme as the target graphic layout optimization scheme.

[0192] When the stability condition of step S159 is met, the dynamic self-adaptive typesetting engine stops adjusting the typesetting parameters and determines the current typesetting scheme as the target graphic and text typesetting optimization scheme. This target graphic and text typesetting optimization scheme can adapt to the real-time display environment and audience intent, achieving the best typesetting effect for online menu graphics and text.

[0193] Step S1511: Output the target graphic layout optimization scheme for digital multimedia content display.

[0194] The target graphic and text layout optimization plan is sent to the user's mobile terminal device via a data interface. The restaurant APP displays the optimized online menu on the screen according to the plan, including the position, size, brightness, contrast, and loading order of graphic and text units, providing users with a good menu browsing experience.

[0195] Figure 2This illustration shows an intelligent text and image layout optimization system 100 integrating digital multimedia, provided in an embodiment of this application. The system includes a processor 1001, a memory 1003, and program code stored in the memory 1003. The processor 1001 executes the program code to implement the steps of the intelligent text and image layout optimization method integrating digital multimedia. The processor 1001 and the memory 1003 are connected, for example, via a bus 1002. Optionally, the intelligent text and image layout optimization system 100 may further include a transceiver 1004. The transceiver 1004 can be used for data interaction between this intelligent text and image layout optimization system and other intelligent text and image layout optimization systems integrating digital multimedia, such as sending and / or receiving data. It should be noted that in actual scheduling, the transceiver 1004 is not limited to one, and the structure of this intelligent text and image layout optimization system 100 integrating digital multimedia does not constitute a limitation on the embodiments of this application.

[0196] The memory 1003 is used to store program code for executing the embodiments of this application, and its execution is controlled by the processor 1001. The processor 1001 is used to execute the program code stored in the memory 1003 to implement the steps shown in the foregoing method embodiments.

[0197] This application provides a computer-readable storage medium storing program code, which, when executed by a processor, can implement the steps and corresponding content of the aforementioned method embodiments.

[0198] The above description is only an optional implementation method for some implementation scenarios of this application. It should be noted that for those skilled in the art, other similar implementation methods based on the technical concept of this application, without departing from the technical concept of this application, also fall within the protection scope of the embodiments of this application.

Claims

1. A method for intelligent optimization of text and image layout combining digital multimedia, characterized in that, The method includes: The system acquires a set of digital multimedia images and texts to be optimized, real-time display environment parameters, and audience behavioral intent data. The set of digital multimedia images and texts to be optimized includes multiple image units and multiple text units. The real-time display environment parameters include device screen characteristics, ambient light intensity, and network transmission rate. The audience behavioral intent data includes interactive behavior sequences, sentiment markers, and content demand expressions. The image units and text units in the digital multimedia image and text set to be optimized are modeled in a spatiotemporal manner. Combined with the sentiment tendency markers in the audience behavior intention data, the spatiotemporal correlation and sentiment synergy relationship between the image units and text units are established to obtain the image and text synergy model. Based on the real-time display environment parameters, the interaction behavior sequence in the audience behavior intent data, and the graphic-text collaboration model, a dynamic self-adaptive typesetting engine is constructed. The dynamic self-adaptive typesetting engine includes an environment parameter mapping module, a behavior intent decoding module, a collaboration relationship application module, and a typesetting parameter evolution module. The dynamic adaptive typesetting engine is used to dynamically calculate the typesetting parameters of the image units and text units in the digital multimedia graphic and text collection to be optimized, and generate an initial typesetting scheme. The typesetting parameter evolution module of the dynamic adaptive typesetting engine is used to perform multi-dimensional conflict pre-simulation and adaptive adjustment on the initial typesetting scheme to obtain an intermediate typesetting scheme. The intermediate layout scheme is deployed to the real-time display environment. Combining the content demand description in the audience behavior intent data and the real-time updated display environment parameters, the dynamic self-adaptive layout engine is driven to continuously evolve the layout scheme, generate a target graphic and text layout optimization scheme that conforms to the real-time environment and audience intent, and output the target graphic and text layout optimization scheme for digital multimedia content display.

2. The intelligent optimization method for graphic and text layout combining digital multimedia as described in claim 1, characterized in that, The step involves performing spatiotemporal collaborative modeling of image and text units in the digital multimedia image and text set to be optimized. This is combined with sentiment markers from the audience's behavioral intent data to establish spatiotemporal correlations and sentiment collaborations between image and text units, resulting in an image-text collaboration model. The model includes: Each text unit is analyzed for reading rhythm, extracting sentence pause markers, changes in word density, and semantic transition positions to determine the reading time segments and key semantic intervals of the text unit, thus forming the spatiotemporal characteristics of the text. Visual presentation timing analysis is performed on each image unit to extract the visual focus switching path, color transition rhythm and detail information hierarchy in the image unit, determine the optimal display duration and visual attention order of the image unit, and form the spatiotemporal characteristics of the image. The spatiotemporal features of text and images are matched in a temporal sequence. Based on the correspondence between reading time segments and optimal display time, and the correlation between key semantic intervals and visual attention order, the spatiotemporal relationship between image units and text units is established. Sentiment extraction is performed on each text unit to analyze the distribution of sentiment words, tone expression, and semantic sentiment intensity in the text unit, and to determine the sentiment feature value of the text unit. For each image unit, sentiment extraction processing is performed to analyze the color sentiment attributes, compositional sentiment expression, and visual element sentiment symbolism in the image unit, and to determine the sentiment feature value of the image unit. From the sentiment tendency tags of the audience's behavioral intention data, the sentiment preference feature values ​​of the audience for similar images and texts are extracted as a reference benchmark for sentiment coordination; The emotional feature values ​​of text units, image units, and audience emotional preference feature values ​​are collaboratively matched to adjust the deviation between the emotional feature values ​​of image units and text units, and to establish an emotional synergy relationship. By integrating spatiotemporal relationships and emotional collaboration, a text-image collaboration model is constructed, which includes temporal mapping rules, emotional matching parameters, and collaboration weights.

3. The intelligent optimization method for graphic and text layout combining digital multimedia as described in claim 1, characterized in that, The construction of a dynamic self-adaptive typesetting engine based on the real-time display environment parameters, the interaction behavior sequence in the audience behavior intent data, and the graphic-text collaboration model includes: The device screen characteristics in the real-time display environment parameters are analyzed and processed to extract the screen size specifications, resolution parameters, screen orientation and pixel density. The initial layout space is calculated based on the screen size specifications, resolution parameters and pixel density, and the layout space elasticity coefficient is determined in combination with the screen orientation. The layout space elasticity coefficient is used to dynamically adjust the size adaptation range of the graphic unit. The ambient light intensity in the real-time display environment parameters is detected and processed, and the visual comfort parameter is determined based on the change in light intensity. The visual comfort parameter is used to adjust the brightness and contrast of the image unit and the font clarity of the text unit. The network transmission rate in the real-time display environment parameters is monitored and processed, and the content loading priority parameter is determined based on the transmission rate fluctuation. The content loading priority parameter is used to sort the loading order of graphic units. Input the layout space flexibility coefficient, visual comfort parameter, and content loading priority parameter into the environment parameter mapping module to establish a dynamic mapping relationship between environment parameters and layout adjustment parameters; The interaction sequence in the audience behavior intent data is decoded, the interaction features of the interaction are extracted, and the audience intent corresponding to the interaction features is analyzed. The interaction features include frequency, duration, operation path and operation intensity. The audience intent classification results and corresponding interaction features are input into the behavior intent decoding module to establish the correspondence between the interaction behavior sequence and the layout strategy. Different audience intents correspond to different graphic layout strategies. The temporal mapping rules, sentiment matching parameters, and collaborative weights in the graphic-text collaboration model are input into the collaborative relationship application module to establish the association logic between collaborative relationships and layout parameters, thereby realizing the effective application of spatiotemporal association and sentiment collaboration during the layout process. A layout parameter evolution module is constructed, which includes a conflict pre-simulation algorithm, adaptive adjustment rules, and an evolution iteration mechanism. It can dynamically update layout parameters based on the outputs of the environment parameter mapping module, the behavior intention decoding module, and the collaborative relationship application module. The environmental parameter mapping module, behavioral intent decoding module, collaborative relationship application module, and typesetting parameter evolution module are connected through a data interaction interface. The data transmission timing and triggering conditions between modules are set to form a dynamic self-adaptive typesetting engine.

4. The intelligent optimization method for graphic and text layout combining digital multimedia as described in claim 1, characterized in that, The application of the dynamic adaptive typesetting engine dynamically calculates the typesetting parameters of image units and text units in the digital multimedia graphic and text set to be optimized, generating an initial typesetting scheme. The typesetting parameter evolution module of the dynamic adaptive typesetting engine then performs multi-dimensional conflict pre-simulation and adaptive adjustment on the initial typesetting scheme to obtain an intermediate typesetting scheme, including: The attribute information of the image units and text units in the digital multimedia graphic and text set to be optimized is input into the collaborative relationship application module of the dynamic self-adaptive typesetting engine. Combined with the temporal mapping rules and sentiment matching parameters in the graphic and text collaboration model, the basic typesetting parameters of each graphic and text unit are determined. The basic typesetting parameters include the initial position, initial size and initial display duration. The real-time display environment parameters are input into the environment parameter mapping module to obtain the layout space flexibility coefficient, visual comfort parameter and content loading priority parameter. Based on the layout space flexibility coefficient, visual comfort parameter and content loading priority parameter, the basic layout parameters are adjusted for environmental adaptation to obtain the environment-adapted layout parameters. The interaction behavior sequence in the audience behavior intent data is input into the behavior intent decoding module to determine the audience intent classification result. Based on the audience intent classification result, the environment adaptation layout parameters are adjusted to obtain the strategy adaptation layout parameters. Integrate the strategy adaptation and layout parameters of all graphic and text units, determine the final position coordinates, display size, display duration and loading order of each graphic and text unit, and generate an initial layout scheme; Input the initial layout scheme into the layout parameter evolution module, start the conflict pre-simulation algorithm, and simulate the layout effect under different display environment changes and different audience interaction behaviors; Based on the pre-simulation results, potential layout conflicts are identified, including overlapping display areas, unbalanced loading delays, decreased visual comfort, and broken collaborative relationships. Based on the adaptive adjustment rules, the layout parameters that have conflict risks are adjusted. The adjustment operations include adjusting the position coordinates of overlapping units, optimizing the loading order to balance the delay, adjusting the brightness and contrast to improve comfort, and correcting the timing mapping rules to repair the collaborative relationship. Repeat the conflict pre-simulation algorithm and adaptive adjustment rules until there are no layout conflicts in the pre-simulation results. Record the adjusted layout parameters and generate an intermediate layout scheme.

5. The intelligent optimization method for graphic and text layout combining digital multimedia as described in claim 2, characterized in that, The process of analyzing the reading rhythm of each text unit involves extracting sentence pause markers, changes in lexical density, and semantic transition locations within the text unit. This determines the reading time segments and key semantic intervals of the text unit, forming the spatiotemporal features of the text, including: Perform sentence structure parsing on each text unit, identify commas, periods, exclamation marks, and question marks as sentence pause markers in the text unit, and count the frequency and interval of different pause markers. The text unit is divided into multiple sentence segments based on the number of characters between the pause marks, with each sentence segment bounded by two adjacent pause marks. For each sentence segment, perform lexical density calculation, count the number of words and characters in each sentence segment, and calculate the lexical density value per unit character length. Analyze the changing trend of word density values, identify sentence fragments whose word density values ​​exceed a preset density threshold, and these sentence fragments correspond to areas that need to be focused on during reading; Semantic analysis is performed on text units to identify semantic transition conjunctions and semantic emphasis words, and to determine the positions of semantic transitions and semantic emphasis. Based on the results of sentence segmentation, the trend of word density change, and the results of semantic analysis, the text unit is divided into multiple reading time segments. The segments with word density values ​​exceeding the preset density threshold and prominent semantic points correspond to segments with reading times exceeding the preset time threshold. Mark the sentence segments where the semantic focus is located to determine the key semantic range of the text unit; By integrating data on reading time segments, key semantic intervals, sentence pause marker distribution, and word density changes, spatiotemporal features of the text are formed.

6. The intelligent optimization method for graphic and text layout combining digital multimedia as described in claim 2, characterized in that, The step involves performing visual presentation timing analysis on each image unit, extracting the visual focus switching path, color transition rhythm, and detail information hierarchy within the image unit, determining the optimal display duration and visual attention sequence for each image unit, and forming the spatiotemporal characteristics of the image, including: Visual focus detection is performed on each image unit, and multiple visual focuses in the image unit are identified based on color contrast, brightness difference, edge complexity and object salience. Analyze the positional relationships, size differences, and attractiveness of visual focal points, simulate the eye movement trajectory of the audience when viewing the image, and determine the path of visual focal point switching. Color analysis is performed on image units to extract the color distribution and color transition patterns in different regions of the image units, identify regions with significant color transitions, calculate their areas, and determine the color transition rhythm characteristics. The image unit is processed for detail information extraction. Based on the image resolution, texture complexity and object detail richness, the image unit is divided into multiple detail information levels. The levels with rich detail elements to the levels with simple detail elements correspond to different visual attention depths. Based on the physical length of the visual focus switching path and the number of visual focuses, calculate the basic time required for the audience to fully browse all visual focuses; The base time is adjusted based on the proportion of the color transition area to the total image area. When the proportion of the color transition area exceeds the preset threshold, the browsing time is increased accordingly. Based on the number and complexity of the detailed information levels, the browsing time is further adjusted. When the proportion of levels with rich detailed elements exceeds the preset proportion threshold, the browsing time is increased accordingly, and the optimal display time of the image unit is finally determined. Based on the visual focus switching path and the level of detail information, the order of visual attention when the audience views the image is determined. First, attention is paid to areas where the visual focus attractiveness is greater than the preset attractiveness threshold and areas with rich detail information levels, and then attention is paid to other areas. By integrating the visual focus switching path, color transition rhythm, detail information hierarchy, optimal display duration, and visual attention order, the spatiotemporal characteristics of the image are formed.

7. The intelligent optimization method for graphic and text layout combining digital multimedia as described in claim 3, characterized in that, The process of decoding the interaction sequence in the audience behavior intent data, extracting the interaction features of the interaction behaviors, and analyzing the audience intent corresponding to the interaction features includes: Perform type identification processing on each interactive operation in the interactive behavior sequence to distinguish different interaction types such as click operation, swipe operation, zoom operation, and long press operation; Count the number of times each type of interaction occurs within a preset time window to obtain the frequency of interaction behavior; Record the duration of each interaction from start to finish to obtain the duration of the interaction behavior; Track the position coordinate changes of continuous interactive operations to form interactive operation paths, and analyze the straightness, tortuosity and coverage of the operation paths; For devices that support pressure sensing, pressure data during interactive operation is collected to determine the intensity of the interactive operation. For devices that do not support pressure sensing, the intensity of the operation is indirectly inferred based on the interaction duration and operation speed. Analyze the combination relationship between interaction frequency, interaction duration, operation path characteristics and operation force. When the swipe operation frequency exceeds the preset frequency threshold and the single swipe duration is less than the preset duration threshold, and at the same time the straightness of the operation path is higher than the preset straightness threshold and the coverage of the operation path is greater than the preset range threshold, the corresponding quick browsing intent is determined. The click frequency is lower than the preset frequency standard and the dwell time after a single click exceeds the preset dwell time standard. The zoom operation ratio exceeds the preset zoom ratio standard. The operation path is concentrated in an area that meets the preset local range standard, which corresponds to a type of audience intent. The frequency of click operations exceeds the preset frequency standard and the dispersion of operation paths meets the preset dispersion standard. The proportion of short-distance swipes in swipe operations exceeds the preset short-swipe proportion standard. Combined with the fluctuation of operation force exceeding the preset force fluctuation standard, it corresponds to a type of audience intent. The correspondence between different combinations of interactive features and audience intent is verified, and misjudged intent classification results are corrected.

8. The intelligent optimization method for graphic and text layout combining digital multimedia as described in claim 4, characterized in that, The adjustment of layout parameters with potential conflict risks according to adaptive adjustment rules includes: To address overlapping conflicts in the display area, the system queries the collaborative weight and audience intent relevance of the conflicting graphic and text units. Graphic and text units with a collaborative weight greater than a preset high weight threshold and an audience intent relevance greater than a preset high relevance threshold are given priority to retain their original position coordinates. Graphic and text units with a collaborative weight less than a preset low weight threshold and an audience intent relevance less than a preset low relevance threshold are moved away from the conflict area. The moving distance is positively correlated with the size of the conflict intersection area, and the moved position conforms to the layout space elasticity coefficient constraint. To address loading delay imbalance conflicts, the loading order of text and image units is reordered based on content loading priority parameters and display duration of text and image units. Text and image units with display duration exceeding the preset display duration threshold and audience intent relevance greater than the preset high relevance threshold are loaded first. For graphic units whose loading speed is lower than the preset loading speed threshold, adjust their size parameters to reduce the amount of data, or adopt a progressive loading method, first loading low-resolution image data to meet basic display requirements, and then loading high-resolution image data to balance loading delay. Adjust the font size, character spacing, and line spacing of the text units to ensure that the font clarity is not lower than the preset clarity threshold under different lighting conditions; To address the conflict caused by the break in the collaborative relationship, the temporal matching relationship between the spatiotemporal features of the text and the spatiotemporal features of the image is recalculated. Based on the current display environment parameters, the temporal mapping rules are corrected, and the deviation between the reading time segments and the optimal display time is adjusted to not exceed the preset corresponding deviation threshold. Adjust the sentiment feature values ​​of image units and text units to reduce the sentiment deviation between them to no more than a preset sentiment deviation threshold, thereby repairing the sentiment synergy relationship. After the adjustment is completed, the conflict pre-simulation algorithm is re-executed to verify whether the conflict has been resolved. If not, the adjustment steps are repeated until the conflict is eliminated.

9. The intelligent optimization method for graphic and text layout combining digital multimedia as described in claim 1, characterized in that, The process of deploying the intermediate layout scheme to the real-time display environment, combining the content demand descriptions in the audience behavior intent data and the real-time updated display environment parameters, drives the dynamic self-adaptive layout engine to continuously evolve the layout scheme, generating a target graphic and text layout optimization scheme that conforms to the real-time environment and audience intent, includes: Deploy the intermediate layout scheme to the digital multimedia real-time display environment and establish a real-time data connection between the layout scheme and the display environment parameter acquisition module and the audience behavior intent acquisition module; By displaying the environmental parameter acquisition module, changes in device screen characteristics, fluctuations in ambient light intensity, and changes in network transmission rate are collected in real time, forming a real-time updated display environmental parameter stream. The audience behavior intent collection module collects new interaction behavior sequences of the audience in real time during the browsing of intermediate layout schemes, and collects the content demand expressions input by the audience to form a real-time updated audience behavior intent data stream. The real-time updated display environment parameter stream is input into the environment parameter mapping module of the dynamic self-adaptive typesetting engine, which updates the dynamic mapping relationship between environment parameters and typesetting adjustment parameters, and outputs new typesetting space elasticity coefficient, visual comfort parameter and content loading priority parameter. The real-time updated audience behavior intent data stream is input into the behavior intent decoding module to update the correspondence between the interaction behavior sequence and the layout strategy. The audience intent classification results are corrected in combination with the content requirement description, and a new layout strategy is output. Based on the new layout strategy and updated environmental parameters, the collaborative relationship application module adjusts the temporal mapping rules, sentiment matching parameters and collaborative weights in the graphic-text collaborative model, and outputs new collaborative application logic. The layout parameter evolution module updates the layout parameters of the intermediate layout scheme in real time based on the new environmental parameter mapping relationship, layout strategy and collaborative application logic, and adjusts the position coordinates, display size, display duration and loading order of graphic units. During the typesetting parameter update process, a conflict pre-simulation algorithm is executed simultaneously, and no new conflicts are generated in the updated typesetting scheme; The process of continuously repeating real-time data collection, module parameter updates, layout parameter adjustments, and conflict rehearsals continues until there are no new content requirements expressed in the audience behavior intent data stream, and the fluctuation range of each parameter in the real-time display environment parameter stream remains stable within its respective preset threshold range. The final stable layout scheme was determined as the target text and image layout optimization scheme.

10. A smart optimization system for graphic and text layout combining digital multimedia, characterized in that, The method includes a processor and a computer-readable storage medium storing machine-executable instructions that, when executed by the processor, implement the intelligent optimization method for graphic and text layout combining digital multimedia as described in any one of claims 1-9.

Citation Information

Patent Citations

  • Exhibition hall content self-adaptive central control display method and exhibition hall content self-adaptive central control display system

    CN120491810A

  • Page structure optimization method and system for PowerPoint

    CN120493879A

  • Automatic content typesetting method based on AI identification

    CN120745555A

  • Commodity multimedia recommendation method and system combining RPA and AI

    CN121052901A

  • Dynamic content control in an information processing system based on cultural characteristics

    US9672537B1