Information Display Method Based on AI Algorithm
Through the information display method based on AI algorithm, combined with voice to text, text enhancement and multimodal alignment models, the problem of image generation in the prior art does not meet user description and does not meet the quality standards, achieving high accuracy and high-quality image generation, and improving user experience.
Patent Information
- Application Number
- CN202510142501.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-10
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2045-02-10
AI Technical Summary
When generating pictures, the prior art cannot guarantee that the picture content meets user description and meets the quality standards, and the traditional image generation scheme relies on matching to sort, resulting in pictures that meet user needs may be ranked last, affecting the user experience.
Through the information display method based on AI algorithm, including converting voice information into text, text enhancement, inputting enhanced text into the AI system to generate pictures, using a multimodal alignment model to calculate the matching degree between pictures and text, and prioritizing the picture quality and popularity to ensure that the generated pictures meet user description and meet the quality standards.
The generated image content is highly consistent with user description, which improves the accuracy and quality of image generation, enhances user experience and satisfaction, and solves the problem of limitations in traditional methods of image sorting.
Smart Images

Figure CN119597949B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of AI technology, and specifically to an information display method based on AI algorithms. Background Art
[0002] In an existing document with the application number 202310484946.9 and the title of "An AR Space Annotation and Display Method Integrating an AI General Assistant", it is pointed out that: First, the intelligent terminal identifies each target object in the scene through the AR engine; based on the target detection algorithm, each target object is identified respectively to obtain the feature information of the target object; the feature information is input into the AI general assistant to obtain the output result; the output result is matched with the preset user requirements to identify the objects related to the user requirements in the scene, and the corresponding feedback information is extracted for feature words and sentences and annotated and displayed in the AR space, so as to achieve the effect of personalized recommendation of relevant content for users; the present invention combines the AR space annotation technology with the AI general assistant, which can not only meet the intuitive, rich and vivid user interaction experience, but also provide more personalized recommended content for users through the AI general assistant, and has been improved both in form and content; however, this invention does not further ensure that the generated content meets the user requirements and the quality is up to standard, and there are certain limitations.
[0003] Combined with the above document and the existing technology:
[0004] Traditionally, in the process of generating pictures from text, after analyzing the user's voice and converting the voice into text, the text content is input into the picture generation function module of the AI large model, and several pictures meeting the user's requirements can be generated simultaneously, but it is not guaranteed to satisfy the user. In the traditional picture generation scheme, only relying on the matching degree for sorting has limitations. The pictures with higher matching degrees are sorted more forward, but in actual use, some situations will be ignored. If only a few pictures can be displayed on the display screen, and the number of pictures generated according to the matching degree is as high as dozens, and the matching degrees of each picture are similar, it is possible that the pictures meeting the user's requirements are displayed at the end, affecting the user experience. At the same time, there are also limitations in the picture display sorting, which further affects the efficiency and satisfaction of picture selection. Summary of the Invention
[0005] (I) Technical Problems to be Solved
[0006] Aiming at the deficiencies of the existing technology, the present invention provides an information display method based on AI algorithms. Through the specific scheme design in the technical solution, the generated picture content meets the user's description and the quality is up to standard, solving the problems raised in the background art.
[0007] (II) Technical Solutions
[0008] To achieve the above objectives, the present invention is realized through the following technical solutions:
[0009] An information display method based on an AI algorithm, comprising the following steps:
[0010] The app terminal obtains the voice information of the picture to be generated and converts the voice information into text information;
[0011] Parse the text information, obtain the grammatical structure and keywords, and perform text enhancement. Detect whether there is size data in the enhanced text. If not detected, trigger the supply adjustment mechanism;
[0012] Input the enhanced text into the AI system to generate several pictures;
[0013] Use a multi-modal alignment model to analyze several pictures, calculate the matching degree between each picture and the original text information, and execute a priority determination mechanism based on the matching degree; Extract the corresponding picture with the maximum matching degree, and obtain the number of pictures whose difference from the maximum matching degree is within S points and S points;
[0014] If the number of pictures does not exceed the maximum value of the same-screen display amount on the app terminal, trigger the first determination strategy;
[0015] If the number of pictures exceeds the maximum value of the same-screen display amount on the app terminal, trigger the second determination strategy;
[0016] According to the determination result of the corresponding determination strategy, sort several pictures according to priority;
[0017] Display several sorted pictures on the app terminal for selection;
[0018] If there is a selected picture, push the selected picture from the app terminal to the device terminal. If there is no selected picture, obtain the feedback voice data, convert the feedback voice data into text information again, and execute the optimization and adjustment strategy.
[0019] Furthermore, the process of converting voice information into text information is as follows:
[0020] Audio preprocessing: The preprocessing includes at least noise reduction, echo removal, and audio format conversion;
[0021] Feature extraction: The preprocessed voice information is converted into feature vectors through a feature extraction algorithm;
[0022] Speech recognition model processing: The extracted feature vectors are sent into a trained HMM-GMM model, and the HMM-GMM model performs speech-to-text conversion according to the input feature vector sequence;
[0023] Text output and post-processing: Post-process the converted text information, which includes error correction and format adjustment.
[0024] Furthermore, text enhancement: perform detail completion operations based on syntactic structures and keywords, and synchronously optimize the text information, including at least: correcting grammar errors and adjusting the order of vocabulary;
[0025] The trigger supply adjustment mechanism is as follows:
[0026] Pre-tune sample pictures of various sizes on the app side for selection; among them, the sample pictures of various sizes are displayed in ascending order of size specifications.
[0027] Furthermore, before text enhancement, it also includes recommended recognition processing:
[0028] When the user uses it for the first time or the historical data volume under the corresponding user has not reached the set threshold, then perform an emotion recognition action to determine the emotional tendency, including positive, negative, and neutral; when the historical data volume under the corresponding user reaches the set threshold, then perform a preference recommendation action to determine the preference options;
[0029] Among them, the process of determining the emotional tendency is as follows:
[0030] Run the emotion orientation library, which records words representing positive, neutral, and negative emotional tendencies as keywords for emotion recognition, used to judge the emotional tendency, and accordingly select a matching picture style; positive emotional tendency corresponds to bright tones; negative emotional tendency corresponds to dull tones; neutral emotional tendency corresponds to tones other than bright and dull;
[0031] The process of determining the preference options is as follows:
[0032] Obtain the historical selection and behavior records of the corresponding user, and use machine learning algorithms for analysis to construct a preference model for the corresponding user.
[0033] Furthermore, after inputting the enhanced text into the AI system, it also includes combining the results of the recommended recognition processing to generate several pictures.
[0034] Furthermore, the value and calculation process of S score are as follows:
[0035] Calculate the matching degree between all pictures and the original text information to obtain a matching degree set. Based on the pre-constructed threshold setting engine, set S based on the standard deviation of the matching degree set, using the standard deviation method: calculate the standard deviation σ of the matching degree set, and set S = σ * Bs, where Bs represents the multiple value.
[0036] Furthermore, the calculation process of the maximum value of the same-screen display quantity on the app side is as follows:
[0037] Determine the display area: Obtain the width W and height H of the app screen size; subtract the layout space, which at least includes the space occupied by the screen edge, navigation bar, and toolbar, to get the actual display area width W_usable and height H_usable;
[0038] Determine the grid cell size: According to the design requirements, determine the width w_unit and height h_unit of each grid cell;
[0039] Calculate the number of displays on the same screen: In the width direction, the number of grid cells placed is: W_usable / w_unit; in the height direction, the number of grid cells placed is: H_usable / h_unit;
[0040] Then, the calculation formula for the maximum value N_max of the number of displays on the same screen is as follows:
[0041] N_max = floor(W_usable / w_unit) * floor(H_usable / h_unit)
[0042] Among them, the floor() function represents rounding down.
[0043] Furthermore, the first triggered decision strategy is as follows:
[0044] The priority is determined according to the matching degree of each picture:
[0045] ;
[0046] Among them, P represents the priority, and Em represents the matching degree;
[0047] The second triggered decision strategy is as follows:
[0048] Calculate and obtain the picture quality evaluation value and picture popularity evaluation value of each picture. Combining the matching degree, build a comprehensive calculation model based on the product and exponential function to generate the comprehensive value of each picture. The priority is determined according to the comprehensive value of each picture:
[0049] ;
[0050] Among them, Zs represents the comprehensive value;
[0051] The process of calculating the picture quality evaluation value and picture popularity evaluation value of each picture is:
[0052] Obtain the clarity and color restoration degree of each picture, and perform weighted calculation to obtain the picture quality evaluation value;
[0053] Obtain the popularity score and time decay score of each picture, and calculate the weighted value to obtain the picture heat value estimation;
[0054] Build a comprehensive calculation model based on the product and exponential functions:
[0055] Normalize the picture quality estimation, picture heat value estimation, and matching degree of each picture. The comprehensive value of each picture is calculated by the following formula:
[0056] ;
[0057] In the formula, λ represents the decay coefficient, and its value range is [0, 1]. Qp represents the picture quality estimation, and Phv represents the picture heat value estimation.
[0058] Furthermore, the corresponding index of clarity is: average gradient; the corresponding index of color restoration degree is: structural similarity index;
[0059] The popularity score is calculated by the number of likes, shares, and comments of the corresponding picture under the current display count. The formula is:
[0060] ;
[0061] In the formula, FS represents the popularity score, h1, h2, and h3 are weight coefficients, and their value ranges are all [0, 1]. dz represents the number of likes, fs represents the number of shares, pl represents the number of comments, and zs represents the display count;
[0062] The time decay score represents the heat decay of the corresponding picture over time after it is released, and is calculated using a time decay function. The formula is:
[0063] ;
[0064] In the formula, Fj represents the time decay score, e represents the natural constant, k represents the cooling coefficient, and its value range is [0, 1]. (T - T0) represents the time interval, where T is the current time and T0 is the picture release time.
[0065] Furthermore, the content of the optimization adjustment strategy is:
[0066] Use keyword extraction technology to obtain the keywords in the feedback text information converted from the feedback voice information, make corresponding parameter optimizations based on the content of the keywords, and the interval of parameter optimization is a preset fixed value. And at least three categories of pictures after parameter optimization are displayed on the app side for secondary selection; in the state where the user selects any picture, record the parameters under the corresponding picture and use these parameters as the standard parameters.
[0067] An information display system based on an AI algorithm. The system includes:
[0068] Voice acquisition and conversion module: The app terminal obtains the voice information of the picture to be generated and converts the voice information into text information;
[0069] Text enhancement module: Parse the text information, obtain the grammatical structure and keywords, and perform text enhancement. Detect whether there is size data in the enhanced text. If not detected, trigger the supply adjustment mechanism;
[0070] Text-to-image module: Input the enhanced text into the AI system to generate several pictures;
[0071] Alignment and screening module: Use the multimodal alignment model to analyze several pictures, calculate the matching degree between each picture and the original text information, and execute the priority judgment mechanism according to the matching degree; Extract the picture corresponding to the maximum matching degree, and obtain the number of pictures whose difference from the maximum matching degree is within S points and S points;
[0072] If the number of pictures does not exceed the maximum value of the app terminal's on-screen display quantity, trigger the first judgment strategy;
[0073] If the number of pictures exceeds the maximum value of the app terminal's on-screen display quantity, trigger the second judgment strategy;
[0074] According to the judgment result of the corresponding judgment strategy, sort several pictures according to the priority;
[0075] Pre-display module: Display several sorted pictures on the app terminal for selection;
[0076] Display adjustment module: If there is a selected picture, push the selected picture from the app terminal to the device terminal. If there is no selected picture, obtain the feedback voice data, convert the feedback voice data into text information again, and execute the optimization adjustment strategy.
[0077] (3) Beneficial effects
[0078] The present invention provides an information display method based on an AI algorithm, having the following beneficial effects:
[0079] (1) By introducing the emotion recognition and preference recommendation mechanisms, this solution can intelligently recommend picture styles that match the user's taste according to the user's emotional tendency and historical preferences. This not only improves the personalization and customization degree of picture generation, but also makes the generated pictures more in line with the user's actual needs and aesthetic preferences, thereby enhancing user satisfaction and ensuring the subsequent picture display effect to a certain extent;
[0080] (2) This solution uses a multi-modal alignment model to calculate the matching degree between each picture and the original text information, ensuring that the generated pictures are highly consistent with the user's description. Then, by setting the maximum value of the S score and the number of pictures shown on the app screen simultaneously, combined with the calculation of picture quality and picture popularity, it solves the limitation of only relying on the matching degree for sorting in traditional picture generation solutions. When the number of pictures shown on the screen does not meet the requirements, it can automatically re-sort the recommended pictures according to the actual situation to ensure that the generated picture content meets the user's description and the quality standard.
[0081] (3) This solution realizes the function of dynamically optimizing picture display parameters according to the user's real-time feedback. By extracting keywords in the user's feedback voice, it accurately adjusts picture parameters such as color and brightness, and shows multiple pictures with optimized parameters on the app for the user to select again. This not only significantly improves the user experience, making it easier for the user to select a satisfactory picture through personalized adjustment, but also effectively solves the problem that the user has to flip through multiple pages to select pictures or cannot select a picture because the picture display does not meet expectations by recording the parameters preferred by the user as the system standard parameters, greatly improving the efficiency and satisfaction of picture selection. Brief Description of the Drawings
[0082] Figure 1 It is the overall step flow chart of the information display method in Embodiment 1 of the present invention;
[0083] Figure 2 It is the overall flow schematic diagram of executing S4 in Embodiment 2 of the present invention;
[0084] Figure 3 It is the modular schematic diagram of the information display system based on the AI algorithm in the present invention. Detailed Embodiments
[0085] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0086] Research and Development Concept:
[0087] With the continuous development of AI technology, there is a need to build an information display application integrating AI technology, with the core being to achieve the automatic generation of high-quality pictures from user voices; the solution will focus on optimizing the accuracy of speech recognition, enhancing the richness of text descriptions, improving the accuracy of AI picture generation, and the effectiveness of matching degree calculation;
[0088] Existing Problem Directions to be Solved:
[0089] 1. Text enhancement requires more refined semantic understanding and completion;
[0090] 2. The expressiveness of AI image generation in complex scenarios and styles;
[0091] 3. The sensitivity of the matching degree calculation model to details and styles, etc.;
[0092] Solution description:
[0093] 1. Introduce a more advanced speech recognition model, combined with filtering technology, to improve recognition accuracy; use deep learning technology to enhance the depth of text understanding, or combine with a knowledge graph to complete description details;
[0094] 2. Upgrade the AI image generation model, introduce more style and scenario data, and improve the diversity and realism of the images;
[0095] 3. Optimize the multi-modal alignment model, such as CLIP, enhance the recognition ability of image details and styles, and improve the accuracy of matching degree calculation; at the same time, consider introducing user behavior data to further personalize the matching results;
[0096] The problems and solutions given above are all in one direction or basic concepts. The specific solution design given later may be implemented based on one direction or several concepts. This solution aims to create an efficient, accurate and user-friendly image generation and display system to meet the diverse image needs of users.
[0097] Example 1
[0098] Please refer to Figure 1 , this example provides an information display method based on AI algorithms. The steps of this method are as follows:
[0099] S1. The app side obtains the voice information of the image to be generated, and uses speech recognition technology to convert the voice information into text information;
[0100] Among them, the app side user speaks out the information such as the content and size of the image to be generated;
[0101] S2. Use natural language processing technology to parse the text information, obtain the grammatical structure and keywords, and perform text enhancement;
[0102] Among them, after the steps given in S1: the app side submits the voice information to the backend, and after the backend converts the voice information into text information, it is handed over to the "text enhancement module" for further processing; after the "text enhancement module" parses the input text information, it obtains the grammatical structure and keywords, completes the description details, and ensures that the enhanced text matches the requirements of the image generation model (i.e., the AI system);
[0103] S3. The "Text Enhancement Module" accesses the AI interface and submits a task of generating images to the AI system;
[0104] S4. After receiving the task, the AI system generates a series of images in a specific style;
[0105] S5. The backend uses CLIP or a similar multimodal alignment model to calculate the matching degree between each image and the original text description (i.e., the text information in S2), and preferentially displays the images with higher matching degrees to the user;
[0106] S6. The user selects from the app end according to the requirements. If selected, the image is pushed out from the app end for display on the target device; if not selected, the operation of S1 is re-executed.
[0107] The above technical solution realizes an efficient and personalized generation process from speech to image, bringing significant beneficial effects; First, through speech recognition technology, users can express information such as the content and size of the images they want to generate in the most natural way - speech input, greatly improving the user experience and convenience; Second, the application of natural language processing technology enables the system to accurately parse the user's intention, refine and enrich the description details through the text enhancement module, ensuring that the generated images are more in line with the user's needs; In addition, the seamless docking with the AI system enables text information to be quickly converted into images in a specific style, not only enriching the diversity and creativity of image generation, but also improving the generation efficiency; By evaluating the matching degree between the image and the original text description through CLIP or a similar model, it is ensured that the image displayed to the user is the one that best matches its description, further improving user satisfaction; Finally, the user can select images on the app end according to actual needs and re-generate if not selected. This flexible interaction mechanism gives users a higher degree of participation and decision-making power, making the entire image generation process more personalized and intelligent; In summary, the technical solution has achieved remarkable results in improving the user experience and enhancing the accuracy and personalization of generated images.
[0108] Embodiment 2:
[0109] Based on Embodiment 1, please refer to Figure 2 , this embodiment also provides an optimized information display method based on AI algorithms (this embodiment elaborates in detail on some content and optimized content in Embodiment 1), and the specific steps of this method are as follows:
[0110] S1. The app end obtains the voice information of the image to be generated and converts the voice information into text information using speech recognition technology;
[0111] Among them, the user speaks out the content and size of the picture to be generated through the app side. The content and size together constitute the original voice information. Before the user speaks out the original voice information, the app side will give information prompts, such as: reminding the user that the content to be spoken should include the content and size. Of course, it is not limited to the content and size. At this time, user A may say:
[0112] Example 1: "I want to generate a landscape painting with a beautiful environment showing a pleasant atmosphere, and the size is larger";
[0113] Or Example 2: "I want to generate a landscape painting, and the size is moderate";
[0114] Or Example 3: "I want to generate a landscape painting with a simple style, and the size is 1920*1080";
[0115] It should be noted that for the size, not all users will say specifically. Most will say words like larger, medium, and smaller, which are vague and colloquial. Therefore, a supply adjustment mechanism needs to be executed in S2 later, specifically referring to the content described in S2;
[0116] The conversion process of converting voice information into text information is as follows:
[0117] Audio preprocessing (i.e., voice information): The preprocessing includes at least: noise reduction, echo removal, and audio format conversion to improve the accuracy of speech recognition; audio segmentation is also performed as needed to cut continuous speech into individual words or phrases for subsequent processing;
[0118] Feature extraction: The preprocessed voice information is converted into a series of feature vectors through feature extraction algorithms (such as MFCC, FBank, etc.). These feature vectors are the representations of the voice signal in the frequency domain and time domain, reflecting the acoustic characteristics of the voice;
[0119] Speech recognition model processing: The extracted feature vectors are fed into a trained speech recognition model (such as a deep learning model, HMM-GMM model, etc.); the model performs speech-to-text conversion based on the input sequence of feature vectors, combining a language model and an acoustic model; the language model is used to predict possible word sequences, and the acoustic model is used to evaluate the matching degree of these sequences with the input speech;
[0120] Text output and post-processing: The text output by the speech recognition model may contain some errors or uncertain words and needs to be post-processed; the post-processing includes error correction, format adjustment, special word replacement, etc. to ensure that the output text is accurate, clear, and in line with the user's intention.
[0121] S2. Use natural language processing technology to parse text information, obtain syntactic structures and keywords, and perform text enhancement; preliminarily detect whether there is dimensional data in the enhanced text, and trigger the supply adjustment mechanism under the condition that no dimensional data is detected;
[0122] Before performing text enhancement, it also includes recommendation recognition processing;
[0123] When the user uses it for the first time or the historical data volume (i.e., voice information volume) under the corresponding user does not reach the set threshold, then perform an emotion recognition action to determine the emotional tendency, including positive, negative, and neutral;
[0124] When the historical data volume under the corresponding user reaches the set threshold, then perform a preference recommendation action to determine the preference options;
[0125] It should be noted that the voice information volume represents the memory occupied by the original voice information spoken by the corresponding user, and the unit is represented by kb. For example: the set threshold is 1000 kb. After the current app end obtains the voice information of the corresponding user, it is detected that the current voice information volume of the corresponding user reaches 1200 kb, which means that the corresponding user has entered the option of performing the recommendation action; but if the corresponding user is using it for the first time, the historical data volume of this user is not considered;
[0126] Of course, if the converted text information from the original voice information already contains a style description, such as the content in Example 3, which has clearly stated that pictures in a simple style are required, but generally users will not or leave the required style description, so the operation of recommendation recognition processing needs to be performed;
[0127] The process of performing the emotion recognition action and determining the emotional tendency is as follows:
[0128] Run the emotion orientation library, which records: The words representing positive emotional tendencies at least include: "happiness", "joy", "excitement", "excitation", "satisfaction", "relief", and "optimism", and these words are usually associated with positive emotions and pleasant feelings; The words representing negative emotional tendencies at least include: "sadness", "frustration", "disappointment", "anxiety", "pain", "anger", "melancholy", and these words are often associated with negative emotions and unpleasant feelings; After excluding the words representing positive and negative emotional tendencies, the remaining words are those representing neutral emotional tendencies; These words can be used as keywords for emotion recognition to help the system judge the user's emotional tendency and select a matching picture style accordingly;
[0129] Match the corresponding picture style according to the emotional tendency;
[0130] The positive emotional tendency corresponds to a bright (lively) color tone;
[0131] Negative emotional tendencies correspond to dull (deep) hues;
[0132] Neutral emotional tendencies correspond to hues other than bright and dull, i.e., conventional or normal hues;
[0133] The process of performing preference recommendation actions and determining preference options is as follows:
[0134] Obtain the historical selections and behavior records of the corresponding user, analyze them using machine learning algorithms, and construct a preference model for the corresponding user. This model is used to reflect the user's strong preferences for different styles; form the personalized preferences of the corresponding user;
[0135] Example: If the user selects pictures in a minimalist style multiple times, the system will preferentially recommend pictures in a minimalist style in subsequent generations; during the subsequent picture generation process, the system will use this preference model to preferentially recommend pictures that suit the user's taste; specifically, the system will adjust the parameters of the generated pictures so that the newly generated pictures are closer to the minimalist style preferred by the user in terms of color, composition, and theme; in this way, when the user browses pictures again, they can more easily find the types they like, thereby enhancing the user experience and satisfaction;
[0136] Generally speaking, after performing recommendation recognition processing, the result obtained is: the determination of the corresponding user's picture style;
[0137] When parsing text information using natural language processing technology, first analyze the grammatical structure of the text information, identify language units such as sentences, phrases, and words, and understand the relationships and hierarchical structures between them; this helps to more accurately grasp the user's intentions and the key points of expression; for the acquisition of keywords, use keyword extraction technology to identify words with key meanings from the text information, including at least nouns, verbs, and adjectives; these keywords are usually closely related to specific information such as the content, style, and size of the pictures that the user wants to generate;
[0138] Text enhancement: Based on the grammatical structure and keywords, perform detail completion operations to synchronously optimize the text information, including at least: correcting grammar errors, adjusting the order of words, and adding necessary modifiers; to ensure that it is more fluent, accurate, and meets the requirements of the image generation model; among them, the detail completion operation means performing detail completion on the text information;
[0139] For example, if the user only mentions "want a landscape painting", based on the grammatical structure and keywords, and according to common landscape painting elements, the detail completion operation will automatically complete it to "want a landscape painting containing blue sky, white clouds, green mountains, green waters, and small bridges over flowing waters";
[0140] An example of integrity is as follows:
[0141] Suppose the user inputs via voice: "I want a picture of a living room decorated in a retro style, with a sofa, a coffee table, and a bookshelf."
[0142] After text enhancement and parsing, the "Text Enhancement Module" may enhance this text to:
[0143] "I want a picture of a living room decorated in a retro style. There is a brown leather sofa placed in the living room. In front of the sofa is a wooden coffee table, on which there are several decorative books and a vase. In a corner of the living room, there is a retro bookshelf filled with books. The bookshelf is made of dark solid wood, which is coordinated with the overall style of the living room. The overall color tone of the living room is mainly warm colors, creating a warm and comfortable atmosphere."
[0144] The enhanced text not only contains the key information of the user's original input but also complements more detailed descriptions, such as the color and material of the sofa, the items on the coffee table, the style and location of the bookshelf, etc., so as to more comprehensively meet the requirements of the image generation model and generate pictures that better meet the user's expectations;
[0145] The content triggering the supply and adjustment mechanism is as follows:
[0146] On the app side, sample pictures of various sizes are pre - tuned for the user to select; among them, the sample pictures of various sizes are displayed in ascending order of size specifications until the user selects one.
[0147] Specifically, by introducing an emotion recognition and preference recommendation mechanism, this solution can intelligently recommend picture styles that match the user's taste according to the user's emotional tendency and historical preferences. This not only improves the personalization and customization of picture generation but also makes the generated pictures more in line with the user's actual needs and aesthetic preferences, thus enhancing user satisfaction and ensuring the display effect of subsequent pictures to a certain extent;
[0148] This solution also solves the problem of ambiguous user input. For ambiguous size descriptions in the user's input, such as colloquial expressions like "a bit larger", "medium", etc., this solution can convert them into specific size parameters by triggering the supply and adjustment mechanism, thus ensuring the accuracy of picture generation and meeting the user's expectations;
[0149] In summary, the above - mentioned solution not only realizes the intelligent generation process from voice to pictures but also improves the personalization, accuracy, and quality of picture generation by introducing technologies such as emotion recognition, preference recommendation, and text enhancement, bringing a more convenient, efficient, and personalized picture generation experience to users.
[0150] S3. Input the enhanced text into the AI system and generate several pictures in combination with the results of recommendation recognition processing;
[0151] Among them, the AI system is a text-to-image tool;
[0152] The AI system specifically refers to a comprehensive system integrating advanced text processing and image generation technologies. Such systems usually utilize deep learning algorithms, especially generative adversarial networks (GANs) and diffusion models (such as Stable Diffusion), etc., to achieve the function of converting enhanced text input into high-quality images;
[0153] Taking AI-Chat as an example, it integrates multiple AI models, including GPT4.0 and the Stable Diffusion painting model. Users can generate various anime and cartoon-style avatars or other artworks by inputting text descriptions; in addition, there are also various AI text-to-image software such as Jasperart, Photosonic, Shutterstock, NightCafe, etc. in the market. They each have their own characteristics and can provide users with diverse image generation experiences; these systems not only improve the efficiency of image creation but also provide powerful auxiliary tools for creative workers such as designers and illustrators;
[0154] The number of pictures produced by this AI system is not just 1. Instead, according to the situation that meets the text description, all the pictures that meet the requirements are output. For some common text descriptions, the number of pictures that meet the requirements can be as high as dozens or hundreds; while for some uncommon or special text descriptions, the number of pictures that meet the requirements is at least 1 and at most 10. The above numbers are just examples, and the specific data can only be obtained according to the actual situation.
[0155] S4. Use a multi-modal alignment model to analyze a number of pictures, calculate the matching degree of each picture with the original text information, and execute a priority determination mechanism based on the matching degree;
[0156] Extract the corresponding picture with the maximum matching degree, obtain the number of pictures whose difference from the maximum matching degree is within S points and S points, and if the number of pictures does not exceed the maximum value of the app-side on-screen display quantity, trigger the first determination strategy;
[0157] If the number of pictures exceeds the maximum value of the app-side on-screen display quantity, trigger the second determination strategy;
[0158] Sort a number of pictures according to the priority based on the determination result of the corresponding determination strategy;
[0159] In this scenario, the system uses CLIP (Contrastive Language–Image Pre-training) or a similar multi-modal alignment model to perform in-depth analysis on each image generated by the AI system; the CLIP model (i.e., the multi-modal alignment model used in this embodiment) is a powerful tool that can learn the association between vision and language, and it can encode images and texts into the same high-dimensional space, enabling the similarity between the two to be measured by the distance between them in this space;
[0160] The calculation of the matching degree is described as follows:
[0161] In the multi-modal alignment model, calculating the matching degree between each image and the original text description is a key step, which is achieved by calculating the distance between the image and the original text information in the encoding space; the smaller the distance, the higher the matching degree between the image and the text description; conversely, the lower the matching degree; therefore, the matching degree represents the reciprocal of the distance value between the corresponding image and the original text information in the encoding space, for example: 1 / distance value;
[0162] In addition, in actual situations, in order to ensure the accuracy of the matching degree, the model needs to comprehensively consider factors such as the content, style, and size of the image; for example, in terms of content, the model needs to determine whether the objects and scenes in the image match the text description; in terms of style, the model needs to evaluate whether the color, lines, composition, etc. of the image are consistent with the style of the text description; in terms of size, the model needs to ensure that the generated image meets the size requirements specified by the user;
[0163] Regarding the value and calculation process of S score are as follows:
[0164] Calculate the matching degrees between all the images generated in S3 and the original text information to obtain a matching degree set. Based on the pre-constructed threshold setting engine, S is set based on the standard deviation of the matching degree set, using the standard deviation method: calculate the standard deviation σ of the matching degree set, and set S = σ * Bs, where Bs represents the multiple value, and the value of Bs is 1 or 2, and the value in this embodiment is 2; this method takes into account the distribution width of the matching degree and is more dynamic and adaptable;
[0165] Among them, the formula for calculating the standard deviation σ of the matching degree set is as follows:
[0166] ;
[0167] In the formula, N represents the number of data points, that is, the number of each matching degree in the matching degree set; xi represents each matching degree value in the matching degree set; μ represents the average value of all matching degree values in the matching degree set;
[0168] Therefore, the formula for calculating the S score is:
[0169] S = 2σ;
[0170] Suppose we have the following set of matching degrees (represented as 1 / distance value):
[0171] [0.1, 0.2, 0.4, 0.5, 0.7, 0.8, 0.9, 1.0];
[0172] Among them, the maximum matching degree is 1.0;
[0173] Using the standard deviation method: Assume that the standard deviation σ of this set is approximately 0.25 (this is an assumed value and should be calculated actually); if we set S to be twice the standard deviation, then S = 2 * 0.25 = 0.5; then, the matching degrees within 0.5 and within the gap from the maximum value are: [0.5, 0.7, 0.8, 0.9, 1.0]; in this example, the number of pictures with the matching degree within S points and within S points from the maximum matching degree is 4;
[0174] The calculation process of the maximum value of the same - screen display quantity on the app side is as follows:
[0175] Determine the available display area:
[0176] Obtain the width W and height H of the app - side screen size; subtract the layout space, including at least the space occupied by interface elements such as the screen edge, navigation bar, and toolbar, to get the actual available display area width W_usable and height H_usable;
[0177] Determine the grid cell size (pictures are displayed on the app side in the form of a grid):
[0178] According to the design requirements, determine the width w_unit and height h_unit of each grid cell; these sizes are determined in advance according to the app's built - in program settings;
[0179] Calculate the same - screen display quantity:
[0180] In the width direction, the number of grid cells placed is W_usable / w_unit (take the integer part);
[0181] In the height direction, the number of grid cells that can be placed is approximately H_usable / h_unit (take the integer part);
[0182] Therefore, the maximum value of the same - screen display quantity is (W_usable / w_unit) * (H_usable / h_unit) (take the integer part); thus, the calculation formula based on which the maximum value N_max of the same - screen display quantity is as follows:
[0183] N_max = floor(W_usable / w_unit) * floor(H_usable / h_unit)
[0184] Among them, the floor() function represents rounding down;
[0185] It should be noted that the purpose of comparing the number of pictures with the maximum number of pictures displayed on the app side at the same time is to: automatically select which judgment strategy to execute. If the number of pictures does not exceed the maximum number of pictures displayed on the app side at the same time, it means that all pictures can be displayed on the app side at the same time (they can be regarded as "no difference or negligible difference" when displayed), and there is no need to scroll down or switch to the next screen (if there is such an operation, it will not only waste time but also affect the user's selection and judgment). At this time, only the matching degree needs to be considered to determine the picture priority; but if the number of pictures exceeds the maximum number of pictures displayed on the app side at the same time, it means that all pictures cannot be completely displayed on one screen. To ensure that users can select the required pictures only on the first screen, at this time, not only the matching degree needs to be considered, but also other factors (i.e., picture quality and popularity) need to be considered to determine the picture priority;
[0186] The content of the first judgment strategy triggered is as follows:
[0187] The priority is determined according to the matching degree of each picture:
[0188] ;
[0189] Among them, P represents the priority, and Em represents the matching degree;
[0190] The content of the second judgment strategy triggered is as follows:
[0191] Calculate and obtain the picture quality estimation value and picture popularity estimation value of each picture, combine the matching degree, build a comprehensive calculation model based on the product and exponential function, generate the comprehensive value of each picture, and the priority is determined according to the comprehensive value of each picture:
[0192] ;
[0193] Among them, Zs represents the comprehensive value;
[0194] Regardless of the first or second judgment strategy, several pictures are sorted according to the priority. For example: they are sorted in turn on the app screen from left to right and from top to bottom;
[0195] Among them, the process of calculating the picture quality estimation value and picture popularity estimation value of each picture is explained as follows:
[0196] Obtain the clarity and color restoration degree of each picture, and the picture quality estimation value can be obtained after weighted calculation;
[0197] Obtain the popularity score and time decay score of each picture, and the picture heat estimation value can be obtained after weighted calculation;
[0198] The explanation is as follows:
[0199] Clarity is an important indicator in picture quality assessment, which reflects the visible degree of picture details; Clarity assessment methods include edge detection-based methods and frequency domain analysis-based methods; Among them, the corresponding index of clarity is: Mean Gradient, which reflects the rate of detail contrast and texture change in the image;
[0200] Color restoration degree refers to the degree of approximation between the colors in the image and the colors in the real scene; Color restoration degree assessment methods include color space-based methods and histogram-based methods; Among them, the corresponding index of color restoration degree is: Structural Similarity Index (SSIM), which reflects the color restoration degree of the image;
[0201] The formula for calculating the picture quality estimation value is as follows:
[0202] ;
[0203] In the formula, Qp represents the picture quality estimation value, a1 and a2 are both weight coefficients, and their value ranges are both [0, 1], MG represents clarity, and SM represents color restoration degree;
[0204] The popularity score is calculated by the number of likes, shares, and comments on the corresponding picture under the current display count. The specific formula is:
[0205] ;
[0206] In the formula, FS represents the popularity score, h1, h2, and h3 are the single score items for their respective behaviors, which are set according to the platform strategy to reflect the contribution degree of different behaviors to the popularity score, that is, the weight coefficients, and their value ranges are both [0, 1], dz represents the number of likes, fs represents the number of shares, pl represents the number of comments, and zs represents the display count;
[0207] The time decay score represents the heat decay of the corresponding picture over time after it is released, and is calculated using a time decay function (such as Newton's law of cooling). The specific formula is:
[0208] ;
[0209] In the formula, Fj represents the time decay fraction, e represents the natural constant, k represents the cooling coefficient, which reflects the rate of heat value decay over time, and its value range is [0, 1], (T - T0) represents the time interval, where T is the current time and T0 is the picture release time;
[0210] The formula for calculating the picture heat value estimate is as follows:
[0211] ;
[0212] In the formula, Phv represents the picture heat value estimate, and b1 and b2 are both weight coefficients, with a value range of [0, 1];
[0213] The weight coefficients are determined by the coefficient of variation method. The coefficient of variation method is a method of assigning weights to each index according to the degree of variation between the current value and the target value of each evaluation index; if the values of a certain index vary greatly and can clearly distinguish each evaluated object, it means that the resolution information of this index is rich, so a larger weight should be given to this index; conversely, if the values of each evaluated object vary little in a certain index, then the ability of this index to distinguish each evaluation object is weak, so a smaller weight should be given to this index; this method directly utilizes the information contained in each index and calculates the weights of the indexes, so it has objectivity;
[0214] Build a comprehensive calculation model based on the product and exponential function:
[0215] Normalize the picture quality estimate, picture heat value estimate, and matching degree of each picture so that their value ranges are restricted to between 0 and 1, and then the comprehensive value of each picture is calculated through the following formula:
[0216] ;
[0217] In the formula, λ represents the attenuation coefficient, which can be adjusted according to the actual situation, and its value range is [0, 1];
[0218] Logical explanation:
[0219] Product part: Em * Qp * Phv reflects the combined effect of the three indicators of matching degree, picture quality estimate, and picture heat value estimate. When any one of these indicators approaches 0, the entire product term will also approach 0, indicating that the comprehensive value will be significantly affected;
[0220] Denominator part: 1 - (1 - Em)(1 - Qp)(1 - Phv) is an adjustment factor used to enhance the interaction between the three indicators; when all three indicators are relatively high, the denominator approaches 1, and the influence of the product term is more significant; when at least one of the three indicators is relatively low, the denominator will increase, thereby reducing the influence of the product term on the comprehensive value;
[0221] Exponential part: (-)^(1 / 3) takes the square root of the product term to smooth out extreme differences among the three metrics, which can ensure that when one metric is extremely high or low, the comprehensive value will not deviate too much from the normal range;
[0222] Time decay part: It is a decay term based on the exponential function, used to consider the decay of the average value of the three metrics. When the average value of Em + QP + Phv approaches 3 (i.e., all three metrics are high), the decay term approaches 1 and has a small impact on the comprehensive value; when the average value is low, the decay term will decrease, thus reducing the comprehensive value;
[0223] Comprehensive consideration: The entire formula comprehensively considers three metrics: matching degree, image quality, and image popularity, and combines them through a non-linear relationship. This combination method can not only reflect the individual effects of each metric but also embody their interaction and overall effect.
[0224] First, this solution uses a multi-modal alignment model (such as CLIP) to calculate the matching degree between each image and the original text information, ensuring that the generated images are highly consistent with the user's description. This not only improves the accuracy of image generation but also enables users to more easily find images that meet their needs; by setting the maximum values of S score and the number of images shown on the app side simultaneously, this solution can flexibly adapt to different quantities and qualities of images, providing more accurate recommendations for users;
[0225] Second, by introducing the calculation of image quality and image popularity, this solution solves the limitation of traditional image generation solutions that only rely on matching degree for sorting. When the number of images shown on the same screen does not meet the requirements, it can automatically select images that, on the basis of considering the matching degree, also comprehensively consider image quality and image popularity. It can not only re-sort the recommended images according to the actual situation but also ensure that the generated image content meets the user's description and the quality standard;
[0226] This solution constructs a comprehensive calculation model based on product and exponential functions, organically combines the three metrics of matching degree, image quality, and image popularity, and generates a comprehensive value to evaluate the overall performance of each image. This comprehensive evaluation method not only considers the individual effects of each metric but also reflects their interaction and overall effect, making the sorting results more comprehensive, objective, and in line with user expectations;
[0227] Finally, by adopting the above solution, this method not only improves the accuracy and quality of image generation, but also enhances the user experience and satisfaction; users can more easily find images that meet their needs, and can view more high-quality and popular images on one screen; at the same time, this solution also has a certain degree of adaptability and scalability, and can adjust parameters and models according to actual situations to meet the needs and scenarios of different users.
[0228] In summary, the above solution realizes the comprehensive evaluation and sorting of AI-generated images by integrating a multi-modal alignment model, image quality evaluation, image popularity calculation, and a comprehensive calculation model, improves the accuracy and quality of image generation, and enhances the user experience and satisfaction.
[0229] S5. Display a number of sorted images on the app side for users to select;
[0230] Among them, a number of images sorted by priority are arranged in a grid on the display screen of the app side. Users select on the display screen. If selected, perform the operations of the subsequent steps. If not selected, continue to flip through the pages to select until selected. However, in the actual selection process, usually users can complete the selection on the home page. If flipping through the pages to select, it is very likely that no selected images will be available;
[0231] S6. If there are selected images, push the selected images from the app side to the device side. If there are no selected images, obtain feedback voice data, re-execute S1, and perform an optimization adjustment strategy according to the output result of S1;
[0232] Among them, the device side refers to a specified user device, such as a large screen, so as to achieve an information display effect;
[0233] The feedback voice data refers to the user's feedback voice. For example: the user feedbacks that the image color is not bright enough; then, when re-executing S1 later, the content of the optimization adjustment strategy performed according to the output result of S1 is as follows:
[0234] Use keyword extraction technology to obtain keywords in the feedback text information converted from the feedback voice information, make corresponding parameter optimizations based on the content of the keywords, and the interval of parameter optimization is a preset fixed value, and display at least three categories of images after parameter optimization on the app side for users to select again; in the state where the user selects any image, record the parameters corresponding to the image, and use this parameter as the standard parameter of the current system;
[0235] Among them, if the user feedbacks that the picture color is not bright enough, the keywords are "color" and "not bright enough". The system optimizes and adjusts the color parameters. The original color parameter is U, and the interval value is 1. Then, on the app side, pictures with three types of color parameters, namely U+1, U+2, and U+3 (the three types of pictures only differ in color parameters), are displayed. If the user selects the picture with the U+3 type of color parameter, the color parameter is recorded. The next time the corresponding user operates again, the color parameter is retrieved. Thus, to a certain extent, it can avoid the related problems of being unable to select the picture at one time and improve work efficiency.
[0236] Specifically, the above scheme realizes the function of dynamically optimizing the picture display parameters according to the user's real-time feedback. By extracting the keywords in the user's feedback voice, the picture parameters such as color and brightness are accurately adjusted, and multiple pictures with optimized parameters are displayed on the APP side for the user to select again. This scheme not only significantly improves the user experience and makes it easier for the user to select a satisfactory picture through personalized adjustment, but also effectively solves the problem that the user has to flip through the pages multiple times or cannot select the picture because the picture display does not meet the expectations by recording the parameters preferred by the user as the system standard parameters, greatly improving the efficiency and satisfaction of picture selection.
[0237] Embodiment 3:
[0238] Please refer to Figure 3 , based on Embodiment 2, this embodiment also provides an information display system based on the AI algorithm. The system includes a voice acquisition and conversion module, a text enhancement module, a text-to-image generation module, an alignment and screening module, a pre-display module, and a display adjustment module that run in sequence;
[0239] Voice acquisition and conversion module: The app side obtains the voice information of the picture to be generated and converts the voice information into text information;
[0240] Text enhancement module: Analyze the text information, obtain the grammatical structure and keywords, and perform text enhancement. Detect whether there is size data in the enhanced text. If not detected, trigger the supply adjustment mechanism;
[0241] Text-to-image generation module: Input the enhanced text into the AI system to generate several pictures;
[0242] Alignment and screening module: Use the multi-modal alignment model to analyze several pictures, calculate the matching degree between each picture and the original text information, and execute the priority determination mechanism based on the matching degree; Extract the picture corresponding to the maximum matching degree, and obtain the number of pictures whose difference from the maximum matching degree is within S points and S points;
[0243] If the number of pictures does not exceed the maximum value of the app side's same-screen display quantity, trigger the first determination strategy;
[0244] If the number of pictures exceeds the maximum value of the same-screen display quantity on the app side, the second determination strategy is triggered;
[0245] According to the determination result of the corresponding determination strategy, several pictures are sorted according to priority;
[0246] Pre-display module: Display several sorted pictures on the app side for selection;
[0247] Display adjustment module: If there are selected pictures, the selected pictures are pushed from the app side to the device side. If there are no selected pictures, feedback voice data is obtained, the feedback voice data is converted back into text information, and the optimization adjustment strategy is executed.
[0248] The above embodiments can be implemented in whole or in part by software, hardware, firmware or any other combination. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product. Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution.
[0249] The units described as separate components may or may not be physically separated. The components shown as units may or may not be physical units. They may be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0250] The above is only the specific implementation manner of this application, but the protection scope of this application is not limited thereto. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed in this application, and all of them should be covered by the protection scope of this application.
Claims
1. The information display method based on AI algorithm is characterized by: The steps of this method are as follows: The app obtains the voice information of the image to be generated and converts the voice information into text information; Parse text information, obtain grammatical structure and keywords, and perform text enhancement. Detect whether there is dimension data in the enhanced text. If not, trigger the adjustment mechanism. Input the enhanced text into the AI system to generate several images; Using the multimodal alignment model, several images are analyzed, the matching degree between each image and the original text information is calculated, and the priority determination mechanism is implemented according to the matching degree; the corresponding image with the maximum matching degree is extracted, and the number of images with a difference of S points or less from the maximum matching degree is obtained; the value and calculation process of S score are as follows: Calculate the matching degree of all images and the original text information to obtain the matching degree set. According to the pre-built threshold setting engine, set S based on the standard deviation of the matching degree set. Use the standard deviation method: calculate the standard deviation σ of the matching degree set and set S=σ*Bs, where Bs represents the multiple value. If the number of pictures does not exceed the maximum number of pictures displayed on the same screen on the app side, the first determination strategy is triggered, and the priority P is determined based on the matching degree Em of each picture: ; If the number of pictures exceeds the maximum number of pictures that can be displayed on the same screen on the app side, the second judgment strategy is triggered to calculate and obtain the picture quality estimation and picture heat estimation of each picture. Combined with the matching degree, a comprehensive calculation model based on the product and exponential function is built to generate the comprehensive value Zs of each picture. The priority P is determined based on the comprehensive value Zs of each picture: ; The process of calculating the image quality estimation and image heat estimation of each image is as follows: Obtain the clarity and color reproduction of each image, and perform weighted calculations to obtain an estimated image quality; Obtain the popularity score and time decay score of each image, and perform weighted calculation to obtain the image heat estimation; Build a comprehensive calculation model based on product and exponential functions: The image quality estimation, image heat estimation, and matching degree of each image are normalized, and the comprehensive value of each image is calculated using the following formula: In the formula, λ represents the attenuation coefficient, which ranges from [0 to 1], Qp represents the image quality estimation, and Phv represents the image heat estimation; According to the determination result of the corresponding determination strategy, a number of images are sorted according to priority; Display a number of sorted pictures on the app for selection; If the selected picture exists, the selected picture will be pushed from the app to the device. If the selected picture does not exist, the feedback voice data will be obtained, the feedback voice data will be converted into text information again, and the optimization adjustment strategy will be executed.
2. The information display method based on AI algorithm according to claim 1, characterized in that: The process of converting voice messages to text messages is as follows: Audio preprocessing: Preprocessing includes at least noise reduction, echo removal, and audio format conversion; Feature extraction: The preprocessed speech information is converted into a feature vector through a feature extraction algorithm; Speech recognition model processing: The extracted feature vectors are fed into the trained HMM-GMM model, which converts speech to text based on the input feature vector sequence; Text output and post-processing: Post-process the converted text information, including error correction and format adjustment.
3. The information display method based on AI algorithm according to claim 1, characterized in that: Text enhancement: Perform detailed completion operations based on grammatical structure and keywords, and simultaneously optimize text information, including at least: correcting grammatical errors and adjusting vocabulary order; The triggering mechanism is: Sample images of various sizes are pre-selected on the app for selection; the sample images of various sizes are displayed in order from small to large size specifications.
4. The information display method based on AI algorithm according to claim 1, characterized in that: Before text enhancement, recommendation recognition processing is also included: When the user uses the app for the first time or the amount of historical data under the corresponding user does not reach the set threshold, the emotion recognition action is performed to determine the emotional tendency, including positive, negative and neutral; when the amount of historical data under the corresponding user reaches the set threshold, the preference recommendation action is performed to determine the preferred option; The process of determining emotional tendency is as follows: Run the emotional index library, which records words representing positive, neutral, and negative emotional tendencies as keywords for emotional recognition, and is used to judge emotional tendencies and select matching picture styles accordingly; positive emotional tendencies correspond to bright tones; negative emotional tendencies correspond to dark tones; neutral emotional tendencies correspond to tones other than bright and dark; The process of determining the preferred option is: Obtain the historical selection and behavior records of the corresponding user, analyze them using machine learning algorithms, and build a preference model for the corresponding user.
5. The information display method based on AI algorithm according to claim 4 is characterized in that: After the enhanced text is input into the AI system, it is also combined with the results of the recommended recognition processing to generate several pictures.
6. The information display method based on AI algorithm according to claim 1, characterized in that: The calculation process of the maximum number of displays on the same screen on the app side is as follows: Determine the display area: Get the width W and height H of the app screen size; subtract the layout space, including at least the space occupied by the screen edge, navigation bar, and toolbar, to get the actual display area width W_usable and height H_usable; Determine the grid unit size: According to the design requirements, determine the width w_unit and height h_unit of each grid unit; Calculate the same-screen display volume: In the width direction, the number of grid units placed is: W_usable / w_unit; in the height direction, the number of grid units placed is: H_usable / h_unit; Then, the calculation formula for the maximum value N_max of the same-screen display volume is as follows: N_max = floor(W_usable / w_unit) * floor(H_usable / h_unit) Among them, the floor() function means rounding down.
7. The information display method based on AI algorithm according to claim 1, characterized in that: The corresponding index of clarity is: average gradient; the corresponding index of color reproduction is: structural similarity index; The popularity score is calculated by the number of likes, shares, and comments on the corresponding image under the current display number. The formula is: Where FS represents the popularity score, h1, h2 and h3 are weight coefficients, all ranging from [0, 1], dz represents the number of likes, fs represents the number of shares, pl represents the number of comments, and zs represents the number of impressions; The time decay score indicates the heat decay of the corresponding image over time after it is released. It is calculated using the time decay function. The formula is: Where Fj represents the time decay fraction, e represents the natural constant, k represents the cooling coefficient, and its value range is [0, 1]. (T-T0) represents the time interval, where T is the current time and T0 is the time when the picture is released.
8. The information display method based on AI algorithm according to claim 1, characterized in that: The content of the optimization and adjustment strategy is as follows: Keyword extraction technology is used to obtain keywords in the feedback voice information conversion feedback text information, and corresponding parameter optimization is made according to the content of the keywords. The interval of parameter optimization is a preset fixed value, and at least three types of pictures with optimized parameters are displayed on the app for secondary selection; when the corresponding user selects any picture, the parameters under the corresponding picture are recorded and used as standard parameters.
Citation Information
Patent Citations
An AR spatial annotation and display method integrating a general AI assistant
CN116186310B
Image display method and device, computer equipment and storage medium
CN116150421A
Image generation method, device and equipment and computer readable storage medium
CN117541683A
Method and system for generating controllable high-quality AI drawing picture description
CN119151793A