Image analysis and beautification method and system

Through the precise matching of images and text tokens and modular design, the problem of the lack of interpretability and multi-task coordination in facial aesthetic evaluation and beautification of deep learning methods is solved, and efficient and precise facial beautification and aesthetic analysis is achieved, enhancing user interactivity and system scalability.

CN120236171APending Publication Date: 2025-07-01THE CHINESE UNIV OF HONG KONG (SHENZHEN)
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510307818.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-17
Publication Date
2025-07-01

AI Technical Summary

Technical Problem

Existing deep learning methods lack interpretability and multi-task coordination capabilities in facial aesthetic evaluation and beautification, and cannot achieve interpretable and interactive facial aesthetic analysis and beautification.

Method used

By extracting image tokens and text tokens, the alignment model is used to generate beautified text, and iterative training is combined with multiple loss functions to achieve accurate matching between images and text, and a modular design image analysis and beautification system is adopted.

Benefits of technology

It improves beautification efficiency and accuracy, enhances the interpretability of aesthetic analysis and beautification tasks, and improves the scalability and user experience of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120236171A_ABST
    Figure CN120236171A_ABST
Patent Text Reader

Abstract

The invention discloses an image analysis and beautification method and system. The method comprises the following steps: extracting a first image token according to an original image; determining a first text token according to the first image token and an alignment model; determining a beautified text according to the first text token; and beautifying the original image according to the beautifying text. The first text token used for aesthetic analysis is generated by using the alignment model, then the first text token is used for generating the beautified text highly matched with the image content, the beautified text is used for guiding the beautification process of the original image, and the image and the text token are accurately matched, so that the beautification efficiency and accuracy are improved, and the image quality is improved. The utilization efficiency of multi-modal information is remarkably improved, aesthetic analysis and beautification tasks are successfully considered, and the interpretability and the model effect of the aesthetic analysis and beautification tasks are enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of computer vision and natural language processing, and particularly relates to an image analysis and beautification method and system. Background Art

[0002] Facial beauty plays an important role in people's daily activities such as social interaction and entertainment. Therefore, objectively evaluating the degree of facial beauty, discovering the common facial attributes of beautiful people, and guiding plastic surgery, daily makeup styling, and the design of beautification algorithms based on this have important significance.

[0003] With the continuous progress of deep learning technology, researchers have begun to use CNN (Convolutional Neural Network) or Transformer models to perform end-to-end aesthetic score prediction on facial images, and these methods have achieved remarkable results in improving the prediction accuracy. However, such deep learning methods have a significant defect, that is, they lack interpretability. They cannot intuitively show the basis and internal mechanism of facial beauty evaluation, which makes it difficult for people to understand and trust these prediction results. In terms of facial beautification, based on the deep learning-based generative model, by selecting beautiful samples as reference images, performing style transfer on the input images, and learning and applying the key attributes and features of the reference images, facial beautification can be achieved.

[0004] However, the existing technologies are often designed for a single task, ignoring the multi-task coordination ability of the model. When designing the model, researchers usually only focus on a certain sub-task in "aesthetic analysis" or "face beautification", lacking the ability to unify the modeling of multiple tasks. This limits the model's ability to perform multi-modal fusion at a higher level (such as text modality and image modality), and thus it is impossible to achieve interpretable and interactive face aesthetic analysis and beautification. Summary of the Invention

[0005] To solve the above problems, the present invention discloses an image analysis and beautification method and system.

[0006] The present invention discloses an image analysis and beautification method, including the following steps: Extract a first image token according to the original image; Determine a first text token according to the first image token and an alignment model; Determine a beautification text according to the first text token; Beautify the original image according to the beautification text.

[0007] Preferably, before determining the first text token according to the first image token and the alignment model, it further includes: Determine a second image token according to the training image, and determine a second text token according to the training text; Input the second image token and the second text token into a large language model for semantic alignment training to obtain an alignment model.

[0008] Preferably, determine the beautified text according to the first text token, specifically: Train a text model according to the first text token and the label beautified text; Determine the beautified text according to the first text token and the text model.

[0009] Preferably, beautify the original image according to the beautified text, specifically: Input the beautified text and the original image into a preset generation model to obtain an initial beautified image; Determine multiple loss functions according to the initial beautified image; Perform multiple iterative trainings on the generation model according to the multiple loss functions to obtain a beautified image.

[0010] Preferably, the loss function includes a semantic matching loss, Correspondingly, determine multiple loss functions according to the initial beautified image, specifically: Determine the third text according to the initial beautified image; Determine a semantic matching loss function according to the beautified text and the third text.

[0011] Preferably, perform multiple iterative trainings on the generation model according to the multiple loss functions to obtain a beautified image, specifically: Perform a weighted sum on the multiple loss functions; Perform multiple iterative trainings on the generation model according to the value of the weighted sum to obtain a beautified image.

[0012] Preferably, the loss function further includes an aesthetic score improvement loss, Correspondingly, determine multiple loss functions according to the initial beautified image, specifically: Determine an original image score value according to the first image token, the alignment model, and a preset classifier; Determine an initial beautified image score value according to the initial beautified image, the alignment model, and a preset classifier; Determine an aesthetic score improvement loss according to the original image score value and the initial beautified image score value.

[0013] Preferably, the loss function further includes an adversarial loss.

[0014] Preferably, the loss function further includes an identity preservation loss, Correspondingly, determine multiple loss functions according to the initial beautified image, specifically: Determine the original identity information corresponding to the original image and the beautified identity information corresponding to the initial beautified image; Determine the identity preservation loss according to the original identity information and the beautified identity information.

[0015] The present invention also discloses an image analysis and beautification system, including: An extraction module for extracting a first image token according to the original image; A graphic-text module for determining a first text token according to the first image token and an alignment model; A beautified text module for determining beautified text according to the first text token; A beautified image module for beautifying the original image according to the beautified text.

[0016] Compared with the prior art, the present invention has the following beneficial effects: (1) The present invention uses an alignment model to generate a first text token for aesthetic analysis. Subsequently, the first text token is used to generate beautified text that highly matches the image content, and the beautified text is used to guide the beautification process of the original image. By precisely matching the image and the text token, the present invention improves the beautification efficiency and accuracy, not only significantly improving the utilization efficiency of multimodal information, but also successfully taking into account aesthetic analysis and beautification tasks; (2) The present invention analyzes the aesthetic features of the original image in detail through the first text token, identifies aspects that need to be improved, such as defective facial attributes, such as age spots, eye bags, facial contours, etc., and then generates specific beautified text to guide the image generation process, achieving a higher-quality beautification result; (3) Further, in the present invention, the user can modify the beautified text before generating the beautified image, improving the control over the final result; and the present invention generates beautified text through the first text token, providing a clear optimization path, and the user can understand the improvement basis for each step, enhancing the interpretability; (4) The system of the present invention adopts a modular design, and can be optimized and additionally trained for a single module, facilitating the expansion and maintenance of the model, while improving the overall processing efficiency. Through the modular design, the system can flexibly adapt to different task requirements and ensure the efficient cooperation between modules, thereby improving the scalability and maintainability of the system; (5) The present invention ensures the effective training of each module (such as the extraction module, the graphic-text module, the beautified text module, the beautified image module) through a specially designed loss function and optimization strategy, achieving accurate generation effects under different tasks. This design not only improves the overall performance of the system, but also significantly improves the user experience, making the image beautification process more efficient and accurate. Description of the Drawings

[0017] Figure 1 It is a process diagram for the present invention to extract the second image token and the second text token; Figure 2 It is a schematic diagram of the training process of each module of the present invention. Detailed implementation manners

[0018] In the following description, specific details such as specific system architectures and technologies are presented for the purpose of illustration rather than limitation, so as to thoroughly understand the embodiments of the present invention. However, those skilled in the art should clearly understand that the present invention can also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid unnecessary details from interfering with the description of the present invention.

[0019] As Figure 2 shown, the present invention discloses an image analysis and beautification method, including the following steps: S1. Extract the first image token according to the original image; Preferably, before S2, it includes: Determine the second image token according to the training image, and determine the second text token according to the training text; The acquisition and processing processes of the training image and the training text are as follows: To obtain rich and high-quality paired image-text data for facial aesthetics analysis, face images can be collected as training images based on open-source datasets such as SCUT-FBP5500, CelebA, FFHQ, etc.

[0020] In order to enable the input training images to be processed in a standardized form, the present invention also performs standardization processing on them. The specific content of the standardization processing includes: (1) Facial key point localization: Locate 68 key points of the training image through Dlib, MTCNN or other key point detection algorithms; (2) Alignment and cropping: Rotate the face region of the training image to a standard frontal face view based on the key point information, and crop out an image of a unified size (such as 224×224 or 256×256). If more background information needs to be retained, a certain amount of blank space can also be set during cropping, but it should be ensured that the proportion of the face in the image is large enough; (3) Data augmentation: Conventional augmentation operations such as random rotation, horizontal mirroring, brightness / contrast adjustment can be appropriately added to improve the robustness of the model. It should be noted that too strong visual transformations will affect the recognition of aesthetic features and should be kept within an appropriate range; The training text contains more hierarchical text information such as beauty scores, facial attribute analysis, plastic surgery suggestions, makeup suggestions, and facial feature summaries, laying a data foundation for subsequent multi-modal learning.

[0021] Specifically, the training text includes: (1) Beauty score: such as an overall beauty score from 1 to 5, or more detailed aesthetic score values for each area, etc.

[0022] (2) Facial attribute analysis: including specific observations and descriptions of eyebrows, eyes, nose, lips, skin condition, face shape contour, etc. For example, "The tip of this user's nose is slightly round and blunt, and the nasal wings are relatively wide", etc. These descriptions need to be operable; (3) Cosmetic / plastic surgery suggestions: Based on clinical medicine and aesthetic experience, give suggestions on the areas that need improvement, such as "Perform moderate nasal plastic surgery to enhance the three-dimensionality of the nose bridge and narrow the nasal wings", etc.; (4) Makeup suggestions: Propose improvement plans from the perspective of daily makeup or stage makeup, such as "Modify the eyebrow shape with an eyebrow pencil, deepen the eyebrow color, and enhance the three-dimensionality of the face", etc. (5) Summary of facial features: Make a comprehensive evaluation of the advantages and disadvantages of the overall face, and can give simple suggestions on hairstyles, accessories, etc.

[0023] In other embodiments, it is also possible to model the user's historical selections and preferences, add personalized text to the training text, so that the system can give different scores and suggestions according to different aesthetic preferences when facing the same face.

[0024] Finally, we will obtain a large amount of data in the form of "binary tuples": (training images, training text), and the training images are standardized cropped images.

[0025] To complete multi-modal pre-training and subsequent fine-tuning, it is necessary to accurately pair and label the training image-training text pairs.

[0026] Specific annotation methods include but are not limited to a hierarchical annotation strategy: first, a professional annotator makes a rough annotation, then a senior expert conducts a review and fine-tuning, and finally, consistency analysis and cleaning are performed to ensure the correctness, effectiveness, and consistency of the annotation.

[0027] After preparing the training images and training text, the present invention also selects an open-source large language model with multi-modal capabilities, such as "Llama3.2-vision". This model has usually been pre-trained on a large number of text and image pairs, and has general alignment capabilities and powerful language generation capabilities.

[0028] The present invention performs fine-tuning on the basis of the large language model, which can significantly reduce the demand for large-scale labeled data and obtain better generalization capabilities and knowledge transfer capabilities.

[0029] The second image token and the second text token are input into a large language model for semantic alignment training to obtain an alignment model.

[0030] The overall structure of the alignment model includes the following four parts: (1) Image Encoder (Vision Encoder): Such as ViT (Vision Transformer) or other CNN-Transformer hybrid structures, which are used to encode the training images into a sequence of visual encoding vectors (Tokens), that is, the second image tokens. The second image tokens should retain as much facial key feature information as possible, such as Figure 1 shown; (2) Text Encoder (Language Encoder): Usually based on the Transformer architecture of the large language model and retaining its original structure during fine-tuning. The advantage of this is that it can make full use of the word vectors and context expression capabilities learned during the pre-training stage, so as to achieve better performance in specific tasks; As Figure 1 shown, in the present invention, the training text (including aesthetic scores, facial feature analysis, suggestions, summaries, etc.) is converted into the second text tokens {w_1, w_2,.., w_m} by the text encoder.

[0031] (3) Cross-modal Interaction Module (Cross-Attention / Optional Image-Text Matching Head): In the multi-modal large language model, through the cross-modal attention mechanism, the second image tokens and the second text tokens interact with each other, and high-level semantic alignment of the training images and training texts is performed based on the mapping learned during pre-training; Preferably, the present invention uses the attention mechanism to learn the semantic correspondence between the second image tokens and the second text tokens. The core idea is to maximize the similarity between the second image tokens representing the same training image and the keywords in the second text tokens, and use 0 as a mask to cover other irrelevant text descriptions in the second text tokens, which can be achieved by contrastive learning loss or maximizing cosine similarity, etc.

[0032] Suppose there is a training image, and its corresponding second text tokens may contain the following descriptions: Keywords: bright eyes, soft skin color, symmetrical face shape (directly related to the aesthetic features of the image). Irrelevant text descriptions: cluttered background, dim light (not related to the aesthetic features of the image).

[0033] The core idea of the alignment training is: Maximize similarity: Make the similarity between the second image tokens of the training images (such as the visual features of "bright eyes", "soft skin tone", "symmetrical face shape") and the keywords in the second text tokens ("bright eyes", "soft skin tone", "symmetrical face shape") as high as possible.

[0034] Minimize similarity: Mask the text descriptions in the second text tokens that are not relevant to the second image tokens (such as "cluttered background", "dim light") with 0.

[0035] (4) Language Decoder: The role of the language decoder is to convert the text tokens generated by the cross-modal interaction module into natural language output, that is, a text sequence, in scenarios where text needs to be output (such as automatically generating an aesthetic report based on a face image).

[0036] The training process of the language decoder is as follows: When we need to make the model "output an explanatory text based on the training image", we will give the second image tokens as conditional input during the training phase, let the language decoder generate a text sequence aligned with it, and calculate the autoregressive language model loss (Cross-Entropy Loss). The training objective is to make the generated text sequence close to the labeled training text in terms of syntax, semantics, and aesthetic description.

[0037] After training, inputting an image into the trained alignment model can generate text corresponding to the image; it is also possible to generate corresponding image tokens based on the input text.

[0038] S2. Determine the first text token according to the first image token and the alignment model; S2 specifically includes: (1) Image feature encoding Perform standardization processing on the original image, and input the standardized original image into the Vision Encoder (such as the ViT model) to obtain the first image tokens {v_1, v_2,.., v_k}. In ViT, each first image token corresponds to the embedding representation of an image patch.

[0039] (2) Cross-modal alignment and matching Generate the first text token according to the first image token; (3) Language generation decoding Furthermore, the language decoder can be used to convert the first text token into the first text, and the first text contains aesthetic analysis content such as aesthetic scores, facial feature analysis, and summary information.

[0040] S3. Determine the beautified text according to the first text token; Preferably, S3 is specifically: S31. Train a text model by beautifying text based on the first text token and labels; the text model is a large language model, which can be the large language model used when training the alignment model or other large language models; in the present invention, the beautified text is the text marked and used for training the text model, which contains makeup suggestions and plastic surgery suggestions corresponding to the first text token, such as semantic information like "raise the nose bridge", "narrow the nasal wings", "make the facial contour softer", etc.

[0041] In the present invention, by fine-tuning on the basis of a pre-trained large language model, the obtained alignment model can fully utilize the aesthetic knowledge in the large-scale corpus. For example, the first text token can more accurately describe the aesthetic features of the image, so that the generated beautified text is more precisely semantically aligned with the image content. This precise semantic alignment ensures that the present invention can generate high-quality beautified images according to the accurate beautified text.

[0042] Furthermore, the model of the present invention can not only output beautified images, but also output aesthetic analysis and beautified text generated based on the first text token, enabling users to understand why a certain beauty score is given and how to improve the face, etc., with better user interactivity and transparency.

[0043] S32. Determine the beautified text according to the first text token and the text model; Specifically, input the first text token into the trained text model to obtain the beautified text.

[0044] S4. Beautify the original image according to the beautified text.

[0045] Preferably, S4 is specifically: S41. Input the beautified text and the original image into a preset generation model to obtain an initial beautified image; S41 is specifically as follows: Convert the beautified text into beautified text tokens. In generation models such as GAN, VAE, or Diffusion-based, first encode the original face into the latent space z. Then adjust the vectors of some dimensions in the z space according to t (beautified text tokens) to generate an initial beautified image.

[0046] In other embodiments, techniques such as Neural Rendering and NeRF (Neural Radiance Field) can also be used to generate the initial beautified image. In this embodiment, the beautification of the original image not only focuses on geometry and texture, but also attaches importance to light, shadow, and texture, and can show natural transitions and real details under indoor lighting and outdoor light conditions, with a more persuasive visual effect.

[0047] S42. Determine multiple loss functions according to the initial beautified image; S43. Iteratively train the generation model multiple times according to multiple loss functions to obtain the beautified image.

[0048] Preferably, S43 is specifically as follows: S431. Perform weighted summation on multiple loss functions; S432. Iteratively train the generation model multiple times according to the value of the weighted summation to obtain the beautified image.

[0049] During the iterative training, the generation model continuously performs backpropagation according to the weighted summation of multiple loss functions, and according to the preset iterative termination condition, strives to find the best balance among multiple loss functions.

[0050] Preferably, the loss function includes a semantic matching loss. Correspondingly, S42 is specifically as follows: S421. Determine the third text according to the initial beautified image; S422. Determine the semantic matching loss function according to the beautified text and the third text.

[0051] In one embodiment, S42 is specifically as follows: Input the initial beautified image into the alignment model. Similar to the generation process of the first text token, obtain the third text token corresponding to the initial beautified image, and then obtain the third text. The third text includes the aesthetic analysis information in the initial beautified image; then compare the third text token with the beautified text token, and use contrastive learning loss or maximize the cosine similarity to minimize the distance between the two in subsequent multiple iterative trainings.

[0052] In other embodiments, an external attribute recognition model can also be used to generate the third text.

[0053] Preferably, the loss function further includes an aesthetic score improvement loss. Correspondingly, S42 is specifically as follows: S421. Determine the original image score value according to the first image token, the alignment model, and a preset classifier; S422. Determine the initial beautified image score value according to the initial beautified image, the alignment model, and a preset classifier; S423. Determine the aesthetic score improvement loss according to the original image score value and the initial beautified image score value.

[0054] In this embodiment, a small regression head or classifier is added to the output end of the alignment model to predict the score values of the original image and the beautified image; Preferably, mean squared error (MSE) or cross-entropy loss can be used when training the regression head or classifier.

[0055] By inputting the first text token and the beautified text token output by the alignment model into the trained regression head or classifier respectively, the original image score value and the beautified image score value can be obtained respectively.

[0056] In other embodiments, if you want to generate discrete scores and distinguish specific details at the same time, you can also jointly train the regression head and the classifier to improve the accuracy and stability of the scores.

[0057] The original image score value and the initial beautified image score value are obtained from the beauty score prediction model, the aforementioned classifier or regression head. After each iteration, calculate the latest beautified image score value, and hope that the latest beautified image score value is as high as possible compared to the original image score value; Furthermore, a relatively high ideal aesthetic score value can be selected, and in subsequent multiple iterative trainings, make the latest beautified image score value as close to or exceed the ideal aesthetic score value as possible; How much higher than the original image score value, how much exceeding the ideal aesthetic score value, and the ideal aesthetic score value can all be set by those skilled in the art themselves.

[0058] Preferably, the loss function further includes an adversarial loss. In the adversarial loss, a discriminator is used to make the generated image more realistic, ensuring that the generated face is more in line with the real face distribution in terms of texture, lighting, details, etc.

[0059] Preferably, the loss function further includes an identity preservation loss. Correspondingly, S42 is specifically as follows: S421. Determine the original identity information corresponding to the original image and the beautified identity information corresponding to the initial beautified image; S422. Determine the identity preservation loss according to the original identity information and the beautified identity information.

[0060] Both the original identity information and the beautified identity information can be extracted through face recognition models such as ArcFace. When using the identity preservation loss for subsequent iterative training, it is necessary to ensure that the cosine similarity of the identity features extracted from the face before and after beautification in face recognition models such as ArcFace is as high as possible. The specific similarity is set by those skilled in the art themselves.

[0061] When beautifying the original image in the present invention, by combining multiple loss functions, it is ensured that the beautified image is still recognizable and the visual beauty is further enhanced.

[0062] In other embodiments, the beautified text can also be manually input. When the user inputs an original image and gives a text description of "wanting to become more beautiful", the system will first extract the identity vector and perform aesthetic scoring on the original image, then encode the text description, fuse it with the existing beautification experience vector, and finally send it to the generation model to obtain the output beautified image. If the user is not satisfied, they can further modify the beautified text description and generate it again, or make further corrections to the system (such as "don't make the chin too pointed"), thus realizing interactive and controllable face beautification.

[0063] The present invention also discloses an image analysis and beautification system, including: An extraction module, configured to extract a first image token according to the original image; A graphic and text module, configured to determine a first text token according to the first image token and an alignment model; A beautified text module, configured to determine the beautified text according to the first text token; A beautified image module, configured to beautify the original image according to the beautified text.

[0064] Multiple modules of the present invention can cooperate with each other. Multiple modules share a set of encoders, decoders, and alignment models, and there is no need to establish multiple models for different subtasks (such as scoring prediction, facial analysis, beautification suggestions, plastic surgery suggestions, etc.), reducing the deployment and maintenance costs.

[0065] Furthermore, through the integrated design of the present invention, the information flow between modules is more coherent, ensuring that the output of each module can be fully utilized by the next module, and the optimization objectives of the system are more consistent. All modules jointly optimize the final effect of image beautification, ensuring that the output of each module can serve the final beautified image, improving the overall efficiency and accuracy of the system. And with the modular design, the system has stronger flexibility and scalability. Each module can be independently optimized and extended while maintaining close cooperation with other modules. For example, if a new aesthetic analysis function needs to be added, only the corresponding sub-module needs to be added to the graphic and text module without modifying other modules; when the user is not satisfied with the beautified image, they can directly modify the beautified text without retraining other modules.

[0066] In one embodiment, when a given original face image is provided, the system of the present invention has the following capabilities: (1) Facial aesthetic scoring After the user uploads the face image, the model can directly output the aesthetic score value and a brief text description, such as "The overall aesthetic score value is 4.2 points, the proportion of facial features is good, but the eyebrows are slightly sparse and the color is light, etc.".

[0067] (2) Facial feature analysis and suggestions The model will automatically generate a detailed aesthetic analysis report, including key elements such as eyebrows, eyes, nose, lips, facial contour, and skin condition, and give corresponding plastic surgery or makeup suggestions. Even on this basis, it will recommend suitable skin care plans and hairstyle combinations.

[0068] (3)Personalized human face beautification image generation If the user has improvement requirements for specific parts such as the nose, jawline, and lip color, they can indicate them in the text input. The model will generate an optimized human face image that not only retains personal identity characteristics but also meets the user's beautification needs based on the multimodal interaction and human face generation module.

[0069] (4)Interactivity and interpretability On the one hand, users can interact with the system in a natural language way to view the text of the human face aesthetic analysis and beautification suggestions given by the system; on the other hand, the system has a certain interpretability for the scoring and generation process of the beautification image, such as "because the bridge of the nose is relatively low, so it is recommended to raise the height of the bridge of the nose in the generated image", etc.

[0070] Compared with the prior art, the present invention has the following beneficial effects: (1)The present invention uses an alignment model to generate the first text token for aesthetic analysis. Subsequently, the first text token is used to generate a beautification text that highly matches the image content, and the beautification text is used to guide the beautification process of the original image; by precisely matching the image with the text token, the present invention improves the beautification efficiency and accuracy, not only significantly improving the utilization efficiency of multimodal information, but also successfully taking into account the aesthetic analysis and beautification tasks, enhancing the interpretability and model effect of the aesthetic analysis and beautification tasks; (2)The present invention uses the first text token to detailedly analyze the aesthetic features of the original image, identify the aspects that need to be improved, such as defective facial attributes, such as age spots, eye bags, facial contour, etc., and then generate specific beautification text to guide the image generation process, achieving a higher-quality beautification result; (3)Furthermore, in the present invention, the user can modify the beautification text before generating the beautification image, improving the control over the final result; and the present invention provides a clear optimization path through the first text token and the beautification text, enabling the user to understand the improvement basis for each step and enhancing the interpretability; (4)The system of the present invention adopts a modular design, which can optimize and additionally train a single module, facilitating the expansion and maintenance of the model, while improving the overall processing efficiency. Through the modular design, the system can flexibly adapt to different task requirements and ensure the efficient cooperation between modules, thereby improving the scalability and maintainability of the system; (5) Through a specially designed loss function and optimization strategy, the present invention ensures that each module (such as the extraction module, the graphic and text module, the beautification text module, and the beautification image module) can be effectively trained to achieve accurate generation effects under different tasks. This design not only improves the overall performance of the system but also significantly enhances the user experience, making the image beautification process more efficient and precise.

[0071] As described above, these are only several embodiments of the present application and do not impose any form of limitation on the present application. Although the present application is disclosed above with preferred embodiments, it is not intended to limit the present application. Any person skilled in the art, without departing from the scope of the technical solution of the present application, making some changes or modifications using the disclosed technical content is equivalent to equivalent implementation cases and all fall within the scope of the technical solution.

Claims

1. An image analysis and beautification method, characterized in that: The following steps are involved: extracting a first image token based on the original image; determining a first text token based on the first image token and the alignment model; Determine the beautified text according to the first text token; The original image is beautified according to the beautification text.

2. The image analysis and beautification method according to claim 1, characterized in that: Before determining the first text token according to the first image token and the alignment model, the method further includes: Determine a second image token based on the training image, and determine a second text token based on the training text; The second image token and the second text token are input into a large language model for semantic alignment training to obtain an alignment model.

3. The image analysis and beautification method according to claim 1, characterized in that: Determine the beautified text according to the first text token, specifically: Training a text model based on the first text token and the beautified text with the tag; A beautified text is determined according to the first text token and the text model.

4. The image analysis and beautification method according to claim 1, characterized in that: Beautify the original image according to the beautification text, specifically: Inputting the beautified text and the original image into a preset generation model to obtain an initial beautified image; Determining a plurality of loss functions according to the initial beautified image; The generation model is iteratively trained multiple times according to the multiple loss functions to obtain a beautified image.

5. The image analysis and beautification method according to claim 4, characterized in that: The loss function includes semantic matching loss, Accordingly, multiple loss functions are determined according to the initial beautified image, specifically: Determining a third text according to the initial beautified image; A semantic matching loss function is determined according to the beautified text and the third text.

6. The image analysis and beautification method according to claim 4, characterized in that: The generation model is iteratively trained multiple times according to the multiple loss functions to obtain a beautified image, specifically: Performing weighted summation on the multiple loss functions; The generation model is iteratively trained multiple times according to the weighted sum value to obtain a beautified image.

7. The image analysis and beautification method according to claim 5, characterized in that: The loss function also includes the aesthetic score improvement loss, Accordingly, multiple loss functions are determined according to the initial beautified image, specifically: Determining an original image score value according to the first image token, the alignment model and a preset classifier; Determining a score value of the initial beautified image according to the initial beautified image, the alignment model and a preset classifier; An aesthetic score improvement loss is determined according to the original image score value and the initial beautified image score value.

8. The image analysis and beautification method according to claim 7, characterized in that: The loss function also includes adversarial loss.

9. The image analysis and beautification method according to claim 8, characterized in that: The loss function also includes an identity preservation loss, Accordingly, multiple loss functions are determined according to the initial beautified image, specifically: Determining original identity information corresponding to the original image and beautified identity information corresponding to the initial beautified image; An identity preservation loss is determined according to the original identity information and the beautified identity information.

10. An image analysis and beautification system, characterized in that: include: An extraction module, configured to extract a first image token based on the original image; An image-text module, configured to determine a first text token according to the first image token and the alignment model; A beautified text module, used for determining the beautified text according to the first text token; The image beautification module is used to beautify the original image according to the beautification text.