Character image generation method and device, electronic equipment and storage medium

By coding semantic feature codes and feature fusion of historically generated image images on the image description text input by the user, the problem of inappropriate generation results is solved, and a more accurate and personalized virtual character image generation is achieved.

CN120374767APending Publication Date: 2025-07-25PING AN TECH (SHENZHEN) CO LTD
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510437489.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-08
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

The prior art is difficult to accurately generate virtual character image images that meet users' personalized needs, resulting in inappropriate results.

Method used

By semantic feature encoding of the image description text input by the user, the image constituent elements are identified, and the feature update and fusion is updated and integrated with historically generated image images and user preference features, the fusion features are generated, and the image generation is finally generated.

Benefits of technology

Improve the accuracy and personalized matching of generated character image images, ensuring that the generated results are more in line with user preferences and historical behaviors.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120374767A_ABST
    Figure CN120374767A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a character image generation method and device, electronic equipment and a storage medium, belongs to the technical field of artificial intelligence, and is suitable for the fields of financial science and technology and medical science and technology. The method comprises the following steps: carrying out semantic feature coding on a current image description text to obtain description text semantic features; and carrying out image element identification on the semantic features of the description text to obtain image components. And updating the semantic features of the description text according to the historically generated image and the image composition elements to obtain the current semantic features. Performing image screening according to the current semantic feature to obtain a reference character image, and performing image feature coding on the reference character image to obtain a reference image feature; and performing feature fusion according to the description text semantic features and the reference image features to obtain guidance fusion features. And performing image generation based on the guidance fusion features. According to the embodiment of the invention, the accuracy of generating the character image can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology, and is applicable to the fields of fintech and medical technology. In particular, it relates to a method and device for generating a character image, an electronic device, and a storage medium. Background Art

[0002] Currently, relevant application programs have opened an image generation function. When a user inputs description text related to an image in the relevant application interface, the image generation model configured in the application program can output the corresponding image according to the description text. In some application scenarios such as online sales of financial products or medical consultations, users expect to generate a virtual character image that meets their preferences by inputting their own description text, and use this virtual character as the dialogue interaction object in subsequent consultation services.

[0003] However, the clarity of the description text input by each user varies, making it difficult for the model to grasp the accurate meaning of the description text, resulting in the finally generated character image not being close enough to the personalized needs of the user. Therefore, how to accurately generate a character image has become a technical problem to be solved urgently. Summary of the Invention

[0004] The main purpose of the embodiments of this application is to propose a method and device for generating a character image, an electronic device, and a storage medium, aiming to improve the accuracy of generating a character image.

[0005] To achieve the above object, in the first aspect of the embodiments of this application, a method for generating a character image is proposed. The method includes: Obtain the current image description text, and perform semantic feature encoding on the current image description text to obtain the semantic features of the description text; Perform image element recognition on the semantic features of the description text to obtain the image composition elements; Obtain the historical generated character image, and update the semantic features of the description text according to the historical generated character image and the image composition elements to obtain the current semantic features; Perform image screening according to the current semantic features to obtain a reference character image, and perform image feature encoding on the reference character image to obtain reference image features; Perform feature fusion according to the semantic features of the description text and the reference image features to obtain a guiding fusion feature; Generate an image based on the guiding fusion feature.

[0006] In some embodiments, the step of updating the semantic features of the description text according to the historical generated character image and the image composition elements to obtain the current semantic features includes: Determine the type of image elements based on the described image composition elements; Select a reference element type different from the image element type from multiple preset reference element types to obtain a missing element type; Conduct user preference analysis on the historically generated image to obtain user preference characteristics corresponding to the missing element type; Perform feature splicing based on the user preference characteristics and the semantic characteristics of the description text to obtain the current semantic characteristics.

[0007] In some embodiments, the conducting user preference analysis on the historically generated image to obtain user preference characteristics corresponding to the missing element type includes: Perform image disassembly on the historically generated image according to the missing element type to obtain an image composition sub-region corresponding to the missing element type; Perform human recognition on the historically generated image to obtain a human image sub-region; Perform image feature analysis based on the image composition sub-region and the human image sub-region to obtain the user preference characteristics.

[0008] In some embodiments, the performing image feature analysis based on the image composition sub-region and the human image sub-region to obtain the user preference characteristics includes: Perform area preference analysis based on the human image sub-region and the historically generated image to obtain a preference characteristic for the proportion of the human area; Perform image semantic recognition on the image composition sub-region to obtain element semantic characteristics; Perform feature correlation analysis based on the element semantic characteristics and the preference characteristic for the proportion of the human area to obtain the user preference characteristics.

[0009] In some embodiments, the performing image generation based on the guidance fusion feature includes: Perform audio screening according to the current semantic characteristics to obtain a reference audio file; Perform audio feature encoding on the reference audio file to obtain reference audio features; Perform feature fusion based on the reference audio features and the guidance fusion feature to obtain a target fusion feature; Perform image generation based on the target fusion feature.

[0010] In some embodiments, the performing image generation based on the guidance fusion feature includes: performing image generation on the guidance fusion feature through a preset target human image generation model, where the target human image generation model is trained through the following training method: Obtain a sample character image description text; wherein, the sample character image description text is used to describe a preset sample character image; Perform image screening based on the sample character image description text and the sample generated image to obtain a guiding image; Perform feature fusion based on the sample character image description text and the guiding image to obtain a sample fusion feature; Generate an initial character image by generating an image of the sample fusion feature through a preset initial image generation model; Perform quality evaluation on the initial character image to obtain a character image quality score; Construct target loss data based on the character image quality score, the sample character image, and the initial character image; Adjust the parameters of the initial image generation model according to the target loss data to obtain the target character image generation model.

[0011] In some embodiments, the performing quality evaluation on the initial character image to obtain a character image quality score includes: Perform rationality evaluation on the initial character image to obtain a rationality sub-score; Perform content richness evaluation on the initial character image to obtain a richness sub-score; Perform user satisfaction evaluation on the initial character image to obtain a satisfaction sub-score; Perform numerical calculation according to the rationality sub-score, the richness sub-score, and the satisfaction sub-score to obtain the character image quality score.

[0012] To achieve the above object, a second aspect of the embodiments of the present application proposes a character image generation device, the device includes: An acquisition module, configured to acquire a current image description text, and perform semantic feature encoding on the current image description text to obtain a description text semantic feature; An image element recognition module, configured to perform image element recognition on the description text semantic feature to obtain image composition elements; A semantic feature update module, configured to acquire a historical generated image, and update the description text semantic feature according to the historical generated image and the image composition elements to obtain a current semantic feature; An image screening module, configured to perform image screening according to the current semantic feature to obtain a reference character image, and perform image feature encoding on the reference character image to obtain a reference image feature; A feature fusion module, configured to perform feature fusion based on the semantic features of the description text and the reference image features to obtain guiding fusion features; An image generation module, configured to generate an image based on the guiding fusion features.

[0013] To achieve the above object, a third aspect of the embodiments of the present application provides an electronic device, where the electronic device includes a memory and a processor, the memory stores a computer program, and when the processor executes the computer program, the method described in the first aspect above is implemented.

[0014] To achieve the above object, a fourth aspect of the embodiments of the present application provides a computer-readable storage medium, where the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method described in the first aspect above is implemented.

[0015] The method, device, electronic device, and storage medium for generating a character image proposed in the present application encode the semantic features of the current image description text input by the user, further identify the image composition elements, and deeply understand the user's description at the semantic level, improving the parsing ability of the description intention. Secondly, the historical generated character image is introduced, and the semantic features of the description text are updated by combining the semantic features of the description text and the image elements, so that the description text can approach the user's historical preferences to obtain the current semantic features. Then, image screening is performed based on the current semantic features to obtain a reference character image, effectively guiding the generation process to approach the user's preferences, thereby improving the personalized matching degree of the generated image. Then, the reference character image is encoded for image features to obtain reference image features, and the features are fused with the description semantic features and the reference image features to generate guiding fusion features. And based on this, the target image is generated, which helps to establish a closer connection between semantic understanding and visual presentation, thereby enhancing the accuracy and expressiveness of the generated image. Therefore, the method of the embodiments of the present application effectively solves the problem that the generated results are not appropriate due to the large differences in user descriptions, realizes the accurate restoration of the character image, and improves the accuracy of generating the character image. Description of the Drawings

[0016] Figure 1 is a flowchart of the method for generating a character image provided by the embodiments of the present application; Figure 2 is Figure 1 a flowchart of step S103 in Figure 3 is Figure 2 a flowchart of step S203 in Figure 4 is Figure 3 a flowchart of step S303 in Figure 5 is Figure 1 the flowchart of step S106 in Figure 6 is Figure 1 the flowchart before step S106 in Figure 7 is the schematic structural diagram of the character image generation device provided by the embodiment of the present application; Figure 8 is the schematic hardware structure diagram of the electronic device provided by the embodiment of the present application. Detailed implementation manners

[0017] In order to make the objectives, technical solutions and advantages of the present application more clear and understandable, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0018] It should be noted that although the functional modules are divided in the device schematic diagram and the logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order from the module division in the device or the order in the flowchart. Terms such as "first" and "second" in the specification, claims and the above-mentioned drawings are used to distinguish similar objects and do not necessarily need to describe a specific order or sequence.

[0019] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which this application belongs. The terms used herein are only for the purpose of describing the embodiments of this application and are not intended to limit this application.

[0020] First, several nouns involved in the present application are analyzed: Artificial intelligence (AI): It is a new technical science that studies, develops theories, methods, technologies and application systems for simulating, extending and expanding human intelligence; artificial intelligence is a branch of computer science. Artificial intelligence attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a way similar to human intelligence. The research in this field includes robots, speech recognition, image recognition, natural language processing and expert systems, etc. Artificial intelligence can simulate the information process of human consciousness and thinking. Artificial intelligence also uses digital computers or machines controlled by digital computers to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain the best results of theories, methods, technologies and application systems.

[0021] Natural Language Processing (NLP): NLP uses computers to process, understand, and apply human languages (such as Chinese, English, etc.). NLP is a branch of artificial intelligence and an interdisciplinary field of computer science and linguistics, often referred to as computational linguistics. Natural language processing includes syntactic analysis, semantic analysis, discourse understanding, etc. Natural language processing is commonly used in technical fields such as machine translation, handwritten and printed character recognition, speech recognition and text-to-speech conversion, information intention recognition, information extraction and filtering, text classification and clustering, public opinion analysis, and opinion mining. It involves data mining, machine learning, knowledge acquisition, knowledge engineering, artificial intelligence research related to language processing, and linguistic research related to language computation, etc.

[0022] Information Extraction: A text processing technology that extracts factual information such as entities, relationships, events, etc. of a specified type from natural language texts and forms structured data for output. Information extraction is a technology for extracting specific information from text data. Text data is composed of some specific units, such as sentences, paragraphs, and texts. Text information is exactly composed of some small specific units, such as characters, words, phrases, sentences, paragraphs, or combinations of these specific units. Extracting noun phrases, personal names, place names, etc. from text data are all text information extraction. Of course, the information extracted by text information extraction technology can be various types of information.

[0023] Image Caption generates a natural language description for an image and uses the generated description to help applications understand the semantics expressed in the visual scene of the image. For example, image captioning can convert image retrieval into text retrieval, be used to classify images, and improve image retrieval results. People can usually describe the details of the visual scene of an image with just a quick glance, while automatically adding a description to an image is a comprehensive and challenging computer vision task that requires converting the complex information contained in the image into a natural language description. Compared with ordinary computer vision tasks, image captioning not only requires identifying objects in the image but also associating the identified objects with natural semantics and describing them in natural language. Therefore, image captioning requires people to extract the deep features of the image, associate them with semantic features, and convert them for generating descriptions.

[0024] In scenarios where it is necessary to generate human images, the user can draw by themselves or ask a painter to draw on a painting software. However, this method is limited by the skills and painting styles of the painters and has a certain operation threshold. Currently, some related applications have opened an image generation function. Users input description texts related to images in the relevant application interfaces, and the image generation models configured in the applications can output corresponding images according to the description texts. In some application scenarios of online sales of financial products or medical consultations, users expect to generate virtual human images that meet their preferences by inputting their own description texts, and use the virtual humans as the dialogue interaction objects in subsequent consultation services. However, the clarity of the description texts input by each user varies, making it difficult for the model to grasp the accurate meaning of the description texts, resulting in the finally generated human images not being close enough to the personalized needs of the users.

[0025] Therefore, how to accurately generate human image pictures has become a technical problem to be solved urgently.

[0026] Based on this, the embodiments of the present application provide a method and device for generating human image pictures, an electronic device, and a storage medium, aiming to improve the accuracy of generating human image pictures.

[0027] The method and device for generating human image pictures, the electronic device, and the storage medium provided by the embodiments of the present application are specifically described through the following embodiments. First, the method for generating human image pictures in the embodiments of the present application is described.

[0028] The embodiments of the present application can obtain and process relevant data based on artificial intelligence technology. Among them, artificial intelligence (AI) is a theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results.

[0029] Artificial intelligence basic technologies generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technologies, operation / interaction systems, and mechatronics. Artificial intelligence software technologies mainly include several major directions such as computer vision technology, robotics, biometric technology, speech processing technology, natural language processing technology, and machine learning / deep learning.

[0030] The method for generating a character image provided by an embodiment of the present application relates to the field of artificial intelligence technology. The method for generating a character image provided by an embodiment of the present application can be applied to a terminal, or to a server side, or can be software running on a terminal or a server side. In some embodiments, the terminal can be a smart phone, a tablet computer, a laptop computer, a desktop computer, etc.; the server side can be configured as an independent physical server, or can be configured as a server cluster or a distributed system composed of multiple physical servers, or can also be configured as a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms; the software can be an application that implements the method for generating a character image, etc., but is not limited to the above forms.

[0031] The present application can be used in many general or special computer system environments or configurations. For example: personal computers, server computers, handheld or portable devices, tablet devices, multi-processor systems, microprocessor-based systems, set-top boxes, programmable consumer electronic devices, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, and so on. The present application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. The present application can also be practiced in a distributed computing environment where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media including storage devices.

[0032] It should be noted that in each specific embodiment of the present application, when it comes to relevant processing that needs to be carried out according to data related to the user's identity or characteristics such as user information, user behavior data, user historical data, and user location information, the user's permission or consent will be obtained first. Moreover, the collection, use, and processing of these data will comply with relevant laws, regulations, and standards. In addition, when an embodiment of the present application needs to obtain sensitive personal information of the user, the user's separate permission or separate consent will be obtained through methods such as pop-up windows or jumping to a confirmation page. After clearly obtaining the user's separate permission or separate consent, the necessary user-related data for enabling the normal operation of the embodiment of the present application will be obtained.

[0033] Figure 1 is an optional flowchart of the method for generating a character image provided by an embodiment of the present application. Figure 1 The method in may include but is not limited to steps S101 to S106.

[0034] Step S101: Obtain the current image description text, and perform semantic feature encoding on the current image description text to obtain the semantic features of the description text.

[0035] Step S102: Identify the image elements from the semantic features of the description text to obtain the image composition elements.

[0036] Step S103: Obtain the historical generated image, and update the semantic features of the description text based on the historical generated image and the image composition elements to obtain the current semantic features.

[0037] Step S104: Perform image screening based on the current semantic features to obtain a reference character image, and perform image feature encoding on the reference character image to obtain reference image features.

[0038] Step S105: Perform feature fusion based on the semantic features of the description text and the reference image features to obtain the guiding fusion features.

[0039] Step S106: Generate an image based on the guiding fusion features.

[0040] Steps S101 to S106 shown in the embodiments of the present application, by performing semantic feature encoding on the current image description text input by the user, and further identifying the image composition elements, deeply understand the user's description at the semantic level, and improve the parsing ability of the description intention. Secondly, introduce the historical generated image, and update the semantic features of the description text by combining the semantic features of the description text and the image elements, so that the description text can approach the user's historical preferences to obtain the current semantic features. And perform image screening based on the current semantic features to obtain a reference character image, effectively guiding the generation process to approach the user's preferences, thereby improving the personalized matching degree of the generated image. Then, perform image feature encoding on the reference character image to obtain reference image features, and fuse the features with the description semantic features and the reference image features to generate guiding fusion features. And generate the target image based on this, which helps to establish a closer connection between semantic understanding and visual presentation, thereby enhancing the accuracy and expressiveness of the generated image. Therefore, the method of the embodiments of the present application effectively solves the problem that the generated results are not appropriate due to the large differences in user descriptions, realizes the accurate restoration of the character image, and improves the accuracy of character image generation.

[0041] In step S101 of some embodiments, the current image description text is a natural language description input by the user based on the expected generated image. The user can be the user who initiates the request for generating a character image, specifically including but not limited to mobile terminal users and web users. Exemplarily, in some image generation scenarios related to online medical consultations, the current image description text can be: "a female doctor in a white coat and black-framed glasses", etc. This description text can be directly obtained after being input by the user, or can be obtained after being converted into text form through speech recognition, and is not limited thereto.

[0042] The semantic features of the description text are the vectorized expression of the content of the current image description text at the semantic level, and are used to represent the key semantic information contained in the description. In the embodiments of the present application, a text encoder is provided. This text encoder can be implemented based on the Transformer architecture, and a multi-head attention mechanism is introduced inside to capture the context dependence and semantic association between different words in the current image description text. During the encoding process, first, the current image description text is tokenized, and each word is mapped to a word vector or an embedding representation. Then, through parallel processing of multiple attention heads, the association degree of each word in different semantic subspaces is modeled, so as to obtain a richer and more detailed context semantic representation, that is, the semantic features of the description text.

[0043] In step S102 of some embodiments, the image composition elements refer to the specific content information that constitutes the character image in the current image description text, specifically including but not limited to facial features, clothing styles, clothing colors, postures, etc., and is not limited thereto. Exemplarily, if the current image description text is: "a female doctor in a white coat and black-framed glasses", then "wearing a white coat", "wearing black-framed glasses", and "female" all belong to the image composition elements.

[0044] The recognition process of the image composition elements can be carried out through natural language processing techniques. First, perform language preprocessing operations on the input current image description text, including tokenization, part-of-speech tagging, and syntactic analysis, to better understand the basic structure of the text. Subsequently, use a pre-trained language model to perform semantic encoding on the current image description text, extract context correlation information, so as to obtain the semantic representation of each word or phrase in the sentence. Finally, through methods such as semantic attention mechanism, sentence structure analysis, or context correlation scoring, the language model can automatically judge which parts are the most valuable information when the user describes the character image.

[0045] In step S103 of some embodiments, the historical generated image is an image sample generated by the user based on other description inputs before, and can be retrieved from the user image generation records pre-stored in the image generation application, or can be actively selected and uploaded by the user.

[0046] For the method of updating the semantic features of the description text according to the historical generated image and the image composition elements, reference can be made to Figure 2 . In some embodiments, step S103 may include but is not limited to steps S201 to S204: Step S201, determining the type of image element based on the image composition elements.

[0047] Step S202, selecting a reference element type different from the image element type from multiple preset reference element types to obtain the missing element type.

[0048] Step S203, performing user preference analysis on the historical generated image to obtain the user preference features corresponding to the missing element type.

[0049] Step S204, performing feature splicing according to the user preference features and the semantic features of the description text to obtain the current semantic features.

[0050] In step S201 of some embodiments, the image element type refers to the identifier after classifying the composition elements in the current image description text according to semantic categories. For example, the human face (such as facial features and expressions), human clothing (such as accessory styles, clothing styles, clothing colors, and clothing materials), human actions (such as gestures and postures), backgrounds (such as scene environments and decorative sticker patterns), and lighting styles, etc., are not limited to this.

[0051] Specifically, the semantic features of the image composition elements can be parsed by constructing a parsing model containing semantic tags. For example, techniques such as named entity recognition models and text structure analysis networks based on attention mechanisms are used to automatically identify and classify the types of image elements actually mentioned in the user input.

[0052] In step S202 of some embodiments, the reference element type refers to the type of constituent elements required to generate an image of a human figure, which can be set according to the needs of the human figure, including but not limited to human face types, human clothing types, human action types, background types, and lighting styles. It can be understood that when the reference element type increases, it indicates that the elements to be considered in generating the image are more comprehensive, thereby enhancing the content richness of the generated image. The missing element type refers to the reference element type not involved in the current image description text. For example, in the scenario of generating human figures related to online medical consultations, the reference element types include human face types, human clothing types, human action types, and background types. If the current image description text input by the user is "I want to generate an image of a male doctor sitting upright with a smile", this description only describes the human action and facial state, but does not mention the information about clothing and background. Then, "human clothing type" and "background type" can be automatically identified as the missing element types. Another example is that in the scenario of generating human figures related to online sales of financial products, if the current image description text input by the user is "I want to generate an image of a young male wearing a white suit and gold-rimmed glasses", this description only describes the human clothing, but does not mention the information about action, human face, and background. Then, "human face type", "human action type", and "background type" can be automatically identified as the missing element types.

[0053] In some other embodiments, if there is no reference element type different from the image element type among multiple preset reference element types, it indicates that the content of the current image description text is comprehensive, and the subsequent steps S203 to S204 do not need to be performed.

[0054] In step S203 of some embodiments, the user preference feature refers to the characteristic representation of the stable visual tendency or style preference reflected by the user in the historical generated image of the human figure in the dimension corresponding to the missing element type.

[0055] Please refer to Figure 3 , in some embodiments, step S203 may include but is not limited to steps S301 to S303: Step S301, disassemble the historical generated image of the human figure according to the missing element type to obtain the image constituent sub-regions corresponding to the missing element type.

[0056] Step S302, perform human figure recognition on the historical generated image of the human figure to obtain the human figure sub-region.

[0057] Step S303, perform image feature analysis based on the image constituent sub-regions and the human figure sub-region to obtain the user preference feature.

[0058] In step S301 of some embodiments, the image composition sub-region refers to the local image region extracted from the historically generated image according to the type of missing elements. For example, in the scenario of generating a character image related to medical consultation, if the type of missing elements has been determined to be the character's clothing category, and the historically generated image shows a nurse in a blue nurse's uniform, the image region corresponding to the character's clothing category in the historically generated image is extracted, and the image region of the blue nurse's uniform is the image composition sub-region in this scenario.

[0059] The image decomposition process can be implemented by using an image processing technology based on semantic segmentation. Specifically, a pre-trained deep semantic segmentation model is used to perform pixel-level semantic annotation on the image, so that each region in the image has corresponding semantic category information, such as people, clothing, background, accessories, etc. After segmentation, combined with the semantic label corresponding to the type of missing elements, the target region is selected from the segmentation result, and then the corresponding image region is extracted from the historically generated image to obtain the image composition sub-region.

[0060] In step S302 of some embodiments, the character image sub-region refers to the image region where the character subject is located in the historically generated image. The extraction of the character image sub-region can be achieved by using a deep learning model to recognize the character image in the image and generate its corresponding bounding box or mask, so as to accurately delineate the pixel coordinates of the character at the pixel level or region level.

[0061] The reason for performing object recognition on the character image is to clarify the user's data on the position and relative size of the character image in the generated image. In the process of generating an image based on text, users often do not describe the size of the character image, but only realize that the size or position of the character image does not meet the expectations after the image is generated, and then need to adjust the prompt words or model parameters multiple times to optimize the image output. Compared with other optional reference image-guided element types, the size of the character image is a fundamental factor that must be considered in each generation, rather than an optional expression component of a certain type. Therefore, separating the size of the character image from the reference element type and explicitly modeling it through recognition and area analysis in historical images can better fit the actual usage behavior and generation logic of users, and help provide accurate size guidance in the early stage of generation, reducing the iteration cost.

[0062] In step S303 of some embodiments, the user preference can be predicted according to the position or area ratio of the character image sub-region in the historically generated image.

[0063] Steps S301 to S303 shown in the embodiments of the present application, by accurately extracting the historical generated image, obtain the image regions related to the missing element types, and combine the spatial distribution information of the character image in the image to construct the preference characteristics of the user in terms of visual composition and semantic expression. By using the image decomposition technology to extract the sub-regions of the image composition, the image content related to the elements not reflected in the current text description can be obtained specifically, for personalizing and complementing the visual dimensions not mentioned by the user. By using the character recognition to extract the sub-regions of the character image, the layout information such as the position and size of the character image in the image can be modeled. Finally, through the joint feature analysis of these two types of regions, the specific expression preferences of the user for the missing elements in the image composition can be inferred. The method of this embodiment can automatically supplement the element characteristics guided by the user's preferences through refined image decomposition and region analysis, so as to improve the visual performance of the finally generated image in terms of style continuity and visual fit.

[0064] Please refer to Figure 4 , in this embodiment, only the prediction of the user's preferences based on the area ratio of the character sub-region in the historical generated image is described, but it does not mean that the generation basis of the user's preference characteristics is limited to only the area ratio of the character sub-region.

[0065] Specifically, step S303 may include but is not limited to steps S401 to S403: Step S401, perform area preference analysis based on the character sub-region and the historical generated image to obtain the preference characteristics of the proportion of the character region.

[0066] Step S402, perform image semantic recognition on the sub-regions of the image composition to obtain the semantic characteristics of the elements.

[0067] Step S403, perform feature correlation analysis based on the semantic characteristics of the elements and the preference characteristics of the proportion of the character region to obtain the user's preference characteristics.

[0068] In step S401 of some embodiments, the human figure sub-region and its recognition frame are detected through a target detection method. Then, the pixel area of the task image sub-region is determined based on the recognition frame, and calculations are performed according to the resolution (at the pixel level) of the historically generated image, so as to obtain the area ratio of the human figure sub-region in the historically generated image. The same operations can be separately performed on multiple recent historically generated images to obtain multiple area ratio data, and based on the statistical characteristics (such as mean, variance, distribution interval, or clustering center) of these data, the preference characteristics of the human figure area ratio are obtained. That is, the preference characteristics of the human figure area ratio refer to the structured feature data formed through statistical analysis of the area ratio of the human figure area in the historically generated image in the overall image, which can represent the preference pattern in terms of the composition of the human figure size.

[0069] In step S402 of some embodiments, it can be performed through image classification, image recognition, or a text-image alignment model. For example, the CLIP model is used to perform semantic mapping on the image region, extract the semantic vector corresponding to the visual content, or convert the recognition result into structured semantic features.

[0070] In step S403 of some embodiments, feature correlation analysis can be achieved through a feature fusion method. For example, vector splicing or an attention mechanism is adopted to integrate the semantic features and the region ratio features into a joint representation, and the co-occurrence relationship between the two is analyzed. Exemplarily, when it is recognized that there is a background description related to the indoor environment in the current image description text, it is often preferred to set the human figure subject to a large area ratio, while in the outdoor natural background, a smaller human figure ratio is preferred.

[0071] Steps S401 to S403 illustrated in the embodiments of the present application can effectively mine the personalized preferences of users in image composition and visual expression by deeply analyzing the relationship between the human figure area in the historically generated image and other image elements, thereby providing more targeted guiding features for subsequent image generation. By performing statistical modeling on the area ratio of the human figure sub-region, the preference tendency of users for the composition of the human figure size can be quantified, and the key information often ignored in the text description can be supplemented. At the same time, combined with the recognition of the semantic features of the image composition sub-region, the system can understand the visual meaning and expression style of various composition elements in the image, and further perform fusion analysis on these semantic features and the spatial preferences of users in the human figure area, so as to establish a preference association pattern between the image structure and the semantic content. Generally speaking, this process not only improves the refinement degree and multi-dimensional expression ability of user preference modeling, but also significantly enhances the complementation and reasoning ability of the image generation system in the absence of a complete text description, helping the generated results to be closer to the aesthetic habits and composition expectations of users, so as to improve the accuracy of the reference human figure images obtained after searching. In step S204 of some embodiments, the user preference features obtained in the foregoing steps are concatenated with the semantic features of the description text containing multiple image element types to form new semantic features, that is, the current semantic features. This concatenation process can be vector-level concatenation, weighted fusion, or information integration through an attention mechanism to construct a comprehensive semantic representation that not only retains the user's clear description intention and complete description dimensions but also integrates the user's historical preference features.

[0072] Steps S201 to S204 illustrated in the embodiments of the present application first identify the types of image composition elements, determine the image composition elements not covered in the current text, and then determine the missing element types therefrom. Subsequently, by analyzing the image regions corresponding to the missing element types in the user's historical generated images, the preference expression forms of the user on these missing element types are extracted, and these preference features are fused with the current text semantic features to form a more complete and more in line with the user's true intention current semantic feature. The method of this embodiment not only enhances the semantic integrity of text-driven image search but also introduces the ability to complement the user's potential needs with historical behavior data, and can still achieve high-quality image generation guidance when facing incomplete or implicit input.

[0073] In step S104 of some embodiments, the reference character image refers to an existing image in the image database that has similar semantic or visual features to the current image description text. The reference character image can include images in a preset image material library and historical generated character images.

[0074] The process of image screening can be based on a text-image alignment model (such as CLIP), projecting the image semantics embedding of all images in the preset database and the current semantic features into the same vector space. Next, a similarity calculation is performed between the current semantic features and the image semantics embedding vectors. Usually, cosine similarity is used to measure the closeness of the two in the semantic space, and the images are sorted according to the similarity scores, and the image that best matches the current description intention is selected as the reference character image.

[0075] The reference image features refer to the feature encoding of the reference character image to extract its multi-level structured expression in the visual dimension. The reference character image can be feature encoded by a deep convolutional neural network (CNN).

[0076] In step S105 of some embodiments, the guided fusion features refer to the features obtained by fusing the reference image features and the description text semantic features. To effectively integrate the feature information of these two modalities, methods such as weighted summation, vector concatenation, and attention mechanism can be used for feature fusion in the fusion layer.

[0077] In step S106 of some embodiments, generating an image based on the guidance fusion feature includes: generating an image of the guidance fusion feature through a preset target character image generation model. It can be understood that in the embodiment of the present application, a text encoder, a fusion layer, and an image generator (target character image generation model) are provided in the model structure. The text encoder encodes the current image description text input by the user, and then obtains the semantic features of the description text. According to the semantic features of the description text and the historical generated image, an image search is performed to obtain a reference character image. In other embodiments, an audio search is also performed according to the semantic features of the description text and the historical generated image to obtain a reference audio file, and its detailed process will be further described in the embodiments of steps S601 to S604 below. Subsequently, the semantic features of the description text, the reference character image, and the reference audio file are cross-modally fused in the fusion layer, and the fused features are input into the target character image generation model, and finally the generated character image is output.

[0078] Specifically, the target character image generation model is trained by the following training method.

[0079] Please refer to Figure 5 , in some embodiments, the training method of the target character image generation model includes but is not limited to steps S501 to S507: Step S501, obtaining a sample character image description text.

[0080] Step S502, performing an image search according to the sample character image description text and the sample generated image to obtain a guidance image.

[0081] Step S503, performing feature fusion according to the sample character image description text and the guidance image to obtain a sample fusion feature.

[0082] Step S504, generating an initial character image by generating an image of the sample fusion feature through a preset initial image generation model.

[0083] Step S505, performing a quality assessment on the initial character image to obtain a character image quality score.

[0084] Step S506, constructing target loss data according to the character image quality score, the sample character image, and the initial character image.

[0085] Step S507, adjusting the parameters of the initial image generation model according to the target loss data to obtain the target character image generation model.

[0086] In step S501 of some embodiments, the sample character image description text is used to describe a preset sample character image. The sample character image refers to a real image sample used to guide the image generation model to learn the target image representation form during the training process. It is usually collected manually, professionally drawn, or extracted from an existing image database, and is the standard reference for the image generation model in the training stage on "what kind of image should be generated". The sample character image description text is a linguistic expression of the image content, describing information such as the identity, actions, clothing, expressions, and scene background of the characters in the image in natural language.

[0087] In step S502 of some embodiments, the sample generated image refers to an image sample used to simulate the "historically generated image" during the training process, aiming to restore the image results that the user has generated during the inference stage. The guiding image refers to a reference image retrieved from a preset image library through an image search operation based on the sample character image description text and the sample generated image, which has a high degree of relevance in semantics and style to the description content.

[0088] The implementation process of the steps in this embodiment is the same as that of steps S101 to S104, so it will not be elaborated here.

[0089] In step S503 of some embodiments, the sample fusion feature refers to a multi-modal semantic representation generated by performing feature extraction and fusion operations on the sample character image description text and the guiding image, and is used as the input of the image generation model during the training stage.

[0090] Perform text feature encoding on the sample character image description text to obtain the corresponding text features. Perform image feature encoding on the guiding image to obtain the corresponding image features, and fuse the image features with the aforementioned text features to obtain the sample fusion features.

[0091] In step S504 of some embodiments, the initial image generation model refers to an image generation model preset at the initial stage of training, which is used to generate a corresponding character image according to the input sample fusion features. The initial character image refers to the image result generated by the initial image generation model according to the sample fusion features.

[0092] In step S505 of some embodiments, in this embodiment, the initial character image will be scored on three evaluation dimensions, namely rationality, content richness, and user satisfaction, and the order of scoring can be adjusted. Exemplarily, first, the rationality of the initial character image is evaluated to obtain a rationality sub-score. Rationality refers to the correctness of the initial character image in terms of structural composition, semantic alignment, and visual logic, mainly evaluating whether there are structural errors, semantic deviations, or cultural violations in the image. For example, whether the generated task image has normal human proportions, whether the posture is natural, and whether it semantically conforms to the input description text. A pre-trained classifier model can be used to determine whether the initial character image is compliant. The rationality sub-score is a quantitative expression of the initial character image in the rationality dimension. In the embodiments of the present application, its value range is [0, 1].

[0093] Next, the content richness of the initial character image is evaluated to obtain a richness sub-score. Among them, the richness evaluates the number of visual elements contained in the generated image, its sense of hierarchy, and the degree of detail performance, etc., aiming to measure whether the image is sufficiently substantial. For example, whether the image has a complete background, and the richness of the image in terms of texture, color matching, and structural complexity. The implementation method can include using an image complexity evaluation algorithm (such as edge density, color distribution entropy), an object detection method (for detecting the number of semantic elements appearing in the image), or a perceptual feature extraction network (such as using a VGG network to count the responses of features at different levels) to score. The richness sub-score is a numerical representation of the evaluation result of the richness of the image content in terms of the number of elements, hierarchical structure, and detail expression. In the embodiments of the present application, its value range is [0, 1].

[0094] Then, the user satisfaction of the initial character image is evaluated to obtain a satisfaction sub-score. During the training process, real user interaction data (such as clicks, saves, selections) or a preference feature model summarized from historical generated images can be used to evaluate the degree of match between the current initial image and user preferences for user satisfaction. The satisfaction sub-score is a quantitative index generated based on the evaluation result of the overall adaptation degree of the current image according to user potential preferences or generation expectations. In the embodiments of the present application, its value range is [0, 1].

[0095] Finally, numerical calculations are performed based on the rationality sub-score, richness sub-score, and satisfaction sub-score to obtain the character image quality score. In this embodiment, the sub-scores of the three dimensions are weighted and calculated, and the weight data of the sub-scores of each dimension can be set according to actual needs. Exemplarily, the weight parameter of the rationality sub-score can be 0.5, the weight parameter of the richness sub-score can be 0.2, and the weight parameter of the satisfaction sub-score can be 0.3. The data obtained after weighted calculation of the sub-scores of multiple dimensions is the character image quality score.

[0096] In step S506 of some embodiments, this embodiment is trained in combination with the adversarial training method of the generative adversarial network (GAN). The initial character image is input into the discriminator to let it determine whether it is a real image (i.e., the sample character image). The target loss data of this embodiment can be shown as the following analytical formula: (1), Where, represents the target loss data, represents the gradient of the traditional GAN loss term of the image generator (initial image generation model) with respect to the parameter θG, specifically expressed as: . Where z represents the random noise vector, T represents the input sample fusion feature, represents the initial character image. represents the output of the discriminator, indicating the probability that the discriminator believes the initial character image is real. p z represents the random noise the probability distribution of z. Ez ~ pz[] represents the expectation of the probability distribution p of the noise variable z z of.

[0097] represents the gradient of the character image quality score with respect to the generator parameter θG. λ is a balancing coefficient used to adjust the relative importance of the adversarial loss and the reinforcement learning reward in training, and takes a positive value, such as 0.1, 1, 5, or 10, and its specific value can be adjusted according to actual needs.

[0098] In step S507 of some embodiments, when the target loss data drops to a preset threshold, the training ends, and finally the target character image generation model is obtained.

[0099] Steps S501 to S507 shown in the embodiments of the present application establish a training basis between semantics and vision by introducing sample character image and its corresponding descriptive text. In the feature fusion stage, through the deep integration of text and image features, precise guidance is provided for image generation. With the help of a multi-dimensional quality evaluation mechanism, quantitative feedback on the image quality is carried out from aspects such as rationality, content richness, and user satisfaction, and target loss data for supervising model optimization is constructed. Training is carried out based on this target loss data, so that the finally obtained target character image generation model has stronger generalization ability and generation effect in aspects such as semantic consistency, visual expressiveness, and cultural adaptability.

[0100] It should be noted that in addition to obtaining the guided fusion feature according to the semantic features of the descriptive text and the reference image features, the embodiments of the present application can also introduce audio features for fusion, so as to obtain a guided fusion feature that integrates more modal data.

[0101] Please refer to Figure 6 , in some embodiments, step S106 may further include but is not limited to steps S601 to S604: Step S601, perform audio screening according to the current semantic features to obtain a reference audio file.

[0102] Step S602, perform audio feature encoding on the reference audio file to obtain reference audio features.

[0103] Step S603, perform feature fusion according to the reference audio features and the guided fusion features to obtain target fusion features.

[0104] Step S604, perform image generation based on the target fusion features.

[0105] In step S601 of some embodiments, the specific implementation process of this step is similar to the missing element completion mechanism adopted in steps S201 to S204 and step S104. First, determine the type of the image element to which the extracted image constituent element belongs, and identify the part not covered by the current text from multiple preset reference element types as the missing element type. Subsequently, perform encoding processing on the image area corresponding to the missing element type in the historically generated image, and extract the user preference features contained in this area. Then, splice the user preference features with the semantic features of the description text. Based on the spliced semantic features, map the audio in the preset audio material library to a unified semantic space, calculate the semantic similarity between the two, and take the audio with the highest semantic similarity as the reference audio file. In the scenario of generating relevant character images for online medical consultations, the reference audio file can include recordings of doctors' disease popular science explanations, voice simulations of common disease consultation scenarios, voice content in medical education videos, or voice data of pre-consultation by a medical guide robot. For another example, in the scenario of generating relevant character images for financial product sales, the reference audio file can include recordings of financial advisors introducing financial products, voice conversations between insurance sales representatives and customers, host explanation audio of financial product promotion meetings, or voice segments of lecturers in investment courses.

[0106] In step S602 of some embodiments, the reference audio file can be processed based on a pre-trained audio encoder to extract audio spectral features, rhythm features, timbre textures, etc., and further encode these audio information into high-level semantic vectors, that is, reference audio features. Thus, it can express the emotional atmosphere contained in the audio, providing perceptually consistent sound guidance information for image generation.

[0107] In step S603 of some embodiments, the target fusion feature refers to a multi-modal feature expression that fuses the three-modal data information of the current image description text, the reference character image, and the reference audio file. This fusion process can be achieved by means such as weighted summation, feature splicing, attention mechanism, or multi-modal fusion network.

[0108] In step S604 of some embodiments, the image generation process will not only be guided by the text semantic content and the reference image style, but also further fuse the emotional or scene information provided by the audio, thereby realizing multi-modal driven image generation. So that the finally generated image is more in line with the description input by the user in terms of content structure, visual style, and overall atmosphere.

[0109] Steps S601 to S604 illustrated in the embodiments of the present application retrieve reference audio related to the generation intention from the audio material library through a semantic matching mechanism, and perform semantic encoding on it, so that the emotional atmosphere, rhythm, or scene perception in the sound can be understood by the model. By further fusing the audio features with the guidance fusion features, a multi-modal fusion feature including the semantics of the current image description text, the reference image features, and the audio features is formed, thereby providing richer and more comprehensive guiding conditions for the image generation model.

[0110] It should be noted that this embodiment can provide a fixed template for the user to select. When the user determines to use the template, the relevant feature vectors can be directly determined, and the feature vectors of the template are fused with the features obtained by fusing other modal data again. The fusion result is input into the target task image generation model for image generation.

[0111] Please refer to Figure 7 , the embodiments of the present application also provide a device for generating a character image, which can implement the above method for generating a character image. The device includes: An acquisition module 701, configured to acquire the current image description text, and perform semantic feature encoding on the current image description text to obtain the semantic features of the description text.

[0112] An image element recognition module 702, configured to perform image element recognition on the semantic features of the description text to obtain the image composition elements.

[0113] A semantic feature update module 703, configured to acquire the historical generated image, and update the semantic features of the description text according to the historical generated image and the image composition elements to obtain the current semantic features.

[0114] An image screening module 704, configured to perform image screening according to the current semantic features to obtain a reference character image, and perform image feature encoding on the reference character image to obtain reference image features.

[0115] A feature fusion module 705, configured to perform feature fusion according to the semantic features of the description text and the reference image features to obtain guidance fusion features.

[0116] An image generation module 706, configured to generate an image based on the guidance fusion features.

[0117] The specific implementation manner of the device for generating a character image is basically the same as the specific embodiments of the above method for generating a character image, and will not be elaborated here.

[0118] The embodiments of the present application also provide an electronic device, which includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the above-mentioned method for generating a character image is implemented. The electronic device can be any intelligent terminal including a tablet computer, an in-vehicle computer, etc.

[0119] Please refer to Figure 8 , Figure 8 which schematically shows the hardware structure of an electronic device in another embodiment. The electronic device includes: A processor 801, which can be implemented in ways such as a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided by the embodiments of the present application; A memory 802, which can be implemented in forms such as a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM). The memory 802 can store an operating system and other application programs. When implementing the technical solutions provided by the embodiments of this specification through software or firmware, the relevant program codes are stored in the memory 802, and the processor 801 is used to call and execute the method for generating a character image in the embodiments of the present application; An input / output interface 803, which is used to implement information input and output; A communication interface 804, which is used to implement communication interaction between this device and other devices. Communication can be achieved through wired means (such as USB, network cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.); A bus 805, which transmits information between various components of the device (such as the processor 801, the memory 802, the input / output interface 803, and the communication interface 804); Among them, the processor 801, the memory 802, the input / output interface 803, and the communication interface 804 are communicatively connected to each other inside the device through the bus 805.

[0120] The embodiments of the present application also provide a computer-readable storage medium, which stores a computer program, and when the computer program is executed by a processor, the above-mentioned method for generating a character image is implemented.

[0121] As a non-transitory computer-readable storage medium, the memory can be used to store non-transitory software programs and non-transitory computer-executable programs. In addition, the memory may include high-speed random access memory, and may also include non-transitory memory, such as at least one magnetic disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some embodiments, the memory optionally includes a memory remotely disposed relative to the processor, and these remote memories can be connected to the processor through a network. Examples of the above networks include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0122] The method for generating a character image, the device for generating a character image, the electronic device, and the storage medium provided by the embodiments of the present application encode semantic features of the current image description text input by the user, and further identify the image composition elements, so as to deeply understand the user's description at the semantic level and improve the parsing ability of the description intention. Secondly, the historical generated image is introduced, and the semantic features of the description text are updated by combining the semantic features of the description text and the image elements, so that the description text can approach the user's historical preference to obtain the current semantic features. Then, the reference character image is screened according to the current semantic features to obtain the reference character image, which effectively guides the generation process to approach the user's preference, thereby improving the personalized matching degree of the generated image. Then, the reference character image is encoded with image features to obtain the reference image features, and the features are fused with the description semantic features and the reference image features to generate the guiding fusion features. And based on this, the target image is generated, which helps to establish a closer connection between semantic understanding and visual presentation, thereby enhancing the accuracy and expressiveness of the generated image. Therefore, the method of the embodiments of the present application effectively solves the problem that the generated result is not appropriate due to the large difference in user descriptions, realizes the accurate restoration of the character image, and improves the accuracy of character image generation.

[0123] The embodiments described in the embodiments of the present application are for more clearly illustrating the technical solutions of the embodiments of the present application, and do not constitute a limitation on the technical solutions provided by the embodiments of the present application. Those skilled in the art can know that with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of the present application are also applicable to similar technical problems.

[0124] Those skilled in the art can understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of the present application, and may include more or fewer steps than shown in the figures, or combine certain steps, or different steps.

[0125] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the objectives of the solutions in this embodiment.

[0126] Those of ordinary skill in the art can understand that all or some of the steps in the methods disclosed above, and the functional modules / units in systems and devices, can be implemented as software, firmware, hardware, and appropriate combinations thereof.

[0127] As used in the specification of this application and the above drawings, the terms "first", "second", "third", "fourth", etc. (if any) are used to distinguish similar objects and do not necessarily describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances so that the embodiments of this application described here can be implemented in an order other than those illustrated or described here. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that comprises a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.

[0128] It should be understood that in this application, "at least one (item)" means one or more, and "a plurality" means two or more. "And / or" is used to describe the association relationship of associated objects and indicates that there can be three relationships. For example, "A and / or B" can mean: only A exists, only B exists, and both A and B exist at the same time. Here, A and B can be singular or plural. The character " / " generally indicates an "or" relationship between the associated objects before and after. "At least one (one)" or similar expressions below refer to any combination of these items, including any combination of single items (ones) or plural items (ones). For example, at least one (one) of a, b, or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.

[0129] In several embodiments provided by the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the above-mentioned units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of the devices or units can be in electrical, mechanical or other forms.

[0130] The units described above as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0131] In addition, in each embodiment of the present application, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.

[0132] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes multiple instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods in each embodiment of the present application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM), random access memories (RAM), magnetic disks, or optical discs that can store programs.

[0133] The preferred embodiments of the embodiments of the present application have been described above with reference to the accompanying drawings. However, this does not limit the scope of the rights of the embodiments of the present application. Any modifications, equivalent replacements, and improvements made by those skilled in the art without departing from the scope and essence of the embodiments of the present application shall be within the scope of the rights of the embodiments of the present application.

Claims

1. A method for generating a character image, characterized in that, The method includes: Obtain the current image description text, and perform semantic feature encoding on the current image description text to obtain the semantic features of the description text; Perform image element recognition on the semantic features of the description text to obtain image composition elements; Obtain the historical generated image, and update the semantic features of the description text according to the historical generated image and the image composition elements to obtain the current semantic features; Perform image screening according to the current semantic features to obtain a reference character image, and perform image feature encoding on the reference character image to obtain reference image features; Perform feature fusion according to the semantic features of the description text and the reference image features to obtain a guidance fusion feature; Generate an image based on the guidance fusion feature.

2. The method according to claim 1, characterized in that, The step of updating the semantic features of the description text according to the historical generated image and the image composition elements to obtain the current semantic features includes: Determine the type of image element based on the image composition elements; Select a reference element type different from the image element type from multiple preset reference element types to obtain a missing element type; Perform user preference analysis on the historical generated image to obtain user preference features corresponding to the missing element type; Perform feature splicing according to the user preference features and the semantic features of the description text to obtain the current semantic features.

3. The method according to claim 2, wherein The step of performing user preference analysis on the historical generated image to obtain user preference features corresponding to the missing element type includes: Decompose the historical generated image according to the missing element type to obtain an image composition sub-region corresponding to the missing element type; Perform character recognition on the historical generated image to obtain a character image sub-region; Perform image feature analysis according to the image composition sub-region and the character image sub-region to obtain the user preference features.

4. The method according to claim 3, wherein The step of performing image feature analysis according to the image composition sub-region and the character image sub-region to obtain the user preference features includes: Perform area preference analysis according to the character image sub-region and the historical generated image to obtain a character area ratio preference feature; Perform image semantic recognition on the image composition sub-region to obtain element semantic features; Perform feature correlation analysis according to the element semantic features and the character area ratio preference feature to obtain the user preference features.

5. The method according to claim 1, characterized in that, The step of generating an image based on the guidance fusion feature includes: Perform audio screening according to the current semantic features to obtain a reference audio file; Perform audio feature encoding on the reference audio file to obtain reference audio features; Perform feature fusion according to the reference audio features and the guidance fusion feature to obtain a target fusion feature; Generate an image based on the target fusion feature.

6. The method according to any one of claims 1 to 5, characterized in that, The step of generating an image based on the guidance fusion feature includes: generating an image of the guidance fusion feature through a preset target character image generation model, where the target character image generation model is trained by the following training method: Obtain a sample character image description text; wherein, the sample character image description text is used to describe a preset sample character image; Perform image screening based on the sample character image description text and the sample generated image to obtain a guiding image; Perform feature fusion based on the sample character image description text and the guiding image to obtain a sample fusion feature; Generate an initial character image by generating an image of the preset initial image model for the sample fusion feature; Perform quality assessment on the initial character image to obtain a character image quality score; Construct target loss data based on the character image quality score, the sample character image, and the initial character image; Adjust the parameters of the initial image generation model according to the target loss data to obtain the target character image generation model.

7. The method according to claim 6, characterized in that, The performing quality assessment on the initial character image to obtain a character image quality score includes: Perform rationality assessment on the initial character image to obtain a rationality sub-score; Perform content richness assessment on the initial character image to obtain a richness sub-score; Perform user satisfaction assessment on the initial character image to obtain a satisfaction sub-score; Perform numerical calculation based on the rationality sub-score, the richness sub-score, and the satisfaction sub-score to obtain the character image quality score.

8. An apparatus for generating a character image, characterized in that, The device includes: An acquisition module, configured to acquire a current image description text and perform semantic feature encoding on the current image description text to obtain a description text semantic feature; An image element recognition module, configured to perform image element recognition on the description text semantic feature to obtain image composition elements; A semantic feature update module, configured to acquire a historical generated image and update the description text semantic feature according to the historical generated image and the image composition elements to obtain a current semantic feature; An image screening module, configured to perform image screening according to the current semantic feature to obtain a reference character image and perform image feature encoding on the reference character image to obtain a reference image feature; A feature fusion module, configured to perform feature fusion according to the description text semantic feature and the reference image feature to obtain a guiding fusion feature; An image generation module, configured to generate an image based on the guiding fusion feature.

9. An electronic device, characterized in that, The electronic device includes a memory and a processor, the memory stores a computer program, and when the processor executes the computer program, the method according to any one of claims 1 to 7 is implemented.

10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, the method according to any one of claims 1 to 7 is implemented.

Citation Information

Cited By

  • Virtual image generation method based on machine learning

    CN120726194A

  • A method for generating a virtual avatar based on machine learning

    CN120726194B

  • Intelligent material recommendation method and system based on image communication

    CN120856940A