Text and Image Generation Method, Apparatus, Device, Storage Medium, and Program Product

By using the image annotation model to automatically label image materials and use the design interactive interface, the problems of low image search efficiency and poor image annotation efficiency in image design are solved, and efficient image design and automatic text generation are achieved.

CN118820503BActive Publication Date: 2025-05-30QINGDAO HAIGAO DESIGN & MANUFACTURING CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411275018.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-09-12
Publication Date
2025-05-30
Estimated Expiration
2044-09-12

AI Technical Summary

Technical Problem

The prior art has problems in image design with low image search efficiency, cumbersome operation, and poor image labeling efficiency and accuracy.

Method used

By automatically labeling image materials with a pre-trained image annotation model, descriptive information is generated, and the integrated design of image search and application is realized through the design of interactive interfaces and correlation with the database.

Benefits of technology

It improves the efficiency and accuracy of image annotation, simplifies user operations, improves the efficiency of image design, and assists users in quickly generating design documents by automatically generating text.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118820503B_ABST
    Figure CN118820503B_ABST
Patent Text Reader

Abstract

The present application provides a method, apparatus, device, storage medium, and program product for generating graphics and texts. The method for generating graphics and texts includes: via an image annotation model, annotating an image material under the guidance of a prompt word corresponding to the image material to obtain description information of the image material, and storing the image material and its description information in a database; displaying a design interaction interface; after detecting the input search information, determining the type hit by the search information; based on the matching result of the search information and the description information of the image materials of the hit type stored in the database, searching for alternative image materials from at least one type of image materials stored in the database; in response to an editing operation of the user on at least one alternative image material, generating a target image, and generating text corresponding to the target image based on the description information of at least one alternative image material. The integrated design of image search and application is realized, and the user operation is simplified.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image processing technologies, and in particular, to a method, apparatus, device, storage medium, and program product for generating graphics and text. Background Art

[0002] During the process of product iteration, a large number of digital assets are generated, including pictures related to product design, reports on product analysis, etc. How to fully release the value of digital assets and improve design efficiency is an urgent problem to be solved.

[0003] In order to achieve the sharing and reuse of digital assets, digital assets often need to be stored in a database and the fast data retrieval and query functions of the database are used to find the data of interest. For designers, when facing image design tasks such as new product design and promotional picture drawing, they can effectively use the image materials stored in the database for design. However, the existing databases only provide the function of image search. Designers need to download the images searched through the relevant web pages of the database to the local and then upload them to the design software for design, which is inefficient and cumbersome in operation. At the same time, when storing images, the existing databases often add labels to the images in the way of manual labeling, and both the labeling efficiency and accuracy are poor, resulting in unsatisfactory image search results. Summary of the Invention

[0004] This application provides a method, apparatus, device, storage medium, and program product for generating graphics and text. By using a model to label image materials to obtain descriptive information, the accuracy is high and the efficiency is high. At the same time, through the association between the design interaction interface and the database, the integrated design of image search and application is realized, reducing user operations and improving the efficiency of generating target images.

[0005] In a first aspect, this application provides a method for generating graphics and text, including:

[0006] For each of multiple image materials to be stored, determine a prompt corresponding to the image material according to the type of the image material;

[0007] Under the guidance of the prompt corresponding to the image material, label the image material via a pre-trained image annotation model to obtain descriptive information of the image material, and store the image material and its descriptive information in a database;

[0008] Display a design interaction interface, where the design interaction interface is used for a user to search for image materials and generate a target image and text corresponding to the target image according to the searched image materials;

[0009] After detecting search information input by the user in the design interaction interface, determine at least one type hit by the search information;

[0010] Search for alternative image materials from the at least one type of image materials stored in the database based on the matching result between the search information and the description information of the at least one type of image materials stored in the database;

[0011] In response to the user's editing operation on at least one of the alternative image materials, generate a target image, and generate text corresponding to the target image based on the description information of the at least one alternative image material.

[0012] Optionally, determine the prompt words corresponding to the image material according to the type of the image material, including:

[0013] Determine the type of the image material based on the objects included in the image material;

[0014] Determine the prompt words corresponding to the image material based on the corresponding relationship between the type of the image material and the prompt words, and the type of the image material.

[0015] Optionally, in the corresponding relationship, the prompt words corresponding to the furniture type of image materials include style prompt words, design key point prompt words, material prompt words, color prompt words, and function prompt words, and the prompt words corresponding to the scene type of image materials include scene prompt words, style prompt words, and object included prompt words.

[0016] Optionally, the design interaction interface includes a first type of search box and multiple second type of search boxes, the type hit by the first type of search box is empty, and the second type of search box hits at least one type;

[0017] Determine the at least one type hit by the search information, including:

[0018] If the search box for inputting the search information is a second type of search box, then determine the type hit by the search box for inputting the search information as the type hit by the search information;

[0019] If the search box for inputting the search information is a first type of search box, then determine the at least one type hit by the search information based on the keywords extracted from the search information.

[0020] Optionally, if the search information is an image, extract the keywords of the search information based on the objects and text recognized from the search information.

[0021] Optionally, after detecting the search information input by the user in the design interaction interface, determine the at least one type hit by the search information, including:

[0022] After detecting the search information input by the user in the design interaction interface, if the search information is text, determine whether the search information is a sentence;

[0023] If not, determine at least one type hit by the search information.

[0024] Optionally, determining whether the search information is a sentence includes:

[0025] Based on at least one of the length, syntactic structure, and semantic analysis result of the search information, determine whether the search information is a sentence.

[0026] Optionally, determining whether the search information is a sentence includes:

[0027] If the length of the search information is shorter than a preset length, or the search information includes multiple delimiters, determine that the search information is not a sentence.

[0028] Optionally, the method further includes:

[0029] For each image material to be stored, based on a pre-trained vectorization model, obtain the image vector of the image material, and store the image vector in the database;

[0030] After detecting the search information input by the user in the design interaction interface, if the search information is an image, based on a pre-trained vectorization model, obtain the image vector of the search information;

[0031] Based on the image vector of the search information and the image vectors stored in the database, search for the alternative image materials from the images stored in the database.

[0032] Optionally, the vectorization model is a multi-modal vectorization model. Based on a pre-trained vectorization model, obtaining the image vector of the image material includes:

[0033] Input the image material and its description information into a pre-trained multi-modal vectorization model to obtain the image vector of the image material.

[0034] Optionally, the method further includes:

[0035] For the text material to be stored, based on a pre-trained vectorization model, obtain the text vector of the text material, and store the text material and its text vector in the database;

[0036] After detecting the search information input by the user in the design interaction interface, if it is determined that the search information is a sentence, based on a pre-trained vectorization model, obtain the text vector of the search information;

[0037] Based on the text vector of the search information and the image vectors and / or text vectors stored in the database, obtain search results from the image materials and / or text materials stored in the database;

[0038] Display the search results on the design interaction interface.

[0039] Optionally, based on the text vector of the search information and the image vectors and / or text vectors stored in the database, obtaining search results from the image materials and / or text materials stored in the database includes:

[0040] Identify the search intention of the search information;

[0041] If the search intention of the search information is an image search intention, based on the text vector of the search information and the image vectors stored in the database, obtain search results from the image materials stored in the database;

[0042] If the search intention of the search information is a text search intention, based on the text vector of the search information and the text vectors stored in the database, obtain search results from the text materials stored in the database;

[0043] If the search intention of the search information is a graphic and text search intention, based on the text vector of the search information and the image vectors and text vectors stored in the database, obtain search results from the image materials and text materials stored in the database.

[0044] Optionally, identifying the search intention of the search information includes:

[0045] If the end of the search information includes a preset keyword, determine that the search intention of the search information is an image search intention;

[0046] Wherein, the preset keyword includes a word associated with a graph.

[0047] Optionally, identifying the search intention of the search information includes:

[0048] Input the search information into a pre-trained search intention recognition model to determine a first search intention;

[0049] Based on the user's search history, determine the user's intention preference;

[0050] Based on the user's role, the user's intention preference, and the interface opened on the screen when inputting the search information, determine a second search intention;

[0051] Determine that the search intention is the union of the first search intention and the second search intention.

[0052] Optionally, the image annotation model includes an image encoder, a modality alignment module, and a large language model layer;

[0053] The image encoder is used to extract the encoded features of the input image;

[0054] The modality alignment module is used to project the encoded features into a preset text space to obtain an aligned feature vector;

[0055] The large language model layer is used to output the description information of the input image under the corresponding prompt word based on the aligned feature vector under the guidance of the prompt word corresponding to the input image.

[0056] In a second aspect, the present application provides a graphic and text generation device, including:

[0057] A prompt word determination module, configured to determine, for each of multiple image materials to be stored, the prompt word corresponding to the image material according to the type of the image material;

[0058] An image material annotation module, configured to annotate the image material under the guidance of the prompt word corresponding to the image material through a pre-trained image annotation model, obtain the description information of the image material, and store the image material and its description information in a database;

[0059] An interface display module, configured to display a design interaction interface, where the design interaction interface is used for a user to search for image materials and generate a target image and the text corresponding to the target image according to the searched image materials;

[0060] A hit type determination module, configured to determine at least one type hit by the search information after detecting the search information input by the user in the design interaction interface;

[0061] An image material search module, configured to search for alternative image materials from the image materials of at least one type stored in the database based on the matching result between the search information and the description information of the image materials of at least one type stored in the database;

[0062] A graphic and text generation module, configured to generate a target image in response to an editing operation of the user on at least one of the alternative image materials, and generate the text corresponding to the target image based on the description information of the at least one alternative image material.

[0063] In a third aspect, the present application provides an electronic device, including: a processor, and a memory communicatively connected to the processor;

[0064] The memory stores computer execution instructions;

[0065] The processor executes the computer-executable instructions stored in the memory to implement the method provided in the first aspect of the present application.

[0066] In a fourth aspect, the present application provides a computer-readable storage medium storing computer-executable instructions, which are used to implement the method provided in the first aspect of the present application when executed by a processor.

[0067] In a fifth aspect, the present application provides a computer program product including a computer program, which implements the method provided in the first aspect of the present application when executed by a processor.

[0068] The graphic generation method, device, equipment, storage medium and program product provided by the present application realize automatic annotation of image materials using a model during the image material storage stage, improving the efficiency of image annotation. Moreover, the prompt words guiding the model to output description information are flexibly determined based on the type of image materials, improving the accuracy of the prompt words, and further improving the accuracy of the model in text standardization of image materials. During the search application stage, the search information input by the user is obtained by designing an interactive interface, and the image materials of the type in the search information command are used as the search scope. Within the search scope, based on the matching result of the search information and the description information of the image tree branches, the search result is obtained, that is, the search for alternative image materials is realized, so that the user can synthesize the required target image through the image materials shown in the search result, improving the efficiency of image design and simplifying the user operation. At the same time, the text corresponding to the target image can be automatically generated using the description information of the image materials. Through the text corresponding to the target image, the user can be assisted in quickly generating a design document, adding annotations to the target image, etc., further improving the convenience of image design. BRIEF DESCRIPTION OF THE DRAWINGS

[0069] The drawings here are incorporated into the specification and form a part of this specification, showing embodiments consistent with the present application, and are used together with the specification to explain the principles of the present application.

[0070] Figure 1 It is a schematic diagram of an application scenario provided by an embodiment of the present application;

[0071] Figure 2 It is a schematic flowchart of a graphic generation method provided by an embodiment of the present application;

[0072] Figure 3 It is a schematic flowchart of another graphic generation method provided by an embodiment of the present application;

[0073] Figure 4 It is a schematic structural diagram of an image annotation model provided by an embodiment of the present application;

[0074] Figure 5 A schematic diagram of image search by sentence provided by an embodiment of the present application;

[0075] Figure 6 A schematic diagram of four search methods provided by an embodiment of the present application;

[0076] Figure 7 A schematic structural diagram of a graphic and text generation device provided by an embodiment of the present application;

[0077] Figure 8 A schematic structural diagram of an electronic device provided by an embodiment of the present application.

[0078] Through the above-mentioned drawings, specific embodiments of the present application have been shown, and there will be more detailed descriptions hereinafter. These drawings and text descriptions are not intended to limit the scope of the concept of the present application in any way, but to illustrate the concept of the present application to those skilled in the art by referring to specific embodiments. Detailed Description of Specific Embodiments

[0079] Here, exemplary embodiments will be described in detail, and examples are shown in the drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present application. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present application as detailed in the appended claims.

[0080] First, some terms related to the present application are explained:

[0081] Prompt: An input to the model, which guides the model to perform a specific output or task through specific instructions or questions.

[0082] Multimodal Large Language Model (MLLM): Based on the large language model, it performs multimodal tasks, such as processing modal data such as images, texts, videos, and audios. In the present application, the input and output of the MLLM have different modalities, the input is an image, and the output is text (description information).

[0083] In the process of enterprise product iteration and R & D, a large number of digital assets will be generated, such as product design documents, design material images, etc. It is necessary to establish a centralized and efficient data management system for these digital assets. Store data assets, such as documents and images, in the database of the data management system. By adding appropriate indexes to the images, documents and other data stored in the database, the search for resources of interest is realized.

[0084] Figure 1 A schematic diagram of an application scenario provided by an embodiment of the present application. As Figure 1 shown, in a data management system, at the data source end, such as the terminals of enterprise designers, mobile terminals, servers, etc., new generated modal data such as images and texts can be uploaded to the database regularly or irregularly; when the text is stored in the database, the keywords of the text can be extracted as indexes; when the image is stored in the database, the image can be labeled with keywords by manual annotation. In the retrieval stage, the user needs to input one or more keywords, such as keyword 1, keyword 2, and keyword 3. Through keyword matching, the search results are recalled from the stored images and texts and sorted by similarity, such as Figure 1 the images img1 to img3, and the texts text1 and text2 in

[0085] The above image search method requires the user to clearly define the keywords of the image. In some scenarios, the user's search intention is relatively vague and the specific keywords cannot be determined, resulting in unsatisfactory search results. At the same time, the foregoing search method is relatively single and cannot meet the diverse needs of users, such as the need to search for images by image, the need to search for images using natural language description statements, etc.

[0086] When the user needs to perform graphic design based on the searched image, operations such as copying and downloading the searched image in the database are required, so as to load the searched image to the design platform for image design. The user's operations are cumbersome and inefficient.

[0087] In order to enrich the image search methods, improve the flexibility of search, and enhance the efficiency of graphic design, the present application provides a graphic generation method. During the image material storage stage, it realizes the automatic generation of description information for image materials based on a model, thereby supporting users to search for image materials in a more diverse way, such as searching for image materials with natural sentences. To improve the comprehensiveness and accuracy of the description information, it realizes the differential customization of prompt words corresponding to image materials based on the type of image materials, so that the model can label the required description information for the image materials under the guidance of the prompt words. During the search application stage, it obtains the search information input by the user, such as natural sentences, keywords, etc., through the design of an interactive interface. When the search information matches the type of image materials, it uses the image materials of the hit type as the search scope, and based on the matching result between the search information and the description information of the image materials, it obtains the search results, that is, alternative image materials. Users can directly operate the alternative image materials to generate the target image, without performing operations such as copying, downloading, and uploading of image materials, improving the efficiency of image design. At the same time, it also supports automatically labeling the text of the target image based on the description information of the operated image materials, and the labeled text can better assist users in tasks such as image annotation and design text generation.

[0088] The following will specifically describe the technical solutions of the present application and how the technical solutions of the present application solve the above technical problems with specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. The embodiments of the present application will be described below with reference to the accompanying drawings.

[0089] Figure 2 It is a schematic flowchart of a graphic generation method provided by an embodiment of the present application. This graphic generation method can be executed by an electronic device with corresponding data processing capabilities, such as a database system, the aforementioned data management system, etc., as Figure 2 shown, this data search method includes the following steps:

[0090] Step S201, for each of the multiple image materials to be stored, determine the prompt words corresponding to the image materials according to the type of the image materials.

[0091] The image materials can be images generated at historical times or images imported through external devices.

[0092] The type of the image materials can be determined based on one or more of the parameters such as the objects included in the image materials, the style of the image, and the scene of the image.

[0093] Multiple prompting words can be designed in advance, and an applicable type can be configured for each prompting word. After determining the type of the image material to be stored, based on the type of the image material, the prompting word corresponding to the image material can be screened from the pre-designed prompting words.

[0094] Exemplarily, for image materials of the furniture type, the corresponding prompting words can be used to prompt the image annotation model to output information such as the design points, materials, textures, and design styles of the furniture in the image materials; for image materials of the human type, the corresponding prompting words can be used to prompt the image annotation model to output information such as the roles, outfits, and genders of the people in the image materials.

[0095] Optionally, determining the prompting word corresponding to the image material according to the type of the image material includes:

[0096] Based on the objects included in the image material, determining the type of the image material; based on the correspondence between the type of the image material and the prompting words, and the type of the image material, determining the prompting word corresponding to the image material.

[0097] If the image material only contains one object, the type of the image material can be determined based on the object or the category to which the object belongs; if the image material contains multiple objects, the type of the image material can be determined based on the categories to which the multiple objects respectively belong.

[0098] Furthermore, the type of the image material can be determined based on the objects included in the image material and the scene recognition result of the image material.

[0099] Exemplarily, if the image material only includes a sofa, the type of the image material can be determined as a sofa or furniture; if the image material includes multiple objects such as a sofa, a coffee table, a person, and a plant, the type of the image material can be determined as an interior decoration type, and the type of the image material can be further subdivided according to the decoration style, such as modern minimalist style interior decoration, Nordic style interior decoration, etc.

[0100] Prompting words corresponding to various types can be configured in advance to obtain the correspondence between the type and the prompting words; after determining the type of the image material, the prompting word corresponding to the type of the image material is searched from this correspondence as the prompting word corresponding to the image material.

[0101] Using the objects included in the image material to classify the type of the image material has a simple classification logic and high accuracy; by using the pre-established correspondence to determine the prompting word corresponding to the image material of a certain type, the calculation complexity is low and the prompting word search efficiency is high.

[0102] Optionally, in the corresponding relationship, the prompt words corresponding to the image materials of furniture category include style prompt words, design key point prompt words, material prompt words, color prompt words and function prompt words, and the prompt words corresponding to the image materials of scene category include scene prompt words, style prompt words and included object prompt words.

[0103] For the image materials of furniture category, the style prompt words in the corresponding prompt words are used to prompt the image annotation model to output the design style of the furniture in the image material, such as modern minimalist style, Chinese style, neo-Chinese style, etc.; the design key point prompt words are used to prompt the image annotation model to output the design key points of the furniture in the image material, such as the main design elements; the material prompt words, color prompt words and function prompt words are respectively used to prompt the image annotation model to output the material, color and main functions of the furniture in the image material, and the main functions can include providing a comfortable / beautiful / convenient / safe rest / sleep / entertainment environment, storage function, etc.

[0104] For the image materials of scene category, the scene prompt words in the corresponding prompt words are used to prompt the image annotation model to output the specific scene in the image material, such as sports scene, sleep scene, etc.; the style prompt words are used to prompt the image annotation model to output the specific style of the scene in the image material, such as leisure and entertainment style, warm and comfortable style, etc.; the included object prompt words are used to prompt the image annotation model to output the main objects included in the image material and their related information, such as the color, material, state, etc. of the object.

[0105] Step S202, under the guidance of the prompt words corresponding to the image material, annotate the image material through a pre-trained image annotation model to obtain the description information of the image material, and store the image material and its description information in the database.

[0106] The image annotation model is a pre-trained multi-modal language model, such as a multi-modal large language model. The image annotation model takes the image material (the modality is image) and its corresponding prompt words as inputs, and takes the description information of the image material (the modality is text) as the output, and is used to perform text annotation on the input image material under the guidance of the prompt words and output the description information of the image material.

[0107] For each image material to be stored, after obtaining the description information of the image material output by the image annotation model, in the database, associate and store the image material and its description information to facilitate subsequent image material search.

[0108] Furthermore, the image vector of the image material can also be stored in the database. Any model for image vectorization can be used to vectorize the image material to obtain the image vector, and associate and store the image material, the image vector of the image material and the description information.

[0109] Furthermore, text materials can also be stored in the database.

[0110] In some embodiments, for the text materials to be stored, the keywords of the text materials can be extracted as the description information of the text materials, or description information such as the abstract or title of the text materials can be generated, and the text materials and their description information are stored in an associated manner.

[0111] In other embodiments, for the text materials to be stored, any text vectorization model can be used to vectorize the text materials to obtain text vectors, and the text materials and their text vectors are stored in an associated manner.

[0112] After obtaining the text vectors and description information of the text materials, the text materials, the text vectors of the text materials, and the description information can be stored in the database in an associated manner.

[0113] By introducing preset prompt words, the richness of the description information output by the model can be improved, and at the same time, the efficiency of model training can be improved and the generality of the model can be enhanced.

[0114] Step S203: Display the design interaction interface.

[0115] Among them, the design interaction interface is used for the user to search for image materials and generate a target image and the text corresponding to the target image based on the searched image materials.

[0116] The database system also provides a front-end interaction interface, including a design interaction interface. In the design interaction interface, the user can search for image materials and design the required image, that is, the target image, based on the search results.

[0117] The design interaction interface can include a search box, a canvas, and a series of controls. The user can enter search information in the search box.

[0118] Step S204: After detecting the search information input by the user in the design interaction interface, determine at least one type hit by the search information.

[0119] Among them, the search information can be text or an image. The text-type search information can be a natural sentence, one or more keywords, etc.

[0120] The user can enter search information through the search box of the design interaction interface.

[0121] The search information can be input through input devices such as a keyboard or a touch screen, or the search voice can be input through a microphone, and the search voice is converted into text through voice recognition, that is, the search information. Other ways can also be used to input the search information, and the present application does not limit this.

[0122] The modality of the search information can be image, sound, or text. If it is sound, it can be converted into text through speech recognition. Therefore, only the two modalities of image and text will be analyzed subsequently.

[0123] Exemplarily, during the training phase or fine-tuning phase of the model, the user can search for image materials, text materials, etc. of interest in the database to form training samples. The user can enter the search information in the search box, and through the data management system or database system, search for data such as image materials and text materials that match the search information from the database.

[0124] Exemplarily, when designing a new product, such as a smart home product, the designer can design the product diagram on the design interaction interface provided by the database system. To assist in the design of the new product, the user can search for historical image materials, trend analysis reports, etc. of the product in the database through this design interaction interface as a reference.

[0125] When the modality of the search information is an image, the search information is recorded as a search image. The type hit by the search image can be determined based on the type of the search image. For example, determining the type of the search image and the type to which the type of the search image belongs is the type hit by the search image. Suppose the type of the search image is a sofa, then the types hit by the search image can include two types: sofa and furniture.

[0126] When the modality of the search information is text, the type hit by the search information can be determined based on the type of the noun in the search information, and the type to which the type of the noun in the search information belongs can also be considered. For example, if the search information includes a washing machine, then the types hit by the search information can include two types: washing machine and household appliance.

[0127] In some embodiments, the design interaction interface can include multiple search boxes. If the search scope or the type hit by different search boxes is different, then the type hit by the search information can be determined based on the search box corresponding to the search information.

[0128] Optionally, the design interaction interface includes a first type of search box and multiple second type of search boxes. The type hit by the first type of search box is empty, and the second type of search box hits at least one type. Determining at least one type hit by the search information includes:

[0129] If the search box for inputting the search information is a second type of search box, then determine the type hit by the search box for inputting the search information as the type hit by the search information. If the search box for inputting the search information is a first type of search box, then determine at least one type hit by the search information based on the keywords extracted from the search information.

[0130] Optionally, if the search information is an image, keywords of the search information are extracted based on the objects and text identified from the search information.

[0131] Step S205: Search for alternative image materials from the at least one type of image materials stored in the database based on the matching result between the search information and the description information of the at least one type of image materials stored in the database.

[0132] The matching result is used to represent the matching degree or similarity between the search information and the description information of the image material. When the search information is text, the similarity between two texts can be measured by the intersection-over-union, cosine similarity, longest common subsequence, etc.

[0133] When the type hit by the search information is not empty, that is, when the search information hits at least one type, the image materials of the type hit by the search information stored in the database are used as the search scope, and based on the matching result between the description information of the image materials in this search scope and the search information, alternative image materials are searched from the image materials in this search scope to obtain the search result.

[0134] When the modality of the search information is an image, the description information of the search information can be obtained in a manner similar to obtaining the description information of the image material. That is, based on the type of the search image, the prompt word corresponding to the search image is determined, and under the guidance of the prompt word corresponding to the search image, the search image is labeled through a pre-trained image annotation model to obtain the description information of the search image.

[0135] The search for alternative image materials is realized through the matching result between the description information of the image materials in the search scope and the description information of the search image.

[0136] Further, when the search for alternative image materials is empty or no alternative image materials are found, the search scope can be expanded. For example, the full amount of image materials stored in the database, that is, all the stored image materials, are used as the search scope, and based on the matching result between the search information and the description information of the image materials, alternative image materials are searched from all the image materials stored in the database.

[0137] If no alternative image materials are found after the search scope is expanded, corresponding prompt information can be displayed to prompt the user that no relevant image materials are found.

[0138] Step S206: Respond to the editing operation of the user on at least one of the alternative image materials, generate a target image, and generate the text corresponding to the target image based on the description information of the at least one alternative image material.

[0139] The searched alternative image materials can be displayed in a container on the design interaction interface. The user can move one or more alternative image materials to the canvas of the design interaction interface by selecting, dragging, and other operations, and generate a target image by adjusting, drawing, and other processing the alternative image materials in the canvas.

[0140] Exemplarily, the target image may be an image synthesized from a plurality of candidate image materials, for example, an interior decoration rendering may be synthesized from a plurality of image materials of furniture, home appliances, etc.

[0141] After the target image is generated, the target image may be annotated with text based on description information of one or more candidate image materials used to generate the target image, and text corresponding to the target image may be obtained and displayed.

[0142] Specifically, the description information of one or more candidate image materials used to generate the target image may be integrated, including operations such as content extraction and redundancy removal, to generate text corresponding to the target image.

[0143] The image-text generation method provided in the present embodiment realizes automatic annotation of image materials by using a model in the image material storage stage, thereby improving the efficiency of image annotation, and guiding the prompt words for outputting the description information of the model to be flexibly determined based on the type of the image material, thereby improving the accuracy of the prompt words, and further improving the accuracy of the model in textual standardization of the image material; in the search application stage, the search information input by the user is obtained by designing an interactive interface, and the image material of the type in the search information command is used as the search scope. Within the search scope, the search results are obtained based on the matching results of the search information and the description information of the image branches, that is, the search for the alternative image materials is realized, so that the user can synthesize the desired target image through the image materials displayed in the search results, thereby improving the efficiency of image design and simplifying user operations; at the same time, the description information of the image material can be used to automatically generate the text corresponding to the target image, and the text corresponding to the target image can be used to assist the user in quickly generating design documents, adding annotations to the target image, etc., thereby further improving the convenience of image design.

[0144] Figure 3 A flow chart of another method for generating images and texts provided in an embodiment of the present application. In the data storage stage, this embodiment also provides steps for extracting and storing image vectors of image materials, and adds steps for extracting and storing text material vectors. At the same time, in the search application stage, multi-layer logical judgments based on the modality of the search information and whether it is a statement are implemented to flexibly determine the search logic scheme to enrich the image and text search method.

[0145] like Figure 3 As shown, the image and text generation method provided in this application may specifically include the following steps:

[0146] Step S301: For each image material to be stored, determine the prompt word corresponding to the image material according to the type of the image material.

[0147] Step S302: Under the guidance of the prompt word corresponding to the image material, label the image material through a pre-trained image annotation model to obtain the description information of the image material.

[0148] Step S303: Based on a pre-trained vectorization model, obtain the image vector of the image material.

[0149] In some embodiments, for distinction, the model that vectorizes the search information for images, including image materials and image modalities, can be denoted as the first vectorization model, while the model that vectorizes the search information for subsequent texts, including text materials and text modalities, can be denoted as the second vectorization model.

[0150] The image vector and description information of the image material can be obtained in parallel. Figure 3 Taking the case of serial operation as an example.

[0151] Step S304: Store the image material, its description information, and the image vector in the database.

[0152] For the convenience of retrieval or search, in the database, the image material, the image vector, and the description information will be stored in the database together, using the image vector and the description information as the index of the image material to facilitate image search.

[0153] Exemplarily, the database can be a vector database.

[0154] Exemplarily, the image annotation model can be a multi-modal large language model, such as RAM (Recognize Anything Model).

[0155] For the image materials to be stored in the database, that is, the image materials to be stored, input the image materials or the pre-processed image materials into the multi-modal large language model, and output multiple labels of the image materials and the description information of each label. The multiple labels output by the multi-modal large language model include at least one label under the prompt word corresponding to the image material. At the same time, the image materials also need to go through a vectorization model for vectorization to obtain image vectors.

[0156] Exemplarily, the image vector can be a 1×1024 vector.

[0157] Optionally, Figure 4 This is a schematic structural diagram of the image annotation model provided by the embodiments of the present application, such as Figure 4As shown in the figure, the image annotation model includes: an image encoder, a modality alignment module, and a large language model layer. The image encoder is used to extract the encoded features of the input image; the modality alignment module is used to project the encoded features into a preset text space to obtain an aligned feature vector; the large language model layer is used to output the description information of the input image under the corresponding prompt word based on the aligned feature vector under the guidance of the prompt word corresponding to the input image.

[0158] Since the modalities of the input and output of the image annotation model are different, after obtaining the encoded features, it is also necessary to map the encoded features to the space where the text is located, that is, the preset text space, through the modality alignment module to build a bridge between the two modalities of image and text. The preset text space is a text space that can be understood by the large language model.

[0159] Exemplarily, the modality alignment module can convert the encoded features of the image to be stored into Tokens (text tokens), and then the large language model layer outputs the description information of the image to be stored based on the Tokens.

[0160] The large language model layer is the core of the image annotation model and has powerful generalization and reasoning capabilities. It is used to analyze and process the output of the modality alignment module and output the description information of the image under the guidance of the corresponding prompt word.

[0161] The large language model layer can include a pre-trained large language model.

[0162] Correspondingly, under the guidance of multiple preset prompt words based on a pre-trained multi-modal large language model, the description information of the image to be stored is obtained, including:

[0163] Input the image into the pre-trained multi-modal large language model. Through the image encoder, extract the encoded features of the image to be stored; through the modality alignment module, project the encoded features into the preset text space to obtain an aligned feature vector; through the large language model, under the guidance of multiple preset prompt words, based on the aligned feature vector, output the description information of the image to be stored under each of the preset prompt words.

[0164] Optionally, the prompt words can include object prompt words, style prompt words, color prompt words, and scene prompt words; the object prompt words are used to prompt the large language model layer to output the object labels and their description information included in the image; the style prompt words are used to prompt the large language model layer to output the style labels and their description information of the image; the color prompt words are used to prompt the large language model layer to output the colors of the object labels; the scene prompt words are used to prompt the large language model layer to output the scene labels and their description information of the image.

[0165] The description information of the object label can be used to describe the attributes of the corresponding object, such as its state and behavior; the description information of the style label is used to describe the specific style presented by the image to be stored; the description information of the scene label can be used to describe the characteristics of the scene in the image to be stored and how to reflect the corresponding characteristics.

[0166] The color prompt can be specifically used to prompt the large language model to output the color or approximate color of the main object (such as a relatively large object) in the image.

[0167] After obtaining the description information of the image material, when vectorizing the image material, the information of both the image material itself and its description information can be comprehensively considered to achieve multi-modal vectorization of the image material, obtaining an image vector, so that the image vector of the image material includes the vector of the description information of the image material.

[0168] Optionally, the vectorization model is a multi-modal vectorization model. Based on the pre-trained vectorization model, the image vector of the image material is obtained, including:

[0169] Input the image material and its description information into the pre-trained multi-modal vectorization model to obtain the image vector of the image material.

[0170] The multi-modal vectorization model can realize the vectorization of information in multiple modalities, specifically the vectorization of information in two modalities, namely image and text. At the same time, through vector mapping, an image vector composed of two vectors in the same vector space is obtained.

[0171] Through the multi-modal vectorization model, the image material and its description information are respectively vectorized to obtain the image vector of the image material. This image vector consists of two parts, one part is the vector obtained by vectorizing the image material, and the other part is the vector obtained by vectorizing the description information of the image material.

[0172] Exemplarily, the image vector of the image material can include two vectors, namely vector V img and vector V txt , vector V img is the vector obtained by vectorizing the image material, vector V txt is the vector obtained by vectorizing the description information of the image material, and the dimensions of vector V img and vector V txt can be the same, such as both being 1×1024.

[0173] Obtaining the image vector through multi-modal information makes the image vector attached with the vector of the description information. When calculating the vector distance, the distance between the text vector of the search information and each vector of the image vector can be calculated respectively, improving the comprehensiveness of image search and thus the accuracy and hit rate of the search results.

[0174] Step S305: For the text material to be stored, based on the pre-trained vectorization model, obtain the text vector of the text material, and store the text material and its text vector in the database.

[0175] For the text material to be stored in the database, that is, the text material to be stored, the text material passes through the text vectorization model, that is, the second vectorization model, and the text vector of the text material is output and stored in the database together with the text material.

[0176] In some embodiments, the vectorization of image materials and text materials can be achieved through the same multi-modal vectorization model to obtain the image vector of the image material and the text vector of the text material.

[0177] In other embodiments, two single-modal vectorization models, that is, the first vectorization model and the second vectorization model, can be used to perform the vectorization of image materials and text materials respectively.

[0178] By repeating steps S301 to S305, the continuous storage of image materials and text materials can be realized, and the addition of database data can be achieved.

[0179] Step S306: Display the design interaction interface.

[0180] Step S307: Obtain the search information input by the user on the design interaction interface. If the search information is an image, execute step S308; if the search information is text, execute step S310.

[0181] Step S308: Based on the pre-trained vectorization model, obtain the image vector of the search information.

[0182] If the modality of the search information is an image, that is, the user searches in the way of searching by image, then use the pre-trained vectorization model such as the first vectorization model or the multi-modal vectorization model to vectorize the search information to obtain the image vector of the search information.

[0183] Before inputting the search information in the image modality, that is, the search image, into the vectorization model, the search information can be pre-processed first, such as noise reduction processing, color space conversion, size adjustment, etc. Then, the pre-processed search information is input into the vectorization model, and the image vector of the search information is output via the vectorization model, such as image encoding.

[0184] The vectorization model for vectorizing an image, such as the first vectorization model, can be an image feature extraction model. By extracting features in the image, such as texture, color, shape and other features, an image vector is obtained. The vectorization model for vectorizing an image can also be a Graph Neural Network (GNN) model, such as a graph autoencoder. Through convolution operations, an image vector is generated.

[0185] The training of the vectorization model for vectorizing an image can be carried out by using a large number of pre-collected images to form positive sample pairs and negative sample pairs, constructing a loss function, and continuously optimizing the parameters of the vectorization model by minimizing the loss function to achieve model training until the training end condition is met, such as the loss value of the loss function being less than a preset threshold. The loss function has a positive correlation with the distance between the image vectors in the positive sample pairs and an inverse correlation with the distance between the image vectors in the negative sample pairs.

[0186] When training the vectorization model, optimization can be carried out by using the stochastic gradient descent algorithm.

[0187] Step S309, based on the image vector of the search information and the image vectors stored in the database, search for the alternative image materials from the images stored in the database.

[0188] In the database establishment stage, for the image materials to be stored, it is necessary to use the vectorization model to generate the image vectors of the image materials to be stored, and associate the image vectors with the image materials and store them in the database.

[0189] In the search stage, if the search information input by the user is an image, after generating the image vector of the search information, calculate the similarity between the image vector and the image vectors of the image materials stored in the database, and recall the image materials stored in the database with higher similarity in descending order of similarity as the alternative image materials to obtain the search result of image search by image.

[0190] The similarity between image vectors can be characterized by the distance between image vectors, such as Euclidean distance, Manhattan distance, etc. The smaller the distance between image vectors, the higher the similarity.

[0191] The image materials stored in the database with a similarity higher than the preset value to the image vector of the search information can be recalled as the search result, or the top N image materials stored in the database with the highest similarity can be used as the search result.

[0192] Step S310, determine whether the search information is a statement. If not, execute step S311; if so, execute step S314.

[0193] Step S311, determine at least one type hit by the search information.

[0194] Step S312, search for alternative image materials based on the matching result between the search information and the description information of the at least one type of image materials stored in the database.

[0195] Step S313, in response to the user's editing operation on at least one of the alternative image materials, generate a target image, and generate text corresponding to the target image based on the description information of the at least one alternative image material.

[0196] Step S314, obtain a text vector of the search information based on a pre-trained vectorization model.

[0197] For search information in text modality, vectorization of the search information can be performed through a vectorization model such as a second vectorization model or multimodal vectorization to obtain a text vector of the search information.

[0198] Optionally, after obtaining the text vector of the search information, the method further includes the following steps:

[0199] Map the text vector of the search information to the vector space where the image vectors stored in the database are located.

[0200] To further improve the flexibility of vector search and implement various search methods such as image search by text, text search by text, and text and image search by text, the image vectors and text vectors stored in the database can be located in the same vector space. Correspondingly, in the search stage, after obtaining the text vector of the search information, the text vector is projected through vector mapping to the vector space where the vectors stored in the database are located, such as the vector space where the image vectors are located.

[0201] In some embodiments, if the vectors stored in the database are all located in the vector space where the text vectors are located, the step of mapping the text vector of the search information can be omitted.

[0202] Step S315, based on the text vector of the search information and the image vectors and / or text vectors stored in the database, obtain search results from the image materials and / or text materials stored in the database, and display the search results on the design interaction interface.

[0203] If the modality of the search information input by the user is text, it is necessary to determine the user's search intent. The search information in text modality is divided into two types: statements and keywords. If it is a keyword, the user's search intent can be determined as an image. If it is a statement, it is necessary to further identify the intent of the search information to determine the user's search intent. That is to say, for the search information in text modality, when determining the user's search intent, it is possible to first determine whether the search information is a statement. If not, that is, the search information input by the user includes one or more keywords, and the keywords can be separated by delimiters such as spaces, commas, and Chinese punctuation marks for separating words in a series, then the user's search intent is determined as searching for images. The keywords in the search information are matched with the description information of the image materials of the type hit by the search information stored in the database. Based on the matching results, the image materials stored in the database are recalled as alternative image materials to obtain the search results of image search by text.

[0204] In the database establishment stage, for the image materials to be stored, it is necessary to use a multimodal large language model to generate the description information (modality is text) of the image materials to be stored, and associate and store the image vector, description information of the same image material with the image material in the database.

[0205] The input modality of the multimodal large language model is image, and the output modality is text. In the training stage of the multimodal large language model, the description information samples of the image samples can be obtained by manual annotation, or the objects, colors of the objects, scenes, etc. included in the image samples can be identified by using existing models to automatically annotate some of the description information in the description information samples, and the remaining description information is manually annotated by experts.

[0206] The multimodal large language model can include a visual encoder, specifically an image encoder, which is used to encode the input image. The encoding of the image is analyzed by the large language model to generate description information. By adjusting the parameters of the multimodal large language model according to the loss value between the description information samples and the description information output by the multimodal large language model until the model training end condition is met.

[0207] In the output description information, there can be multiple tags, and each tag corresponds to the description information of the image material under that tag. The tags can include tags of the objects included in the image material, scene tags of the image material, style tags of the image material, color tags of the main objects included in the image material, etc.

[0208] When matching, the keywords in the search information input by the user can be matched with the tags of the image materials stored in the database, or with the tags and their description information, and the image materials in the database to be recalled are determined by parameters such as the number of matching keywords and the similarity of the matching keywords.

[0209] Exemplarily, if the search information input by the user is "dog, grassland, running", then through tag matching, from the image materials of the types hit by the search information stored in the database, such as the image materials of the outdoor sports type, the image materials whose object tags include both "dog" and "grassland" and the description information of the object tag "dog" includes words similar to "running" are preferentially recalled as alternative image materials.

[0210] On the premise that the search information is text, to determine whether the search information is a sentence, it can be determined through information such as the composition, semantics, and length of the search information. If the search information includes a delimiter and has a short length, it is determined that the search information is not a sentence.

[0211] Exemplarily, if the search information includes colloquial words, it is determined that the search information is a sentence.

[0212] Optionally, determining whether the search information is a sentence includes:

[0213] If the length of the search information is shorter than a preset length, or the search information includes multiple delimiters, it is determined that the search information is not a sentence.

[0214] Among them, the delimiter can be a space, a comma, a semicolon, etc.

[0215] Optionally, determining whether the search information is a sentence includes:

[0216] Based on at least one of the length, syntactic structure, and semantic analysis result of the search information, determine whether the search information is a sentence.

[0217] For the search information in text modality, if the length of the search information is shorter than the preset length, it is determined that the search information is not a sentence. It can also be determined whether the search information is a sentence by whether the syntactic structure of the search information is complete.

[0218] The syntactic structure of the search information can be analyzed through natural language processing algorithms.

[0219] Exemplarily, for the search information whose length is not shorter than the preset length, it can be further determined whether the syntactic structure of the search information is complete. If it is complete, the search information is a sentence; if it is incomplete, the search information is not a sentence.

[0220] It is possible to perform semantic analysis on the search information, identify information such as entities in the search information and the relationships between entities, and then determine whether the search information is a sentence based on the results of semantic analysis such as the identified entities and the relationships between entities.

[0221] In some embodiments, for search information with a length not shorter than a preset length, semantic analysis can be performed on the search information, and based on the results of the semantic analysis, it can be determined whether the search information is a sentence.

[0222] If it is determined through sentence determination that the search information is a sentence, then first, a vectorization model needs to be used to convert the search information into a text vector. The vectorization model is a text vectorization model. For example, an N - Gram model, a Word2vec (Word to Vector) model, etc.

[0223] After obtaining the text vector of the search information, taking the text vectors or image vectors stored in the database as the scope, through vector matching, search results can be obtained from the text materials or image materials stored in the database; or taking the text vectors and image vectors stored in the database as the scope, through vector matching, search results can be obtained from the text materials and image materials stored in the database.

[0224] Vector matching can specifically be calculating the similarity between text vectors, or between a text vector and an image vector.

[0225] The text vectors and image vectors stored in the database are in the same vector space. Before storage, through vector mapping, the text vectors can be mapped to the space where the image vectors are located, or the image vectors can be mapped to the space where the text vectors are located.

[0226] After obtaining the text vector of the search information, through vector matching, the text vector or image vector that matches the text vector of the search information can be determined from the text vectors and image vectors stored in the database, and the text materials stored in the database corresponding to the matching text vectors and the image materials stored in the database corresponding to the matching image vectors are returned as search results.

[0227] Furthermore, when the search results are empty, the search scope can be expanded. If the search scope cannot be expanded, corresponding prompt information is displayed to indicate that no relevant content has been searched.

[0228] Expanding the search scope can be: when the search results for searching with text materials as the search scope are empty, through the keyword or image vector of the search information and the matching results of the description information and image vectors stored in the database, search results are found from the image materials stored in the database; when the search results for searching with image materials as the search scope are empty, through the matching results such as the similarity between the text vector of the search information and the text vectors stored in the database, search results are found from the text materials stored in the database.

[0229] Through the statement judgment of the text search information, the preliminary search intention recognition is realized. By delimiting the search scope according to the search intention, such as the image materials and text materials stored in the database, etc., the search efficiency is improved. At the same time, by combining multi-dimensional features such as length, syntactic structure, and semantic analysis results for statement judgment, the accuracy of statement judgment is improved, and then the accuracy of search intention recognition is improved, making the search results more in line with the user's expectations.

[0230] In the graphic generation method provided in this embodiment, the image vectors of the image materials and the description information of the image materials output by the multi-modal large language model under the guidance of multiple prompt words are stored in the database. During image search, it supports users to search for images by image and search for images by text. When searching for images by image, the search information input by the user is an image, and image retrieval is realized through the matching result of the image vector and the image vectors stored in the database; when searching for images by text, the user can input one or more keywords, and image retrieval is realized through the matching result of the keywords and the description information of the images stored in the database. The data search method provided in this application supports searching for images by image and searching for images by text, enriches the image retrieval methods, and improves the flexibility of image search. At the same time, it also supports searching for text by text, and realizes the recall of the stored text materials through vector matching. The search methods are flexible and diverse, improving the user's search experience.

[0231] Figure 5 It is a schematic diagram of searching for images by statement provided in the embodiment of this application, as Figure 5 shown, when searching for images or text, the user can input a statement described in natural language. This statement can be described in colloquial language. For example, "I want a picture of a puppy running happily on the grassland". After the front-end device of the database system obtains this statement, it sends this statement to the search device. The search device converts this statement into a vector through a vectorization model to obtain a text vector; calculates the similarity between this text vector and the image vectors of each image material stored in the database, and recalls the image materials stored in the database (that is, the search alternative image materials) in the order of the similarity between the vectors from high to low, realizing image search in natural language and improving the flexibility of image search. Figure 5 Taking the recall of 3 images, namely img51 to img53, as an example.

[0232] Searching for images in the database by using statements enables users to avoid refining accurate and comprehensive keywords, simplifies user operations, and improves the flexibility of image search.

[0233] In some embodiments, the vectorization model can be a multimodal vectorization model. The input modalities can be text or images, and the output can be an image vector, or both an image vector and a text vector. When the search information is a statement, the search intent of the search information can be identified first. If the search intent is an image search intent, then based on the pre-trained multimodal vectorization model, an image vector of the search information can be obtained. Thus, through the image vector of the search information and the image vectors stored in the database, search results can be obtained from the image materials stored in the database. If the search intent is a text search intent, then based on the pre-trained multimodal vectorization model, a text vector of the search information can be obtained. Thus, through the text vector of the search information and the text vectors stored in the database, search results can be obtained from the text materials stored in the database. If the search intent is a text-and-image search intent, then based on the pre-trained multimodal vectorization model, an image vector and a text vector of the search information can be obtained respectively. Thus, through the image vector and the text vector of the search information, and the image vectors and text vectors stored in the database, search results can be obtained from the image materials and text materials stored in the database.

[0234] Optionally, when the search information is a statement, the search results can be obtained through the following steps:

[0235] Identify the search intent of the search information; if the search intent of the search information is an image search intent, then based on the text vector of the search information and the image vectors stored in the database, search results can be obtained from the image materials stored in the database; if the search intent of the search information is a text search intent, then based on the text vector of the search information and the text vectors stored in the database, search results can be obtained from the text materials stored in the database; if the search intent of the search information is a text-and-image search intent, then based on the text vector of the search information and the image vectors and text vectors stored in the database, search results can be obtained from the image materials and text materials stored in the database.

[0236] In some embodiments, the text materials stored in the database can be abbreviated as text, and the image materials can be abbreviated as images.

[0237] If the search information is a statement, the search intent of the search information can be determined based on the keywords included in the search information. The search intent of the search information can also be determined based on the semantic analysis result of the statement corresponding to the search information.

[0238] Search intents can be classified into three types according to the search scope: image search intent, text search intent, and image-text search intent. The search scope of the image search intent is the image materials stored in the database, the search scope of the text search intent is the text materials stored in the database, and the search scope of the image-text search intent is the image materials and text materials stored in the database.

[0239] Optionally, identifying the search intent of the search information includes:

[0240] Inputting the search information into a pre-trained search intent recognition model, and outputting the search intent of the search information via the search intent recognition model.

[0241] The search intent recognition model can be any model for classifying text, such as a deep learning model, a support vector machine, a large language model, etc.

[0242] Using the model for search intent recognition improves the accuracy of intent recognition, and thus improves the accuracy of determining the search scope, taking into account both search efficiency and accuracy.

[0243] Optionally, identifying the search intent of the search information includes:

[0244] If the end of the search information includes a preset keyword, determining the search intent of the search information as an image search intent; wherein, the preset keyword includes words associated with images.

[0245] Exemplarily, words associated with images include, but are not limited to, illustration, line drawing, real scene picture, image, drawing, schematic diagram, left view, top view, etc., words for representing pictures or synonyms of these words.

[0246] The search information can be regarded as a text sequence, and the end of the search information can be the word at the end of the text sequence, such as the last word in the search information that is not a punctuation mark.

[0247] Exemplarily, if the search information is "Nordic-style home improvement pictures", since the last word of this search information is "home improvement pictures", which is a word associated with images, the search intent of this search information is an image search intent.

[0248] Identifying the image search intent through the keyword at the end of the search information has a simple logic, is easy to implement, and improves the efficiency of search intent recognition.

[0249] Optionally, identifying the search intent of the search information includes:

[0250] Input the search information into a pre-trained search intent recognition model to determine the first search intent; based on the user's search history, determine the user's intent preference; based on the user's role, the user's intent preference, and the interface opened on the screen when the search information is input, determine the second search intent; determine that the search intent is the union of the first search intent and the second search intent.

[0251] The user's role can be determined by the user's position, the product the user is responsible for, etc.

[0252] The user's intent preference can be the search intent with the highest frequency of occurrence in the search history, or it can also be the proportion of various search intents.

[0253] The interface opened when the user inputs search information can include the interfaces of other software except the interface of the software corresponding to the database system.

[0254] If the interface opened by the user is a text editing interface, such as the interface of a text editing software, then the second search intent can be determined as a text search intent; if the interface opened by the user is an image editing interface, such as the interface of an image drawing software, then the second search intent is determined as an image search intent; if the interface opened by the user is a graphic and text editing interface, then the second search intent cannot be determined based on the opened interface, and it is necessary to further consider the user's intent preference to determine the second search intent.

[0255] If the first search intent is an image search intent and the second search intent is a text search intent, then the finally determined search intent is a graphic and text search intent.

[0256] In this embodiment, in the database establishment stage, the automatic tagging of images by the large model improves the efficiency, comprehensiveness, and accuracy of image annotation; at the same time, the large model is guided to perform image annotation through differentially set prompt words, improving the diversity and flexibility of annotation; using the feature data of two modalities, image vectors and description information, as the index of image materials enriches the index of image materials and provides a basis for the implementation of various search methods; at the same time, different strategies are adopted to process the data to be stored in the database in different modalities, avoiding an excessive number of indexes stored in the database while ensuring high-accuracy search. In the search stage, it supports the user to search for images by image, by label, and by statement in multiple search methods, the search methods are rich and flexible, and it supports the user to perform graphic and text search through natural description language. Through the accurate recognition of search intent, the accuracy of determining the search scope is improved, taking into account both search efficiency and accuracy.

[0257] Figure 6 Schematic diagrams of 4 search methods provided by the embodiments of this application, as Figure 6As shown in the figure, when a user searches for data of interest in a vector database, there are four search methods available. The search information for the first search method is tags. The search information for the second search method (searching for images with natural language) and the fourth search method (searching for text with natural language) is natural language. The search information for the third search method (searching for images with images) is pictures.

[0258] For the first search method, the user can enter one or more tags in the search box. By matching the entered tags with the tags in the description information of the image materials stored in the vector database, the images that match the tags can be obtained.

[0259] For the second and third search methods, since the search intents are both image search intents, the image vector of the statement input in the second search method and the image vector of the input image in the third search method can be extracted through a multimodal vectorization model (an example of the first vectorization model). Through the matching result of the image vector and the image vectors stored in the vector database, the image materials that match the image vector of the search information can be obtained.

[0260] For the fourth search method, whose search intent is a text search intent, the text vector of the search information can be obtained through a text vectorization model (i.e., the above-mentioned second vectorization model); through the matching result of the text vector and the text vectors stored in the vector database, the text materials that match the text vector of the search information can be obtained.

[0261] It is also possible to search for images and text with natural language. The search process of this search method can be obtained by combining the search processes of the aforementioned second and fourth search methods, which will not be elaborated here.

[0262] The embodiment of the present application also provides a data storage method, which includes:

[0263] For the image materials to be stored, according to the type of the image materials, determine the corresponding prompt words; through a pre-trained multimodal large language model, obtain the description information of the image materials under the guidance of the corresponding prompt words; based on a pre-trained vectorization model, obtain the image vectors of the image materials; in the database, store the image materials, as well as the image vectors and description information of the image materials, in an associated manner.

[0264] Optionally, the data storage method further includes:

[0265] For the text materials to be stored, based on a pre-trained vectorization model, obtain the text vectors of the text materials; in the database, store the text materials and their text vectors in an associated manner.

[0266] The embodiment of the present application further provides a data search method, including:

[0267] Obtain the search information input by the user; determine whether the search information is a statement; if so, determine at least one type hit by the search information; based on the matching result between the search information and the description information of the at least one type of image material stored in the database, search for alternative image materials from the at least one type of image material stored in the database.

[0268] Optionally, determining whether the search information is a statement includes:

[0269] Based on at least one of the length, syntactic structure, and semantic analysis result of the search information, determine whether the search information is a statement.

[0270] Optionally, determining whether the search information is a statement includes:

[0271] If the length of the search information is shorter than a preset length, or the search information includes multiple delimiters, determine that the search information is not a statement.

[0272] Optionally, the search method further includes:

[0273] If the search information is an image, based on a pre-trained vectorization model, obtain the image vector of the search information; based on the image vector of the search information and the image vectors stored in the database, search for the alternative image materials from the images stored in the database.

[0274] Optionally, the vectorization model is a multi-modal vectorization model. Based on a pre-trained vectorization model, obtaining the image vector of the image material includes:

[0275] Input the image material and its description information into a pre-trained multi-modal vectorization model to obtain the image vector of the image material.

[0276] Optionally, the search method further includes:

[0277] If it is determined that the search information is a statement, based on a pre-trained vectorization model, obtain the text vector of the search information; based on the text vector of the search information and the image vectors and / or text vectors stored in the database, obtain search results from the image materials and / or text materials stored in the database.

[0278] Optionally, based on the text vector of the search information and the image vectors and / or text vectors stored in the database, obtaining search results from the image materials and / or text materials stored in the database includes:

[0279] Identify the search intent of the search information; if the search intent of the search information is an image search intent, then based on the text vector of the search information and the image vectors stored in the database, obtain search results from the image materials stored in the database; if the search intent of the search information is a text search intent, then based on the text vector of the search information and the text vectors stored in the database, obtain search results from the text materials stored in the database; if the search intent of the search information is a text and image search intent, then based on the text vector of the search information and the image vectors and text vectors stored in the database, obtain search results from the image materials and text materials stored in the database.

[0280] Optionally, identifying the search intent of the search information includes:

[0281] If the end of the search information includes a preset keyword, then determine that the search intent of the search information is an image search intent; wherein, the preset keyword includes a word associated with a figure.

[0282] Optionally, identifying the search intent of the search information includes:

[0283] Input the search information into a pre-trained search intent recognition model to determine a first search intent; based on the user's search history, determine the user's intent preference; based on the user's role, the user's intent preference, and the interface opened on the screen when inputting the search information, determine a second search intent; determine that the search intent is the union of the first search intent and the second search intent.

[0284] Corresponding to the text and image generation method provided in the foregoing embodiment, the embodiment of the present application further provides a corresponding device. Figure 7 As shown in the structural schematic diagram of a text and image generation device provided by the embodiment of the present application, Figure 7 as shown, the text and image generation device includes: a prompt word determination module 710, an image material annotation module 720, an interface display module 730, a hit type determination module 740, an image material search module 750, and a text and image generation module 760.

[0285] The prompt word determination module 710 is configured to determine, for each of multiple image materials to be stored, a prompt word corresponding to the image material according to the type of the image material; the image material annotation module 720 is configured to annotate the image material under the guidance of the prompt word corresponding to the image material via a pre-trained image annotation model, obtain description information of the image material, and store the image material and its description information in a database; the interface display module 730 is configured to display a design interaction interface, wherein the design interaction interface is used for a user to search for image materials and generate a target image and text corresponding to the target image according to the searched image materials; the hit type determination module 740 is configured to determine at least one type hit by the search information after detecting the search information input by the user in the design interaction interface; the image material search module 750 is configured to search for alternative image materials from the image materials of at least one type stored in the database based on a matching result between the search information and the description information of the image materials of at least one type stored in the database; the graphic and text generation module 760 is configured to generate a target image in response to an editing operation of the user on at least one of the alternative image materials, and generate text corresponding to the target image based on the description information of the at least one alternative image material.

[0286] Optionally, the prompt word determination module 710 includes:

[0287] Determine the type of the image material based on the objects included in the image material; determine the prompt word corresponding to the image material based on the corresponding relationship between the type of the image material and the prompt word, and the type of the image material.

[0288] Optionally, the design interaction interface includes a first type of search box and multiple second type of search boxes, the type hit by the first type of search box is empty, and the second type of search box hits at least one type; the hit type determination module 740 is specifically configured to:

[0289] After detecting the search information input by the user in the design interaction interface, if the search box for inputting the search information is a second type of search box, determine the type hit by the search box for inputting the search information as the type hit by the search information; if the search box for inputting the search information is a first type of search box, determine at least one type hit by the search information based on the keywords extracted from the search information.

[0290] Optionally, the hit type determination module 740 includes:

[0291] A statement determination unit, configured to, after detecting search information input by a user in the design interaction interface, if the search information is text, determine whether the search information is a statement; a hit type determination unit, configured to, if it is not a statement, determine at least one type hit by the search information.

[0292] Optionally, the statement determination unit is specifically configured to:

[0293] After detecting search information input by a user in the design interaction interface, if the search information is text, based on at least one of the length, syntactic structure, and semantic analysis result of the search information, determine whether the search information is a statement.

[0294] Optionally, the statement determination unit is specifically configured to:

[0295] After detecting search information input by a user in the design interaction interface, if the search information is text, and the length of the search information is shorter than a preset length, or the search information includes multiple delimiters, determine that the search information is not a statement.

[0296] Optionally, the graphic and text generation device further includes:

[0297] An image vector storage module, configured to, for each image material to be stored, based on a pre-trained vectorization model, obtain an image vector of the image material, and store the image vector in the database; an image search module, configured to, after detecting search information input by a user in the design interaction interface, if the search information is an image, based on a pre-trained vectorization model, obtain an image vector of the search information, and based on the image vector of the search information and the image vectors stored in the database, search for the alternative image materials from the images stored in the database.

[0298] Optionally, the image vector storage module is specifically configured to:

[0299] For each image material to be stored, input the image material and its description information into a pre-trained multi-modal vectorization model, obtain an image vector of the image material, and store the image vector in the database.

[0300] Optionally, the graphic and text generation device further includes:

[0301] A text vector storage module, which is used to obtain the text vector of the text material to be stored based on a pre-trained vectorization model, and store the text material and its text vector in the database; A statement search module, which is used to, after detecting the search information input by the user on the design interaction interface, if it is determined that the search information is a statement, obtain the text vector of the search information based on a pre-trained vectorization model, and based on the text vector of the search information and the image vectors and / or text vectors stored in the database, obtain search results from the image materials and / or text materials stored in the database; A result display module, which is used to display the search results on the design interaction interface.

[0302] Optionally, when the statement search module performs a search, it is specifically used for:

[0303] Identify the search intent of the search information; if the search intent of the search information is an image search intent, obtain search results from the image materials stored in the database based on the text vector of the search information and the image vectors stored in the database; if the search intent of the search information is a text search intent, obtain search results from the text materials stored in the database based on the text vector of the search information and the text vectors stored in the database; if the search intent of the search information is a graphic and text search intent, obtain search results from the image materials and text materials stored in the database based on the text vector of the search information and the image vectors and text vectors stored in the database.

[0304] Optionally, when the statement search module identifies the search intent, it is specifically used for:

[0305] If the end of the search information includes a preset keyword, determine that the search intent of the search information is an image search intent; where the preset keyword includes a word associated with a graph.

[0306] Optionally, when the statement search module identifies the search intent, it is specifically used for:

[0307] Input the search information into a pre-trained search intent recognition model to determine a first search intent; determine the user's intent preference based on the user's search history; determine a second search intent based on the user's role, the user's intent preference, and the interface opened on the screen when the search information is input; determine the search intent as the union of the first search intent and the second search intent.

[0308] The graphic and text generation device provided in the embodiments of the present application can be used to execute the technical solutions of the graphic and text generation methods provided in any of the above embodiments of the present application. The implementation principles and technical effects are similar, and will not be elaborated here in this embodiment.

[0309] Figure 8 This is a schematic structural diagram of an electronic device provided by an embodiment of the present application. As Figure 8 shown, the electronic device of this embodiment may include: at least one processor 801; and a memory 802 communicatively connected to the at least one processor; wherein, the memory 802 stores instructions executable by the at least one processor 801, and the instructions are executed by the at least one processor 801 to cause the electronic device to execute the method described in any of the foregoing embodiments.

[0310] Optionally, the memory 802 may be either independent or integrated with the processor 801.

[0311] When the memory 802 is independently provided, the device further includes a bus for connecting the memory 802 and the processor 801.

[0312] The implementation principle and technical effects of the electronic device provided by this embodiment may be referred to the foregoing embodiments, and will not be elaborated herein.

[0313] An embodiment of the present application further provides a computer-readable storage medium, in which computer-executable instructions are stored, and when the computer-executable instructions are executed by a processor, the method provided in any of the foregoing embodiments can be implemented.

[0314] An embodiment of the present application further provides a computer program product, including a computer program, and when the computer program is executed by a processor, the method provided in any of the foregoing embodiments is implemented.

[0315] In several embodiments provided by the present application, it should be understood that the disclosed devices and methods may be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the modules is only a logical function division, and there may be other division methods in actual implementation. For example, multiple modules may be combined or integrated into another system, or some features may be ignored or not executed.

[0316] The integrated modules implemented in the form of software function modules as described above may be stored in a computer-readable storage medium. The software function modules are stored in a storage medium, including several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) or a processor to execute some steps of the methods described in various embodiments of the present application.

[0317] It should be understood that the above-mentioned processor can be a Central Processing Unit (CPU), or other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), etc. The general-purpose processor can be a microprocessor or any conventional processor, etc. The steps of the method disclosed in combination with the application can be directly implemented by a hardware processor, or implemented by a combination of hardware and software modules in the processor.

[0318] The memory may include high-speed memory and may also include non-volatile memory, such as at least one disk memory, and can also be a USB flash drive, a mobile hard disk, a read-only memory, a magnetic disk, or an optical disc, etc.

[0319] The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, an Extended Industry Standard Architecture (EISA) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of convenience in representation, the buses in the drawings of this application are not limited to only one bus or one type of bus.

[0320] The above-mentioned storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory, electrically erasable programmable read-only memory, erasable programmable read-only memory, programmable read-only memory, read-only memory, magnetic memory, flash memory, a magnetic disk, or an optical disc. The storage medium can be any available medium that can be accessed by a general-purpose or special-purpose computer.

[0321] An exemplary storage medium is coupled to the processor, enabling the processor to read information from the storage medium and write information to the storage medium. Of course, the storage medium can also be a component of the processor. The processor and the storage medium can be located in an application specific integrated circuit. Of course, the processor and the storage medium can also exist as discrete components in an electronic device or a control device of a vehicle sentry mode.

[0322] It should be noted that in this text, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, such that a process, method, article or device comprising a series of elements not only includes those elements but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising one..." does not exclude the presence of additional identical elements in the process, method, article or device comprising such element.

[0323] The serial numbers of the embodiments of the present application above are only for description and do not represent the superiority or inferiority of the embodiments.

[0324] Through the description of the above embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, it can also be implemented by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to enable a terminal device (which can be a mobile phone, computer, server, or network device, etc.) to execute the methods provided by the various embodiments of the present application.

[0325] After considering the specification and practicing the invention disclosed herein, those skilled in the art will readily conceive of other embodiments of the present application. The present application is intended to cover any variations, uses, or adaptations of the present application that follow the general principles of the present application and include known common knowledge or conventional technical means in the technical field not disclosed in the present application. The specification and embodiments are only regarded as exemplary, and the true scope and spirit of the present application are pointed out by the following claims.

[0326] It should be understood that the present application is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present application is only limited by the appended claims.

Claims

1. A method for generating images and texts, characterized in that: include: For each image material among the plurality of image materials to be stored, determining the type of the image material based on an object contained in the image material; Determining the prompt word corresponding to the image material based on a preset correspondence between the type of the image material and the prompt word and the type of the image material; Using a pre-trained image annotation model, under the guidance of the prompt words corresponding to the image material, annotate the image material to obtain description information of the image material, and store the image material and its description information in a database; Displaying a design interaction interface, wherein the design interaction interface is used for users to search for image materials, and to generate a target image and text corresponding to the target image based on the searched image materials; After detecting the search information input by the user in the design interaction interface, when the search information is an image or the search information is text and the text is not a sentence, determining that the user's search intention is an image, and determining at least one of the types hit by the search information; Based on a matching result between the search information and the description information of the at least one image material of the type stored in the database, searching for candidate image materials from the at least one image material of the type stored in the database; In response to a user's editing operation on at least one of the candidate image materials, a target image is generated, and based on description information of the at least one candidate image material, a text corresponding to the target image is generated; When the search information is text and the text is a sentence, the search information is input into a pre-trained search intent recognition model to determine a first search intent; based on the search history of the user, the user's intent preference is determined; based on the user's role, the user's intent preference and the interface opened on the screen when the search information is input, a second search intent is determined; the search intent is determined to be the union of the first search intent and the second search intent; and according to the search intent, search results are obtained from the image material and / or text material stored in the database; The design interaction interface includes a first-type search box and multiple second-type search boxes, the type hit by the first-type search box is empty, and different second-type search boxes hit different types; determining at least one type hit by the search information includes: If the search box in which the search information is input is a second type search box, determining the type of the search box hit in which the search information is input is the type of the search information hit; If the search box in which the search information is input is a first type search box, at least one type of hits of the search information is determined based on keywords extracted from the search information.

2. The method according to claim 1, characterized in that In the corresponding relationship, the prompt words corresponding to the furniture-type image materials include style prompt words, design point prompt words, material prompt words, color prompt words and function prompt words, and the prompt words corresponding to the scene-type image materials include scene prompt words, style prompt words and included object prompt words.

3. The method according to claim 1, characterized in that If the search information is an image, keywords of the search information are extracted based on the objects and texts recognized from the search information.

4. The method according to claim 1, characterized in that: Also includes: Based on at least one of the length, grammatical structure and semantic analysis result of the search information, it is determined whether the search information is a sentence.

5. The method according to claim 1, characterized in that Also includes: If the length of the search information is shorter than a preset length, or the search information includes a plurality of separators, it is determined that the search information is not a sentence.

6. The method according to claim 1, characterized in that The method further comprises: For each image material to be stored, an image vector of the image material is obtained based on a pre-trained vectorization model, and the image vector is stored in the database, wherein the image vector includes a vector of the description information and a vector of the image material.

7. The method according to claim 6, characterized in that If the search information is an image, searching for candidate image materials from the at least one image material of the type stored in the database based on a matching result between the search information and the description information of the at least one image material of the type stored in the database includes: Obtaining an image vector of the search information based on a pre-trained vectorization model; Based on the image vector of the search information and the image vector of the at least one image material of the type stored in the database, a candidate image material is searched from the at least one image material of the type stored in the database.

8. The method according to claim 6, characterized in that The vectorization model is a multimodal vectorization model, and the image vector of the image material is obtained based on a pre-trained vectorization model, including: The image material and its description information are input into a pre-trained multimodal vectorization model to obtain an image vector of the image material.

9. The method according to claim 1, characterized in that: The method further comprises: For the text material to be stored, a text vector of the text material is obtained based on a pre-trained vectorization model, and the text material and its text vector are stored in the database.

10. The method according to claim 9, characterized in that Also includes: If it is determined that the search information is text and the text is a sentence, obtaining a text vector of the search information based on a pre-trained vectorization model; Obtaining search results from the image materials and / or text materials stored in the database according to the search intent includes: obtaining search results from the image materials and / or text materials stored in the database based on the text vector of the search information and the image vector and / or text vector stored in the database according to the search intent; The search results are displayed on the design interaction interface.

11. The method according to claim 9, characterized in that According to the search intent, based on the text vector of the search information and the image vector and / or text vector stored in the database, a search result is obtained from the image material and / or text material stored in the database, including: If the search intent of the search information is an image search intent, obtaining search results from image materials stored in the database based on the text vector of the search information and the image vector stored in the database; If the search intent of the search information is a text search intent, obtaining search results from text materials stored in the database based on the text vector of the search information and the text vector stored in the database; If the search intent of the search information is a picture and text search intent, then based on the text vector of the search information and the image vector and text vector stored in the database, search results are obtained from the image materials and text materials stored in the database.

12. The method according to claim 11, characterized in that Identifying the search intent of the search information, including: If the end of the search information includes a preset keyword, determining that the search intent of the search information is an image search intent; The preset keywords include words associated with the image.

13. The method according to any one of claims 1 to 12, characterized in that: The image annotation model includes an image encoder, a modality alignment module, and a large language model layer; The image encoder is used to extract encoding features of an input image; The modal alignment module is used to project the encoding feature into a preset text space to obtain an alignment feature vector; The large language model layer is used to output description information of the input image under the corresponding prompt word based on the aligned feature vector under the guidance of the prompt word corresponding to the input image.

14. A device for generating images and texts, characterized in that: include: A prompt word determination module, used for determining the type of each image material among a plurality of image materials to be stored based on an object contained in the image material; Determining the prompt word corresponding to the image material based on a preset correspondence between the type of the image material and the prompt word and the type of the image material; An image material annotation module is used to annotate the image material through a pre-trained image annotation model under the guidance of a prompt word corresponding to the image material, obtain description information of the image material, and store the image material and its description information in a database; An interface display module, used to display a design interaction interface, wherein the design interaction interface is used for users to search for image materials, and to generate a target image and text corresponding to the target image based on the searched image materials; A hit type determination module, configured to detect that after the search information input by the user in the design interaction interface is an image or the search information is text and the text is not a sentence, determine that the user's search intention is an image, and determine at least one of the types hit by the search information; an image material search module, configured to search for candidate image materials from the at least one image material of the type stored in the database based on a matching result between the search information and the description information of the at least one image material of the type stored in the database; An image and text generation module, configured to generate a target image in response to a user's editing operation on at least one of the candidate image materials, and generate text corresponding to the target image based on description information of the at least one candidate image material; A sentence search module, used for inputting the search information into a pre-trained search intent recognition model to determine a first search intent when the search information is text and the text is a sentence; determining the user's intention preference based on the user's search history; determining a second search intent based on the user's role, the user's intention preference and an interface opened on the screen when the search information is input; determining the search intent as the union of the first search intent and the second search intent; The image material search module is further used to obtain search results from the image materials and / or text materials stored in the database according to the search intention; The design interaction interface includes a first-category search box and multiple second-category search boxes, the hit type of the first-category search box is empty, and different second-category search boxes hit different types; the hit type determination module is specifically used to determine the hit type of the search box in which the search information is input as the type of search information hit if the search box in which the search information is input is a second-category search box; if the search box in which the search information is input is a first-category search box, then determining at least one type of the search information hit based on keywords extracted from the search information.

15. An electronic device, characterized in that: include: A processor, and a memory communicatively connected to the processor; The memory stores computer-executable instructions; The processor executes the computer-executable instructions stored in the memory to implement the method according to any one of claims 1 to 13.

16. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer-executable instructions, which are used to implement the method according to any one of claims 1 to 13 when executed by a processor.

17. A computer program product, characterized in that The method comprises a computer program, which implements the method according to any one of claims 1 to 13 when being executed by a processor.

Citation Information

Patent Citations

  • Image searching method and device

    CN118503473A

  • Event searching method and device, electronic equipment and computer storage medium

    CN118626674A