Data search method and device, product and equipment

By intent identification of search text from multiple intent understanding dimensions and similarity matching with text labels in the emoticon gallery, the problem of inaccurate expression map search in the prior art is solved, and higher search accuracy and user experience are achieved.

CN120067358APending Publication Date: 2025-05-30TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510123604.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-24
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

Existing emoticon image search methods rely on the degree of overlap between the input text and the emoticon image text, resulting in inaccurate searches and it is difficult to find emoticon images that fully meet user needs.

Method used

By intent identification of search text from multiple intent understanding dimensions, multiple intent texts are generated, and similarity matched with text labels in the emoticon gallery, a more accurate emoticon search is achieved.

Benefits of technology

Improve the accuracy of searching for emoticons in the emoticon gallery that match the search text, and enhances the relevance and user experience of search results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120067358A_ABST
    Figure CN120067358A_ABST
Patent Text Reader

Abstract

The invention discloses a data search method and device, a product and equipment, and the method comprises the steps: obtaining a search text which is used for triggering the search of an emoticon; intention recognition is conducted on the search text from N intention understanding dimensions, N intention texts of the search text are generated, one intention text corresponds to one intention understanding dimension, and N is a positive integer; obtaining a text label of each emoticon in an emoticon library, wherein the text label of any emoticon comprises intention labels of any emoticon on N intention understanding dimensions; and searching a matched first emoticon for the search text in the emoticon library according to the similarity between the N intention texts and the text labels of the emoticons in the emoticon library. By adopting the method and the device, the accuracy of searching the emoticon matched with the search text in the emoticon library can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of data search, and in particular, to a data search method, apparatus, product, and device. Background Art

[0002] In the scenario of emoji search, users often face a pain point, that is, when a user has a sentence to express, it is difficult to find an emoji that exactly matches the meaning of this sentence. In existing applications, the emojis are searched and recalled for the user based on the degree of overlap (such as the proportion of overlapping characters) between the input text used by the user to search for emojis and the text contained in the emojis themselves. However, this simple emoji search method often cannot search for emojis that exactly meet the user's needs for the user, resulting in inaccurate emoji search. Therefore, how to accurately search for the emojis that the user wants is a hot issue. Summary of the Invention

[0003] This application provides a data search method, apparatus, product, and device, which can improve the accuracy of searching for emojis that match the search text in an emoji library.

[0004] On the one hand, this application provides a data search method, which includes:

[0005] Obtain a search text, where the search text is used to trigger the search for emojis;

[0006] Perform intent recognition on the search text from N intent understanding dimensions to generate N intent texts of the search text, where one intent text corresponds to one intent understanding dimension, and N is a positive integer;

[0007] Obtain the text labels of each emoji in the emoji library, where the text label of any emoji includes the intent labels of any emoji in N intent understanding dimensions respectively;

[0008] Search for the first emoji that matches the search text in the emoji library through the similarity between the N intent texts and the text labels of the emojis in the emoji library.

[0009] On the one hand, this application provides a data search apparatus, which includes:

[0010] A first acquisition module, configured to obtain a search text, where the search text is used to trigger the search for emojis;

[0011] A generation module, configured to perform intent recognition on the search text from N intent understanding dimensions to generate N intent texts of the search text, where one intent text corresponds to one intent understanding dimension, and N is a positive integer;

[0012] A second acquisition module, configured to acquire the respective text labels of each expression graph in the expression graph library, where the text label of any expression graph includes the intention labels of the any expression graph in N intention understanding dimensions respectively;

[0013] A search module, configured to search for a first expression graph that matches the search text in the expression graph library by the similarity between N intention texts and the text labels of the expression graphs in the expression graph library.

[0014] In one implementation manner, the generating module performs intention recognition on the search text from N intention understanding dimensions, and the manner of generating N intention texts of the search text includes:

[0015] Acquire a trained intention recognition model, where the trained intention recognition model includes a text encoder and N text decoders, and one text decoder corresponds to one intention understanding dimension;

[0016] Call the text encoder to perform feature encoding processing on the search text to generate the encoded features of the search text;

[0017] Call N text decoders to perform feature decoding processing on the encoded features respectively in the corresponding intention understanding dimensions to generate N intention texts.

[0018] In one implementation manner, the above data search device further includes a training module, and the training module is configured to:

[0019] Acquire a first sample text and an intention recognition model to be trained, where the intention recognition model to be trained includes a text encoder to be trained and N text decoders to be trained, and one text decoder to be trained corresponds to one intention understanding dimension, and the first sample text has corresponding reference intention texts in N intention understanding dimensions;

[0020] Call the text encoder to be trained to perform feature encoding processing on the first sample text to generate the sample encoded features of the first sample text;

[0021] Call N text decoders to be trained to perform feature decoding processing on the sample encoded features respectively in the corresponding intention understanding dimensions to generate N sample intention texts of the first sample text;

[0022] Based on the differences between the sample intention texts and the reference intention texts in the intention understanding dimensions corresponding to each text decoder to be trained respectively, generate the feature decoding losses of each text decoder to be trained for the sample encoded features respectively;

[0023] Based on the feature decoding losses of N text decoders to be trained respectively, correct the model parameters of the intention recognition model to be trained to obtain a trained intention recognition model.

[0024] In one implementation, there are multiple first sample texts, and the multiple first sample texts have the same reference intention text in N intention understanding dimensions. Any text decoder to be trained is a target text decoder, and the target text decoder includes multiple first network layers. The training module corrects the model parameters of the intention recognition model to be trained based on the feature decoding losses of the N text decoders to be trained respectively, and the method for obtaining the trained intention recognition model includes:

[0025] During the process of the target text decoder performing feature decoding processing on the sample encoding features of each first sample text, obtain the decoding features generated by each first network layer except the last first network layer in the multiple first network layers for each first sample text respectively;

[0026] Generate the first clustering loss of the target text decoder for the multiple first sample texts based on the decoding features of the multiple first sample texts. The first clustering loss is used to reflect the differences between the decoding features generated by the same first network layer for the multiple first sample texts;

[0027] Correct the model parameters of the intention recognition model to be trained through the first clustering losses and feature decoding losses of the N text decoders to be trained respectively, and obtain the trained intention recognition model.

[0028] In one implementation, the method for the search module to search for a matching first expression image for the search text in the expression image library through the similarity between the N intention texts and the text labels of the expression images in the expression image library includes:

[0029] Obtain the search matching degrees between the search text and each expression image in the expression image library through the similarities between the search text and the N intention texts and the text labels of each expression image in the expression image library respectively;

[0030] Filter out the first expression image from the expression image library based on the search matching degrees between the search text and each expression image in the expression image library.

[0031] In one implementation, any expression image in the expression image library is a target expression image. The method for the search module to obtain the search matching degrees between the search text and each expression image in the expression image library through the similarities between the search text and the N retrieval texts and the text labels of each expression image in the expression image library respectively includes:

[0032] Call the trained feature embedding model to perform feature embedding processing on the search text, the N intention texts, and each text label of the target expression image respectively, and generate the representation features of the search text, the representation features of each intention text, and the representation features of each text label of the target expression image;

[0033] Obtain M feature similarities between the representation features of the search text, the representation features of N intent texts, and the representation features of the text labels of the target emoticon, where the feature similarity between the representation features of texts is used to reflect the similarity between texts, and M is a positive integer;

[0034] Take the maximum feature similarity among the M feature similarities as the search matching degree between the search text and the target emoticon.

[0035] In one implementation, the text label of any emoticon includes the general label of any emoticon. The target emoticon has K + N text labels, and the K + N text labels include K general labels of the target emoticon and N intent labels, where K is a positive integer;

[0036] The manner in which the search module obtains M feature similarities between the representation features of the search text, the representation features of N intent texts, and the representation features of the text labels of the target emoticon includes:

[0037] Calculate K + N feature similarities between the representation feature of the search text and the representation features of K + N text labels respectively. There is one feature similarity between the representation feature of the search text and the representation feature of a text label of the target emoticon;

[0038] Calculate N×K feature similarities between the representation features of N intent texts and the representation features of K general labels respectively. There is one feature similarity between the representation feature of an intent text and the representation feature of a general label of the target emoticon;

[0039] Calculate N feature similarities between the representation features of N intent texts and the representation features of N intent labels respectively. There is one feature similarity between the representation feature of any intent text and the representation feature of an intent label of the target emoticon in the intent understanding dimension corresponding to any intent text;

[0040] Determine the K + N feature similarities, N×K feature similarities, and N feature similarities as the M feature similarities.

[0041] In one implementation, the above training module is further used for:

[0042] Obtain a sample pair and a feature embedding model to be trained. The sample pair includes two second sample texts. The sample pair has a semantic matching label, and the semantic matching label indicates whether the text semantics between the two second sample texts in the sample pair are matching or not;

[0043] Call the feature embedding model to be trained to perform feature embedding processing on each second sample text in the sample pair, and generate the representation features of each second sample text in the sample pair;

[0044] Generate a feature embedding loss for the feature embedding model to be trained based on the representation features and semantic matching labels of each second sample text in the sample pair.

[0045] Modify the model parameters of the feature embedding model to be trained based on the feature embedding loss to obtain the trained feature embedding model.

[0046] In one implementation, there are multiple sample pairs, and multiple second sample texts with similar text semantics are included in the multiple sample pairs. The feature embedding model to be trained includes multiple second network layers. The way for the training module to modify the model parameters of the feature embedding model to be trained based on the feature embedding loss to obtain the trained feature embedding model includes:

[0047] During the process of the feature embedding model to be trained performing feature embedding processing on the second sample texts in multiple sample pairs, obtain the embedding features generated by each of the multiple second network layers for each second sample text in the multiple second sample texts.

[0048] Generate a second clustering loss for the feature embedding model to be trained for the multiple second sample texts based on the embedding features of the multiple second sample texts. The second clustering loss is used to reflect the differences between the embedding features generated by the same second network layer for the multiple second sample texts.

[0049] Modify the model parameters of the feature embedding model to be trained based on the second clustering loss and the feature embedding loss to obtain the trained feature embedding model.

[0050] In one implementation, the above training module is further used for:

[0051] Obtain the original sample text, and perform key information detection on the original sample text to obtain G key information in the original sample text, where G is a positive integer.

[0052] Perform G times of antonym replacement processing on the G key information in the original sample text to generate G antonym replacement texts of the original sample text. One antonym replacement processing is used to perform antonym replacement on one key information in the original sample text to obtain one antonym replacement text.

[0053] Perform pairwise combination processing between the original sample text and the G antonym replacement texts to obtain at least one sample pair.

[0054] In one implementation, the way for the search module to screen out the first emoji from the emoji library based on the search matching degree between the search text and each emoji in the emoji library includes:

[0055] Sort the expression images in the expression image library in descending order according to the search matching degree between the search text and each expression image in the expression image library, and obtain the sorted expression images;

[0056] Take the first Q expression images sorted in the sorted expression images as the first expression images, where Q is a positive integer and Q is less than the total number of expression images in the expression image library.

[0057] In one implementation, the search text is sent by the communication client; the search module is further configured to:

[0058] Perform expression expansion processing based on the search text to obtain the expanded second expression images;

[0059] Return the first expression images and the second expression images to the communication client.

[0060] In one implementation, the manner in which the search module performs expression expansion processing based on the search text to obtain the expanded second expression images includes:

[0061] Call the trained language rewriting model to perform language rewriting processing on the search text to generate an expression description text corresponding to the search text;

[0062] Call the text-to-image model to generate L candidate expression images based on the expression description text, and call the image-text matching model to predict the image-text relevance between each candidate expression image and the expression description text, where L is a positive integer;

[0063] Sort the L candidate expression images in descending order according to the image-text relevance between each candidate expression image and the expression description text, and obtain the sorted L candidate expression images;

[0064] Take the first P candidate expression images sorted in the sorted L candidate expression images as the second expression images, where P is a positive integer and P is less than L.

[0065] In one implementation, the above training module is further configured to:

[0066] Obtain a third sample text and the language rewriting model to be trained, where the third sample text has a reference rewritten text;

[0067] Call the language rewriting model to be trained to perform language rewriting processing on the third sample text to generate a sample rewritten text corresponding to the third sample text;

[0068] Based on the difference between the sample rewritten text and the reference rewritten text, correct the model parameters of the language rewriting model to be trained to obtain the trained language rewriting model.

[0069] In one embodiment, the method for the search module to perform expression expansion processing based on the search text to obtain the expanded second expression graph includes:

[0070] Obtain general expression graphs;

[0071] Perform synthesis processing on the search text and the general expression graphs to generate second expression graphs.

[0072] In one embodiment, the first expression graph includes an expression template graph. The process for the search module to return the first expression graph to the communication client includes:

[0073] Perform synthesis processing on the search text and the expression template graph to generate a synthesized expression graph;

[0074] Return the synthesized expression graph to the communication client.

[0075] On the one hand, the present application provides a computer device, including a memory and a processor. The memory stores a computer program. When the computer program is executed by the processor, the processor executes the method in one aspect of the present application.

[0076] On the one hand, the present application provides a computer-readable storage medium that stores a computer program. When the computer program is executed by a processor, the processor executes the method in the above-mentioned one aspect.

[0077] On the one hand, the present application provides a computer program product that includes a computer program stored in a computer-readable storage medium. The processor of the computer device reads the computer program from the computer-readable storage medium, and the processor executes the computer program, enabling the computer device to execute the methods provided in the above-mentioned various optional manners such as one aspect.

[0078] This application can obtain a search text, which is used to trigger the search for emoticons; and can perform intent recognition on the search text from N intent understanding dimensions to generate N intent texts of the search text, where one intent text corresponds to one intent understanding dimension, and N is a positive integer; moreover, it can also obtain the text labels of each emoticon in the emoticon library, and the text label of any emoticon includes the intent labels of any emoticon in N intent understanding dimensions respectively; thus, the first emoticon matching the search text can be searched in the emoticon library through the similarity between the N intent texts and the text labels of the emoticons in the emoticon library. It can be seen that the method proposed in this application can perform multi-dimensional intent recognition on the search text for searching emoticons from multiple intent understanding dimensions, and can also configure the intent labels of each emoticon in the emoticon library in the multiple intent understanding dimensions respectively. Thus, through the similarity between the multiple intent texts generated by performing intent recognition on the search text and the text labels (including intent labels) of the emoticons in the emoticon library, the multi-dimensional search for the emoticon matching the search text can be realized from the multiple intent understanding dimensions, thereby improving the accuracy of searching for the emoticon matching the search text in the emoticon library. BRIEF DESCRIPTION OF THE DRAWINGS

[0079] In order to more clearly illustrate the technical solutions in this application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of this application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0080] Figure 1 It is a schematic structural diagram of a network architecture for searching emoticons provided by an embodiment of this application;

[0081] Figure 2 It is a schematic diagram of a scenario for searching for an emoticon matching the search text provided by an embodiment of this application;

[0082] Figure 3 It is a schematic flowchart of a data search method provided by an embodiment of this application;

[0083] Figure 4 It is a schematic diagram of the effect of a chat interface provided by an embodiment of this application;

[0084] Figure 5 It is a schematic flowchart of a process for searching and obtaining emoticons provided by an embodiment of this application;

[0085] Figure 6 It is a schematic flowchart of a method for training an intent recognition model provided by an embodiment of this application;

[0086] Figure 7 It is a schematic structural diagram of an intention recognition model provided by an embodiment of the present application;

[0087] Figure 8 It is a schematic flowchart of a method for searching for a first expression graph in an expression library provided by an embodiment of the present application;

[0088] Figure 9 It is a schematic diagram of the principle for obtaining the search matching degree between a search text and a target expression graph provided by an embodiment of the present application;

[0089] Figure 10 It is a schematic diagram of the principle for constructing an antonym sample pair provided by an embodiment of the present application;

[0090] Figure 11 It is a schematic framework diagram of a data processing system provided by an embodiment of the present application;

[0091] Figure 12 It is a schematic structural diagram of a data search device provided by an embodiment of the present application;

[0092] Figure 13 It is a schematic structural diagram of a computer device provided by an embodiment of the present application. Detailed implementation manners

[0093] Next, the technical solutions in the present application will be clearly and completely described in conjunction with the accompanying drawings in the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without making creative efforts belong to the scope of protection of the present application.

[0094] All data collected in the present application (such as relevant data such as search texts and text labels of expression graphs) are collected with the consent and authorization of the owner of the data (such as users, institutions or enterprises), and the collection, use and processing of relevant data need to comply with relevant laws, regulations and standards in relevant regions.

[0095] Here, relevant technical concepts involved in the present application are described:

[0096] Expression template graph: An expression image obtained by processing an original expression picture or a gif (a kind of image file format) animated picture, which does not contain explicit text and only contains the main image, and can be synthesized with any text to produce new expressions with different meanings.

[0097] Synthetic expression graph (which can be abbreviated as synthetic expression): A new expression created by superimposing any text on an expression template graph, which can expand the usage range of existing expressions and improve the accuracy and interestingness of the expressions presented to users.

[0098] AIGC: That is, AI (Artificial Intelligence) generated content. In this application, it can specifically refer to the technology of generating images through text, including generating static pictures through text, generating dynamic gif pictures or videos through text, and the generated content can be used as an expression template graph.

[0099] Please refer to Figure 1 , Figure 1 FIG. is a schematic structural diagram of a network architecture for searching expression graphs provided by an embodiment of the present application. As Figure 1 shown, the network architecture may include a server 200 and a cluster of terminal devices. The cluster of terminal devices may include one or more terminal devices, and the number of terminal devices will not be limited here. As Figure 1 shown, the multiple terminal devices may specifically include terminal device 1, terminal device 2, terminal device 3,..., terminal device n; As Figure 1 shown, terminal device 1, terminal device 2, terminal device 3,..., terminal device n can all be network-connected to the server 200, so that each terminal device can perform data interaction with the server 200 through the network connection.

[0100] As Figure 1 shown, the server 200 can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms. The terminal device can be: a smart phone, a tablet computer, a notebook computer, a desktop computer, a smart TV, a vehicle-mounted terminal, a smart home and other intelligent terminals. Here, taking the communication between terminal device 1 and the server 200 as an example, the specific description of the embodiment of the present application will be carried out.

[0101] The terminal device 1 may include a communication client, and in this communication client, it supports users to communicate through emoticons. The server 200 may be the background server of this communication client. When a user chats with a friend in the communication client, the user can input a search text for triggering the search for emoticons. The communication client may send the search text input by the user to the server 200, so that the server 200 can search for emoticons matching the search text in the background. The server may return the searched emoticons to the communication client, so that the communication client can display the received emoticons in the chat interface between the user and the friend for the user to view and select. Among them, the process of the server 200 searching for emoticons matching the search text can be referred to the following description.

[0102] Please refer to Figure 2 , Figure 2 which is a schematic diagram of a scenario for searching for emoticons matching a search text provided by an embodiment of the present application. As Figure 2 shown, the above Figure 1 server 200 in can perform intent recognition on the search text from N intent understanding dimensions to generate N intent texts (including intent text 1 to intent text N) of the search text in these N intent understanding dimensions.

[0103] The server 200 can also obtain an emoticon library, and this emoticon library may include multiple emoticons. Each emoticon in the emoticon library has a text label, and the text label of each emoticon can include intent labels of each emoticon in N intent understanding dimensions respectively.

[0104] Therefore, the server 200 can search for emoticons matching the search text in the emoticon library by the similarity between the N intent texts of the search text and the text labels of the emoticons in the emoticon library. Among them, the more similar the text label of an emoticon is to the intent text of the search text, the more matching the emoticon is to the search text. This specific process can also be referred to the relevant descriptions in the following embodiments.

[0105] By using the method provided by the present application, fine-grained semantic understanding can be performed on the search text provided by the user from N intent understanding dimensions to obtain N intent texts of the search text in these N intent understanding dimensions. Thus, through the similarity between the N intent texts and the text labels of the emoticons in the emoticon library, emoticons matching the search text can be searched more finely grained and accurately from these N intent understanding dimensions.

[0106] Please refer to Figure 3 , Figure 3It is a schematic flowchart of a data search method provided by an embodiment of the present application. The execution subject in the embodiment of the present application may be a data search device (which may be abbreviated as a search device). The search device may be a computer device or a computer device cluster composed of multiple computer devices. The computer device may be a server or other devices, and the present application does not limit this. As Figure 3 shown, the method may include:

[0107] Step S101: Obtain a search text, where the search text is used to trigger a search for emoticons.

[0108] Specifically, the search device may obtain a search text, which is used to trigger a search for emoticons. The search text may be text input by a user (which may be referred to as an object) for searching emoticons. For example, the search text may be input by the user in a communication client, specifically, text input in the chat interface of the communication client for searching emoticons for chat conversations. The communication client may be any client that supports communication through emoticons.

[0109] Among them, the search device may be a background device (such as a background server) of the communication client. The search text obtained by the search device may be sent by the communication client to the search device. For example, the communication client may send an emoticon search request to the search device, and the emoticon search request may carry the search text. The search device may search for emoticons through the search text carried by the emoticon search request.

[0110] Exemplarily, if the above search text is "Hello", then the search text is used to search for emoticons related to the text "Hello"; again, if the search text is "Good night", then the search text is used to search for emoticons related to the text "Good night"; also, if the search text is "Hahaha", then the search text is used to search for emoticons related to the text "Hahaha".

[0111] Step S102: Perform intent recognition on the search text from N intent understanding dimensions to generate N intent texts of the search text. One intent text corresponds to one intent understanding dimension, and N is a positive integer.

[0112] Specifically, the search device can perform intent recognition on the search text from N intent understanding dimensions to generate N intent texts for the search text. Performing intent recognition on the search text from one intent understanding dimension can generate a corresponding intent text. Therefore, one intent text corresponds to one intent understanding dimension, and one intent text is used to reflect the search intent of the search text in the intent understanding dimension corresponding to the intent text. This search intent is the intent to search for emoticons, such as the intent to search for emoticons expressing the emotion of "happy". N is a positive integer, and the specific value of N can be determined according to the actual application scenario.

[0113] The N intent understanding dimensions can be N preset dimensions for the search intent of emoticons, and the N intent understanding dimensions can be flexibly set according to the actual search requirements for emoticons. This application does not limit the specific content of the N intent understanding dimensions. Here, the N intent understanding dimensions are described exemplarily. For example, the N intent understanding dimensions can include at least one of the following: the understanding dimension of keywords, the understanding dimension of emotions, and the understanding dimension of buzzwords.

[0114] Among them, the above-mentioned understanding dimension of keywords is the dimension that needs to understand the keywords in the search text. Therefore, the intent text generated by performing intent recognition on the search text from the understanding dimension of keywords can be the keywords extracted from the search text, and this intent text can include at least one keyword extracted from the search text. For example, the search text can be "postures for paying New Year greetings", and the keyword extracted from this search text can be "paying New Year greetings".

[0115] The above-mentioned understanding dimension of emotions is the dimension that needs to understand the type of emotion expressed by the search text. Therefore, the intent text generated by performing intent recognition on the search text from the understanding dimension of emotions can be an emotion word used to reflect the type of emotion expressed by the search text. For example, if the search text is "angry", then the intent text in the understanding dimension of emotions can be the emotion word "angry"; another example, if the search text is "bursting out laughing", then the intent text in the understanding dimension of emotions can be the emotion word "happy".

[0116] The above-mentioned understanding dimension of buzzwords is the dimension that needs to understand the buzzwords associated with the search text. Therefore, the intent text generated by performing intent recognition on the search text from the understanding dimension of buzzwords can be the buzzword represented (or associated) by the search text. For example, the buzzword can be a popular term on the Internet. For example, if the search text is "so tired from work", then the intent text in the understanding dimension of buzzwords can be the buzzword "work fatigue".

[0117] In one implementation, the search device can perform intent recognition on the search text from N intent understanding dimensions to generate N intent texts of the search text, which may include:

[0118] The search device can obtain a trained intent recognition model. The trained intent recognition model can be a trained model capable of recognizing N search intents of the search text from the above N intent understanding dimensions. For the specific training process of this intent recognition model, please refer to the relevant descriptions in the following Figure 6 corresponding embodiments. The trained intent recognition model in this application may include a text encoder and N text decoders. One text decoder corresponds to one intent understanding dimension, and one text decoder can be used to perform intent recognition on the search text in the corresponding intent understanding dimension. Therefore, these N text decoders can perform intent recognition on the search text in the above N intent understanding dimensions.

[0119] The search device can input the search text into the trained intent recognition model to call the text encoder in the trained intent recognition model to perform feature encoding processing on the search text, thereby generating the encoded feature of the search text. The encoded feature can be the feature vector encoded by the text encoder for the search text.

[0120] The search device can call the N text decoders in the trained intent recognition model to perform feature decoding processing on the above encoded features respectively in the corresponding intent understanding dimensions to generate the above N intent texts. Among them, one text decoder can perform feature decoding processing on the encoded feature of the search text in the corresponding intent understanding dimension to generate one intent text of the search text.

[0121] In one implementation, a generation range can be set for each of the above text decoders respectively. One text decoder can decode and generate the intent text of the search text within the set generation range. For example, for the text decoder corresponding to the understanding dimension of the above keywords, the set generation range can include multiple keywords available for selection and generation, such as each keyword in the thesaurus. For the text decoder corresponding to the understanding dimension of the above emotions, the set generation range can include multiple emotion words available for selection and generation. For the text decoder corresponding to the understanding dimension of the above buzzwords, the set generation range can include multiple buzzwords available for selection and generation.

[0122] Moreover, the elements within the generation range corresponding to each text decoder in this application support dynamic addition and deletion, which not only increases the interpretability of the intent text generated by the text decoder, but also makes the content of the intent text generated by the text decoder controllable and modifiable.

[0123] Among them, the intent text of the search text in certain intent understanding dimensions can be empty (i.e., "null"), that is, in the present application, the intent text of the search text in various intent understanding dimensions can have the option of "empty". That is to say, when the above-mentioned text decoder performs feature decoding processing to generate intent text, there can be an option of generating "empty". For example, for a search text that does not express buzzwords, its intent text in the above-mentioned buzzword understanding dimension can be empty.

[0124] By decoupling various intent understanding dimensions in the present application, with one intent understanding dimension corresponding to one text decoder, the flexibility of maintaining, adding, and deleting various intent understanding dimensions can be improved. For example, when a certain intent understanding dimension is not needed, the text encoder corresponding to that intent understanding dimension can be directly removed.

[0125] The present application models the matching process from the user Query (i.e., the search text) to the emoticon as a multi-dimensional and open short text generation task (such as the task of generating intent text), and can implement this short text generation task through an intent recognition model including a multi-task text decoder to generate multiple intent texts of the search text (i.e., the above-mentioned N intent texts). This intent recognition model belongs to a generation model from sequence (such as the text sequence of the search text) to sequence (the text sequence of the intent text).

[0126] Step S103: Obtain the respective text labels of each emoticon in the emoticon library. The text label of any emoticon includes the intent labels of the any emoticon in N intent understanding dimensions respectively.

[0127] Specifically, the emoticon library can contain many emoticons available for selection and search. The search device can obtain the respective text labels of each emoticon in the emoticon library. Any emoticon can be set with multiple text labels, and the present application can implement the search for emoticons through the text labels of the emoticons. In the present application, the text label of any emoticon can include the intent labels of the any emoticon in the above-mentioned N intent understanding dimensions respectively, that is, the text label of any emoticon can include N intent labels of the any emoticon in the N intent understanding dimensions. An emoticon can have one intent label in one intent understanding dimension. The N intent labels of the emoticon can all correspond to the N intent understanding dimensions for intent recognition of the search text.

[0128] Among them, the intent label of the emoticon in certain intent understanding dimensions can also be empty "null", that is, in the present application, the intent labels of the emoticon in various intent understanding dimensions can have the option of "empty". For example, for an emoticon that has no associated buzzwords, its intent label in the above-mentioned buzzword understanding dimension can be empty.

[0129] For example, for an emoticon of an angry expression, its intention label in the above-mentioned dimension of emotion understanding can be the emotion word "angry"; for an emoticon of being tired from work, its intention label in the above-mentioned dimension of catchphrases understanding can be the catchphrase "work flavor"; for an emoticon containing the text "Hello", its intention label in the above-mentioned dimension of keywords understanding can be the text "Hello" it contains, or in other scenarios, it can also be the keyword extracted from the text contained in the emoticon.

[0130] Furthermore, the text labels of the emoticons in the emoticon library can also include general labels, that is, the text label of any emoticon can include the general label of that any emoticon, and this general label can be a text label related to the expression attribute that the emoticon itself possesses. For example, the general label of any emoticon can include the expression type label of that any emoticon (used to describe the expression type of the emoticon, such as the expression type label can be the text "greeting type", the text "apology type", the text "farewell type", etc.), the label of the text contained in the emoticon (used to describe the text contained in the emoticon, such as "Hello"), the emoticon pack to which the emoticon belongs (used to indicate which set of emoticon packs the emoticon belongs to, such as "cute little yellow duck"), and so on. Specifically which text labels are included in the general label of the emoticon can be set according to actual application requirements.

[0131] Step S104, search for the first emoticon that matches the search text in the emoticon library based on the similarity between the N intention texts and the text labels of the emoticons in the emoticon library.

[0132] Specifically, the search device can search for the emoticon that matches the search text in the emoticon library based on the similarity between the N intention texts and the text labels of the emoticons in the emoticon library, such as through the text similarity between the N intention texts and the text labels of the emoticons in the emoticon library. The emoticon found to match the search text here can be called the first emoticon. The text labels of the first emoticon found are similar to the N intention texts, and there can be multiple first emoticons.

[0133] Among them, for the specific process of how to search for the first emoticon in the emoticon library, reference can be made to the description in the following Figure 8 corresponding embodiment. In this application, the search matching degree between each first emoticon found and the search text (the calculation process can be referred to in the following Figure 8(corresponding to the relevant descriptions in the embodiments), sort and display each first expression graph, that is, when displaying the searched first expression graphs in the communication client, the first expression graphs with higher search matching degrees with the search text can be displayed more in the front, and conversely, the first expression graphs with lower search matching degrees with the search text can be displayed more in the back.

[0134] In one implementation manner, the search device can also perform expression expansion processing on the above search text to obtain expanded expression graphs, and the expanded expression graphs can be called second expression graphs. This application can also return the expanded second expression graphs to the communication client for display. Among them, when displaying expression graphs in the communication client, the first expression graphs can be displayed entirely in front of the second expression graphs, or the first expression graphs and the second expression graphs can also be displayed in different regions on the client interface (such as the chat interface). The display order among the second expression graphs can be random or other set orders, which can be specifically determined according to the actual application scenario.

[0135] The above search text can be sent by the communication client to the search device. Therefore, the search device can return both the searched first expression graphs and the expanded second expression graphs to the communication client, so that the communication client can output the received first expression graphs and second expression graphs for the user to view and select.

[0136] Exemplarily, the process of performing expression expansion processing on the search text to obtain expanded expression graphs can include: the search device can obtain a trained language rewriting model, and the trained language rewriting model can be a trained model capable of rewriting the text used for expression graph search into a text more specifically describing the specific style of the expression graph to be generated. For example, the trained language rewriting model can be a trained large language model (LLM model).

[0137] Therefore, the search device can call the trained language rewriting model to perform language rewriting processing on the search text to generate an expression description text corresponding to the search text. The expression description text is to rewrite the search text into a description text of the specific expression graph it specifically wants to search for. For example, the search text can be "My waist hurts so much tonight", and the expression description text corresponding to the search text can be "The expression of a cartoon kitten with its paw pressing on the waist and lying on the ground crying".

[0138] In one implementation, the process of training the trained language rewriting model may include: The search device can obtain the third sample text and the language rewriting model to be trained. The third sample text can be a sample for training the language rewriting model to be trained. The third sample text can be the text that the user's search emoticon map may input during the training phase. The nature of the third sample text is similar to the nature of the above-mentioned search text, that is, the third sample text can be the text for searching emoticon maps.

[0139] The third sample text can have a reference rewritten text. The reference rewritten text can be the description text that ideally rewrites the third sample text into the specific emoticon map to be generated. That is, we hope that the language rewriting model can rewrite the third sample text into its reference rewritten text. The reference rewritten text can be understood as the sample label of the third sample text.

[0140] Therefore, the search device can call the language rewriting model to be trained to perform language rewriting processing on the third sample text, and can generate a sample rewritten text corresponding to the third sample text. The sample rewritten text is the text for specifically describing the emoticon map to be generated obtained by rewriting through the language rewriting model to be trained.

[0141] The search device can correct the model parameters of the language rewriting model to be trained through the difference between the sample rewritten text and the reference rewritten text to obtain the trained language rewriting model. For example, the search device can generate the generation loss (i.e., the loss function) of the language rewriting model to be trained for the sample rewritten text through the sample rewritten text and the reference rewritten text. For example, the generation loss can be the cross-entropy loss between the sample rewritten text and the reference rewritten text. The generation loss can reflect the deviation of the language rewriting model to be trained in performing language rewriting processing on the third sample text. The generation loss can reflect the difference between the sample rewritten text and the reference rewritten text. The larger the generation loss, the greater the difference between the sample rewritten text and the reference rewritten text. On the contrary, the smaller the generation loss, the smaller the difference between the sample rewritten text and the reference rewritten text. The search device can correct the model parameters of the language rewriting model to be trained through the generation loss to obtain the trained language rewriting model.

[0142] The objective of modifying the model parameters of the language rewriting model to be trained through the generation loss is to modify the model parameters of the language rewriting model to be trained so that the generation loss tends to a minimum value (such as tending to 0). There can be multiple third sample texts as described above, and each third sample text can have its own reference rewritten text. The search device can, according to the above principle, perform multiple rounds of iterative training on the language rewriting model to be trained through these multiple third sample texts. When the training of the language rewriting model to be trained is completed, the trained language rewriting model can be obtained. Exemplarily, the completion of the training of the language rewriting model to be trained can mean: modifying the model parameters of the language rewriting model to be trained to a convergent state, or the number of rounds of iterative training of the language rewriting model to be trained being equal to the set round threshold.

[0143] The search device can also obtain a text-to-image model, which can be a trained model that can generate an image that conforms to the text content described by the input text through the input text. Therefore, the search device can call the text-to-image model to generate L expression images that conform to the text content described by the expression description text corresponding to the search text through the search text. These L expression images can be referred to as L candidate expression images. L is a positive integer and is the number of images that the text-to-image model can generate at one time. The value of L can be the default setting of the text-to-image model, or the value of L can also be set for the text-to-image model according to the actual application scenario.

[0144] Through the above process, it can be understood that since the text-to-image model usually cannot intuitively understand the generation requirements of the user's search text for the expression images, therefore, in this application, the search text can be rewritten as an expression description text, and then the text-to-image model can be called to accurately generate the expression images required by the user through the rewritten expression description text.

[0145] The search device can obtain a text-image matching model, which can be a trained model that can predict the relevance (which can be called text-image relevance or text-image matching degree) between the input text and the input image. The higher the text-image relevance between the input text and the image, the closer the text content described by the text is to the image content in the image (that is, the more relevant or the more matching). Therefore, the search device can call the text-image matching model to predict the text-image relevance between each candidate expression image and the above expression description text respectively. There can be a text-image relevance between a candidate expression image and the expression description text.

[0146] In one implementation, the search device may sort the above-mentioned L candidate emoji images in descending order of the graphic-text relevance between each candidate emoji image and the emoji description text, and the sorted L candidate emoji images can be obtained. Among the sorted L candidate emoji images, the candidate emoji image with a greater graphic-text relevance to the emoji description text can be arranged in the front, and conversely, the candidate emoji image with a smaller graphic-text relevance to the emoji description text can be arranged in the back.

[0147] The search device may generate an extended second emoji image through the first P candidate emoji images arranged in the front among the sorted L candidate emoji images. For example, the P candidate emoji images can be used as P emoji template images, and the search device may perform a synthesis process on the search text and each of the P emoji template images respectively to generate P extended second emoji images, and the P extended second emoji images are the new emoji images generated by the text-to-image model. P is a positive integer and P is less than L, and the value of P can be set according to the actual application scenario.

[0148] Alternatively, the search device may also generate an extended second emoji image through the candidate emoji images among the above-mentioned L candidate emoji images whose graphic-text relevance to the emoji description text is greater than or equal to a relevance threshold (which can be preset). For example, the search device may also perform a synthesis process on the search text and each of the candidate emoji images whose graphic-text relevance to the emoji description text is greater than or equal to the relevance threshold to generate an extended second emoji image.

[0149] This application generates new emoji images through a text-to-image model, which can improve the richness, interestingness, and adaptability of the searched emoji images to the user's query, enabling the user to view a more diverse range of emoji images and enhancing the user experience of emoji search and use.

[0150] Furthermore, the process of performing an emoji expansion process on the search text to obtain an extended second emoji image may further include:

[0151] The search device may obtain general emoji images, which can be general emoji template images that do not look out of place when synthesized with any text. There can be multiple such general emoji images, and they can be pre-configured. Therefore, the search device may perform a synthesis process on the search text and the general emoji images to generate extended second emoji images. Performing a synthesis process on the search text and one general emoji image can generate a corresponding second emoji image. Among them, performing a synthesis process on the search text and the general emoji image may mean adding the search text to the general emoji image.

[0152] Exemplarily, when adding search text to a general emoji, the specific style of the added search text (such as font, text size, text color, text stroke, etc.) can be random or pre-default set.

[0153] In one implementation, if the emoji synthesized from the general emoji returned to the user is clicked and used by the user, then subsequently, when other users also recall the emoji synthesized from the general emoji and also recall the emoji synthesized from the emoji generated by the text-to-image model, the emoji synthesized from the general emoji can be displayed in front of the emoji synthesized from the emoji generated by the text-to-image model.

[0154] Among them, the above emoji library can include two types, one is an emoji template library and the other is a standard emoji library. The standard emoji library can include multiple standard emojis, which can be directly used emojis, and explicit text can be added to the standard emojis. The emoji template library can include multiple emoji template images, which can be emoji template images that are not directly used and do not have explicit text added. The emoji template images can be synthesized with other texts (i.e., words) to generate standard emojis. This application can, according to the same principle above, search for the above first emoji from the emoji template library and the standard emoji library respectively. That is, the first emoji can include the emoji template image searched from the emoji template library that matches the search text, or can also include the emoji searched from the standard emoji library that matches the search text.

[0155] Therefore, the first emoji can include an emoji template image, which can be the emoji template image searched from the emoji template library that matches the search text. The process of the search device returning the first emoji to the communication client can include: the search device can perform synthesis processing on the search text and the emoji template image to generate a synthesized emoji, which is the emoji obtained after synthesizing (such as adding) the search text to the emoji template image. The search device can return the synthesized emoji to the communication client.

[0156] For the first emoji searched from the standard emoji library that matches the search text, it can be directly returned to the communication client without being synthesized with the search text. It is very likely that the first emoji already contains the search text or other texts semantically similar to the search text.

[0157] Among them, in some actual application scenarios, when users search for emoticons, it is necessary to ensure the real-time nature of their search. However, there may be a delay when generating the augmented second emoticons through the above text-to-image model. Therefore, the present application can also use the method of synthesizing the above general emoticons and the search text as a fallback solution, and use the synthesized emoticons obtained by synthesizing the general emoticons and the search text as the augmented second emoticons. The emoticons generated by the text-to-image model (such as the candidate emoticons selected from L candidate emoticons through the above sorting or relevance threshold) can be supplemented to the above emoticon template library for subsequent use when other users input the same or similar text as the above search text to search for emoticons. Moreover, when adding the emoticons generated by the text-to-image model to the emoticon template library, corresponding text tags can also be added to the emoticons, and the text tags can include the intent tags of the emoticons in N intent understanding dimensions. Exemplarily, the intent tags to be added to the emoticons generated by the text-to-image model can be generated by a large language model (the large language model here can be different from the above language rewriting model) through language understanding of the above emoticon description text from N intent understanding dimensions.

[0158] Please refer to Figure 4 , Figure 4 is a schematic diagram of the effect of a chat interface provided by an embodiment of the present application. Figure 4 The shown chat interface can be a chat interface between the user and the friend "Xiaoming". The search text input in this chat interface can be "crying", and the emoticons matching the search text searched by the method provided by the present application can be displayed in the interface at the bottom part of the chat interface.

[0159] The method of the present application can make the acquisition of the emoticon library easy and its expansion easy: The present application can make full use of the emoticons (such as emoticon template images) in the existing emoticon library, and can automatically retrieve the corresponding emoticon template images according to the text semantics of the given input text (i.e., the search text) (which can be reflected by the intent text), and can synthesize more emoticons through the retrieved emoticon template images, without the need to produce a large number of emoticons through manual creation or human intervention, greatly reducing the difficulty of emoticon acquisition and maintenance. In addition, the method provided by the present application has strong interpretability and is easy to control for the searched emoticons: The present application realizes the decoupling of multi-dimensional semantic understanding (i.e., the intent recognition of multiple intent understanding dimensions), which is convenient for adding or deleting intent understanding dimensions according to actual needs, and can recall emoticons for users from multiple channels according to these intent understanding dimensions, making it convenient to control the sources and respective display orders of the recall results of each channel, and having high flexibility in controlling the recall process and recall results of multiple-channel emoticons.

[0160] By adopting the method provided in the present application, multi-dimensional intent recognition can be performed on the search text used to search for emoticons from multiple intent understanding dimensions, and each emoticon in the emoticon can be respectively configured with its own intent label on the multiple intent understanding dimensions. Thus, by comparing the similarity between multiple intent texts generated by intent recognition of the search text and the text labels (including intent labels) of the emoticons in the emoticon library, a multi-dimensional search for emoticons matching the search text from the multiple intent understanding dimensions can be achieved, thereby improving the accuracy of searching the emoticon library for emoticons matching the search text.

[0161] See also Figure 5 , Figure 5 is a schematic diagram of a process of searching and obtaining emoticons provided in an embodiment of the present application. Figure 5 As shown, the process may include:

[0162] S1: The search device can obtain the search text.

[0163] S2: The search device may parse the search text (the parsing may be intent recognition) to obtain fine-grained semantics of the search text. The fine-grained semantics may be the N intent texts generated by intent recognition of the search text.

[0164] S3: The search device can search for the expression in the expression library (the expression library here can be the expression template library mentioned above) through the fine-grained semantics. Figure 1 In one embodiment, the emoticon image can be searched by a preset relevance threshold (also called a matching threshold). If the relevance between the retrieved emoticon image and the search text (which can be the following) is Figure 8 If the search matching degree in the corresponding embodiment is greater than or equal to the correlation threshold, it indicates that the currently retrieved emoticon image matches the search text, and the following step S4 can be executed; on the contrary, if the correlation between the retrieved emoticon image and the search text is less than the correlation threshold, it indicates that the currently retrieved emoticon image does not match the search text, and the following steps S6 and S8 can be executed. Among them, here it can also be that when the number of emoticons in the emoticon image library whose correlation with the search text is greater than or equal to the correlation threshold is less than or equal to the set number threshold, the following steps S6 and S8 are executed.

[0165] S4: The search device may synthesize the search text and the currently retrieved expression template image to generate a synthesized expression image. The synthesized expression image may be obtained by adding the search text to the expression template image.

[0166] S5: The search device may use the synthesized emoji image as the search result and return it to the front end (such as a communication client) for display, so that users can view and select it.

[0167] S6: The search device may generate an image generation prompt based on the search text. For example, the search device may perform language rewriting on the search text to generate the image generation prompt, which may be the emoji description text corresponding to the above search text.

[0168] S7: The search device may input the image generation prompt into the text-to-image model to call the text-to-image model to generate an AIGC emoji template image. The AIGC emoji template image may be an emoji image selected from the above L candidate emoji images generated by the text-to-image model, such as the above P candidate emoji images selected, or a candidate emoji image with a text-image relevance greater than or equal to the relevance threshold between the selected image and the image generation prompt. The relevance threshold here may be different from the relevance threshold in the above step S3. Since it is necessary to ensure the real-time performance of the user's emoji image search, and the method of generating emoji images by AIGC may have delays, therefore, the above steps S6 and this step S7 may be executed in the background, and the AIGC emoji template image generated in this step S7 may be supplemented and added to the emoji library for use when performing emoji image search again later.

[0169] S8: The search device may obtain a universal template image, which may be the above general emoji image, and there may be no explicit text in the general emoji image.

[0170] S9: The search device may perform synthesis processing on the search text and the universal template image to generate a synthesized emoji image, which may be obtained by adding the search text to the universal template image. The synthesized emoji image may also be used as the search result and returned to the front end (such as a communication client) for display, so that users can view and select it.

[0171] Since the speed of performing synthesis processing on the search text and the universal template image is extremely fast and the delay during this process can be ignored, therefore, the method of performing synthesis processing on the search text and the universal template image to obtain a synthesized emoji image can be used as a fallback solution and executed online.

[0172] Please refer to Figure 6 , Figure 6 which is a schematic flowchart of a method for training an intent recognition model provided by an embodiment of the present application. As Figure 6 shown, the method may include:

[0173] Step S201: Obtain the first sample text and the intent recognition model to be trained. The intent recognition model to be trained includes a text encoder to be trained and N text decoders to be trained. One text decoder to be trained corresponds to one intent understanding dimension, and the first sample text has corresponding reference intent texts in N intent understanding dimensions.

[0174] Specifically, the search device can obtain the first sample text and the intent recognition model to be trained. Similar to the above-mentioned trained intent recognition model, the intent recognition model to be trained can include a text encoder to be trained and N text decoders to be trained. One text decoder to be trained can correspond to one intent understanding dimension, that is, one text decoder to be trained can be used to perform intent recognition on the input text in the corresponding intent understanding dimension.

[0175] Among them, the first sample text can have corresponding reference intent texts in the above-mentioned N intent understanding dimensions. The first sample text can have a corresponding reference intent text in one intent understanding dimension. The reference intent text is the ideal intent text that is expected to be decoded by the text decoder for the first sample text. That is, the reference intent text of the first sample text in one intent understanding dimension can be the ideal intent text that is expected to be decoded by the text decoder corresponding to this intent understanding dimension for the first sample text. The reference intent texts of the first sample text in N intent understanding dimensions can be understood as the sample labels carried by the first sample text.

[0176] Step S202: Invoke the text encoder to be trained to perform feature encoding processing on the first sample text to generate the sample encoding features of the first sample text.

[0177] Specifically, the search device can invoke the above-mentioned text encoder to be trained to perform feature encoding processing on the first sample text, and the encoding features of the first sample text can be generated. The encoding features of the first sample text can be called sample encoding features. The sample encoding features can be the feature vectors generated by the text encoder to be trained for embedding processing (also called feature extraction processing) on the first sample text.

[0178] Step S203: Invoke N text decoders to be trained to perform feature decoding processing on the sample encoding features in the corresponding intent understanding dimensions respectively to generate N sample intent texts of the first sample text.

[0179] Specifically, the search device can invoke the above-mentioned N text decoders to perform feature decoding processing on the sample encoding features of the first sample text in the corresponding intent understanding dimensions respectively, and N sample intent texts of the first sample text can be generated.

[0180] Among them, by invoking a text encoder to be trained to perform feature decoding processing on the sample encoding features of the first sample text in the corresponding intention understanding dimension, a sample intention text of the first sample text in this intention understanding dimension can be generated. That is, a sample intention text corresponds to one intention understanding dimension.

[0181] Step S204: Based on the differences between the sample intention texts and the reference intention texts in the respective corresponding intention understanding dimensions of each text decoder to be trained, generate the feature decoding losses of each text decoder to be trained for the sample encoding features respectively.

[0182] Specifically, the search device can generate the feature decoding losses (i.e., loss functions) of each text decoder to be trained for the sample encoding features respectively through the differences between the sample intention texts and the reference intention texts in the respective corresponding intention understanding dimensions of each text decoder to be trained. One text decoder to be trained corresponds to one feature decoding loss. For example, the feature decoding loss can be the cross-entropy loss between the sample intention text and the reference intention text.

[0183] Among them, through the differences between the sample intention texts and the reference intention texts in the respective corresponding intention understanding dimensions of a text decoder to be trained, the feature decoding loss of this text decoder to be trained for the sample encoding features can be generated. This feature decoding loss reflects the deviation of this text decoder to be trained in performing feature decoding processing on the sample encoding features. If the feature decoding loss is larger, it indicates that the deviation of this text decoder to be trained in performing feature decoding processing on the sample encoding features is larger, and the difference between the sample intention text and the reference intention text in the corresponding intention understanding dimension of this text decoder to be trained is larger; on the contrary, the smaller the feature decoding loss, it indicates that the deviation of this text decoder to be trained in performing feature decoding processing on the sample encoding features is smaller, and the difference between the sample intention text and the reference intention text in the corresponding intention understanding dimension of this text decoder to be trained is smaller.

[0184] Step S205: Based on the feature decoding losses of the N text decoders to be trained respectively, correct the model parameters of the intention recognition model to be trained to obtain the trained intention recognition model.

[0185] Specifically, the search device can use the N feature decoding losses of the above N text decoders to be trained to correct the model parameters of the intention recognition model to be trained, including correcting the model parameters of the above text encoder to be trained and the model parameters of the N text decoders to be trained, so as to obtain the above trained intention recognition model, that is, the trained intention recognition model.

[0186] Based on the above-mentioned feature decoding loss, the present application can further introduce a clustering loss in which a text decoder performs feature decoding on first sample texts that are semantically similar or identical to train the intent recognition model to be trained. This clustering loss can be referred to as the first clustering loss, and this first clustering loss can be used to reflect the difference (or deviation) in feature decoding of the first sample texts that are semantically similar or identical by the text decoder. The larger this first clustering loss is, the greater the difference in feature decoding of the first sample texts that are semantically similar or identical by the text decoder; the smaller this first clustering loss is, the smaller the difference in feature decoding of the first sample texts that are semantically similar or identical by the text decoder. By introducing this first clustering loss, the present application can make the text decoder more unified when performing feature decoding on the first sample texts that are semantically similar or identical, thereby improving the accuracy of the text decoder in performing feature decoding. The specific process of generating this first clustering loss is described below, as described in the following content.

[0187] There can be multiple such first sample texts, and these multiple first sample texts can have the same reference intent text in N intent understanding dimensions. The first sample texts that have the same reference intent text in N intent understanding dimensions can be regarded as first sample texts that are semantically similar or identical. That is, the reference intent texts of each of these multiple first sample texts in the same intent understanding dimension can be the same. Any one of the above N text decoders to be trained can be referred to as the target text decoder. Since the principle of generating the first clustering loss for each text decoder to be trained is the same, the following will take the principle of generating the first clustering loss of this target text decoder as an example for specific explanation.

[0188] The target text decoder can sequentially include multiple network layers. The network layers included in the target text decoder can be referred to as the first network layers, that is, the target text decoder sequentially includes multiple first network layers, and the input of the subsequent first network layer can be the output of the previous first network layer. During the process of the target text decoder performing feature decoding processing on the sample encoding features of the first sample text, the first network layers among these multiple first network layers except the last first network layer (which can be called the last first network layer) can generate vectors for performing feature decoding processing on the sample encoding features of the first sample text (that is, intermediate vectors, which can be referred to as decoding features), and the last first network layer can decode and generate the sample intent text of the first sample text through the vector generated by the second-to-last first network layer. That is, the last first network layer finally generates the sample intent text, rather than a vector.

[0189] Therefore, during the process of the target text decoder performing feature encoding processing on the sample encoding features of each first sample text, the search device can obtain the decoding features generated by each first network layer except the last first network layer in the multiple first network layers for each first sample text. One first network layer can generate one decoding feature for one first sample text.

[0190] The search device can generate a first clustering loss of the target text decoder for the multiple first sample texts through the decoding features of the multiple first sample texts that are semantically similar or the same. This first clustering loss can be used to reflect the differences between the decoding features generated by the same first network layer for the multiple first sample texts. In this application, by using the decoding features generated by the multiple first network layers included in the target text decoder for the first sample texts to generate the first clustering loss of the target text decoder, it can make each first network layer in the target text decoder tend to be consistent and unified during the process of performing feature decoding processing on the first sample texts that are semantically the same or similar. That is, each first network layer in the target text decoder generates decoding features that are consistent and unified for each first sample text during the process of performing feature decoding processing on the first sample texts that are semantically the same or similar. This can further improve the accuracy of training the intent recognition model to be trained and enhance the unity and consistency of the intent recognition model to be trained in processing texts that are semantically similar or the same. Among them, the formula principle of this first clustering loss can be the same as that of the following second clustering loss, except that when obtaining this first clustering loss, the generation result of the last first network layer is not required.

[0191] The search device can correct the model parameters of the intent recognition model to be trained through the respective first clustering losses and feature decoding losses of the above-mentioned N text decoders to be trained. For example, the search device can perform an addition process (i.e., summation) on the N first clustering losses and N feature decoding losses of the N text decoders to be trained to obtain the overall loss of the intent recognition model to be trained, and can correct the model parameters of the intent recognition model to be trained through this overall loss to obtain the trained intent recognition model. Among them, the goal of correcting the model parameters of the intent recognition model to be trained through this overall loss is to correct the model parameters of the intent recognition model to be trained so that this overall loss tends to the minimum value (such as tending to 0).

[0192] The above-mentioned multiple first sample texts can be a batch of first sample texts for one round of training of the intention recognition model to be trained. According to the principle described above, this application can perform multiple rounds of iterative training on the intention recognition model to be trained through multiple batches of first sample texts. Among them, in order to obtain the above-mentioned first clustering loss, the first sample texts in the same batch used for training the intention recognition model to be trained can be semantically similar or the same, that is, the reference intention texts of the first sample texts in the same batch in N intention understanding dimensions can be the same. And the semantics between different batches of first sample texts can be different, that is, the reference intention texts of different batches of first sample texts in N intention understanding dimensions can be different.

[0193] Exemplarily, when the model parameters of the intention recognition model to be trained are trained to a converged state, or the number of rounds of iterative training on the intention recognition model to be trained is equal to the specified round threshold, it can be considered that the training of the intention recognition model to be trained is completed, and the intention recognition model obtained at this time can be used as the trained intention recognition model.

[0194] Through the above process, this application can perform end-to-end training on the intention recognition model for multiple generation tasks (i.e., multi-intention recognition tasks). Each text decoder can learn a level of understanding (such as a level of an intention understanding dimension) of the user Query, and the levels of various intention understanding dimensions are decoupled from each other. Therefore, the intention recognition model proposed in this application is a model structure that saves computational costs, has strong interpretability, and is pluggable and scalable (such as). For example, a text decoder corresponding to a certain intention understanding dimension that is not required subsequently can be flexibly deleted from the intention recognition model, or more text decoders corresponding to intention understanding dimensions can be added.

[0195] Please refer to Figure 7 , Figure 7 which is a schematic structural diagram of an intention recognition model provided by an embodiment of this application. As Figure 7 shown, the intention recognition model can include a text encoder (Encoder) and N text decoders (Decoder). Here, the N text decoders can include text decoder 1 to text decoder N. One text decoder can correspond to one intention understanding dimension, and the above-mentioned N intention understanding dimensions can include intention understanding dimension 1 to intention understanding dimension N. For example, text decoder 1 can correspond to intention understanding dimension 1, text decoder 2 can correspond to intention understanding dimension 2,... text decoder N can correspond to intention understanding dimension N. Each text decoder can respectively perform intention recognition on the search text in its corresponding intention understanding dimension to generate the intention text in the intention understanding dimension corresponding to the search text.

[0196] Among them, both the text encoder and each text decoder can be composed of multiple model units, and one model unit can be a Transformer block (the basic building block of the model). The input of each text decoder can be the output of the text encoder.

[0197] By adopting the above process of this application, the trained intent recognition model can be accurately obtained. The trained intent recognition model can perform decoupled intent recognition on the input text from N intent understanding dimensions, and thus can recognize the accurate intent text of the input text in these N intent understanding dimensions.

[0198] Please refer to Figure 8 , Figure 8 which is a schematic flowchart of a method for searching for a first emoticon in an emoticon library provided by an embodiment of this application. As Figure 8 shown, the method may include:

[0199] Step S301: Obtain the search matching degrees between the search text and each emoticon in the emoticon library by respectively calculating the similarities between the search text and N intent texts and the text labels of each emoticon in the emoticon library.

[0200] Specifically, the search device may also use the original search text as a text for searching for emoticons. The search device can obtain the search matching degrees between the search text and each emoticon in the emoticon library by respectively calculating the similarities between the search text and its N intent texts and the text labels of each emoticon in the emoticon library. There can be a search matching degree between the search text and an emoticon in the emoticon library.

[0201] Any emoticon in the emoticon library can be referred to as a target emoticon. Since the principle of obtaining the search matching degrees between the search text and each emoticon in the emoticon library is the same, therefore, this application takes the process of obtaining the search matching degree between the search text and this target emoticon as an example for specific description, as described in the following content.

[0202] The search device can obtain the trained feature embedding model. The trained feature embedding model can be a trained model that can embed the input text into its corresponding representation features, and the representation features can be feature vectors generated for the input text. The specific process of training to obtain this trained feature embedding model can be referred to the description of the following process.

[0203] Therefore, the search device can call the trained feature embedding model to perform feature embedding processing on each text label of the search text, the N intent texts, and the target emoticon map respectively, so as to generate the representation features of the search text, the representation features of each intent text, and the representation features of each text label (including the general label and the intent label of the target emoticon map) of the target emoticon map. Among them, one text label of the target emoticon map can correspond to one representation feature, or, for the efficiency of real-time search for emoticon maps, the representation features of each text label of the target emoticon map can also be pre-generated by the trained feature embedding model.

[0204] The search device can obtain M feature similarities between the representation features of the search text and the N intent texts and the representation features of the text labels of the target emoticon map, where M is a positive integer. There can be one feature similarity between the representation feature of one text in the search text and its N intent texts and the representation feature of one text label of the target emoticon map. For example, the feature similarity between two such representation features can be the cosine similarity between the two representation features. Among them, the feature similarity between the representation features of texts can be used to reflect the similarity between texts. For example, the feature similarity between the representation features of texts can be directly used as the text similarity between texts.

[0205] The search device can use the maximum value (i.e., the largest feature similarity) among the M feature similarities as the search matching degree between the search text and the target emoticon map. The process of obtaining the M feature similarities is described in detail below.

[0206] The above-mentioned target emoticon map can have (i.e., be set with) K + N text labels in total. The K + N text labels can include K general labels and N intent labels of the target emoticon map, that is, the target emoticon map can have K general labels, where K is a positive integer.

[0207] The search device can calculate the feature similarities between the representation features of the search text and the representation features of the K+N text labels respectively, and obtain K+N feature similarities. There can be one feature similarity between the representation feature of the search text and the representation feature of a text label of the target emoticon. The search device can also calculate the feature similarities between the representation features of the N intent texts and the representation features of the K general labels of the target emoticon respectively, and obtain N×K feature similarities. There can be one feature similarity between the representation feature of an intent text and the representation feature of a general label of the target emoticon. Moreover, the search device can also calculate the feature similarities between the representation features of the N intent texts and the representation features of the N intent labels of the target emoticon, and obtain N feature similarities. There can be one feature similarity between the representation feature of any intent text and the representation feature of an intent label of the target emoticon in the intent understanding dimension corresponding to the any intent text; that is, when calculating the feature similarity between the representation feature of an intent text and the representation feature of an intent label, it is necessary to calculate the feature similarity between the intent text and the intent label corresponding to the same intent understanding dimension.

[0208] The search device can use the above calculated K+N feature similarities, N×K feature similarities, and N feature similarities as the above M feature similarities, that is, M can be equal to (K+N)+N×K+N.

[0209] The search device can, according to the same principle described above, obtain the search matching degrees between the search text and each emoticon in the emoticon library.

[0210] Please refer to Figure 9 , Figure 9 which is a schematic diagram of the principle for obtaining the search matching degree between the search text and the target emoticon provided by an embodiment of the present application. As Figure 9 shown, here N can be equal to 3 and K can be equal to 4. The N intent texts of the search text can include intent text 1, intent text 2, and intent text 3. The text labels of the target emoticon can include general label 1, general label 2, general label 3, general label 4, intent label 1, intent label 2, and intent label 3.

[0211] Among them, the above N intent understanding dimensions can include intent understanding dimension 1, intent understanding dimension 2, and intent understanding dimension 3. Intent text 1 and intent label 1 can both correspond to intent understanding dimension 1, intent text 2 and intent label 2 can both correspond to intent understanding dimension 2, and intent text 3 and intent label 3 can both correspond to intent understanding dimension 3.

[0212] Therefore, the search device can calculate the feature similarity between the representation features of the search text and the representation features of each text label of the target emoticon, calculate the feature similarity between the representation features of the intent text 1 and the representation features of each general label of the target emoticon, calculate the feature similarity between the representation features of the intent text 2 and the representation features of each general label of the target emoticon, calculate the feature similarity between the representation features of the intent text 3 and the representation features of each general label of the target emoticon, calculate the feature similarity between the representation features of the intent text 1 and the representation features of the intent label 1, calculate the feature similarity between the representation features of the intent text 2 and the representation features of the intent label 2, calculate the feature similarity between the representation features of the intent text 3 and the representation features of the intent label 3, and thus the above M feature similarities can be obtained. Therefore, the maximum feature similarity among the M feature similarities can be used as the search matching degree between the search text and the target emoticon.

[0213] Exemplarily, the specific process of training the above-mentioned trained feature embedding model in this application may include: The search device can obtain a sample pair and a feature embedding model to be trained. The sample pair may include two second sample texts. The sample pair may have a semantic matching label (belonging to the sample label). The semantic matching label can be used to indicate whether the text semantics between the two second sample texts in the sample pair are matching or not matching, that is, the semantic matching label is used to indicate whether the two second sample texts in the sample pair are matching. Whether they are matching may refer to that the semantics of the two second sample texts are matching (such as the semantics are the same or similar).

[0214] In other words, if the semantic matching label indicates that the text semantics between the two second sample texts in the sample pair are matching, it means that the text semantics between the two second sample texts are similar or the same; if the semantic matching label indicates that the text semantics between the two second sample texts in the sample pair are not matching, it means that the text semantics between the two second sample texts are not similar or not the same.

[0215] The search device can call the feature embedding model to be trained to perform feature embedding processing on each second sample text in the sample pair respectively to generate the representation features of each second sample text in the sample pair. One second sample text can have one representation feature.

[0216] Among them, the representation feature of the second sample text can be a feature vector generated for the second sample text.

[0217] The search device can generate a feature embedding loss of the second sample text in the sample pair for the feature embedding model to be trained based on the representation features and semantic matching labels of each second sample text in the sample pair. This feature embedding loss can reflect the difference between the matching of the representation features of the two second sample texts in the sample pair (such as the similarity between the representation features) and the actual matching indicated by the semantic matching label of this sample pair for these two second sample texts (such as the similarity between the second sample texts). The larger this feature embedding loss is, the greater the difference between these two types of matching is. Conversely, the smaller this feature embedding loss is, the smaller the difference between these two types of matching is.

[0218] Among them, if the semantic matching label of the sample pair indicates that the text semantics between the two second sample texts in the sample pair are matched, it means that the two second sample texts in the sample pair are extremely similar. In this case, we hope that the representation features generated by the feature embedding model to be trained for these two second sample texts should also be extremely similar. Conversely, if the semantic matching label of the sample pair indicates that the text semantics between the two second sample texts in the sample pair are not matched, it means that the two second sample texts in the sample pair are extremely dissimilar. In this case, we hope that the representation features generated by the feature embedding model to be trained for these two second sample texts should also be extremely dissimilar.

[0219] Therefore, the above-mentioned feature embedding loss can also be called a similarity loss. This similarity loss can be used to reflect the deviation between the difference between the representation features generated by the feature embedding model to be trained for the two second sample texts in the sample pair and the difference we actually expect. For example, this feature embedding loss can be a cross-entropy loss or a cosine similarity loss generated by the representation features of the two second sample texts in the sample pair and the semantic matching label of the sample pair, and so on.

[0220] This application can correct the model parameters of the feature embedding model to be trained through the feature embedding loss to obtain the above-mentioned trained feature embedding model, as described below.

[0221] In addition to the above-mentioned feature embedding loss, the present application can also introduce a clustering loss for feature embedding of second sample texts with similar or identical semantics by a feature embedding model. When training the feature embedding model to be trained, this clustering loss can be called the second clustering loss, and this second clustering loss can be used to reflect the difference (or deviation) in feature embedding of the feature embedding model for second sample texts with similar or identical semantics. The larger this second clustering loss is, the greater the difference in feature embedding of the feature embedding model for second sample texts with similar or identical semantics; the smaller this second clustering loss is, the smaller the difference in feature embedding of the feature embedding model for second sample texts with similar or identical semantics. By introducing this second clustering loss, the present application can make the feature embedding model more unified when performing feature embedding on second sample texts with similar or identical semantics, and can better distinguish second sample texts with different semantics, thereby improving the accuracy of the feature embedding model in performing feature embedding. The specific process of generating this second clustering loss is described below, as described in the following content.

[0222] There can be multiple such sample pairs, and the multiple sample pairs can include multiple second sample texts with similar text semantics, that is, the multiple second sample texts with similar text semantics (i.e., similar meanings expressed by the texts) can be derived from the multiple sample pairs. The feature embedding model to be trained can sequentially include multiple network layers, and the network layers included in the feature embedding model to be trained can be called the second network layers, that is, the feature embedding model to be trained can sequentially include multiple second network layers, and the input of the subsequent second network layer can be the output of the previous second network layer. During the process of the feature embedding model to be trained performing feature embedding processing on the second sample texts, these multiple second network layers can respectively generate embedding features for the second sample texts. The embedding features generated by the second network layers other than the last second network layer among these multiple second network layers for the second sample texts can be intermediate features (i.e., features in the intermediate process), and the embedding features generated by the last second network layer among these multiple second network layers for the second sample texts are the representation features of the second sample texts.

[0223] Therefore, the search device can, during the process of the feature embedding model to be trained performing feature embedding processing on multiple second sample texts with similar text semantics, obtain the embedding features respectively generated by these multiple second network layers for each of the multiple second sample texts. One second network layer can generate one embedding feature for one second sample text, and this embedding feature can be a feature vector.

[0224] The search device can generate a second clustering loss of the feature embedding model to be trained for the multiple second sample texts through the embedding features of the multiple second sample texts. The second clustering loss can be used to reflect the differences between the embedding features generated by the same second network layer for the multiple second sample texts. By generating the second clustering loss of the feature embedding model through the embedding features generated by the multiple second network layers included in the feature embedding model for the second sample texts, it can be ensured that each second network layer in the feature embedding model has a consistent and unified process of performing feature embedding processing on second sample texts with the same or similar semantics. That is, for each second network layer in the feature embedding model, during the process of performing feature embedding processing on second sample texts with the same or similar semantics, the embedding features generated for each second sample text are consistent and unified. This can further improve the accuracy of training the feature embedding model and enhance the unity and consistency of the feature embedding model to be trained in processing texts with similar or the same semantics. As shown in the following formula, the second clustering loss can be the KL loss (KL Divergence Loss, a method for measuring the difference between two probability distributions. Here, the probability distribution can be understood as the distribution composed of each element within the embedding feature), and the second clustering loss can be expressed as Loss cluster , the second clustering loss Loss cluster can be:

[0225]

[0226] Among them, B batch represents the number of the multiple second sample texts with similar text semantics in the current training batch, and D dim represents the number of layers of the multiple second network layers in the feature embedding model to be trained. Since the last second network layer can have softmax (an activation function), and the second network layers before the last second network layer may not have softmax, therefore, in the above formula, the last second network layer and the second network layers before the last second network layer are represented separately.

[0227] The multiple second sample texts can be sorted in the order of the samples to which they belong before and after being input into the feature embedding model to be trained. i represents the i-th second sample text, and j represents the j-th second network layer. i can be greater than or equal to 1 and less than or equal to B batch , similarly, j can be greater than or equal to 1 and less than or equal to D dim . Therefore, (logits i ) j represents the embedding feature generated by the j-th second network layer for the i-th second sample text, (logits 1) j Denotes the embedding features generated by the j-th second network layer for the first second sample text, logits 1 is the second sample text that is the first to be input into the feature embedding model to be trained among the multiple second sample texts. Similarly, denotes the representation features generated by the last second network layer for the i-th second sample text, denotes the embedding features generated by the last second network layer for the i-th second sample text, denotes the embedding features generated by the last second network layer for the first second sample text. This application can use the embedding features of the first second sample text within the current training batch as a benchmark to evaluate the differences between the embedding features of the second sample texts with similar text semantics, that is, it can be trained so that the embedding features of the multiple second sample texts with similar text semantics are all similar to the embedding features of the first second sample text.

[0228] For the above first clustering loss, it is not necessary to use the partial formula where the last second network layer is located in the above formula (i.e., the partial formula where the softmax function is located), that is, the part of the formula after the plus sign in the above formula.

[0229] This application can use the above second clustering loss and feature embedding loss together to correct the model parameters of the feature embedding model to be trained, so as to obtain the trained feature embedding model. For example, the search device can perform an addition operation (i.e., summation) on the second clustering loss and the feature embedding loss to obtain the overall loss of the feature embedding model to be trained, and can use this overall loss to correct the model parameters of the feature embedding model to be trained, thereby obtaining the trained feature embedding model. Among them, the goal of correcting the model parameters of the feature embedding model to be trained through this overall loss can be to correct the model parameters of the feature embedding model to be trained so that this overall loss tends to the minimum value (such as tending to 0).

[0230] In one implementation, this application can also generate sample pairs (which can be called antonym sample pairs) with opposite text semantics between the two second sample texts it contains through the way of antonym replacement, that is, the above sample pairs can include antonym sample pairs. Of course, the above sample pairs can also include sample pairs with similar text semantics between the two second sample texts it contains. Exemplarily, the process of obtaining the antonym sample pairs can include:

[0231] The search device can obtain the original sample text. The search device can perform key information detection (such as keyword detection) on the original sample text to obtain G key information (such as G keywords) in the original sample text. G is a positive integer, and the specific value of G can be determined according to the number of key information detected in the original sample text. For example, the search device can perform key information detection on the original sample text through a trained key information detection model to detect G key information in the original sample text. The key information detection model can be any model that can detect key information with antonym information (such as antonyms) in the text.

[0232] The search device can perform G times of antonym replacement processing on the G key information in the original sample text to generate G antonym replacement texts of the original sample text. Among them, one antonym replacement processing is used to perform antonym replacement on one key information in the original sample text to obtain one antonym replacement text, and different antonym replacement processes can be used to perform antonym replacement on different key information in the original sample text. Among them, performing antonym replacement on the key information can refer to replacing the key information with its antonym information. For example, the key information can be a keyword, and performing antonym replacement on the key information can refer to replacing the key information with its antonym. Or, the antonym replacement processing here can also be replacing the key information with other words that have the same word type but completely different semantics.

[0233] For example, "you" in the original sample text can be replaced with "me", "you" can be replaced with "him", "me" can be replaced with "him", "is" can be replaced with "is not", "happy" can be replaced with "sad", "not have" can be replaced with "null" (that is, delete the word "not have"), and so on. For example, the original sample text can be "You haven't talked to me for a long time", and the antonym replacement texts of this original sample text can include "He hasn't talked to me for a long time", "You haven't talked to him for a long time", "You have talked to me for a long time", etc.

[0234] The search device can perform pairwise combination processing between the original sample text and the G antonym replacement texts to obtain at least one antonym sample pair. For example, an antonym sample pair can be composed of the original sample text and one antonym replacement text, or an antonym sample pair can also be composed of two antonym replacement texts.

[0235] There can be multiple above-mentioned original sample texts. The search device can perform antonym replacement on the key information in each original sample text according to the above principle to construct a large number of antonym sample pairs.

[0236] In this application, by constructing the above-mentioned antonym sample pairs and using these antonym sample pairs to train the feature embedding model to be trained, the discrimination ability of the feature embedding model to be trained for texts with different (even opposite) text semantics can be improved, thereby improving the accuracy of the feature embedding model to be trained in performing feature embedding processing on the input text.

[0237] In addition, the above-mentioned feature embedding model to be trained can be iteratively trained for multiple rounds. Therefore, after any iterative training of the feature embedding model to be trained, inaccurate antonym sample pairs (i.e., antonym sample pairs with poor generation effects) generated by the feature embedding model to be trained can be screened out from the above-mentioned multiple antonym sample pairs. Thus, when the next iterative training of the feature embedding model to be trained is performed, the screened-out antonym sample pairs can be used to focus on training the feature embedding model to be trained again, so as to achieve self-check of the feature embedding model to be trained and strengthen the learning of the feature embedding model to be trained for the inaccurate antonym sample pairs with poor generation effects. Exemplarily, the inaccurate antonym sample pairs generated may include: antonym sample pairs in which the two representation features generated for the two second sample texts included are very similar. For example, the two representation features being very similar may mean that the feature similarity (such as cosine similarity) between the two representation features is greater than or equal to a set similarity threshold.

[0238] Please refer to Figure 10 , Figure 10 which is a schematic diagram of the principle of constructing antonym sample pairs provided by an embodiment of this application. As Figure 10 shown, here G can be equal to 3. The search device performs antonym replacement on the key information in the original sample text, and 3 antonym replacement texts, including antonym replacement text 1 to antonym replacement text 3, can be obtained. The search device can combine the original sample text and the 3 antonym replacement texts in pairs to construct multiple antonym sample pairs. Here, 6 antonym sample pairs, including antonym sample pair 1 to antonym sample pair 6, are constructed.

[0239] For example, here, antonym sample pair 1 can be constructed through the original sample text and antonym replacement text 1, antonym sample pair 2 can be constructed through the original sample text and antonym replacement text 2, antonym sample pair 3 can be constructed through the original sample text and antonym replacement text 3, antonym sample pair 4 can be constructed through antonym replacement text 1 and antonym replacement text 2, antonym sample pair 5 can be constructed through antonym replacement text 1 and antonym replacement text 3, and antonym sample pair 6 can be constructed through antonym replacement text 2 and antonym replacement text 3.

[0240] The above-mentioned multiple sample pairs can be a batch of sample pairs for one round of training of the feature embedding model to be trained. According to the principle described above, the present application can perform multiple rounds of iterative training on the feature embedding model to be trained through multiple batches of sample pairs. Among them, in order to better obtain the above-mentioned second clustering loss, a large number (such as greater than or equal to a quantity threshold) of second sample texts with similar text semantics can be included in the multiple sample pairs of the same batch used for training the feature embedding model to be trained in the present application, and the present application can identify which second sample texts are semantically similar among the sample pairs of a batch used for one round of training of the feature embedding model to be trained.

[0241] Exemplarily, when the model parameters of the feature embedding model to be trained are trained to a convergence state, or the number of rounds of iterative training on the feature embedding model to be trained is equal to a specified round threshold, it can be considered that the training of the feature embedding model to be trained is completed, and the feature embedding model obtained at this time of training can be used as the above-mentioned trained feature embedding model.

[0242] Through the above process of the present application, excellent training of the feature embedding model to be trained is achieved. The trained feature embedding model can accurately distinguish texts with different text semantics, and can also ensure unified embedding processing of texts with similar text semantics. Therefore, the representation features of the input text can be accurately generated through the trained feature embedding model.

[0243] Step S302: Based on the search matching degree between the search text and each emoji in the emoji library, filter out the first emoji from the emoji library.

[0244] Specifically, the search device can filter out the first emoji from the emoji library according to the search matching degree between the search text and each emoji in the emoji library, as described below.

[0245] In one implementation, the search device can sort the emojis in the emoji library in descending order of the search matching degree between the search text and each emoji in the emoji library to obtain the sorted emojis. Among the sorted emojis, the emoji with a greater search matching degree with the search text can be arranged in the front, and the emoji with a smaller search matching degree with the search text can be arranged in the back.

[0246] The search device can use the first Q emojis (i.e., the first Q emojis in the front) arranged in the sorted emojis as the filtered first emojis. Q is a positive integer, and Q is less than the total number of emojis in the emoji library. The specific value of Q can be set according to the actual application scenario.

[0247] If there are multiple (more than one) expression libraries, the search device can, according to the same principle as above, separately screen out the first expression image from each expression library.

[0248] By adopting the above method of the present application, that is, by searching for the similarity between the search text and its intended text and the text label of the expression image, the first expression image is comprehensively searched from multiple levels (such as the level including the search text itself and the level of the intended text of the search text), so that the searched first expression image can be richer and more in line with the user's search requirements for the expression image.

[0249] The method provided by the present application can significantly improve the accuracy of recommending expression images: through multi-dimensional intention understanding of the user's Query and an expression retrieval algorithm based on text semantics, for the short text (i.e., the search text) with colloquial and daily expressions input by the user, the expression template image corresponding to the emotion and intention of the short text can be accurately retrieved, and the expressiveness of the expression image can be further improved through the synthesis between the text and the expression image, so that an expression image extremely meeting the user's needs can be returned to the user.

[0250] Please refer to Figure 11 , Figure 11 which is a schematic framework diagram of a data processing system provided by an embodiment of the present application. As Figure 11 shown, the data processing system can generally include a peripheral operation part, an online part, and an offline part. Among them, the peripheral operation part includes a manual intervention module, an information tracing module, a whiteboard system module, a log reporting module, etc. The manual intervention module supports the ability to perform real-time incremental processing (such as addition and deletion processing) on expression images. For example, when there are non-compliant or low-quality expression images, these expression images can be deleted from the expression library through manual intervention. The information tracing module can debug the system according to the user's feedback. The whiteboard system module can provide some statistical information for the operation personnel to view and refer to. For example, the statistical information can include some statistical information on the user's use of expression images (such as the number of times of using expression images and feedback, etc.). The log reporting module can be used to report some operation data of the user, such as operation data on the exposure and click of expression images by the user.

[0251] The online part mentioned above supports the ability to retrieve and synthesize emoticons in real time by searching text (which can be denoted as Query). The online part may include modules for business logic, scoring logic, vector retrieval engine, other retrieval engines, and online feature storage. The module for business logic may include "Query content understanding" (such as intent recognition), "retrieval & sorting and distribution" (such as searching for emoticons and sorting and distributing the emoticons), "synthesized emoticon processing" (such as the synthesis processing between text and emoticon template images), "uploading results to CDN" (such as uploading the searched emoticons to CDN to return the emoticons to the front end through CDN), and "packet return" (such as packing the searched emoticons and returning them to the front end), etc.

[0252] The module for scoring logic may include "multi-channel retrieval" (such as retrieving emoticons by means of intent text, text-to-image, or emoticon synthesis, etc.), "feature supplementation" (such as collecting some features of newly added emoticons (such as usage popularity, etc.)), and "sorting and scoring" (such as sorting and scoring the retrieved emoticons), etc.

[0253] The module for the above-mentioned vector retrieval engine may include parts such as "text vector retrieval" (such as retrieving the representation features generated by text) and "image vector retrieval" (such as retrieving the representation features generated by the text labels of emoticons). The module for the above-mentioned other retrieval engines may be a reserved part for supplementing the required processing engines later. The module for online feature storage can be used to store "original content" (such as original emoticons, and the emoticons in the emoticon library are usually obtained by processing the original emoticons (such as size regularization or format regularization, etc.)), and "mined features" (such as the usage popularity of emoticons, etc.).

[0254] The offline part mentioned above supports the capabilities of making, accessing, preprocessing, and index building of the emoticon template library. The offline part may include a data access module (as a unified data entry), a feature calculation module, an index building module, and a data storage module. The data access module may include parts such as "access to original emoticon content" (such as access to original emoticons) and "asynchronous feature access" (such as the representation features of text). The feature calculation module may include "emoticon vector calculation" (such as calculating the representation features of the text labels of emoticons) and the part of "emoticon popularity statistics". For example, when sorting the retrieved emoticons, not only the search match degree between the emoticon and the search text can be referred to, but also the usage popularity of the emoticon can be referred to, and the two can be weighted and summed as the final sorting score of the emoticon.

[0255] The module for index construction may include "vector format index construction" (such as constructing an index from the text labels of emoticons to the representation features of the text labels), and "text format index construction" (such as constructing an index from emoticons to their text labels, or constructing an index from search text to its intent text, etc.). The module for data storage may include "permanent storage of original content / feature vectors" (such as for the emoticons that have been stored in the database, their original emoticons and the representation features of their text labels can be permanently stored), and "temporary storage of original content / feature vectors" (such as for the emoticons that are still under review and not yet stored in the database, their original emoticons and the representation features of their text labels can be temporarily stored).

[0256] Through the collaborative work of the above-mentioned various modules, the present application can achieve efficient search and processing of emoticons.

[0257] Please refer to Figure 12 , Figure 12 which is a schematic structural diagram of a data search device provided by an embodiment of the present application. As Figure 12 shown, the data search device 120 may include: a first acquisition module 1201, a generation module 1202, a second acquisition module 1203, and a search module 1204.

[0258] The first acquisition module 1201 is configured to acquire search text, and the search text is used to trigger the search for emoticons;

[0259] The generation module 1202 is configured to perform intent recognition on the search text from N intent understanding dimensions, and generate N intent texts of the search text. One intent text corresponds to one intent understanding dimension, and N is a positive integer;

[0260] The second acquisition module 1203 is configured to acquire the text labels of each emoticon in the emoticon library. The text label of any emoticon includes the intent labels of the any emoticon in N intent understanding dimensions respectively;

[0261] The search module 1204 is configured to search for a first emoticon matching the search text in the emoticon library through the similarity between the N intent texts and the text labels of the emoticons in the emoticon library.

[0262] In one implementation manner, the manner in which the generation module 1202 performs intent recognition on the search text from N intent understanding dimensions and generates N intent texts of the search text includes:

[0263] Acquire a trained intent recognition model, where the trained intent recognition model includes a text encoder and N text decoders, and one text decoder corresponds to one intent understanding dimension;

[0264] Call the text encoder to perform feature encoding processing on the search text to generate the encoded features of the search text;

[0265] Call N text decoders to perform feature decoding processing on the encoded features respectively in the corresponding intention understanding dimensions to generate N intention texts.

[0266] In one implementation, the above data search device 120 further includes a training module 1205, and the training module 1205 is used for:

[0267] Obtain the first sample text and the intention recognition model to be trained. The intention recognition model to be trained includes a text encoder to be trained and N text decoders to be trained. One text decoder to be trained corresponds to one intention understanding dimension, and the first sample text has corresponding reference intention texts in N intention understanding dimensions;

[0268] Call the text encoder to be trained to perform feature encoding processing on the first sample text to generate the sample encoded features of the first sample text;

[0269] Call N text decoders to be trained to perform feature decoding processing on the sample encoded features respectively in the corresponding intention understanding dimensions to generate N sample intention texts of the first sample text;

[0270] Based on the differences between the sample intention texts and the reference intention texts of each text decoder to be trained in their respective corresponding intention understanding dimensions, generate the feature decoding losses of each text decoder to be trained for the sample encoded features respectively;

[0271] Based on the feature decoding losses of the N text decoders to be trained respectively, correct the model parameters of the intention recognition model to be trained to obtain the trained intention recognition model.

[0272] In one implementation, there are multiple first sample texts, and the multiple first sample texts have the same reference intention texts in N intention understanding dimensions. Any text decoder to be trained is a target text decoder, and the target text decoder includes multiple first network layers. The manner in which the training module 1205 corrects the model parameters of the intention recognition model to be trained based on the feature decoding losses of the N text decoders to be trained respectively to obtain the trained intention recognition model includes:

[0273] During the process of the target text decoder performing feature decoding processing on the sample encoded features of each first sample text, obtain the decoding features generated by each of the first network layers except the last first network layer in the multiple first network layers for each first sample text;

[0274] Generating a first clustering loss of a target text decoder for multiple first sample texts based on decoding features of the multiple first sample texts, where the first clustering loss is used to reflect the differences between the decoding features generated by the same first network layer for the multiple first sample texts;

[0275] By using the first clustering loss and the feature decoding loss of each of the N text decoders to be trained, correcting the model parameters of the intent recognition model to be trained, and obtaining the trained intent recognition model.

[0276] In one implementation, the manner in which the search module 1204 searches for a first expression map that matches the search text in the expression map library by using the similarity between the N intent texts and the text labels of the expression maps in the expression map library includes:

[0277] Obtaining the search matching degrees between the search text and each expression map in the expression map library by using the similarities between the search text and the N intent texts and the text labels of each expression map in the expression map library respectively;

[0278] Based on the search matching degrees between the search text and each expression map in the expression map library, screening out the first expression map from the expression map library.

[0279] In one implementation, any expression map in the expression map library is a target expression map; the manner in which the search module 1204 obtains the search matching degrees between the search text and each expression map in the expression map library by using the similarities between the search text and the N retrieval texts and the text labels of each expression map in the expression map library respectively includes:

[0280] Invoking the trained feature embedding model to perform feature embedding processing on the search text, the N intent texts, and each text label of the target expression map respectively, and generating the representation features of the search text, the representation features of each intent text, and the representation features of each text label of the target expression map;

[0281] Obtaining M feature similarities between the representation features of the search text and the N intent texts and the representation features of the text labels of the target expression map, where the feature similarity between the representation features of texts is used to reflect the similarity between texts, and M is a positive integer;

[0282] Taking the maximum feature similarity among the M feature similarities as the search matching degree between the search text and the target expression map.

[0283] In one implementation, the text label of any expression map includes the general label of any expression map, the target expression map has K + N text labels, the K + N text labels include K general labels of the target expression map and N intent labels, and K is a positive integer;

[0284] The way for the search module 1204 to obtain M feature similarities between the representation features of the search text, the representation features of N intent texts, and the representation features of the text labels of the target emoticon diagram includes:

[0285] Calculate K + N feature similarities between the representation features of the search text and the representation features of K + N text labels respectively. There is one feature similarity between the representation feature of the search text and the representation feature of a text label of the target emoticon diagram;

[0286] Calculate N×K feature similarities between the representation features of N intent texts and the representation features of K general labels respectively. There is one feature similarity between the representation feature of an intent text and the representation feature of a general label of the target emoticon diagram;

[0287] Calculate N feature similarities between the representation features of N intent texts and the representation features of N intent labels. There is one feature similarity between the representation feature of any intent text and the representation feature of an intent label in the intent understanding dimension corresponding to any intent text of the target emoticon diagram;

[0288] Determine the K + N feature similarities, N×K feature similarities, and N feature similarities as M feature similarities.

[0289] In one implementation, the above training module 1205 is further configured to:

[0290] Obtain a sample pair and a feature embedding model to be trained. The sample pair includes two second sample texts, and the sample pair has a semantic matching label, where the semantic matching label indicates whether the text semantics between the two second sample texts in the sample pair are matching or not;

[0291] Call the feature embedding model to be trained to perform feature embedding processing on each second sample text in the sample pair, and generate the representation features of each second sample text in the sample pair;

[0292] Generate a feature embedding loss of the feature embedding model to be trained based on the representation features of each second sample text in the sample pair and the semantic matching label;

[0293] Modify the model parameters of the feature embedding model to be trained based on the feature embedding loss to obtain a trained feature embedding model.

[0294] In one implementation, there are multiple sample pairs, and the multiple sample pairs include multiple second sample texts with similar text semantics. The feature embedding model to be trained includes multiple second network layers. The way for the training module 1205 to modify the model parameters of the feature embedding model to be trained based on the feature embedding loss to obtain a trained feature embedding model includes:

[0295] During the process of the feature embedding model to be trained performing feature embedding processing on the second sample texts in multiple sample pairs, obtain the embedding features generated by multiple second network layers for each of the second sample texts in the multiple second sample texts;

[0296] Generate a second clustering loss of the feature embedding model to be trained for the multiple second sample texts based on the embedding features of the multiple second sample texts. The second clustering loss is used to reflect the differences between the embedding features generated by the same second network layer for the multiple second sample texts;

[0297] Modify the model parameters of the feature embedding model to be trained based on the second clustering loss and the feature embedding loss to obtain the trained feature embedding model.

[0298] In one implementation, the above training module 1205 is further configured to:

[0299] Obtain the original sample text, and perform key information detection on the original sample text to obtain G key information in the original sample text, where G is a positive integer;

[0300] Perform G times of antonym replacement processing on the G key information in the original sample text to generate G antonym replacement texts of the original sample text. One antonym replacement processing is used to perform antonym replacement on one key information in the original sample text to obtain one antonym replacement text;

[0301] Perform pairwise combination processing between the original sample text and the G antonym replacement texts to obtain at least one sample pair.

[0302] In one implementation, the method for the search module 1204 to screen out the first expression image from the expression image library based on the search matching degree between the search text and each expression image in the expression image library includes:

[0303] Sort the expression images in the expression image library in descending order according to the search matching degree between the search text and each expression image in the expression image library to obtain the sorted expression images;

[0304] Use the first Q expression images arranged at the front in the sorted expression images as the first expression images, where Q is a positive integer and Q is less than the total number of expression images in the expression image library.

[0305] In one implementation, the search text is sent by the communication client; the search module 1204 is further configured to:

[0306] Perform expression expansion processing based on the search text to obtain the expanded second expression images;

[0307] Return the first expression images and the second expression images to the communication client.

[0308] In one implementation, the way for the search module 1204 to perform expression augmentation processing based on the search text to obtain the augmented second expression graph includes:

[0309] Call the trained language rewriting model to perform language rewriting processing on the search text to generate an expression description text corresponding to the search text;

[0310] Call the text-to-image model to generate L candidate expression graphs based on the expression description text, and call the image-text matching model to predict the image-text correlation between each candidate expression graph and the expression description text, where L is a positive integer;

[0311] Sort the L candidate expression graphs in descending order according to the image-text correlation between each candidate expression graph and the expression description text to obtain the sorted L candidate expression graphs;

[0312] Take the first P candidate expression graphs in the sorted L candidate expression graphs as the second expression graph, where P is a positive integer and P is less than L.

[0313] In one implementation, the above training module 1205 is further configured to:

[0314] Obtain a third sample text and the language rewriting model to be trained, where the third sample text has a reference rewritten text;

[0315] Call the language rewriting model to be trained to perform language rewriting processing on the third sample text to generate a sample rewritten text corresponding to the third sample text;

[0316] Based on the difference between the sample rewritten text and the reference rewritten text, correct the model parameters of the language rewriting model to be trained to obtain the trained language rewriting model.

[0317] In one implementation, the way for the search module 1204 to perform expression augmentation processing based on the search text to obtain the augmented second expression graph includes:

[0318] Obtain a general expression graph;

[0319] Perform synthesis processing on the search text and the general expression graph to generate the second expression graph.

[0320] In one implementation, the first expression graph includes an expression template graph; the process for the search module 1204 to return the first expression graph to the communication client includes:

[0321] Perform synthesis processing on the search text and the expression template graph to generate a synthesized expression graph;

[0322] Return the synthesized expression graph to the communication client.

[0323] According to an embodiment of the present application, Figure 3 the steps involved in the data search method shown can be performed by Figure 12 each module in the data search device 120 shown. For example, Figure 3 the step S101 shown in can be performed by Figure 12 the first acquisition module 1201 in, Figure 3 the step S102 shown in can be performed by Figure 12 the generation module 1202 in; Figure 3 the step S103 shown in can be performed by Figure 12 the second acquisition module 1203 in, Figure 3 the step S104 shown in can be performed by Figure 12 the search module 1204 in.

[0324] The method proposed in the present application can perform multi-dimensional intent recognition on the search text for searching for emoticons from multiple intent understanding dimensions, and can also configure respective intent labels for each emoticon in the emoticon library in these multiple intent understanding dimensions. Thus, through the similarity between the multiple intent texts generated by performing intent recognition on the search text and the text labels (including intent labels) of the emoticons in the emoticon library, multi-dimensional search for emoticons matching the search text can be achieved from these multiple intent understanding dimensions, thereby improving the accuracy of searching for emoticons matching the search text in the emoticon library.

[0325] According to an embodiment of the present application, Figure 12 each module in the data search device 120 shown can be separately or all combined into one or several units to form, or a certain one (or some) of the units can be further split into multiple smaller sub-units in terms of function, and the same operations can be achieved without affecting the realization of the technical effects of the embodiments of the present application. The above modules are divided based on logical functions. In actual applications, the function of one module can also be realized by multiple units, or the functions of multiple modules can be realized by one unit. In other embodiments of the present application, the data search device 120 can also include other units. In actual applications, these functions can also be assisted by other units and can be realized by the cooperation of multiple units.

[0326] In the embodiments of the present application, the term "module" or "unit" refers to a computer program with a predetermined function or a part of a computer program, which works together with other related parts to achieve a predetermined goal, and can be fully or partially implemented by using software, hardware (such as a processing circuit or a memory), or a combination thereof. Similarly, one processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be a part of an overall module or unit that includes the function of that module or unit.

[0327] According to an embodiment of the present application, a computer program capable of executing the respective steps involved in the corresponding methods shown in the embodiments of the present application can be run on a general-purpose computer device (which may include processing elements and storage elements such as a central processing unit (CPU), a random access storage medium (RAM), a read-only storage medium (ROM), etc.) to construct a data search device 120 as shown in Figure 12 shown. The above computer program can be recorded on a computer-readable recording medium, and can be loaded into the above computer device through the computer-readable recording medium and run therein.

[0328] Please refer to Figure 13 , Figure 13 which is a schematic structural diagram of a computer device provided by an embodiment of the present application. As shown in Figure 13 shown, the computer device 1000 may include: a processor 1001, a network interface 1004, and a memory 1005. In addition, in some embodiments, the computer device 1000 may further include: a user interface 1003, and at least one communication bus 1002. Among them, the communication bus 1002 is used to realize the connection and communication between these components. Among them, the user interface 1003 may include a display screen (Display), a keyboard (Keyboard), and optionally the user interface 1003 may further include a standard wired interface and a wireless interface. The network interface 1004 may optionally include a standard wired interface and a wireless interface (such as a WI-FI interface). The memory 1005 may be a high-speed RAM memory, or a non-volatile memory, such as at least one disk memory. Optionally, the memory 1005 may also be at least one storage device located far from the aforementioned processor 1001. As shown in Figure 13 shown, the memory 1005, as a computer storage medium, may include an operating system, a network communication module, a user interface module, and a device control application program.

[0329] In Figure 13In the computer device 1000 shown, the network interface 1004 can provide network communication functions; the user interface 1003 is mainly used to provide an interface for users to input; and the processor 1001 can be used to call the device control application program stored in the memory 1005 to achieve:

[0330] Obtain a search text, which is used to trigger the search for expression graphs;

[0331] Perform intent recognition on the search text from N intent understanding dimensions to generate N intent texts of the search text, where one intent text corresponds to one intent understanding dimension, and N is a positive integer;

[0332] Obtain the text labels of each expression graph in the expression graph library. The text label of any expression graph includes the intent labels of any expression graph on N intent understanding dimensions respectively;

[0333] Search for a first expression graph matching the search text in the expression graph library based on the similarity between the N intent texts and the text labels of the expression graphs in the expression graph library.

[0334] It should be understood that the computer device 1000 described in the embodiments of the present application can execute the description of the above data search method in the embodiments of the present application, and can also execute the description of the above data search device 120 in the corresponding embodiments described above, which will not be elaborated here. In addition, the description of the beneficial effects of using the same method will not be elaborated either. Figure 12 For the description of the data search device 120 in the corresponding embodiments described above, it will not be elaborated here. In addition, the description of the beneficial effects of using the same method will not be elaborated either.

[0335] In addition, it should be pointed out here that: the present application also provides a computer-readable storage medium, and a computer program is stored in the computer-readable storage medium. When the processor executes the computer program, it can execute the description of the data search method in the embodiments of the present application. Therefore, it will not be elaborated here. In addition, the description of the beneficial effects of using the same method will not be elaborated either. For the technical details not disclosed in the embodiments of the computer storage medium involved in the present application, please refer to the description of the method embodiments of the present application.

[0336] As an example, the above computer program can be deployed to be executed on a computer device, or be deployed to be executed on multiple computer devices located at one location, or, be executed on multiple computer devices distributed at multiple locations and interconnected through a communication network. The multiple computer devices distributed at multiple locations and interconnected through a communication network can form a blockchain network.

[0337] The above computer-readable storage medium may be an internal storage unit of the above computer device, such as the hard disk or memory of the computer device. The computer-readable storage medium may also be an external storage device of the computer device, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. equipped on the computer device. Further, the computer-readable storage medium may also include both the internal storage unit and the external storage device of the computer device. The computer-readable storage medium is used to store the computer program and other programs and data required by the computer device. The computer-readable storage medium may also be used to temporarily store the data that has been output or is to be output.

[0338] The present application provides a computer program product, which includes a computer program stored in a computer-readable storage medium. The processor of the computer device reads the computer program from the computer-readable storage medium, and the processor executes the computer program, so that the computer device executes the description of the above data search method in the embodiments of the present application. Therefore, the description will not be repeated here. In addition, the description of the beneficial effects of using the same method will not be repeated either. For the technical details not disclosed in the embodiments of the computer-readable storage medium involved in the present application, please refer to the description of the method embodiments of the present application.

[0339] The terms "first", "second", etc. in the description, claims and drawings of the embodiments of the present application are used to distinguish different objects, rather than to describe a specific order. In addition, the term "comprising" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, device, product or equipment that includes a series of steps or units is not limited to the listed steps or modules, but may optionally further include steps or modules not listed, or may optionally further include other step units inherent to these processes, methods, devices, products or equipment.

[0340] Those of ordinary skill in the art can realize that the units and algorithm steps of the examples described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the components and steps of the examples have been generally described according to their functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.

[0341] The above disclosure is only a preferred embodiment of the present application. Of course, it cannot be used to limit the scope of rights of the present application. Therefore, equivalent changes made according to the claims of the present application still fall within the scope covered by the present application.

Claims

1. A data search method, characterized in that: The method comprises: Obtaining search text, where the search text is used to trigger a search for emoticons; Performing intent recognition on the search text from N intent understanding dimensions to generate N intent texts of the search text, where one intent text corresponds to one intent understanding dimension, and N is a positive integer; Obtaining a text label for each emoticon in the emoticon library, wherein the text label of any emoticon includes the intent label of any emoticon in the N intent understanding dimensions; The emoticon library is searched for a first emoticon matching the search text based on the similarity between the N intended texts and the text labels of the emoticons in the emoticon library.

2. The method according to claim 1, characterized in that The performing intent recognition on the search text from N intent understanding dimensions to generate N intent texts of the search text includes: Obtaining a trained intent recognition model, wherein the trained intent recognition model includes a text encoder and N text decoders, and one text decoder corresponds to one intent understanding dimension; Calling the text encoder to perform feature encoding processing on the search text to generate encoding features of the search text; The N text decoders are called to perform feature decoding processing on the encoded features at the corresponding intent understanding dimensions respectively to generate the N intent texts.

3. The method according to claim 2, characterized in that The method further comprises: Obtaining a first sample text and an intent recognition model to be trained, wherein the intent recognition model to be trained includes a text encoder to be trained and N text decoders to be trained, wherein one text decoder to be trained corresponds to one intent understanding dimension, and the first sample text has corresponding reference intent texts in each of the N intent understanding dimensions; Calling the text encoder to be trained to perform feature encoding processing on the first sample text to generate sample encoding features of the first sample text; Calling the N text decoders to be trained to perform feature decoding processing on the sample encoding features at corresponding intent understanding dimensions respectively, to generate N sample intent texts of the first sample text; Based on the difference between the sample intent text and the reference intent text on the intent understanding dimension corresponding to each text decoder to be trained, generating a feature decoding loss for each text decoder to be trained for the sample encoding feature; The model parameters of the intention recognition model to be trained are corrected based on the feature decoding losses of each of the N text decoders to be trained to obtain the trained intention recognition model.

4. The method according to claim 3, characterized in that There are multiple first sample texts, and the multiple first sample texts have the same reference intent text in the N intent understanding dimensions. Any text decoder to be trained is a target text decoder, and the target text decoder includes multiple first network layers; the model parameters of the intent recognition model to be trained are corrected based on the feature decoding losses of each of the N text decoders to be trained to obtain the trained intent recognition model, including: In the process of the target text decoder performing feature decoding processing on the sample encoding feature of each first sample text, obtaining decoding features generated by each first network layer except the last first network layer in the plurality of first network layers for each first sample text; Generating a first clustering loss of the target text decoder for the plurality of first sample texts based on the decoding features of the plurality of first sample texts, wherein the first clustering loss is used to reflect the difference between the decoding features generated by the same first network layer for the plurality of first sample texts; The model parameters of the intention recognition model to be trained are corrected by using the first clustering loss and feature decoding loss of each of the N text decoders to be trained, so as to obtain the trained intention recognition model.

5. The method according to claim 1, characterized in that The step of searching the emoticon library for a first emoticon image that matches the search text based on the similarity between the N intended texts and the text labels of the emoticon images in the emoticon library comprises: Obtaining a search matching degree between the search text and each emoticon in the emoticon library by comparing the similarities between the search text and the N intended texts and the text labels of each emoticon in the emoticon library; The first emoticon image is selected from the emoticon image library based on a search matching degree between the search text and each emoticon image in the emoticon image library.

6. The method according to claim 5, characterized in that Any emoticon in the emoticon library is a target emoticon; obtaining the search matching degree between the search text and each emoticon in the emoticon library by comparing the similarity between the search text and the N search texts and the text label of each emoticon in the emoticon library, comprises: Calling the trained feature embedding model to perform feature embedding processing on the search text, the N intended texts, and each text label of the target emoticon respectively, to generate representation features of the search text, representation features of each intended text, and representation features of each text label of the target emoticon; Obtaining M feature similarities between the representation features of the search text and the representation features of the N intended texts and the representation features of the text label of the target emoticon, wherein the feature similarity between the representation features of the texts is used to reflect the similarity between the texts, and M is a positive integer; The maximum feature similarity among the M feature similarities is used as the search matching degree between the search text and the target expression image.

7. The method according to claim 6, characterized in that The text label of any expression image includes the common label of any expression image, the target expression image has K+N text labels, and the K+N text labels include K common labels and N intention labels of the target expression image, where K is a positive integer; The obtaining of M feature similarities between the representation feature of the search text and the representation features of the N intended texts and the representation feature of the text label of the target emoticon image includes: Calculating K+N feature similarities between the representative features of the search text and the representative features of the K+N text labels, respectively, wherein there is one feature similarity between the representative feature of the search text and the representative feature of a text label of the target emoticon; Calculate N×K feature similarities between the representation features of the N intended texts and the representation features of the K general tags, respectively, where a representation feature of an intended text and a representation feature of a general tag of the target emoticon have one feature similarity; Calculate N feature similarities between the representation features of the N intention texts and the representation features of the N intention labels, where the representation feature of any intention text and the representation feature of an intention label of the target emoticon image on the intention understanding dimension corresponding to any intention text have one feature similarity; The K+N feature similarities, the N×K feature similarities, and the N feature similarities are determined as the M feature similarities.

8. The method according to claim 6, characterized in that The method further comprises: Obtaining a sample pair and a feature embedding model to be trained, wherein the sample pair includes two second sample texts, the sample pair has a semantic matching label, and the semantic matching label indicates whether the text semantics between the two second sample texts in the sample pair are matched or not matched; Calling the feature embedding model to be trained to perform feature embedding processing on each second sample text in the sample pair to generate a representative feature of each second sample text in the sample pair; Generating a feature embedding loss of the feature embedding model to be trained based on the representation features of each second sample text in the sample pair and the semantic matching label; The model parameters of the feature embedding model to be trained are corrected based on the feature embedding loss to obtain the trained feature embedding model.

9. The method according to claim 8, characterized in that There are multiple sample pairs, and the multiple sample pairs include multiple second sample texts with similar text semantics. The feature embedding model to be trained includes multiple second network layers; and the model parameters of the feature embedding model to be trained are corrected based on the feature embedding loss to obtain the trained feature embedding model, including: In the process of the feature embedding model to be trained performing feature embedding processing on the second sample texts in the multiple sample pairs, obtaining the embedding features generated by the multiple second network layers for each second sample text in the multiple second sample texts respectively; Generate a second clustering loss of the feature embedding model to be trained for the plurality of second sample texts based on the embedding features of the plurality of second sample texts, wherein the second clustering loss is used to reflect the difference between the embedding features generated by the same second network layer for the plurality of second sample texts; Based on the second clustering loss and the feature embedding loss, the model parameters of the feature embedding model to be trained are corrected to obtain the trained feature embedding model.

10. The method according to claim 8, characterized in that The method further comprises: Obtaining an original sample text, and performing key information detection on the original sample text to obtain G key information in the original sample text, where G is a positive integer; Performing G antonym replacement processing on G key information in the original sample text to generate G antonym replacement texts of the original sample text, wherein one antonym replacement processing is used to perform antonym replacement on one key information in the original sample text to obtain one antonym replacement text; The original sample text and the G antonym replacement texts are processed in pairs to obtain at least one sample pair.

11. The method according to claim 5, characterized in that The step of selecting the first emoticon from the emoticon library based on the search matching degree between the search text and each emoticon in the emoticon library includes: Sorting the emoticon images in the emoticon image library in descending order of the search matching degree between the search text and each emoticon image in the emoticon image library to obtain sorted emoticon images; The first Q expression pictures in the sorted expression pictures are used as the first expression pictures, where Q is a positive integer and is smaller than the total number of expression pictures in the expression picture library.

12. The method according to claim 1, characterized in that The search text is sent by a communication client; the method further comprises: Performing expression expansion processing based on the search text to obtain an expanded second expression image; The first emoticon image and the second emoticon image are returned to the communication client.

13. The method according to claim 12, characterized in that The step of performing expression expansion processing based on the search text to obtain an expanded second expression image includes: Calling the trained language rewriting model to perform language rewriting processing on the search text to generate an expression description text corresponding to the search text; Calling the text-based graph model to generate L candidate expression graphs based on the expression description text, and calling the graph-text matching model to predict the graph-text relevance between each candidate expression graph and the expression description text, where L is a positive integer; Sorting the L candidate expression images in descending order of the image-text relevance between each candidate expression image and the expression description text to obtain sorted L candidate expression images; The second expression image is generated based on the first P candidate expression images among the sorted L candidate expression images, where P is a positive integer and P is less than L.

14. The method according to claim 13, characterized in that The method further comprises: Acquire a third sample text and a language rewriting model to be trained, wherein the third sample text has a reference rewriting text; Calling the language rewriting model to be trained to perform language rewriting processing on the third sample text to generate a sample rewritten text corresponding to the third sample text; Based on the difference between the sample rewritten text and the reference rewritten text, the model parameters of the language rewriting model to be trained are modified to obtain the trained language rewriting model.

15. The method according to claim 12, characterized in that The step of performing expression expansion processing based on the search text to obtain an expanded second expression image includes: Get the universal emoticon image; The search text and the general emoticon are synthesized to generate the second emoticon.

16. The method according to claim 12, characterized in that The first emoticon image includes an emoticon template image; and the process of returning the first emoticon image to the communication client includes: Performing synthesis processing on the search text and the expression template image to generate a synthetic expression image; The synthesized expression image is returned to the communication client.

17. A data search device, characterized in that: The device comprises: A first acquisition module, used to acquire a search text, wherein the search text is used to trigger a search for emoticons; A generating module, configured to perform intent recognition on the search text from N intent understanding dimensions, and generate N intent texts of the search text, wherein one intent text corresponds to one intent understanding dimension, and N is a positive integer; A second acquisition module is used to acquire a text label of each emoticon in the emoticon library, wherein the text label of any emoticon includes the intent label of any emoticon in the N intent understanding dimensions; The search module is used to search for a first emoticon image that matches the search text in the emoticon image library based on similarities between the N intended texts and text labels of the emoticons in the emoticon image library.

18. A computer program product, characterized in that The computer program product comprises a computer program, which is stored in a computer-readable storage medium. The computer program is suitable for being read and executed by a processor, so that a computer device having the processor executes the method according to any one of claims 1 to 16.

19. A computer device, characterized in that: The invention comprises a memory and a processor, wherein the memory stores a computer program, and when the computer program is executed by the processor, the processor executes the steps of the method according to any one of claims 1 to 16.

20. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and the computer program is suitable for being loaded by a processor and executing the steps of the method according to any one of claims 1 to 16.

Citation Information

Cited By

  • Data search method and apparatus, and product and device

    WO2026158095A1