Picture retrieval method and device, terminal equipment and storage medium

By using the joint semantic prediction model to predict the search elements in the image retrieval system, the problem of high complexity in the image retrieval operation in the prior art is solved, and a more efficient and user-friendly image retrieval experience is achieved.

CN120045730APending Publication Date: 2025-05-27SHENZHEN INTELLIFUSION TECHNOLOGIES CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311553073.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-20
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

The prior art image retrieval method requires the user to complete multiple steps of key and mouse operations, resulting in high workload and operational complexity of the user, low efficiency and poor user experience.

Method used

By obtaining the search statements of the image to be retrieved and the preset picture database, the built joint semantic prediction model predicts the search elements corresponding to the search statement, including the search intention, operator, the attribute category of the image to be retrieved or the attribute values ​​corresponding to the attribute category of the attribute category, and the image to be retrieved is determined in the preset picture database based on these search elements.

Benefits of technology

It reduces the workload of users and improves the efficiency and user experience of image retrieval. Users do not need to frequently perform keyboard and mouse operations, and can quickly find the required pictures.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120045730A_ABST
    Figure CN120045730A_ABST
Patent Text Reader

Abstract

The invention is applicable to the technical field of picture retrieval, and provides a picture retrieval method and device, terminal equipment and a storage medium, and the method comprises the following steps: obtaining a retrieval statement of a to-be-retrieved picture and a preset picture database; on the basis of the retrieval statement, retrieval elements corresponding to the retrieval statement are predicted through a constructed joint semantic prediction model, and the retrieval elements comprise at least one of a retrieval intention, an operator, an attribute category of the to-be-retrieved picture or an attribute value corresponding to the attribute category; according to the method, the to-be-retrieved picture is determined based on the retrieval elements and the preset picture database, and the to-be-retrieved picture can be rapidly determined by adopting the retrieval elements including the retrieval intention, the operators, the attribute category of the to-be-retrieved picture or the attribute value corresponding to the attribute category, so that a user does not need to perform frequent keyboard and mouse operation, the workload of the user is greatly reduced, and the user experience is improved. And the picture retrieval efficiency and the user experience are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application belongs to the field of image retrieval technology, and in particular, relates to an image retrieval method, apparatus, terminal device and storage medium. Background Art

[0002] Currently, image retrieval methods are used in various fields such as life knowledge, social media, and intelligence analysis to help users quickly search and identify images of specific topics, people, or places.

[0003] The image retrieval method of the prior art requires the user to search in various image databases through frequent keyboard and mouse operations. For example, first select the image database corresponding to the image type according to the image type to be found, then enter the search keyword and click to search, including filtering the query location through the drop-down box, selecting the corresponding date and time period through the date and time pop-up box, etc. The image retrieval method of the prior art requires the user to complete multiple steps of keyboard and mouse operations. When there are many filtering conditions, the user's workload and operation complexity for retrieving images are greatly increased, the efficiency of image retrieval is reduced, and the user experience is reduced.

[0004] The existing technology has the problem of greatly increasing the workload and operation complexity of users' image retrieval, reducing the efficiency of image retrieval, and reducing the user experience. Summary of the invention

[0005] The embodiments of the present application provide a method, apparatus, terminal device and storage medium for image retrieval, which can solve the problem of greatly increasing the workload and operation complexity of user image retrieval, reducing the efficiency of image retrieval, and reducing the user experience.

[0006] A first aspect of an embodiment of the present application provides a method for image retrieval, including:

[0007] Obtaining a search statement for the image to be retrieved and a preset image database;

[0008] Based on the search statement, predicting the search elements corresponding to the search statement through the constructed joint semantic prediction model, the search elements including at least one of the search intent, the operator, the attribute category of the image to be searched, or the attribute value corresponding to the attribute category;

[0009] Based on the search elements and the preset image database, the image to be searched is determined.

[0010] In one of the embodiments, the joint semantic prediction model is an improved intent recognition and slot filling model, and the joint semantic prediction model includes a Bert model, an intent classifier, an attribute classifier, and an operator classifier.

[0011] In one embodiment, based on the search statement, predicting the search elements corresponding to the search statement by using the constructed joint semantic prediction model includes:

[0012] The search statement obtains the predicted search elements corresponding to the search statement through the Bert model, the intent classifier, the attribute classifier and the operator classifier.

[0013] In one embodiment, constructing the joint semantic prediction model includes:

[0014] Input any training sentence into the Bert model, and obtain an intention hidden vector corresponding to the training sentence and a plurality of slot value hidden vectors corresponding to the training sentence;

[0015] The joint semantic prediction model is trained based on the intention hidden vector, each slot value hidden vector, the intention classifier, the attribute classifier and the operator classifier until the retrieval factor loss value meets the preset conditions, thereby obtaining the trained joint semantic prediction model.

[0016] In one embodiment, the joint semantic prediction model is trained based on the intent hidden vector, each slot value hidden vector, the intent classifier, the attribute classifier, and the operator classifier until the retrieval factor loss value meets a preset condition, and the trained joint semantic prediction model is obtained, including:

[0017] Based on the intent hidden vector, the intent classifier, and the true intent of the intent hidden vector, obtaining a first loss value, where the first loss value is a loss value of a first loss function for intent classification;

[0018] Based on each of the slot value hidden vectors, the attribute classifier, and the true attribute of each of the slot value hidden vectors, obtaining a second loss value, where the second loss value is a loss value of a second loss function for attribute category classification;

[0019] Based on each of the slot value hidden vectors, the operator classifier, and the real operator of each of the slot value hidden vectors, obtaining a third loss value, where the third loss value is a loss value of a third loss function of operator classification;

[0020] Based on the first loss value, the second loss value, the third loss value and a preset weight, obtaining a search factor loss value;

[0021] If the retrieval factor loss value meets the preset condition, the trained joint semantic prediction model is obtained.

[0022] In one of the embodiments, obtaining a first loss value based on the intent hidden vector, the intent classifier, and the true intent of the intent hidden vector includes:

[0023] Inputting the intention hidden vector into the intention classifier to obtain a predicted intention corresponding to the intention hidden vector;

[0024] Obtaining a first loss value based on the predicted intent and the true intent of the intent hidden vector;

[0025] The obtaining of a second loss value based on each of the slot value hidden vectors, the attribute classifier, and the true attribute of each of the slot value hidden vectors includes:

[0026] Inputting each slot value hidden vector into the attribute classifier to obtain a predicted attribute corresponding to each slot value hidden vector;

[0027] Obtaining a second loss value based on the predicted attribute and the true attribute of each of the slot value hidden vectors;

[0028] The obtaining a third loss value based on each of the slot value hidden vectors, the operator classifier, and the real operator of each of the slot value hidden vectors includes:

[0029] Inputting each of the slot value hidden vectors into the operator classifier to obtain a prediction operator corresponding to each of the slot value hidden vectors;

[0030] A third loss value is obtained based on the predicted operator and the true operator of each slot value hidden vector.

[0031] In one embodiment, the preset weights include a first preset weight corresponding to the first loss value, a second preset weight corresponding to the second loss value, and a third preset weight corresponding to the third loss value;

[0032] The acquiring the search factor loss value based on the first loss value, the second loss value, the third loss value and the preset weight includes:

[0033] Based on the first loss value, the second loss value, the third loss value, the first preset weight, the second preset weight and the third preset weight, the retrieval factor loss value is obtained by using a retrieval factor loss function calculation formula.

[0034] A second aspect of an embodiment of the present application provides a device for image retrieval, including:

[0035] An acquisition module is used to acquire a search statement of a picture to be retrieved and a preset picture database;

[0036] A prediction module, configured to predict, based on the search statement, a search element corresponding to the search statement through a constructed joint semantic prediction model, wherein the search element includes at least one of a search intent, an operator, an attribute category of the image to be searched, or an attribute value corresponding to the attribute category;

[0037] A determination module is used to determine the image to be retrieved based on the search elements and the preset image database.

[0038] A third aspect of an embodiment of the present application provides a terminal device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the image retrieval method as described above when executing the computer program.

[0039] A fourth aspect of an embodiment of the present application provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, the image retrieval method as described above is implemented.

[0040] Compared with the prior art, the embodiments of the present invention have the following beneficial effects:

[0041] The method for image retrieval of the embodiment of the present application obtains a search statement of the image to be retrieved and a preset image database; based on the search statement, predicts the search elements corresponding to the search statement through the constructed joint semantic prediction model, the search elements including the search intent, the operator, the attribute category of the image to be retrieved or at least one of the attribute values ​​corresponding to the attribute category; based on the search elements and the preset image database, determines the image to be retrieved, because the search elements including the search intent, the operator, the attribute category of the image to be retrieved or at least one of the attribute values ​​corresponding to the attribute category are adopted, the preset image database can be searched for the specified search elements according to the user's search statement of the image to be retrieved, and the image to be retrieved can be quickly determined, so the user does not need to perform frequent keyboard and mouse operations, which greatly reduces the user's workload and improves the efficiency of image retrieval and user experience.

[0042] It can be understood that the beneficial effects of the second, third and fourth aspects mentioned above can be found in the relevant description of the first aspect mentioned above, and will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS

[0043] In order to more clearly illustrate the specific implementation methods of the present application or the technical solutions in the prior art, the drawings required for use in the specific implementation methods or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are some implementation methods of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0044] Figure 1A flowchart of a method for image retrieval provided in an embodiment of the present application;

[0045] Figure 2 A schematic diagram of the structure of the joint semantic prediction model provided in the embodiment of the present application;

[0046] Figure 3 A schematic diagram of a process for constructing a joint semantic prediction model provided in an embodiment of the present application;

[0047] Figure 4 A schematic diagram of a structure of inputting a training sentence into a Bert model to obtain an intention hidden vector and multiple slot value hidden vectors corresponding to the training sentence provided in an embodiment of the present application;

[0048] Figure 5 A flowchart of the joint semantic prediction model trained based on the intent hidden vector, each slot value hidden vector, intent classifier, attribute classifier and operator classifier provided in the embodiment of the present application until the retrieval factor loss function value meets the preset conditions;

[0049] Figure 6 A schematic diagram of a process for training a joint semantic prediction model provided in an embodiment of the present application;

[0050] Figure 7 A schematic diagram of a framework flow of image retrieval using a search statement provided in an embodiment of the present application;

[0051] Figure 8 A schematic diagram of the structure of an image retrieval device provided in an embodiment of the present application;

[0052] Fig. 9 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0053] In the following description, specific details such as specific system structures, technologies, etc. are provided for the purpose of illustration rather than limitation, so as to provide a thorough understanding of the embodiments of the present application. However, it should be clear to those skilled in the art that the present application may also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to prevent unnecessary details from obstructing the description of the present application.

[0054] It should be understood that when used in the present specification and the appended claims, the term "comprising" indicates the presence of described features, wholes, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components and / or combinations thereof.

[0055] It should also be understood that the term “and / or” used in the specification and appended claims refers to any and all possible combinations of one or more of the associated listed items, and includes these combinations.

[0056] As used in the specification and appended claims of this application, the term "if" can be interpreted as "when" or "uponce" or "in response to determining" or "in response to detecting", depending on the context. Similarly, the phrase "if it is determined" or "if [described condition or event] is detected" can be interpreted as meaning "uponce it is determined" or "in response to determining" or "uponce [described condition or event] is detected" or "in response to detecting [described condition or event]", depending on the context.

[0057] In addition, in the description of the present application specification and the appended claims, the terms "first", "second", "third", etc. are only used to distinguish the descriptions and cannot be understood as indicating or implying relative importance.

[0058] References to "one embodiment" or "some embodiments" etc. described in the specification of this application mean that one or more embodiments of the present application include specific features, structures or characteristics described in conjunction with the embodiment. Therefore, the statements "in one embodiment", "in some embodiments", "in some other embodiments", "in some other embodiments", etc. that appear in different places in this specification do not necessarily refer to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized in other ways. The terms "including", "comprising", "having" and their variations all mean "including but not limited to", unless otherwise specifically emphasized in other ways. "Multiple" means "two or more".

[0059] Currently, image retrieval methods are used in various fields such as life knowledge, social media, and intelligence analysis to help users quickly search and identify images of specific topics, people, or places.

[0060] The image retrieval method of the prior art requires the user to search in various image databases through frequent keyboard and mouse operations. For example, first select the image database corresponding to the image type according to the image type to be found, then enter the search keyword and click to search, including filtering the query location through the drop-down box, selecting the corresponding date and time period through the date and time pop-up box, etc. The image retrieval method of the prior art requires the user to complete multiple steps of keyboard and mouse operations. When there are many filtering conditions, the user's workload and operation complexity for retrieving images are greatly increased, the efficiency of image retrieval is reduced, and the user experience is reduced.

[0061] The existing technology has the problem of greatly increasing the workload and operation complexity of users' image retrieval, reducing the efficiency of image retrieval, and reducing the user experience.

[0062] In response to the above problems, an embodiment of the present application provides an image retrieval method, which obtains a retrieval statement of an image to be retrieved and a preset image database; based on the retrieval statement, predicts the retrieval elements corresponding to the retrieval statement through a constructed joint semantic prediction model, the retrieval elements including retrieval intent, operators, attribute categories of the image to be retrieved, or at least one of the attribute values ​​corresponding to the attribute categories; based on the retrieval elements and the preset image database, determines the image to be retrieved, because the retrieval elements including the retrieval intent, operators, attribute categories of the image to be retrieved, or at least one of the attribute values ​​corresponding to the attribute categories are adopted, the preset image database can be searched for the specified retrieval elements according to the user's retrieval statement of the image to be retrieved, and the image to be retrieved can be quickly determined, so the user does not need to perform frequent keyboard and mouse operations, which greatly reduces the user's workload and improves the efficiency of image retrieval and user experience.

[0063] The image retrieval method provided by this application is exemplarily described below in conjunction with specific embodiments.

[0064] First, as Figure 1 As shown, this embodiment provides a method for image retrieval, including:

[0065] S100, obtaining a search statement for a picture to be searched and a preset picture database.

[0066] In one embodiment, the user inputs the voice content of the image to be retrieved by voice through a mobile phone, computer or other mobile terminal, such as "find snapshots of men in the Civic Square", and then converts the input voice content into text through a voice-to-text conversion module, wherein the recognition accuracy of the voice-to-text conversion module is greater than or equal to 95%, thereby obtaining an accurate search sentence for the image to be retrieved. The text search sentence for the image to be retrieved is obtained by the voice-to-text conversion module, so that the user can input the image search instruction only by voice, without the need for the user to manually input the instruction every time the image is retrieved, and the use experience on mobile phones, computers or other mobile terminals is better, eliminating the need for the user to click and filter multiple times on the screen, improving the input efficiency of the image search instruction, and improving the user's experience of image retrieval.

[0067] In one embodiment, the preset picture database includes preset picture databases of multiple categories, for example, the preset picture database includes at least one picture database of a portrait snapshot database, a vehicle snapshot database, a plant and animal material database, a life knowledge material, a science and technology knowledge material, and a film and television knowledge material; the preset picture database is a structured picture database, that is, the preset picture database analyzes and processes each picture, extracts key information such as various attributes, features, tags or other metadata related to the picture content, and represents and stores each picture in a structured form; for example, the key information is the gender of the portrait snapshot, the color of the clothes, or the license plate number of the vehicle; the structured storage of key information related to the picture is conducive to quickly determining the picture to be retrieved according to the search statement, thereby improving the efficiency of picture retrieval, reducing the user's waiting time, and improving the user's experience of picture retrieval.

[0068] S200, based on the search sentence, predict the search elements corresponding to the search sentence through the constructed joint semantic prediction model.

[0069] In one embodiment, based on the search statement, the search elements corresponding to the search statement are predicted by the constructed joint semantic prediction model. The search elements predicted by the joint semantic prediction model reduce the scope of image retrieval, thereby improving the efficiency of image retrieval, reducing the user's waiting time, and improving the accuracy of image retrieval. It improves the user's retrieval efficiency, reduces the user's retrieval time, and further enhances the user's experience of image retrieval.

[0070] In one embodiment, the search elements include at least one of a search intent, an operator, an attribute category of a picture to be searched, or an attribute value corresponding to an attribute category, wherein the search intent represents the purpose of the user's image search, that is, which attribute categories the user wants to retrieve, which attribute values ​​corresponding to each attribute category, or pictures of at least one information in a range corresponding to the attribute values; the attribute category represents one or more information categories of the user's picture to be searched, such as information such as the gender, age, clothing color, or location of the person in the picture; the attribute value represents the numerical value corresponding to any attribute category; the operator represents the operation on the attribute category or attribute value to find a picture that better meets the search intent; through the search elements corresponding to the search statement, the scope of the image search is further reduced, thereby further improving the efficiency of the image search, reducing the user's waiting time, and further improving the accuracy of the image search, improving the user's search efficiency, reducing the user's search time, and further enhancing the user's experience of image search. For example, if the attribute category is age, the attribute operator can be "age is greater than / exceeds", "age is equal to / is", "age is less than / is not more than", and the attribute values ​​corresponding to age include "30 years old", "40 years old" and other specific ages; if the attribute category is gender, the operator can be "gender is", and the attribute values ​​corresponding to gender include "male" or "female"; if the attribute category is location, the operator can be "location is", "location is", "XX location" and so on, and the attribute values ​​corresponding to location include "Citizen Square", "People's Park" and other specific locations.

[0071] In one embodiment, the search statement "find snapshots of men in Civic Square" is intended to search for snapshots of portraits, and the attribute categories, operators, and attribute values ​​include: gender = male, location = Civic Square.

[0072] In one embodiment, the joint semantic prediction model is used to improve the intent recognition and slot filling model, such as Figure 2As shown, the joint semantic prediction model includes a Bert model, an intent classifier, an attribute classifier and an operator classifier, wherein the output end of the Bert model is connected to the intent classifier, the attribute classifier and the operator classifier respectively, and the output end of the intent classifier, the output end of the attribute classifier and the output end of the operator classifier are commonly connected to the output end of the joint semantic prediction model. Among them, the Bert model (Bidirectional Encoder Representations from Transformers) is a pre-trained language model based on the Transformer structure. It was proposed by Google in 2018. The Bert model enhances the semantic representation of word vectors by pre-training on large-scale unlabeled text data and fully considering the relationship features at the character level, word level, sentence level and between sentences. The learned semantic knowledge is applied to various downstream natural language processing (NLP) tasks through transfer learning, such as text classification, named entity recognition, sentence relationship judgment and other tasks, which enables the model to better mine the feature information of domain texts; the Transformer structure is a sequence-to-sequence model architecture based on the self-attention mechanism, which is used to process sequence data, such as text data in natural language processing. It was proposed in 2017 and has made major breakthroughs in machine translation tasks. The transformer structure abandons the traditional recurrent neural network (RNN) structure and adopts the self-attention mechanism. The self-attention mechanism enables the transformer structure to consider all positions in the input sequence at the same time, without the need to process step by step like the recurrent neural network, so that the transformer structure can be calculated in parallel, which speeds up the model training. Among them, intent recognition is a technology in the field of natural language processing, which is used to identify the user's intention or purpose in text or voice input, and is used to classify and understand the user's language request for further response and processing; slot filling is a method for filling specific information slots in natural language processing. It is used in dialogue systems or voice assistants to extract specific information from the user's input and fill it into predefined slots for subsequent processing and response. Compared with the Bert-based joint intent classification and slot filling model (i.e., the JointBert model), in this embodiment, the joint semantic prediction model includes a Bert model, an intent classifier, an attribute classifier, and an operator classifier. Due to the addition of the operator classifier, it can search within a smaller range of the preset image database according to the user's search statement for the image to be retrieved, further reducing the time for determining the image to be retrieved and improving the efficiency of image retrieval and user experience.

[0073] In one embodiment, based on a retrieval statement, a retrieval element corresponding to the retrieval statement is predicted through a constructed joint semantic prediction model, including: the retrieval statement obtains a retrieval element corresponding to the predicted retrieval statement through a Bert model, an intent classifier, an attribute classifier, and an operator classifier. Since a retrieval element including at least one of a retrieval intent, an operator, an attribute category of a picture to be retrieved, or an attribute value corresponding to the attribute category is adopted, it is possible to search for a specified retrieval element in a preset picture database according to the retrieval statement of the picture to be retrieved by the user, further reducing the time for determining the picture to be retrieved. The user does not need to perform frequent keyboard and mouse operations, greatly reducing the workload of the user and improving the efficiency of picture retrieval and the user experience.

[0074] In one embodiment, as Figure 3 shown, a joint semantic prediction model is constructed, including:

[0075] S210, input any training statement into the Bert model to obtain an intent hidden vector corresponding to the training statement and multiple slot value hidden vectors corresponding to the training statement.

[0076] In one embodiment, input any training statement into the Bert model. The Transformer structure of the Bert model performs bidirectional encoding on the training statement to obtain an intent hidden vector corresponding to the training statement and multiple slot value hidden vectors corresponding to the training statement, facilitating intent classification of the intent hidden vector and attribute category classification and operator classification of the multiple slot value hidden vectors, thereby improving the efficiency of picture retrieval.

[0077] In one embodiment, since one attribute value can continuously associate multiple text slot values (tokens), data preprocessing is performed on the training statement. According to the slot type annotation information given in the training statement, the text sequence in the training statement is converted into a BIO-formatted annotation sequence. The BIO form means that each element in the text sequence is annotated as "B-X", "I-X", or "O", where "B-X" indicates that the segment where this element is located belongs to the X type and this element is at the beginning of this segment; "I-X" indicates that the segment where this element is located belongs to the X type and this element is in the middle position of this segment; "O" indicates that it does not belong to any type. For example, four slot values (tokens) "city", "people", "square", and "field" form an attribute value, and when performing data annotation data preprocessing, they are respectively marked as B-location, I-location, I-location, and I-location.

[0078] In one embodiment, as Figure 4As shown, the training sentences are re-segmented according to the input requirements of Bert, and the words outside the built-in dictionary of Bert are split to obtain a new sentence sequence. For example, the training sentence is "Find snapshots of men in the Civic Square", and the training sentence is processed by word segmentation and slot value, that is, the training sentence "Find snapshots of men in the Civic Square" is segmented, and each word becomes a slot value (token). Then the new sentence sequence is annotated with intent and slot type. Intent labeling refers to numbering all intent categories according to predefined intent categories, and then marking all sentences with corresponding intent numbers according to the intent numbers. Similarly, all slot types need to be numbered, and then each word of the training sentence is numbered using the slot type number to construct a corresponding slot labeling sequence.

[0079] In one embodiment, Figure 4 As shown, a global classification slot value ([CLS]token) is added to the beginning of the training sentence text, which is generally used to represent the sentence information of the content text of the entire training sentence, and a separator slot value ([SEP]token) is added to the end of each sentence, and then each slot value (token) is mapped to a slot value code (tokenID) to identify the content text, and then the slot value code is input into the transformer structure of the Bert model. The transformer structure reads the feature representation corresponding to each slot value (token), encodes each slot value of the input content, and interactively processes and outputs the feature vector (features) corresponding to each slot value; for example, the global classification slot value ([CLS]token) outputs the corresponding intent hidden vector C, and multiple slot value hidden vectors corresponding to each slot value (token) of the training sentence, such as T1 to Tn, where n is a positive integer.

[0080] S220, training the joint semantic prediction model based on the intent hidden vector, each slot value hidden vector, the intent classifier, the attribute classifier and the operator classifier until the retrieval factor loss function value meets the preset conditions, thereby obtaining a trained joint semantic prediction model.

[0081] In one embodiment, the joint semantic prediction model is trained based on the intent hidden vector, each slot value hidden vector, the intent classifier, the attribute classifier and the operator classifier until the retrieval factor loss value meets the preset conditions, thereby obtaining a trained joint semantic prediction model. Since the operator classifier is added to the original intent recognition and slot filling model (JointBert model), it can search in a smaller range of the preset image database according to the user's retrieval statement of the image to be retrieved. At the same time, the joint semantic prediction model can simultaneously perform intent recognition, attribute category classification and operator classification, which reduces the space occupied by the joint semantic prediction model, improves the response speed of the joint semantic prediction model, further reduces the time to determine the image to be retrieved, and improves the efficiency of image retrieval of the joint semantic prediction model.

[0082] In one embodiment, the intent classifier, attribute classifier and operator classifier are all multi-layer perceptrons (MLP). The multi-layer perceptron is an artificial neural network structure in deep learning and is also the most basic feedforward neural network model. It consists of multiple fully connected layers (also called dense layers or fully connected layers). The neurons in each fully connected layer are connected to all neurons in the previous layer. The multi-layer perceptron includes an input layer, a hidden layer and an output layer, wherein the hidden layer includes one or more layers of fully connected layers. The intent classifier, attribute classifier and operator classifier are all multi-layer perceptrons, which are beneficial to improving the accuracy of intent classification, attribute category classification and operator classification. In other embodiments, the intent classifier, attribute classifier and operator classifier can also be other structures that can achieve classification, which are not specifically limited here.

[0083] In one embodiment, Figure 5 As shown, the joint semantic prediction model is trained based on the intent hidden vector, each slot value hidden vector, the intent classifier, the attribute classifier and the operator classifier until the retrieval factor loss value meets the preset conditions, and the trained joint semantic prediction model is obtained, including:

[0084] S221, obtaining a first loss value based on the intent hidden vector, the intent classifier and the true intent of the intent hidden vector.

[0085] In one embodiment, based on the intent hidden vector C, the intent classifier and the true intent of the intent hidden vector, a first loss value is obtained, and the first loss value is the loss value of the first loss function of intent classification. The intent classifier is screened according to the first loss value of the first loss function of intent classification, which can improve the accuracy of the output result of the intent classifier.

[0086] In one embodiment, a first loss value is obtained based on an intent hidden vector, an intent classifier, and the true intent of the intent hidden vector, including: inputting the intent hidden vector into the intent classifier to obtain the predicted intent corresponding to the intent hidden vector; and obtaining the first loss value based on the predicted intent and the true intent of the intent hidden vector.

[0087] S222, obtaining a second loss value based on each slot value hidden vector, the attribute classifier, and the true attribute of each slot value hidden vector.

[0088] In one embodiment, based on each slot value hidden vector, an attribute classifier and the true attribute of each slot value hidden vector, a second loss value is obtained, and the second loss value is the loss value of the second loss function of attribute category classification. The attribute classifier is screened according to the second loss value of the second loss function of attribute category classification, which can improve the accuracy of the output result of the attribute classifier.

[0089] In one embodiment, based on each slot value hidden vector, an attribute classifier and the true attributes of each slot value hidden vector, a second loss value is obtained, including: inputting each slot value hidden vector into the attribute classifier to obtain the predicted attributes corresponding to each slot value hidden vector; based on the predicted attributes and the true attributes of each slot value hidden vector, a second loss value is obtained.

[0090] S223, obtaining a third loss value based on each slot value hidden vector, the operator classifier, and the real operator of each slot value hidden vector.

[0091] In one embodiment, based on each slot value hidden vector, an operator classifier and the real operator of each slot value hidden vector, a third loss value is obtained, and the third loss value is the loss value of the third loss function of the operator classification. The operator classifier is screened according to the loss value of the third loss function of the operator classification, which can improve the accuracy of the output result of the operator classifier.

[0092] In one embodiment, a third loss value is obtained based on each slot value hidden vector, an operator classifier and the real operator of each slot value hidden vector, including: inputting each slot value hidden vector into the operator classifier to obtain a predicted operator corresponding to each slot value hidden vector; and obtaining a third loss value based on the predicted operator and the real operator of each slot value hidden vector.

[0093] S224, based on the first loss value, the second loss value, the third loss value and the preset weight, obtain the search factor loss value.

[0094] In one embodiment, based on the first loss value, the second loss value, the third loss value and the preset weights, it is helpful to set the preset weights according to various needs of image retrieval, thereby obtaining the loss value of the retrieval factor and improving the accuracy of the joint semantic prediction model in predicting the retrieval factor.

[0095] In one embodiment, the preset weights include a first preset weight corresponding to the first loss value, a second preset weight corresponding to the second loss value, and a third preset weight corresponding to the third loss value. By setting corresponding preset weights for the first loss value, the second loss value, and the third loss value, it is helpful to give different levels of attention to different classifiers according to the needs of image retrieval, thereby improving classification results with higher levels of attention.

[0096] In one embodiment, based on the first loss value, the second loss value, the third loss value and the preset weight, the retrieval factor loss value is obtained, including: based on the first loss value, the second loss value, the third loss value, the first preset weight, the second preset weight and the third preset weight, the retrieval factor loss value is obtained by calculating the retrieval factor loss function, which is beneficial to give preset weights to different classifiers according to the needs of image retrieval, thereby improving the matching degree between the prediction results of the joint semantic prediction model and user needs, thereby improving the user experience.

[0097] In one embodiment, the retrieval factor loss function calculation formula includes a first retrieval factor loss function calculation formula, and the first retrieval factor loss function calculation formula is:

[0098] L total =a×A 1 +b×B 2 +c×C 3

[0099] Among them, L total Represents the retrieval factor loss value, A 1 represents the first loss value, a represents the first preset weight, B 2 represents the second loss value, b represents the second preset weight, C 3 represents the third loss value, and c represents the third preset weight.

[0100] In another embodiment, the first loss function, the second loss function, and the third loss function are all cross entropy loss functions, and the optimal model can be selected according to the cross entropy loss function value, which is conducive to improving the prediction accuracy of the joint semantic prediction model. In other embodiments, the first loss function, the second loss function, and the third loss function may also be other loss functions, such as L1loss, MSELoss, etc. The first loss function, the second loss function, and the third loss function may be the same or different, and are not limited here.

[0101] In another embodiment, the retrieval factor loss function calculation formula also includes a second retrieval factor loss function calculation formula; based on the first loss function, the second loss function, the third loss function, the first preset weight, the second preset weight and the third preset weight, a composite retrieval factor loss function is obtained, and the optimized retrieval factor loss value is obtained through the second retrieval factor loss function calculation formula. Since the composite retrieval factor loss function integrates the classification results of the three classifiers, it has better compatibility with unevenly distributed prediction results. At the same time, through the sharing and complementarity of feature information of the three related classification tasks, the accuracy of the prediction results of the trained joint semantic prediction model is improved.

[0102] In another embodiment, the second search factor loss function is calculated as:

[0103] L total =a×L 1 +b×L 2 +c×L 3

[0104] Among them, L total represents the retrieval factor loss function, L 1 represents the first loss function, a represents the first preset weight, L 2 represents the second loss function, b represents the second preset weight, L 3 represents the third loss function, and c represents the third preset weight.

[0105] S225, if the retrieval factor loss value meets the preset conditions, a trained joint semantic prediction model is obtained.

[0106] In one embodiment, if the retrieval factor loss value meets the preset conditions, that is, the prediction results of intent classification, attribute category classification and operator classification have met the training requirements, a relatively accurate predicted retrieval factor corresponding to the training sentence is obtained, thereby obtaining a trained joint semantic prediction model.

[0107] In one embodiment, the preset condition is that the fluctuation range of the retrieval factor loss function value after a preset number of loop steps is less than or equal to a preset value (for example, 0.01, without restriction), wherein the preset number of loop steps is greater than or equal to a preset number of steps (for example, 100, without restriction), which is beneficial to improving the training efficiency of the joint semantic prediction model while improving the accuracy of the prediction results of the trained joint semantic prediction model.

[0108] In another embodiment, the preset condition is that the loss function value of the retrieval element satisfies the preset condition, for example, the loss function value is less than a preset value or the loss function value fluctuation is less than a preset value, which is not limited here.

[0109] In another embodiment, the preset condition is that the number of steps in the training meets the preset condition, for example, the number of cycle training steps reaches 100 steps. The preset condition can be adjusted according to actual conditions and is not limited here.

[0110] In one embodiment, Figure 6 As shown, Figure 6 The flowchart of training the joint semantic prediction model is as follows: after inputting the training sentence, the intent classification, attribute classification and operator classification are performed respectively after the Bert model; the predicted intent and the real intent obtained through the intent classification obtain a first loss value through the first loss function; the predicted attribute and the real attribute obtained through the attribute classification obtain a second loss value through the second loss function; the predicted operator and the real operator obtained through the operator classification obtain a third loss value through the third loss function; the first loss value (or first function), the first preset weight corresponding to the first loss value, the second loss value (or second function), the second preset weight corresponding to the second loss value, the third loss value (or third function), and the third preset weight corresponding to the third loss value are substituted into the retrieval factor function calculation formula to obtain the retrieval factor loss value, until the retrieval factor loss value meets the preset conditions, and the trained joint semantic prediction model is obtained.

[0111] S300, determining the image to be retrieved based on the retrieval elements and the preset image database.

[0112] In one embodiment, based on the retrieval elements and the preset image database, the type of the preset image database is determined by the retrieval intent, the images of the same attribute category of the preset image database are determined by the attribute category, the images of the range of attribute values ​​corresponding to each attribute category are determined by operator classification, and the images with the same attribute values ​​are determined by the attribute values ​​corresponding to the attribute categories, and finally the user's image to be retrieved is determined.

[0113] In one embodiment, Figure 7As shown in the figure, a structured preset image database is constructed in advance. Taking the user's voice "find a snapshot of a man in the Civic Square" as an example, the search sentence of the image to be retrieved is obtained, and each search element is predicted by the trained joint semantic prediction model. For example, the predicted search intention is to find a portrait image, the predicted attribute categories are gender and location, the predicted operators are "gender is" and "location is", the predicted attribute value corresponding to the attribute category "gender" is "male" and the attribute value corresponding to the attribute category "location" is "Citizen Square". Therefore, the image to be retrieved in the preset image database with the attribute categories of "gender" and "location", the attribute operators are "gender is" and "location is", and the attribute values ​​corresponding to each attribute category are "male" and "Citizen Square" are determined as the target image that the user wants to retrieve. After the target image is determined, the target image can be output to the user, so that the user obtains the image he wants to retrieve, and the user does not need to perform frequent keyboard and mouse operations, which greatly reduces the user's workload and improves the efficiency of image retrieval and user experience.

[0114] Compared with the prior art, the embodiments of the present invention have the following beneficial effects:

[0115] The image retrieval method of the present embodiment obtains a retrieval statement of the image to be retrieved and a preset image database; based on the retrieval statement, predicts the retrieval elements corresponding to the retrieval statement through the constructed joint semantic prediction model, the retrieval elements including the retrieval intent, the operator, the attribute category of the image to be retrieved or at least one of the attribute values ​​corresponding to the attribute category; based on the retrieval elements and the preset image database, determines the image to be retrieved, because the retrieval elements including the retrieval intent, the operator, the attribute category of the image to be retrieved or at least one of the attribute values ​​corresponding to the attribute category are adopted, the preset image database can be searched for the specified retrieval elements according to the user's retrieval statement of the image to be retrieved, and the image to be retrieved can be quickly determined, so the user does not need to perform frequent keyboard and mouse operations, which greatly reduces the user's workload and improves the efficiency of image retrieval and user experience.

[0116] The image retrieval device provided in this application is exemplarily described below with reference to the accompanying drawings.

[0117] In the second aspect, corresponding to the image retrieval method described in the above embodiment, Figure 8 As shown, this embodiment provides a device for image retrieval, and the image retrieval device 100 includes:

[0118] The acquisition module 110 is used to acquire the search sentence of the image to be retrieved and the preset image database.

[0119] The prediction module 120 is used to predict the search elements corresponding to the search statement based on the search statement through the constructed joint semantic prediction model. The search elements include at least one of the search intent, operators, attribute categories of the image to be searched, or attribute values ​​corresponding to the attribute categories.

[0120] The determination module 130 is used to determine the pictures to be retrieved based on the retrieval elements and the preset picture database.

[0121] It should be noted that the information interaction, execution process, etc. between the above-mentioned modules / units are based on the same concept as the method embodiment of the present application. Their specific functions and technical effects can be found in the method embodiment part and will not be repeated here.

[0122] The technicians in the relevant field can clearly understand that for the convenience and simplicity of description, only the division of the above-mentioned functional units and modules is used as an example for illustration. In practical applications, the above-mentioned function allocation can be completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiment can be integrated in a processing unit, or each unit can exist physically separately, or two or more units can be integrated in one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of software functional units. In addition, the specific names of the functional units and modules are only for the convenience of distinguishing each other, and are not used to limit the scope of protection of this application. The specific working process of the units and modules in the above-mentioned system can refer to the corresponding process in the aforementioned method embodiment, which will not be repeated here.

[0123] In a third aspect, this embodiment provides an electronic device 900, such as Fig. 9 As shown, it includes a memory 901, a processor 902, and a computer program 903 stored in the memory 901 and executable on the processor 902. When the processor 902 executes the computer program 903, it implements the method described in any one of the above-mentioned first aspects.

[0124] The image retrieval method provided in the embodiment of the present application can be applied to terminal devices such as mobile phones, tablet computers, wearable devices, vehicle-mounted devices, augmented reality (AR) / virtual reality (VR) devices, laptop computers, ultra-mobile personal computers (UMPC), netbooks, personal digital assistants (PDA), etc. The embodiment of the present application does not impose any restrictions on the specific type of terminal devices.

[0125] In applications, the electronic device may include, but is not limited to, a processor and a memory. Figure 8 It is only an example of a terminal device and does not constitute a limitation on the terminal device. It may include more or fewer components than shown in the figure, or combine certain components, or different components, such as input and output devices, network access devices, etc. Input and output devices may include cameras, audio acquisition / playback devices, display screens, etc. The network access device may include a network module for wirelessly communicating with external devices.

[0126] In applications, the processor may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or the processor may be any conventional processor, etc.

[0127] In applications, the memory may be an internal storage unit of a terminal device in some embodiments, such as a hard disk or memory of the terminal device. In other embodiments, the memory may also be an external storage device of the terminal device, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. equipped on the terminal device. The memory may also include both an internal storage unit and an external storage device of the terminal device. The memory is used to store operating systems, applications, boot loaders, data, and other programs, such as program codes of computer programs. The memory may also be used to temporarily store data that has been output or is to be output.

[0128] In a fourth aspect, this embodiment provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method as described in any one of the contents of the first aspect is implemented.

[0129] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. The present application implements all or part of the processes in the above-mentioned embodiment method, which can be completed by instructing the relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium, and the computer program can implement the steps of the above-mentioned method embodiments when executed by the processor. Among them, the computer program includes computer program code, and the computer program code can be in source code form, object code form, executable file or some intermediate form, etc.

[0130] The computer-readable medium may include at least: any entity or device capable of carrying the computer program code to a terminal device, a recording medium, a computer memory, a read-only memory (ROM), a random access memory (RAM), an electric carrier signal, a telecommunication signal, and a software distribution medium, such as a USB flash drive, a mobile hard disk, a magnetic disk, or an optical disk.

[0131] In the above embodiments, the description of each embodiment has its own emphasis. For parts that are not described or recorded in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0132] Those of ordinary skill in the art will appreciate that the devices and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.

[0133] In the embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interface, indirect coupling or communication connection of devices, which can be electrical, mechanical or other forms.

[0134] The embodiments described above are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, a person skilled in the art should understand that the technical solutions described in the aforementioned embodiments may still be modified, or some of the technical features may be replaced by equivalents. Such modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application, and should all be included in the protection scope of the present application.

Claims

1. A method for image retrieval, It is characterized in that include: Obtaining a search statement for the image to be retrieved and a preset image database; Based on the search statement, predicting the search elements corresponding to the search statement through the constructed joint semantic prediction model, the search elements including at least one of the search intent, the operator, the attribute category of the image to be searched, or the attribute value corresponding to the attribute category; Based on the search elements and the preset image database, the image to be searched is determined.

2. The method according to claim 1, It is characterized in that The joint semantic prediction model is an improved intent recognition and slot filling model, and the joint semantic prediction model includes a Bert model, an intent classifier, an attribute classifier and an operator classifier.

3. The method according to claim 2, It is characterized in that Based on the search statement, predicting the search elements corresponding to the search statement by using the constructed joint semantic prediction model includes: The search statement obtains the predicted search elements corresponding to the search statement through the Bert model, the intent classifier, the attribute classifier and the operator classifier.

4. The method according to claim 2, It is characterized in that Constructing the joint semantic prediction model includes: Input any training sentence into the Bert model, and obtain an intention hidden vector corresponding to the training sentence and a plurality of slot value hidden vectors corresponding to the training sentence; The joint semantic prediction model is trained based on the intention hidden vector, each slot value hidden vector, the intention classifier, the attribute classifier and the operator classifier until the retrieval factor loss value meets the preset conditions, thereby obtaining the trained joint semantic prediction model.

5. The method according to claim 4, It is characterized in that The method of training the joint semantic prediction model based on the intent hidden vector, each slot value hidden vector, the intent classifier, the attribute classifier, and the operator classifier until the retrieval factor loss value meets a preset condition to obtain the trained joint semantic prediction model includes: Based on the intent hidden vector, the intent classifier, and the true intent of the intent hidden vector, obtaining a first loss value, where the first loss value is a loss value of a first loss function for intent classification; Based on each of the slot value hidden vectors, the attribute classifier, and the true attribute of each of the slot value hidden vectors, obtaining a second loss value, where the second loss value is a loss value of a second loss function for attribute category classification; Based on each of the slot value hidden vectors, the operator classifier, and the real operator of each of the slot value hidden vectors, obtaining a third loss value, where the third loss value is a loss value of a third loss function of operator classification; Based on the first loss value, the second loss value, the third loss value and a preset weight, obtaining a search factor loss value; If the retrieval factor loss value meets the preset condition, the trained joint semantic prediction model is obtained.

6. The method according to claim 5, It is characterized in that The obtaining a first loss value based on the intention hidden vector, the intention classifier, and the true intention of the intention hidden vector includes: Inputting the intention hidden vector into the intention classifier to obtain a predicted intention corresponding to the intention hidden vector; Obtaining a first loss value based on the predicted intent and the true intent of the intent hidden vector; The obtaining of a second loss value based on each of the slot value hidden vectors, the attribute classifier, and the true attribute of each of the slot value hidden vectors includes: Inputting each slot value hidden vector into the attribute classifier to obtain a predicted attribute corresponding to each slot value hidden vector; Obtaining a second loss value based on the predicted attribute and the true attribute of each of the slot value hidden vectors; The obtaining a third loss value based on each of the slot value hidden vectors, the operator classifier, and the real operator of each of the slot value hidden vectors includes: Inputting each slot value hidden vector into the operator classifier to obtain a prediction operator corresponding to each slot value hidden vector; A third loss value is obtained based on the predicted operator and the true operator of each slot value hidden vector.

7. The method according to claim 5, It is characterized in that The preset weights include a first preset weight corresponding to the first loss value, a second preset weight corresponding to the second loss value, and a third preset weight corresponding to the third loss value; The acquiring the search factor loss value based on the first loss value, the second loss value, the third loss value and the preset weight includes: Based on the first loss value, the second loss value, the third loss value, the first preset weight, the second preset weight and the third preset weight, the retrieval factor loss value is obtained by using a retrieval factor loss function calculation formula.

8. A device for image retrieval, It is characterized in that include: An acquisition module is used to acquire a search statement of a picture to be retrieved and a preset picture database; A prediction module, configured to predict, based on the search statement, a search element corresponding to the search statement through a constructed joint semantic prediction model, wherein the search element includes at least one of a search intent, an operator, an attribute category of the image to be searched, or an attribute value corresponding to the attribute category; A determination module is used to determine the image to be retrieved based on the search elements and the preset image database.

9. A terminal device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, It is characterized in that When the processor executes the computer program, the method according to any one of claims 1 to 7 is implemented.

10. A computer-readable storage medium storing a computer program. It is characterized in that When the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.