Image-text retrieval method and device, server and computer readable storage medium

By obtaining the search keywords with the highest similarity to the natural language search statements entered by the user, and using the dynamic semantic keyword association relationship of natural language and pictures to expand and rewrite the search statements, it solves the problem of low image search recall caused by users' difficulty in accurately describing professional vocabulary, and achieves higher image search recall and better user experience.

CN120045731APending Publication Date: 2025-05-27GUANGZHOU AUTOMOBILE GROUP CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510020850.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-06
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

When users conduct natural language searches, it is difficult for users to accurately describe vocabulary in their professional fields, resulting in the inability to be hit and recalled by natural language search statements entered by users. The recall rate of image search is not high and cannot meet business needs.

Method used

By obtaining the search keywords with the highest similarity to the natural language search statements currently input by the user, based on the association relationship between the natural language search keywords and the dynamic semantic keywords of the picture, obtaining the dynamic semantic keywords associated with the target search keywords, and expanding and rewriting the natural language search statements currently input by the user, and finally searching the picture based on the expanded and rewrite statements.

Benefits of technology

It improves the recall rate of image retrieval, can better meet business needs, and improves user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120045731A_ABST
    Figure CN120045731A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides an image-text retrieval method and device, a server and a computer readable storage medium, and relates to the technical field of application of the Internet of Vehicles. According to the method, the retrieval keyword which has the highest similarity with the natural language retrieval statement currently input by the user can be obtained to serve as the target retrieval keyword; obtaining a dynamic semantic keyword associated with the target retrieval keyword based on an association relationship between the retrieval keyword of the natural language and the dynamic semantic keyword of the picture, and taking the dynamic semantic keyword as a target dynamic semantic keyword; according to the target dynamic semantic keyword, a natural language retrieval statement currently input by the user is expanded and rewritten; and retrieving the picture based on the expanded and rewritten statement. By means of the image-text retrieval method, the recall rate of image retrieval can be increased, and the user experience is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of vehicle networking applications, and more specifically, to a method, device, server, and computer-readable storage medium for graphic and text retrieval. Background Art

[0002] Currently, for the intelligent driving cloud data platform business, it is necessary to retrieve the vehicle-end collected picture data stored in the warehouse to obtain eligible pictures for simulation verification and perception training.

[0003] In the field of machine learning, multi-modal deep learning models are often used to process different modal data, such as pictures and natural language texts. Multi-modal deep learning models can capture the relationships between different modal data. For example, whether the text description matches the picture. A multi-modal pre-training model based on contrastive learning can convert text and pictures into low-dimensional vectors and predict the correlation of the two modal data through the similarity between the vectors. A multi-modal model pre-trained with a large-scale text-picture pair can be directly applied to downstream tasks. For example, in the intelligent driving business, natural language is used to retrieve pictures.

[0004] As Figure 1 shown, the prior art uses a multi-modal deep learning model to extract feature vectors from the vehicle-end collected pictures stored in the warehouse and store them in a vector database. When a user inputs a natural language retrieval statement, after using the multi-modal deep learning model to extract the natural language retrieval statement into a feature vector, vector approximate retrieval is performed in the vector database, so as to obtain driving pictures semantically related to the natural language retrieval statement input by the user.

[0005] However, when users perform natural language retrieval, it is often difficult to accurately describe the vocabulary in the professional field, resulting in some picture data not being hit and recalled by the natural language retrieval statements input by users, resulting in a low recall rate of picture retrieval and unable to well meet the business requirements. Summary of the Invention

[0006] Embodiments of this application propose a method, device, server, and computer-readable storage medium for graphic and text retrieval to solve the above problems.

[0007] In a first aspect, an embodiment of this application provides a method for graphic and text retrieval. The method includes: obtaining the retrieval keyword with the highest similarity to the natural language retrieval statement currently input by the user as the target retrieval keyword; obtaining the dynamic semantic keyword associated with the target retrieval keyword as the target dynamic semantic keyword based on the association relationship between the natural language retrieval keyword and the dynamic semantic keyword of the picture; expanding and rewriting the natural language retrieval statement currently input by the user according to the target dynamic semantic keyword; retrieving pictures based on the statement after expansion and rewriting.

[0008] In a second aspect, an embodiment of the present application provides a graphic and text retrieval device, which includes: a retrieval keyword acquisition module, configured to acquire a retrieval keyword with the highest similarity to the natural language retrieval statement currently input by the user as the target retrieval keyword; a dynamic semantic keyword acquisition module, configured to acquire a dynamic semantic keyword associated with the target retrieval keyword as the target dynamic semantic keyword based on the association relationship between the natural language retrieval keyword and the dynamic semantic keyword of the picture; an expansion and rewriting module, configured to expand and rewrite the natural language retrieval statement currently input by the user according to the target dynamic semantic keyword; and a picture retrieval module, configured to retrieve pictures based on the statement after expansion and rewriting.

[0009] In a third aspect, an embodiment of the present application provides a server, which includes: a memory and a processor, where an application program is stored in the memory, and the application program is configured to cause the processor to execute the method provided by the embodiment of the present application when being called by the processor.

[0010] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, on which program code is stored, and the program code is configured to cause the processor to execute the method provided by the embodiment of the present application when being called by the processor.

[0011] The graphic and text retrieval method provided by the embodiment of the present application has the following technical effects: by acquiring a retrieval keyword with the highest similarity to the natural language retrieval statement currently input by the user as the target retrieval keyword; acquiring a dynamic semantic keyword associated with the target retrieval keyword as the target dynamic semantic keyword based on the association relationship between the natural language retrieval keyword and the dynamic semantic keyword of the picture; and expanding and rewriting the natural language retrieval statement currently input by the user according to the target dynamic semantic keyword, so that dynamic semantic keywords highly associated with the natural language retrieval statement currently input by the user can be extracted, and the natural language retrieval statement currently input by the user can be rewritten in real time during picture retrieval, improving the recall rate of picture retrieval, well meeting the business requirements, and improving the user experience. Description of the Drawings

[0012] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present application, rather than all embodiments. Based on the embodiments of the present application, all other embodiments and drawings obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.

[0013] Figure 1 The structural schematic diagram of an existing picture retrieval device provided by an embodiment of the present application is shown;

[0014] Figure 2 Shows a schematic structural diagram of a graphic and text retrieval device provided in an embodiment of the present application;

[0015] Figure 3 Shows a schematic diagram of a retrieval - click graph provided in an embodiment of the present application;

[0016] Figure 4 Shows a schematic flowchart of a graphic and text retrieval method provided in an embodiment of the present application;

[0017] Figure 5 Shows a schematic flowchart of a graphic and text retrieval method provided in another embodiment of the present application;

[0018] Figure 6 Shows a schematic flowchart of step S220 provided in an embodiment of the present application;

[0019] Figure 7 Shows a schematic flowchart of step S224 provided in an embodiment of the present application;

[0020] Figure 8 Shows a schematic structural diagram of a graphic and text retrieval device provided in another embodiment of the present application;

[0021] Figure 9 Shows a schematic structural diagram of a server provided in an embodiment of the present application. Detailed implementation manners

[0022] In order to enable those skilled in the art to better understand the solutions of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present application.

[0023] Please refer to Figure 2 , Figure 2 which shows a schematic structural diagram of a graphic and text retrieval device provided in an embodiment of the present application. As Figure 2 shown, the graphic and text retrieval device 100 may include a picture storage module 110, a log recording module 120, a log mining module 130, and a retrieval module 140.

[0024] The picture storage module 110 provides a picture feature mining function, including: picture verification, using a multi - modal model to extract features from the picture and storing the feature vectors in a vector database after obtaining the feature vectors, and using dynamic semantic mining rules to perform dynamic semantic mining on the picture and storing the dynamic semantic data in a graph database after obtaining the dynamic semantic data.

[0025] The log recording module 120 provides a log recording function. When a user performs an image search, the user's input request, natural language or graph database language (such as cypher language), and the corresponding image click record after obtaining the search results are stored in the log database.

[0026] The log mining module 130 provides a log mining function for mining the search click records in the log database to obtain Figure 3 The search-click diagram shown. Figure 3 As shown, the black circle represents the natural language search record of the user in the log, and the middle white circle connected to the black circle by a black arrow represents the picture clicked by the user after entering the natural language search statement, indicating that the two are semantically related. The gray circle on the right represents the graph database language (such as cypher language) search record of the user in the log, and the middle white circle connected to it by a gray arrow represents the picture clicked by the user after entering the graph database language, indicating that the two are semantically related. With the picture as the intermediate node, the association relationship between the natural language search statement and the graph database search statement is obtained. Based on the search-click graph, the log mining module 130 executes steps S210 to S220 to establish a rewrite index.

[0027] The retrieval module 140 provides a user retrieval function, which executes steps S110 to S140 or steps S230 to S260 in the image retrieval stage, thereby extracting dynamic semantic keywords that are highly associated with the natural language retrieval statements (text) input by the user based on the association between the natural language retrieval keywords and the dynamic semantic keywords of the image, and rewriting the natural language retrieval statements input by the user in real time, thereby improving the recall rate of image retrieval.

[0028] The image and text retrieval method in the embodiment of the present application can be applied to the image and text retrieval device 100 or the image and text retrieval device 200 or the server 300 mentioned below. The image and text retrieval device 100 and the image and text retrieval device 200 can be integrated in the server 300. Among them, the server can be connected to the vehicle to realize data exchange. The server can be an independent physical server, or it can be a server cluster or distributed system composed of multiple physical servers, or it can be a cloud server that provides cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (Content Delivery Network, referred to as CDN) and big data and artificial intelligence platforms and other basic cloud computing services. Next, the image and text retrieval method of the present application will be introduced.

[0029] See also Figure 4 , Figure 4The flowchart of the graphic and text retrieval method provided by an embodiment of the present application is shown. As Figure 4 shown, the graphic and text retrieval method may include steps S110 to S140.

[0030] Step S110: Obtain the retrieval keyword with the highest similarity to the natural language retrieval statement currently input by the user as the target retrieval keyword.

[0031] The user can use natural language for image retrieval. According to different user requirements, the user can input a natural language retrieval statement (retrieval statement) to the retrieval module. The retrieval module can retrieve images associated with the natural language retrieval statement in the database according to steps S110 to S140 based on the natural language retrieval statement.

[0032] After the user inputs a natural language retrieval statement to the retrieval module, the retrieval module can obtain the natural language retrieval statement currently input by the user; extract features from the natural language retrieval statement currently input by the user to obtain the current retrieval keyword; obtain the retrieval keyword with the highest similarity to the current retrieval keyword as the target retrieval keyword. The specific implementation manner of obtaining the target retrieval keyword can refer to step S230 below.

[0033] Step S120: Based on the association relationship between the retrieval keyword of natural language and the dynamic semantic keyword of the image, obtain the dynamic semantic keyword associated with the target retrieval keyword as the target dynamic semantic keyword.

[0034] The retrieval keyword of natural language may include: the feature vector obtained by extracting features from the image using a multimodal model. The dynamic semantic keyword of the image may include: using dynamic semantic mining rules to perform dynamic semantic mining on the image to obtain dynamic semantic data. The association relationship between the retrieval keyword of natural language and the dynamic semantic keyword of the image can be established in advance based on the above-mentioned retrieval-click graph, and specifically can refer to steps S210 to S220 below. Among them, dynamic semantics refers to the semantics during the operation of the image retrieval program, representing the logical meaning.

[0035] Each retrieval keyword may correspond to at least one dynamic semantic keyword of the image. Based on the association relationship between the retrieval keyword of natural language and the dynamic semantic keyword of the image, the dynamic semantic keyword highly associated with the natural language retrieval statement currently input by the user can be extracted, and the natural language retrieval statement currently input by the user can be expanded and rewritten according to the extracted dynamic semantic keyword, so as to improve the recall rate of image retrieval.

[0036] Step S130: Expand and rewrite the natural language retrieval statement currently input by the user according to the target dynamic semantic keyword.

[0037] There can be at least one target dynamic semantic keyword. When there are multiple target dynamic semantic keywords, the similarity between each target dynamic semantic keyword and the current retrieval keyword can be calculated; at least one target dynamic semantic keyword with the highest similarity to the current retrieval keyword can be obtained (for the specific implementation of obtaining at least one target dynamic semantic keyword, refer to step S250 below); according to the at least one target dynamic semantic keyword, the natural language retrieval statement currently input by the user is expanded and rewritten. Among them, the expansion and rewriting can expand and rewrite the natural language retrieval statement currently input by the user according to certain preset rules, or randomly insert the target dynamic semantic keyword into the natural language retrieval statement currently input by the user to expand and rewrite the natural language retrieval statement currently input by the user. The present application does not limit the specific manner of expansion and rewriting.

[0038] Step S140: Retrieve pictures based on the statement after expansion and rewriting.

[0039] After expanding and rewriting the natural language retrieval statement currently input by the user based on the rewriting of the target dynamic semantic keyword, a statement after expansion and rewriting is obtained. Retrieving pictures based on the statement after expansion and rewriting can improve the recall rate of picture retrieval.

[0040] Steps S110 to S140 have the following technical effects: by obtaining the retrieval keyword with the highest similarity to the natural language retrieval statement currently input by the user as the target retrieval keyword; based on the association relationship between the natural language retrieval keyword and the dynamic semantic keyword of the picture, obtaining the dynamic semantic keyword associated with the target retrieval keyword as the target dynamic semantic keyword; according to the target dynamic semantic keyword, expanding and rewriting the natural language retrieval statement currently input by the user, so that the dynamic semantic keyword highly associated with the natural language retrieval statement currently input by the user can be extracted, and the natural language retrieval statement currently input by the user is rewritten in real time during picture retrieval, improving the recall rate of picture retrieval, being able to well meet the business requirements, and improving the user experience.

[0041] Please refer to Figure 5 , Figure 5 which shows the flowchart of the picture and text retrieval method provided by another embodiment of the present application. As Figure 5 shown, the picture and text retrieval method further includes steps S210 to S260.

[0042] Step S210: Obtain the picture retrieval record logs of multiple users, and at least the natural language retrieval records and the graph database language retrieval records of the users are recorded in the picture retrieval record logs.

[0043] The natural language retrieval records of users include: the natural language retrieval statements input by users and the pictures clicked by users after inputting the natural language retrieval statements. After the users input the natural language retrieval statements, the retrieval module will display picture retrieval results associated with the natural language retrieval statements to the users, and the users can select the pictures that meet their needs by clicking on the pictures.

[0044] The graph database language (such as Cypher language) retrieval records of users include: the graph database language retrieval statements input by users who are proficient in using the graph database language and the pictures clicked by users after inputting the graph database language retrieval statements. After the users input the graph database language statements, the retrieval module will display picture retrieval results associated with the graph database language statements to the users, and the users can select the pictures that meet their needs by clicking on the pictures.

[0045] The picture retrieval record logs of all users are stored in the log database, and the log mining module can mine the picture retrieval record logs of users to establish the association relationship between the retrieval keywords in natural language and the dynamic semantic keywords of pictures. According to the actual picture retrieval accuracy requirements and hardware conditions, the log mining module can mine all the currently stored picture retrieval record logs of users, or can select the picture retrieval logs of multiple users from all the currently stored picture retrieval logs of users for mining. This application does not limit the specific number of users' picture retrieval record logs obtained.

[0046] Step S220: Establish the association relationship between the retrieval keywords in natural language and the dynamic semantic keywords of pictures according to the natural language retrieval records and the graph database language retrieval records.

[0047] Please refer to Figure 6 , Figure 6 which shows the flow schematic diagram of step S220 provided by an embodiment of this application. As Figure 6 shown, step S220 may include step S221 to step S224.

[0048] Step S221: Obtain the natural language retrieval statements and the pictures clicked by users after inputting the natural language retrieval statements from the natural language retrieval records, and obtain the association relationship between the natural language retrieval statements and the pictures.

[0049] As Figure 3 shown, the natural language retrieval statements can be represented by the black circles in Figure 3 , and the pictures clicked by users after inputting the natural language retrieval statements can be represented by the white circles in Figure 3 . The black arrows pointed by the black circles and pointing to the white circles can represent the association relationship between the natural language retrieval statements and the pictures.

[0050] Step S222: From the records retrieved by the graph database language, obtain the graph database retrieval statement and the picture clicked by the user after inputting the graph database retrieval statement, and obtain the association relationship between the graph database retrieval statement and the picture.

[0051] As Figure 3 shown, the graph database retrieval statement can be represented by the gray circle in Figure 3 , and the picture clicked by the user after inputting the graph database retrieval statement can be represented by the white circle in Figure 3 . The gray arrow pointed by the gray circle and pointing to the white circle can represent the association relationship between the graph database retrieval statement and the picture.

[0052] Step S223: According to the association relationship between the natural language retrieval statement and the picture and the association relationship between the graph database retrieval statement and the picture, establish the association relationship between the natural language retrieval statement and the graph database retrieval statement through the picture as an intermediate node.

[0053] As Figure 3 shown in the retrieval-click graph, through the picture as an intermediate node, the association relationship between the natural language retrieval statement and the graph database retrieval statement can be established. As Figure 3 shown, taking the middle white circle as the intermediate node, the connection relationship formed by the black arrow pointed by the black circle and pointing to the white circle - this white circle - and the gray arrow pointed by this white circle and pointing to the gray circle can represent the association relationship between the natural language retrieval statement and the graph database retrieval statement.

[0054] Step S224: According to the association relationship between the natural language retrieval statement and the graph database retrieval statement, establish the association relationship between the retrieval keyword of the natural language and the dynamic semantic keyword of the picture.

[0055] Please refer to Figure 7 , Figure 7 which shows the schematic flow chart of step S224 provided by an embodiment of the present application. As Figure 7 shown, step S224 may include steps S2241 to S2246.

[0056] Step S2241: Extract features from the natural language retrieval statement to obtain the retrieval keyword of the natural language.

[0057] The log mining module can use the text encoder in the multimodal model to extract features from the natural language retrieval statement to obtain the natural language retrieval feature vector, that is, the retrieval keyword of the natural language.

[0058] Step S2242: Cluster the retrieval keywords of natural language to obtain a set of clusters, where each cluster has a cluster center and a set of natural language retrieval statements.

[0059] The log mining module can use the IVF (Inverted File System) indexing algorithm to build a clustering index for the natural language retrieval feature vectors (i.e., the retrieval keywords of natural language). This index can be represented as a set of partitions (i.e., a set of clusters) P = {p 1 , p 2 , …, p n}, and each partition (cluster) corresponds to a cluster center (vector). All cluster centers form a cluster center set O = {o 1 , o 2 , …, o n}.

[0060] After obtaining the cluster center set, the association relationship between the cluster center and the dynamic semantic keywords of the picture can be established according to the association relationship between the natural language retrieval statement and the graph database retrieval statement. Among them, the specific implementation of establishing the association relationship between each cluster center and the dynamic semantic keywords of the picture can include Step S2243 to Step S2246.

[0061] Step S2243: For each cluster center, according to the association relationship between the natural language retrieval statement and the graph database retrieval statement, obtain the set of graph database retrieval statements associated with the set of natural language retrieval statements corresponding to this cluster center.

[0062] For each partition p i in the clustering index, each partition has a cluster center and a set of natural language retrieval statements p i = {q i1 , q i2 , …, q in}, and this cluster center and the set of natural language retrieval statements have a corresponding relationship. For each cluster center, through the association relationship between the natural language retrieval statement and the graph database retrieval statement in the above retrieval - click graph, the set of graph database retrieval statements (such as cypher retrieval statements) C i = {c i1 , c i2 , …, c in} associated with the set of natural language retrieval statements corresponding to this cluster center can be found.

[0063] Step S2244: Extract the keywords in the set of graph database retrieval statements to obtain the set of dynamic semantic keywords of the picture.

[0064] The log mining module can process the set of graph database retrieval statements C i = {ci1 , c i2 , …, c in Perform keyword extraction processing and deduplication processing on the graph database retrieval statements (such as Cypher retrieval statements) in}, and obtain the dynamic semantic keyword set T of the picture i = {t i1 , t i2 , …, t in}.

[0065] Step S2245: Vectorize the dynamic semantic keyword set to obtain a dynamic semantic keyword vector set.

[0066] The log mining module can use the text encoder in the multimodal model to vectorize the dynamic semantic keyword set T i = {t i1 , t i2 , …, t in}, and obtain the dynamic semantic keyword vector set

[0067] Step S2246: Establish an association relationship between the cluster center and the dynamic semantic keyword vector set.

[0068] After obtaining the dynamic semantic keyword vector set , an association relationship between the cluster center o i and the dynamic semantic keyword vector set V i can be established (o i , V i ).

[0069] According to steps S2243 to S2246, for each cluster center in the cluster center set, a corresponding dynamic semantic keyword vector set can be found. The dynamic semantic keyword vector set includes at least one dynamic semantic keyword vector. After obtaining all cluster centers and their corresponding dynamic semantic keyword vector sets, a correspondence set W = {(o 1 , V 1 ), (o 2 , V 2 ), …, (o n , V n )} can be obtained.

[0070] Step S230: Obtain the retrieval keyword with the highest similarity to the natural language retrieval statement currently input by the user as the target retrieval keyword.

[0071] The retrieval module can obtain the natural language retrieval statement currently input by the user; use the text encoder in the multimodal model to extract features from the natural language retrieval statement currently input by the user to obtain a retrieval feature vector (i.e., the current retrieval keyword). The retrieval module can use the retrieval feature vector to compare the similarity with each cluster center in the cluster center set O, and select the cluster center o i (i.e., the target retrieval keyword) that is most similar to the retrieval feature vector u.

[0072] Step S240: Based on the association relationship between the retrieval keyword in natural language and the dynamic semantic keyword of the picture, obtain the dynamic semantic keyword associated with the target retrieval keyword as the target dynamic semantic keyword.

[0073] The association relationship between the retrieval keyword in natural language and the dynamic semantic keyword of the picture is the correspondence set W. The retrieval module can find in the correspondence set W the dynamic semantic keyword vector set V i (i.e., the target retrieval keyword) corresponding to i (i.e., the target dynamic semantic keyword).

[0074] Step S250: Expand and rewrite the natural language retrieval statement currently input by the user according to the target dynamic semantic keyword.

[0075] There is at least one target dynamic semantic keyword. When there are multiple target dynamic semantic keywords, calculate the similarity between each target dynamic semantic keyword and the current retrieval keyword; obtain at least one target dynamic semantic keyword with the highest similarity to the current retrieval keyword; expand and rewrite the natural language retrieval statement currently input by the user according to the at least one target dynamic semantic keyword.

[0076] Specifically, the retrieval module can calculate the retrieval feature vector (i.e., the current retrieval keyword) and each dynamic semantic keyword vector in the dynamic semantic keyword vector set V i respectively, select k (k is greater than or equal to 1) dynamic semantic keyword vectors that are most similar to the retrieval feature vector, and use the corresponding keywords to expand and rewrite the natural language retrieval statement currently input by the user.

[0077] Step S260: Retrieve pictures based on the statement after expansion and rewriting.

[0078] Steps S210 to S260 further have the following technical effects: By mining the picture retrieval record logs of users, a semantic association relationship between natural language retrieval and graph database (such as Cypher) language retrieval is established. By establishing an index, during the natural language retrieval stage, dynamic semantic keywords are supplemented in real time, improving the retrieval recall rate. When users cannot clearly express dynamic semantic information, the recall rate of picture retrieval is increased, which can better meet business requirements and improve the user experience.

[0079] Please refer to Figure 8 , Figure 8 which shows a schematic structural diagram of a graphic and text retrieval device provided by another embodiment of the present application. As Figure 8 shown, the graphic and text retrieval device 200 may include: a retrieval keyword acquisition module 210, a dynamic semantic keyword acquisition module 220, an expansion and rewriting module 230, and a picture retrieval module 240.

[0080] The retrieval keyword acquisition module 210 is configured to acquire a retrieval keyword with the highest similarity to the natural language retrieval statement currently input by the user as the target retrieval keyword.

[0081] The dynamic semantic keyword acquisition module 220 is configured to acquire a dynamic semantic keyword associated with the target retrieval keyword as the target dynamic semantic keyword based on the association relationship between the retrieval keyword in natural language and the dynamic semantic keyword of the picture.

[0082] The expansion and rewriting module 230 is configured to expand and rewrite the natural language retrieval statement currently input by the user according to the target dynamic semantic keyword.

[0083] The picture retrieval module 240 is configured to retrieve pictures based on the statement after expansion and rewriting.

[0084] In some embodiments, the graphic and text retrieval device 200 may further include a log mining module. The log mining module may be configured to: acquire picture retrieval record logs of multiple users, where the picture retrieval record logs at least record the natural language retrieval records and graph database language retrieval records of the users; establish an association relationship between the retrieval keywords in natural language and the dynamic semantic keywords of the pictures according to the natural language retrieval records and graph database language retrieval records.

[0085] In some embodiments, the log mining module can also be used to: obtain, from natural language retrieval records, natural language retrieval statements and the pictures clicked by the user after inputting the natural language retrieval statements, so as to obtain the association relationship between the natural language retrieval statements and the pictures; obtain, from graph database language retrieval records, graph database retrieval statements and the pictures clicked by the user after inputting the graph database retrieval statements, so as to obtain the association relationship between the graph database retrieval statements and the pictures; according to the association relationship between the natural language retrieval statements and the pictures and the association relationship between the graph database retrieval statements and the pictures, establish the association relationship between the natural language retrieval statements and the graph database retrieval statements by using the pictures as intermediate nodes; according to the association relationship between the natural language retrieval statements and the graph database retrieval statements, establish the association relationship between the retrieval keywords of the natural language and the dynamic semantic keywords of the pictures.

[0086] In some embodiments, the log mining module can also be used to: extract features from natural language retrieval statements to obtain the retrieval keywords of the natural language; cluster the retrieval keywords of the natural language to obtain a set of clusters, and each cluster has a cluster center and a set of natural language retrieval statements; according to the association relationship between the natural language retrieval statements and the graph database retrieval statements, establish the association relationship between the cluster center and the dynamic semantic keywords of the pictures.

[0087] In some embodiments, the log mining module can also be used to: for each cluster center, according to the association relationship between the natural language retrieval statements and the graph database retrieval statements, obtain the set of graph database retrieval statements associated with the set of natural language retrieval statements corresponding to the cluster center; extract the keywords in the set of graph database retrieval statements to obtain a set of dynamic semantic keywords of the pictures; vectorize the set of dynamic semantic keywords to obtain a set of dynamic semantic keyword vectors; establish the association relationship between the cluster center and the set of dynamic semantic keyword vectors.

[0088] In some embodiments, the retrieval keyword acquisition module 210 is further configured to obtain the natural language retrieval statement currently input by the user; extract features from the natural language retrieval statement currently input by the user to obtain the current retrieval keyword; obtain the retrieval keyword with the highest similarity to the current retrieval keyword as the target retrieval keyword.

[0089] In some embodiments, when there are multiple target dynamic semantic keywords, the expansion and rewriting module 230 is further configured to calculate the similarity between each target dynamic semantic keyword and the current retrieval keyword; obtain at least one target dynamic semantic keyword with the highest similarity to the current retrieval keyword; and expand and rewrite the natural language retrieval statement currently input by the user according to the at least one target dynamic semantic keyword.

[0090] In some embodiments, the graphic and text retrieval device 200 may further include the picture storage module in the graphic and text retrieval device 100.

[0091] Those skilled in the art can clearly understand that the above devices provided by the embodiments of the present application can implement the methods provided by the embodiments of the present application. For the specific working processes of the above-described devices and modules, reference may be made to the corresponding processes of the methods in the embodiments of the present application, which will not be elaborated herein.

[0092] In the embodiments provided by the present application, the coupling, direct coupling, or communication connection between the modules shown or discussed with each other may be indirect coupling or communication coupling through some interfaces, devices, or modules, and may be in electrical, mechanical, or other forms. The embodiments of the present application do not make specific limitations thereto.

[0093] In addition, in the embodiments of the present application, the various functional modules may be integrated into one processing module, or each module may exist physically alone, or two or more modules may be integrated into one module. The above integrated modules may be implemented in the form of hardware or in the form of software functional modules.

[0094] Please refer to Figure 9 , Figure 9 which shows a schematic structural diagram of a server provided by an embodiment of the present application. As Figure 9 shown, the server 300 may include a memory 310 and a processor 320. An application program is stored in the memory 310, and the application program is configured to cause the processor 320 to execute the method provided by the embodiment of the present application when called by the processor 320.

[0095] The processor 320 may include one or more processing cores. The processor 320 uses various interfaces and lines to connect various parts within the entire server 300, and is used to run or execute instructions, programs, code sets, or instruction sets stored in the memory 310, and to call and run or execute data stored in the memory 310, so as to execute various functions of the server 300 and process data.

[0096] The processor 320 can be implemented in at least one hardware form of digital signal processing (DSP), field-programmable gate array (FPGA), or programmable logic array (PLA). The processor 320 can integrate one or a combination of several of a central processing unit (CPU), a graphics processing unit (GPU), and a modem. Among them, the CPU mainly processes the operating system, user interface, application programs, etc.; the GPU is responsible for rendering and drawing display content; the modem is used to process wireless communication. It can be understood that the above-mentioned modem may not be integrated in the processor 320 and can be implemented separately through a communication chip.

[0097] The memory 310 can include random access memory (RAM) and can also include read-only memory (ROM). The memory 310 can be used to store instructions, programs, codes, code sets, or instruction sets. The memory 310 can include a program storage area and a data storage area. Among them, the program storage area can store instructions for implementing the operating system, instructions for implementing at least one function, instructions for implementing the above method embodiments, etc. The data storage area can store data created by the server 300 during use.

[0098] The embodiments of the present application also provide a computer-readable storage medium, on which program code is stored, and the program code is configured to execute the method provided by the embodiments of the present application when called by a processor.

[0099] The computer-readable storage medium can be an electronic memory such as flash memory, electrically-erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), a hard disk, or ROM.

[0100] In some embodiments, the computer-readable storage medium includes a non-transitory computer-readable medium (Non-Transitory Computer-Readable Storage Medium, abbreviated as Non-TCRSM). The computer-readable storage medium has a storage space for program codes that execute any method steps in the above methods. These program codes can be read out from one or more computer program products or written into these one or more computer program products. The program codes can be compressed in a suitable form.

[0101] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application and are not intended to limit them. Although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments or equivalently replace some of the technical features. These modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present application.

Claims

1. A method for image and text retrieval, characterized in that: include: Obtain the search keyword with the highest similarity to the natural language search sentence currently input by the user as the target search keyword; Based on the association relationship between the natural language search keyword and the dynamic semantic keyword of the image, the dynamic semantic keyword associated with the target search keyword is obtained as the target dynamic semantic keyword; Expand and rewrite the natural language search sentence currently input by the user according to the target dynamic semantic keywords; Retrieve images based on the expanded and rewritten statement.

2. The method according to claim 1, characterized in that The method further comprises: Obtaining a plurality of user's image retrieval record logs, wherein the image retrieval record logs at least record the user's natural language retrieval record and the graph database language retrieval record; According to the natural language search records and the graph database language search records, an association relationship between the natural language search keywords and the dynamic semantic keywords of the image is established.

3. The method according to claim 2, characterized in that The step of establishing an association relationship between natural language search keywords and dynamic semantic keywords of images based on natural language search records and graph database language search records includes: From the natural language search record, a natural language search sentence and a picture clicked by a user after the natural language search sentence is input, and an association relationship between the natural language search sentence and the picture is obtained; From the graph database language search record, obtain the graph database search statement and the picture clicked by the user after inputting the graph database search statement, and obtain the association relationship between the graph database search statement and the picture; According to the association relationship between natural language search statements and images, and the association relationship between graph database search statements and images, the association relationship between natural language search statements and graph database search statements is established by using images as intermediate nodes; According to the association between natural language search statements and graph database search statements, an association between natural language search keywords and dynamic semantic keywords of images is established.

4. The method according to claim 3, characterized in that: The step of establishing an association relationship between a natural language search keyword and a dynamic semantic keyword of a picture according to the association relationship between the natural language search sentence and the graph database search sentence includes: Extract features from natural language search sentences to obtain natural language search keywords; Cluster the natural language search keywords to obtain a set of clusters, each cluster has a cluster center and a set of natural language search sentences; According to the association between natural language search statements and graph database search statements, the association between cluster centers and dynamic semantic keywords of images is established.

5. The method according to claim 4, characterized in that The step of establishing an association relationship between a cluster center and a dynamic semantic keyword of a picture according to an association relationship between a natural language search statement and a graph database search statement includes: For each cluster center, according to the association relationship between the natural language search statement and the graph database search statement, obtain the graph database search statement set associated with the natural language search statement set corresponding to the cluster center; Extract keywords from the graph database search statement set to obtain the dynamic semantic keyword set of the image; Vectorizing the dynamic semantic keyword set to obtain a dynamic semantic keyword vector set; An association relationship is established between the cluster center and the dynamic semantic keyword vector set.

6. The method according to any one of claims 1 to 5, characterized in that: The step of obtaining the search keyword with the highest similarity to the natural language search sentence currently input by the user as the target search keyword includes: Get the natural language search sentence currently input by the user; Extract features of the natural language search sentence currently input by the user to obtain the current search keywords; Get the search keyword with the highest similarity to the current search keyword as the target search keyword.

7. The method according to any one of claims 1 to 5, characterized in that: The step of expanding and rewriting the natural language search sentence currently input by the user according to the target dynamic semantic keyword includes: When there are multiple target dynamic semantic keywords, calculate the similarity between each target dynamic semantic keyword and the current search keyword; Obtain at least one target dynamic semantic keyword with the highest similarity to the current search keyword; The natural language search sentence currently input by the user is expanded and rewritten according to the at least one target dynamic semantic keyword.

8. A graphic and text retrieval device, characterized in that: include: A search keyword acquisition module is used to acquire the search keyword with the highest similarity to the natural language search sentence currently input by the user as the target search keyword; A dynamic semantic keyword acquisition module, based on the association between the natural language search keyword and the dynamic semantic keyword of the image, acquires the dynamic semantic keyword associated with the target search keyword as the target dynamic semantic keyword; An expansion and rewriting module is used to expand and rewrite the natural language search sentence currently input by the user according to the target dynamic semantic keywords; The image retrieval module is used to retrieve images based on the expanded and rewritten sentences.

9. A server, characterized in that: include: A memory and a processor, wherein an application is stored in the memory, and the application is used to enable the processor to execute the method according to any one of claims 1 to 7 when called by the processor.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores program codes, and when the program codes are called by a processor, the processor executes the method according to any one of claims 1 to 7.