Picture retrieval method and electronic equipment

The image tags are directly extracted and corrected from the search instructions through a large language model, which solves the problem of inaccurate tag extraction in the prior art, and improves the accuracy and user experience of image retrieval.

CN120336565APending Publication Date: 2025-07-18SHENZHEN SHULIAN TIANXIA INTELLIGENT TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510250365.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-04
Publication Date
2025-07-18

Smart Images

  • Figure CN120336565A_ABST
    Figure CN120336565A_ABST
Patent Text Reader

Abstract

The embodiment of the invention relates to the technical field of picture retrieval, and discloses a picture retrieval method and electronic equipment. On one hand, a search instruction is obtained, information extraction is conducted on the search instruction based on a large language model to obtain an initial label, the initial label is analyzed and corrected to obtain a corresponding target label, and a picture set is retrieved based on the target label to obtain at least one target picture; according to the method, the initial label can be directly extracted from the search instruction through the large language model and is corrected, so that the accuracy and comprehensiveness of label extraction are improved, and the accuracy of picture retrieval is further improved; and on the other hand, the arrangement sequence of the target pictures is determined based on the target labels, and the target pictures are displayed according to the arrangement sequence, so that the arrangement sequence of the target pictures can be optimized according to the target labels, the picture display logicality and orderliness are improved, and the search experience of the user is optimized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present application relate to the technical field of image retrieval, and in particular, to an image retrieval method and an electronic device. Background Art

[0002] Image retrieval refers to the process of finding images that match the query conditions from a database. Image retrieval includes a tag-based image retrieval method, which means retrieving images through the tags of the images.

[0003] Currently, it is usually achieved by extracting keywords from the search instruction, performing tag matching on the keywords in the tag library to obtain the matching tags, and then searching for the corresponding images in the image library according to these tags, so as to implement tag-based image retrieval.

[0004] In the process of implementing the present application, the inventors found that there are at least the following problems in the prior art: the existing method can only indirectly obtain the tags required for retrieval by extracting keywords, and cannot directly extract tags from the search instruction; moreover, in the process of keyword extraction, the tag system is not considered, which easily leads to problems such as missing keywords or keyword ambiguity, thereby reducing the accuracy and comprehensiveness of tag extraction, and further reducing the accuracy of image retrieval. Summary of the Invention

[0005] The embodiments of the present application aim to provide an image retrieval method and an electronic device to improve the accuracy of image retrieval.

[0006] The embodiments of the present application provide the following technical solutions:

[0007] In a first aspect, the embodiments of the present application provide an image retrieval method, which includes:

[0008] Obtain a search instruction;

[0009] Extract information from the search instruction based on a large language model to obtain initial tags;

[0010] Parse and correct the initial tags to obtain corresponding target tags;

[0011] Retrieve an image set based on the target tags to obtain at least one target image;

[0012] Determine the arrangement order of the target images based on the target tags, and display the target images in the arrangement order.

[0013] In a second aspect, the embodiments of the present application provide an electronic device, including:

[0014] At least one processor, and

[0015] A memory communicatively connected to at least one processor, wherein,

[0016] the memory stores instructions executable by at least one processor, and the instructions are executed by at least one processor to enable the at least one processor to execute the picture retrieval method of the first aspect.

[0017] In a third aspect, an embodiment of the present application provides a non-volatile computer-readable storage medium storing computer-executable instructions for causing an electronic device to execute the picture retrieval method of the first aspect.

[0018] Advantageous effects of the embodiments of the present application: Different from the prior art, an embodiment of the present application provides a picture retrieval method, which includes: obtaining a search instruction; extracting information from the search instruction based on a large language model to obtain an initial label; parsing and correcting the initial label to obtain a corresponding target label; retrieving a picture set based on the target label to obtain at least one target picture; determining an arrangement order of the target pictures based on the target label, and displaying the target pictures in the arrangement order.

[0019] On the one hand, by obtaining a search instruction, extracting information from the search instruction based on a large language model to obtain an initial label, parsing and correcting the initial label to obtain a corresponding target label, and retrieving a picture set based on the target label to obtain at least one target picture, the present application can directly extract the initial label from the search instruction through the large language model and correct it, improving the accuracy and comprehensiveness of label extraction, and further improving the accuracy of picture retrieval; on the other hand, by determining the arrangement order of the target pictures based on the target label and displaying the target pictures in the arrangement order, the present application can optimize the arrangement order of the target pictures according to the target label, improve the logic and coherence of picture display, and optimize the user's search experience. Description of the Drawings

[0020] One or more embodiments are exemplarily illustrated by pictures in the corresponding drawings. These exemplary illustrations do not limit the embodiments. Elements with the same reference numerals in the drawings are represented as similar elements, unless otherwise stated, and the drawings in the drawings do not constitute a proportional limitation.

[0021] Figure 1 is a schematic diagram of an application environment provided by an embodiment of the present application;

[0022] Figure 2 is a schematic flowchart of a picture retrieval method provided by an embodiment of the present application;

[0023] Figure 3It is a schematic flowchart of a process for parsing and correcting an initial label provided by an embodiment of the present application;

[0024] Figure 4 It is a schematic flowchart of a process for label clarification provided by an embodiment of the present application;

[0025] Figure 5 It is a schematic structural diagram of a picture retrieval device provided by an embodiment of the present application;

[0026] Figure 6 It is a schematic structural diagram of an electronic device provided by an embodiment of the present application. Detailed implementation manners

[0027] The present application will be described in detail below with reference to specific embodiments. The following embodiments will help those skilled in the art to further understand the present application, but do not limit the present application in any form. It should be noted that those of ordinary skill in the art can make several modifications and improvements without departing from the concept of the present application. These all fall within the protection scope of the present application.

[0028] In order to make the purpose, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0029] It should be noted that if there is no conflict, the various features in the embodiments of the present application can be combined with each other, and all are within the protection scope of the present application. In addition, although functional module division is performed in the device schematic diagram and the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order from the module division in the device or the flowchart. In addition, the terms "first", "second", "third", etc. used herein do not limit the data and execution order, but only distinguish the same items or similar items with basically the same function and role.

[0030] Unless otherwise defined, all technical and scientific terms used in this specification have the same meaning as commonly understood by those skilled in the technical field to which the present application belongs. The terms used in this specification in the description of the present application are only for the purpose of describing specific implementation manners and are not used to limit the present application. The term "and / or" used in this specification includes any and all combinations of one or more of the related listed items.

[0031] In addition, the technical features involved in the various embodiments of the present application described below can be combined with each other as long as they do not conflict with each other.

[0032] Before introducing the embodiments of the present application, a brief introduction to the image retrieval method known to the inventors of the present application will be given first, so as to facilitate the subsequent understanding of the embodiments of the present application.

[0033] Image retrieval refers to the process of finding images that match the query conditions from a database. Image retrieval includes tag-based image retrieval and content-based image retrieval. Tag-based image retrieval means retrieving images through the tags of the images; content-based image retrieval means calculating the similarity between images by extracting texture, color, gradient or other high-level semantic features of the images as image features to achieve image retrieval.

[0034] In the scenario of retrieving target images by inputting text, for example, when retrieving target images by inputting text in an album business system, in order to meet the requirements of high precision, parsability and precise matching, tag-based image retrieval is usually adopted.

[0035] The tag-based image retrieval method mainly describes the images in text, extracts tag information such as keywords, and searches for corresponding images by retrieving keywords during retrieval. Specifically, it usually adopts the method of extracting keywords from the search instruction, performing tag matching on the keywords in the tag library to obtain the matching tags, and then searching for the corresponding images in the image library according to these tags to achieve tag-based image retrieval.

[0036] However, this method can only indirectly obtain the tags required for retrieval by extracting keywords, and cannot directly extract tags from the search instruction; moreover, in the process of keyword extraction, the tag system is not considered, which easily leads to problems such as missed keyword extraction or keyword ambiguity, thus reducing the accuracy and comprehensiveness of tag extraction, and further reducing the accuracy of image retrieval.

[0037] Based on this, the embodiments of the present application provide an image retrieval method. On the one hand, the initial tags are directly extracted from the search instruction through a large language model and corrected to improve the accuracy and comprehensiveness of tag extraction, and further improve the accuracy of image retrieval; on the other hand, the arrangement order of the target images is optimized through the target tags to improve the logic and coherence of image display and optimize the user's search experience.

[0038] The technical solution of the present application will be specifically described below with reference to the accompanying drawings of the specification.

[0039] Please refer to Figure 1 , Figure 1 which is a schematic diagram of an application environment provided by the embodiments of the present application;

[0040] As Figure 1As shown, the application environment 100 includes: a terminal 10 and a server 20. Among them, the terminal 10 is communicatively connected to the server 20 via a network, where the network includes a wired network and / or a wireless network. It can be understood that the network includes wireless networks such as 2G, 3G, 4G, 5G, Wi-Fi, and Bluetooth, and can also include wired networks such as serial cables and Ethernet cables.

[0041] In some embodiments, the terminal 10 is configured to execute the picture retrieval method in any of the following embodiments. For example: obtain a search instruction; perform information extraction on the search instruction based on a large language model to obtain an initial label; perform parsing and correction processing on the initial label to obtain a corresponding target label; retrieve a picture set based on the target label to obtain at least one target picture; determine the arrangement order of the target pictures based on the target label, and display the target pictures in the arrangement order.

[0042] In some embodiments, the terminal 10 includes an input interface. The user inputs a search instruction through the input interface, and after the input is completed, the terminal 10 automatically obtains the search instruction.

[0043] In some embodiments, the terminal 10 includes a display interface. The terminal 10 is further configured to display the target information corresponding to the target label on the display interface when there is label ambiguity in the target label, and update the target label in response to the user's selection operation on the target information.

[0044] In some embodiments, the terminal 10 includes a display interface. The terminal 10 is further configured to correct the search instruction input by the user, generate a recommended search instruction, and display the recommended search instruction on the display interface.

[0045] In some embodiments, the picture set is stored in the server 20. The terminal 10 is configured to send the search instruction input by the user to the server 20, so that the server 20 executes the picture retrieval method in any of the following embodiments.

[0046] Among them, the terminals in the embodiments of the present application include, but are not limited to: various terminals with computing and processing capabilities such as laptop computers, desktop computers, or mobile devices.

[0047] In some embodiments, the server 20 is communicatively connected to the terminal 10 and is configured to obtain a search instruction input by a user on the terminal 10 sent by the terminal 10, retrieve a matching target image from an image set based on the search instruction, and send the target images to the terminal 10 in a sorted order. For example, information extraction is performed on the search instruction based on a large language model to obtain an initial label; the initial label is parsed and corrected to obtain a corresponding target label; the image set is retrieved based on the target label to obtain at least one target image, the sorting order of the target images is determined based on the target label, and then the target images are sent to the terminal 10 in the sorted order, so that the terminal 10 displays the target images on its own display interface in the sorted order.

[0048] In some embodiments, the server 20 is further configured to, when there is label ambiguity in the target label, send the target information corresponding to the target label to the terminal 10, so that the terminal 10 displays the target information on the display interface. Thus, after the terminal 10 responds to a user's selection operation on the target information, the terminal 10 sends the selected target information to the server 20, and the server 20 updates the target label based on the target information.

[0049] In some embodiments, the server 20 is further configured to correct the search instruction, generate a recommended search instruction, and send the recommended search instruction to the terminal 10, so that the terminal 10 displays the recommended search instruction on its own display interface.

[0050] Wherein, the number of the servers 20 may also be multiple, and the multiple servers may form a server cluster. For example, the server cluster includes: a first server, a second server,..., an Nth server, or the server cluster may be a cloud computing service center, and the cloud computing service center includes several servers. The servers in the embodiments of the present application include but are not limited to: tower servers, rack servers, blade servers, and cloud servers. Preferably, the server is a cloud server (Elastic Compute Service, ECS).

[0051] Please refer to Figure 2 , Figure 2 which is a schematic flowchart of an image retrieval method provided by an embodiment of the present application;

[0052] Wherein, the image retrieval method is applied to an electronic device, and the electronic device includes a terminal or a server. Specifically, the execution subject of the image retrieval method is one or at least two processors in the electronic device.

[0053] The image retrieval method can be applied to various scenarios that require retrieving target images by inputting text. Here, the specific implementation of the image retrieval method is illustrated by taking the image retrieval scenario of a kindergarten photo album as an example. The specific implementation of the image retrieval method in other scenarios is similar and will not be elaborated here.

[0054] As Figure 2 shown, the image retrieval method includes:

[0055] Step S201: Obtain a search instruction;

[0056] Specifically, obtain the search instruction input by the user. Among them, the search instruction is used to instruct the electronic device to execute the image retrieval method to retrieve the target image that matches the search instruction from the image set and display it.

[0057] Among them, the search instruction includes at least one of time information, geographical location information, image scene information, season information, person information, and school record information. Among them, the time information includes but is not limited to year, month, day, or week. The image scene information includes but is not limited to birthday party, group photo, stage performance, teacher-child interaction, interest learning, theme activity, large-scale activity, or parent-child sports meeting. The person information includes but is not limited to name or age. The school record information includes but is not limited to grade or class.

[0058] It can be understood that the specific content of the search instruction can be set according to the specific scenario and the information of the images stored in the image set, which is not limited here. For example: in the image retrieval scenario of a kindergarten photo album, the search instruction is "On December 4, 2034, Zhang San from Class 2, Grade 2 took a group photo".

[0059] In some embodiments, the electronic device includes an input interface. The user inputs the search instruction through the input interface. After the input is completed, the electronic device automatically obtains the search instruction.

[0060] Step S202: Perform information extraction on the search instruction based on a large language model to obtain an initial label;

[0061] Among them, the large language model (LLM) is an artificial intelligence model based on deep learning, used to understand and generate natural language text. Specifically, the LLM consists of an encoder and a decoder. The encoder is used to receive the input text (such as the search instruction) and convert it into a context representation, and the decoder is used to generate a structured information extraction result (such as the initial label) according to the context representation generated by the encoder.

[0062] Specifically, the large language model (LLM) extracts information from the search instruction and outputs an initial label. The initial label is generated by the LLM extracting information from the search instruction, and the format of the initial label is a preset structure. The preset structure is a kind of label format, which can be set by those skilled in the art according to the type or content of the label, and no limitation is made here.

[0063] In the embodiments of the present application, the LLM generates an initial label with a preset structure through the function call function. Specifically, the LLM includes an information extraction function, and the LLM generates an initial label by calling the information extraction function. The information extraction function is used to extract information from the search instruction to generate an initial label with a preset structure.

[0064] In some embodiments, the LLM includes at least one function, and each function corresponds to a function function. After the user inputs a search instruction, the electronic device automatically generates a request instruction. The LLM generates an initial label with a preset structure by identifying the request instruction and calling an appropriate function (i.e., the information extraction function) based on the understanding of the function function of each function. The request instruction is used to request the LLM to extract information from the search instruction.

[0065] Before the LLM calls the information extraction function, the method further includes: defining the information extraction function. The definition of the function includes the function description and the definition of the input parameters of the function. Each input parameter of the information extraction function corresponds to a label name. The definition of the input parameters of the function includes defining the label name corresponding to the input parameter, the type of the label value, and the preset range value of the multi-value label.

[0066] Among them, the label can be divided into a multi-value label and a single-value label. The multi-value label is a label with at least two label values corresponding to one label name, and the single-value label is a label with one label value corresponding to one label name. The preset range value is the optional range of the label value of the multi-value label, which can be set by those skilled in the art according to the specific scenario and the information of the pictures stored in the picture set, and no limitation is made here.

[0067] For example: the information extraction function contains 6 input parameters, and the label names corresponding to the 6 input parameters are respectively: theme label, year, month, day, class, and name. Among them, the label corresponding to the theme label is a multi-value label, and the labels corresponding to the other label names are single-value labels. The preset range value of the theme label is: highlight moment, birthday party, group photo, stage performance, teacher-child interaction, or theme activity.

[0068] In the embodiments of the present application, each initial tag includes a tag name and a tag value, and one tag name corresponds to at least one tag value. Step S201 specifically includes: parsing the search instruction based on the information extraction function, extracting the tag value corresponding to the input parameter from the search instruction, and combining each tag name with the corresponding tag value into an initial tag with a preset structure.

[0069] Specifically, the large language model LLM calls the information extraction function to parse the search instruction, extracts the tag value corresponding to the input parameter from the search instruction, and combines each tag name with the corresponding tag value into an initial tag with a preset structure.

[0070] For example: the search instruction is "On December 4, 2034, Zhang San in the small class 2 took a group photo". According to the information extraction function in the above example, 6 initial tags are obtained, which are respectively: an initial tag with the tag name "theme tag" and the tag value "group photo"; an initial tag with the tag name "year" and the tag value "2034"; an initial tag with the tag name "month" and the tag value "12"; an initial tag with the tag name "day" and the tag value "4"; an initial tag with the tag name "class" and the tag value "small class 2"; an initial tag with the tag name "name" and the tag value "Zhang San".

[0071] In the embodiments of the present application, by performing information extraction on the search instruction based on the large language model to obtain the initial tag, compared with the existing method that can only indirectly obtain the tags required for retrieval by extracting keywords, the present application can rely on the powerful information extraction ability of the large language model to directly extract the initial tag from the search instruction, improving the accuracy of tag extraction.

[0072] In some embodiments, the method further includes: defining the configuration information of the information extraction function in an extensible markup language file; when adding a new tag name to any picture in the picture set, adding the input parameter corresponding to the new tag name to the file.

[0073] Wherein, the configuration information includes input parameters, and the configuration information further includes the tag name corresponding to the input parameter, the type of the tag value, and the preset range value of the multi-value tag. The picture set is used to store a plurality of pictures, each picture corresponds to at least one picture tag, and the picture tag is a tag representing the real information of the picture. The picture tag includes the tag name and the tag value of the picture.

[0074] Specifically, the configuration information of the information extraction function is defined in an Extensible Markup Language (XML) file. When a new tag name is added to the image tag of any image in the image set, the input parameters corresponding to the tag name are added to the XML file. For example, when a new tag name is added to the image tag of an existing image in the image set, the input parameters corresponding to the tag name are added to the XML file. Alternatively, when a new image is added to the image set and any tag name of the new image is not defined in the XML file, the input parameters corresponding to the tag name are added to the XML file. It can be understood that when the image tag corresponding to the tag name is a multi-value tag, the preset range value of the multi-value tag is also added to the XML file.

[0075] Compared with the label extraction method based on machine learning modeling, when the label system is updated frequently, it is necessary to retrain and adjust the model frequently to support new label types, and the scalability is poor. In this application, the configuration information of the information extraction function is uniformly defined in XML. When a new tag name is added, only the corresponding input parameters need to be added to the file, and the continuously expanding and updated label system can be supported by simply modifying the configuration without retraining the model.

[0076] In some embodiments, the large language model LLM generates initial tags with a preset structure through a prompting strategy. For example, through templates or format requirements, the model is guided to fill in the content according to the preset structure to generate initial tags with a preset structure; or, through field names and content requirements, the model is guided to generate initial tags with a preset structure. Among them, the prompting strategy is a prior art, and this application does not limit its specific implementation.

[0077] Step S203: Parse and correct the initial tags to obtain the corresponding target tags;

[0078] Specifically, after the preliminary label extraction of the search instruction is completed based on the large language model LLM, considering the occasional instability in the extraction format, incorrect extraction content (including understanding deviation and label value generation error), or missed extraction problems of the large language model LLM, this application further parses and corrects the initial tags to obtain the corresponding target tags.

[0079] Among them, the calibration process is used to solve the problems of incorrect extraction content and missed extraction. Incorrect extraction content includes three situations: label content creation, label position error, and label ambiguity. Label content creation means that the label value is not within the preset range value corresponding to the label name; label position error means that the label value is placed under the wrong label name; label ambiguity means that the label value uses a broad or vague expression, which does not meet the accuracy requirements of the label value of the picture label. Missed extraction refers to the omission of extracting words containing relative concepts in the search instruction. For example, the relative concept words are "the past two days", "this month", "last query", "this month".

[0080] The calibration process is a processing method for correcting and optimizing the initial label. The calibration process includes at least one of deleting the label value that is not within the preset range value corresponding to any label name, performing position replacement on the label value, performing mapping conversion on the label value, or extracting the missing label. The target label is the label obtained after parsing and calibrating the initial label.

[0081] In the embodiment of the present application, step S203 specifically includes steps S231 - S235:

[0082] Step S231: Based on the format of each initial label, parse the initial label according to the preset parsing method corresponding to the format to obtain the parsed initial label;

[0083] Among them, each format corresponds to a preset parsing method. The preset parsing method can be set by those skilled in the art according to the format of the initial label, and is not limited here. The parsed initial label includes a label name and a label value, and each label name corresponds to at least one label value.

[0084] Specifically, based on the format of each initial label, parse the initial label according to the preset parsing method corresponding to this format, so as to obtain the parsed initial label.

[0085] In the embodiment of the present application, by parsing the initial label based on the format of each initial label according to the preset parsing method corresponding to the format, the present application can consider the format compatibility when parsing the initial label, solve the problem that the large language model LLM is occasionally unstable in the extraction format, and reduce the probability of parsing anomalies.

[0086] It can be understood that in most cases, the large language model LLM outputs the initial label according to the preset structure. In a small probability case, the format of the output initial label will change. At this time, parsing the initial label according to the preset parsing method corresponding to the changed format can reduce the probability of parsing anomalies.

[0087] Further, after obtaining the parsed initial tag, it is successively determined whether the parsed initial tag meets the first condition, the second condition, the third condition, and the fourth condition. When the parsed initial tag meets any of the above conditions, corresponding processing is performed to generate a target tag. Among them, the target tag includes a first target tag, a second target tag, a third target tag, or a fourth target tag.

[0088] Step S232: When any parsed initial tag meets the first condition, delete the first tag value in the parsed initial tag to generate a target tag;

[0089] Among them, the first condition includes: any tag value corresponding to any tag name is not within the preset range value corresponding to the tag name, and the first tag value is the tag value that is not within the preset range value corresponding to any tag name. The parsed initial tag meets the first condition, that is, the parsed initial tag belongs to the tag content creation in the problem of incorrect extraction content.

[0090] Specifically, for any parsed initial tag, compare any tag value it contains with the preset range value corresponding to the tag name to which the tag value belongs. If each tag value is within the preset range value corresponding to the tag name to which the tag value belongs, it is determined that the parsed initial tag does not meet the first condition, and continue to determine whether the parsed initial tag meets the second condition.

[0091] If any tag value is not within the preset range value corresponding to the tag name to which the tag value belongs, it is determined that the parsed initial tag meets the first condition, and delete the tag value in the parsed initial tag to obtain the first target tag. Among them, the first target tag is the tag obtained after deleting the first tag value in the parsed initial tag.

[0092] For example: The tag name of a parsed initial tag is "theme tag", the tag value is "diet health" or "highlight moment", the preset range value of "theme tag" is: "highlight moment", "birthday party", "group photo", or "stage performance", and "diet health" is not within the preset range value, then it is determined that the parsed initial tag meets the first condition, and delete "diet health" in the parsed initial tag.

[0093] In some embodiments, if after deleting the first tag value in the parsed initial tag, the tag value of the parsed initial tag is empty, then delete the parsed initial tag.

[0094] In an embodiment of the present application, when any parsed initial tag meets the first condition, the first tag value is deleted from the parsed initial tag. The present application can make the tag value of the target tag within the preset range values corresponding to the tag name, thereby improving the accuracy of retrieving the target picture based on the target tag subsequently.

[0095] Step S233: When any parsed initial tag meets the second condition, perform a position replacement on the tag value to generate a target tag.

[0096] Among them, the second condition includes: when matching each tag value based on the first regular expression corresponding to each tag name, the tag name corresponding to the first regular expression matched by any tag value is different from the tag name corresponding to this tag value. The first regular expression is a regular expression used to determine whether the parsed initial tag meets the second condition, and the first regular expression can be set by those skilled in the art according to the tag name, and no limitation is made here.

[0097] Position replacement means moving the tag value to the tag name corresponding to the first regular expression matched by this tag value. The parsed initial tag meets the second condition, that is, the parsed initial tag belongs to the tag position error in the problem of incorrect extraction content.

[0098] Specifically, match each tag value of any parsed initial tag based on the first regular expression corresponding to each tag name. If the tag name corresponding to the first regular expression matched by each tag value is the same as the tag name corresponding to this tag value, it is determined that the parsed initial tag does not meet the second condition, and continue to determine whether the parsed initial tag meets the third condition.

[0099] If the tag name corresponding to the first regular expression matched by any tag value is different from the tag name corresponding to this tag value, it is determined that the parsed initial tag meets the second condition, move this tag value to the tag name corresponding to the first regular expression matched by this tag value, and obtain a second target tag. Among them, the second target tag is the tag obtained after performing a position replacement on the tag value, and the tag value included in the second target tag is located under the tag name corresponding to the first regular expression matched by this tag value.

[0100] For example: the tag name of a parsed initial tag is "month" and the tag value is "Tuesday". When the tag name is "week", the first regular expression is "week X" or "X of the week", and "Tuesday" matches this first regular expression. The tag name corresponding to this first regular expression is "week", while the tag name corresponding to "Tuesday" is "month", and the two are different. It is determined that the parsed initial tag meets the second condition, and move "Tuesday" to the tag name of "week".

[0101] In some embodiments, when any of the parsed initial tags does not meet the first condition, it is determined whether the parsed initial tag meets the second condition. When the parsed initial tag meets the second condition, step S233 is performed on the parsed initial tag to obtain a corresponding second target tag.

[0102] In some embodiments, when any of the parsed initial tags meets the first condition and after obtaining the first target tag through step S232, it is determined whether the first target tag meets the second condition. When the first target tag meets the second condition, step S233 is performed on the first target tag to obtain a corresponding second target tag.

[0103] For example: Based on the first regular expression corresponding to each tag name, each tag value of any one of the first target tags is matched. If the tag names corresponding to the first regular expressions matched by each tag value are the same as the tag name corresponding to the tag value, it is determined that the first target tag does not meet the second condition, and it continues to be determined whether the first target tag meets the third condition.

[0104] If the tag name corresponding to the first regular expression matched by any one tag value is different from the tag name corresponding to the tag value, it is determined that the first target tag meets the second condition, and the tag value is moved under the tag name corresponding to the first regular expression matched by the tag value to obtain a second target tag.

[0105] In the embodiments of the present application, by performing a position replacement on the tag value when any of the parsed initial tags meets the second condition, the present application can make the tag value of the target tag under the correct tag name, thereby improving the accuracy of retrieving the target picture based on the target tag subsequently.

[0106] Step S234: When any of the parsed initial tags meets the third condition, perform a mapping conversion on the tag value to generate a target tag;

[0107] Among them, the third condition includes: when matching the tag value corresponding to the tag name based on the second regular expression corresponding to each tag name, any one tag value does not match the second regular expression. The second regular expression is a regular expression used to determine whether the parsed initial tag meets the third condition, and the second regular expression can be set by those skilled in the art according to the accuracy requirements of the tag name and the tag value of the picture tag, and is not limited here.

[0108] The mapping conversion means converting the tag value into a matching form that conforms to the second regular expression corresponding to its corresponding tag name. The parsed initial tag meets the third condition, that is, the parsed initial tag belongs to the problem of vague and general tags in the incorrect extraction content.

[0109] Specifically, based on the second regular expression corresponding to each tag name, each tag value of any parsed initial tag under the tag name is matched. If each tag value matches the second regular expression corresponding to its tag name, it is determined that the parsed initial tag does not meet the third condition, and it is continued to determine whether the parsed initial tag meets the fourth condition.

[0110] If any tag value does not match the second regular expression corresponding to its tag name, it is determined that the parsed initial tag meets the third condition, and the tag value is converted into a matching form that conforms to the second regular expression of its corresponding tag name through a preset mapping rule to obtain a third target tag. The preset mapping rule is used to convert the tag value into a matching form that conforms to the second regular expression of its corresponding tag name, and the preset mapping rule can be set by those skilled in the art according to the tag name and the second regular expression, which is not limited herein. The third target tag is the tag obtained after mapping and converting the tag value, and the tag value included in the third target tag matches the second regular expression of the corresponding tag name.

[0111] For example: the tag name of a parsed initial tag is "month" and the tag value is "last month". When the tag name is "month", the precision requirement for the tag value of the picture tag is a specific month value, and the matching form of the second regular expression is: January - December. Since "last month" does not match this second regular expression, it is determined that the parsed initial tag meets the third condition, and "last month" is converted into a specific month value through the preset mapping rule. For example, the preset mapping rule requires mapping "last month" to the current month minus one. If the current month is December, then "last month" is converted to November.

[0112] In some embodiments, when any parsed initial tag does not meet the first condition and the second condition, it is determined whether the parsed initial tag meets the third condition. When the parsed initial tag meets the third condition, step S234 is executed on the parsed initial tag to obtain the corresponding third target tag.

[0113] In some embodiments, when any first target tag does not meet the second condition, it is determined whether the first target tag meets the third condition. When the first target tag meets the third condition, step S234 is executed on the first target tag to obtain the corresponding third target tag.

[0114] In some embodiments, when any parsed initial tag or first target tag satisfies the second condition, and after obtaining the second target tag through step S233, it is determined whether the second target tag satisfies the third condition. When the second target tag satisfies the third condition, step S234 is performed on the second target tag to obtain the corresponding third target tag.

[0115] Among them, the specific implementation of performing step S234 on the first target tag or the second target tag is similar to the specific implementation of performing step S234 on the parsed initial tag, and will not be elaborated here.

[0116] In the embodiments of the present application, by performing mapping conversion on the tag value when any parsed initial tag satisfies the third condition, the present application can make the tag value of the target tag meet the tag value requirements of the picture tag, thereby improving the accuracy of retrieving the target picture based on the target tag subsequently.

[0117] Step S235: When each parsed initial tag satisfies the fourth condition, extract the tag value matched by the third regular expression from the search instruction, and combine the tag value with the corresponding tag name into a target tag.

[0118] Among them, the fourth condition includes: matching the search instruction based on the third regular expression, there is a preset tag value in the search instruction that matches the third regular expression, and each parsed initial tag does not contain the preset tag value. The third regular expression is a regular expression used to determine whether the parsed initial tag satisfies the fourth condition. The number of the third regular expressions can be multiple, and the third regular expressions can be set by those skilled in the art according to the search instruction containing relative concept words, and are not limited here. The preset tag value is a tag value that matches the third regular expression. For example, the preset tag value is "the past two days", "this month", "last query", "this month". Each parsed initial tag satisfies the fourth condition, that is, the large language model LLM misses extracting some words in the search instruction.

[0119] Specifically, match the search instruction based on the third regular expression to determine whether there is a preset tag value in the search instruction that matches the third regular expression. If there is no tag value in the search instruction that matches the third regular expression, it is determined that there is no problem of missed extraction. At this time, the target tag is the tag obtained by performing steps S232 - S234, and the target tag includes the first target tag, the second target tag or the third target tag.

[0120] For example: when any first target label does not meet the second condition and the third condition, use the first target label as the target label; or, when any second target label does not meet the third condition, use the second target label as the target label; or, when any parsed initial label or first target label or second target label meets the third condition, and after obtaining the third target label through step S307, use the third target label as the target label.

[0121] In some embodiments, when any parsed initial label does not meet the first condition, the second condition, and the third condition, use the parsed initial label as the target label.

[0122] If there is a preset label value in the search instruction that matches the third regular expression, determine whether each parsed initial label does not contain the preset label value. If any parsed initial label contains the preset label value, at least one parsed initial label does not meet the fourth condition, and it is determined that there is no problem of missing extraction. At this time, the target label is the label obtained by executing steps S232 - S234.

[0123] If each parsed initial label does not contain the preset label value, it is determined that each parsed initial label meets the fourth condition. Extract the label value that matches the third regular expression from the search instruction, and combine the label value with the corresponding label name into a fourth target label. The label value included in the fourth target label matches the third regular expression.

[0124] For example: the search instruction contains relative concept words: "the past two days", "this month", "last query", or "this month". When the large language model LLM extracts information from the search instruction, there is a high probability of missing these words. Use the corresponding third regular expression to extract these words from the search instruction as label values, and combine the label values with the label names corresponding to the third regular expression into fourth target labels.

[0125] It can be understood that when each parsed initial label meets the fourth condition, the target label includes the label obtained by executing steps S232 - S234 and the fourth target label obtained by executing step S235.

[0126] In the embodiments of the present application, when each parsed initial label meets the fourth condition, the label value that matches the third regular expression is extracted from the search instruction, and the label value is combined with the corresponding label name into a target label. The present application can extract the missing target labels from the search instruction, thereby improving the accuracy of retrieving target pictures based on the target labels subsequently.

[0127] In the embodiments of the present application, there is no limitation on the specific step sequence for calibrating the initial label. The step sequence of steps S232 - S235 is only an example and can be adjusted by those skilled in the art.

[0128] Please refer to Figure 3 , Figure 3 which is a schematic flowchart of a process for parsing and calibrating an initial label provided by an embodiment of the present application;

[0129] As Figure 3 shown, the process for parsing and calibrating the initial label includes:

[0130] Step S301: Parse the initial label to obtain the parsed initial label;

[0131] Specifically, based on the format of each initial label, the initial label is parsed according to the preset parsing method corresponding to the format to obtain the parsed initial label. This step is the same as step S231 and will not be elaborated here.

[0132] Step S302: Determine whether the parsed initial label meets the first condition;

[0133] Specifically, for any parsed initial label, any label value included therein is compared with the preset range value corresponding to the label name to which the label value belongs, and it is determined whether the label value is within the preset range value corresponding to the label name to which the label value belongs.

[0134] If each label value is within the preset range value corresponding to the label name to which the label value belongs, it is determined that the parsed initial label does not meet the first condition, and step S304 is entered; if any label value is not within the preset range value corresponding to the label name to which the label value belongs, it is determined that the parsed initial label meets the first condition, and step S303 is entered.

[0135] Step S303: Delete the first label value;

[0136] Specifically, when any parsed initial label meets the first condition, the first label value is deleted from the parsed initial label to obtain the first target label.

[0137] Step S304: Determine whether the label meets the second condition;

[0138] Specifically, when any parsed initial label does not meet the first condition, it is determined whether the parsed initial label meets the second condition; or, when any parsed initial label meets the first condition and after obtaining the first target label through step S303, it is determined whether the first target label meets the second condition.

[0139] Specifically, based on the first regular expression corresponding to each tag name, each tag value of the tag input in step S304 is matched to determine whether the tag name corresponding to the first regular expression matched by any tag value is the same as the tag name corresponding to this tag value.

[0140] If the tag name corresponding to the first regular expression matched by each tag value is the same as the tag name corresponding to this tag value, it is determined that the tag input in step S304 does not meet the second condition, and step S306 is entered; if the tag name corresponding to the first regular expression matched by any tag value is different from the tag name corresponding to this tag value, it is determined that the tag input in step S304 meets the second condition, and step S305 is entered.

[0141] Step S305: Perform position replacement on the tag value;

[0142] Specifically, when the tag input in step S304 meets the second condition, perform position replacement on the tag value in this tag to obtain the second target tag. This step is similar to step S233 and will not be elaborated here.

[0143] Step S306: Determine whether this tag meets the third condition;

[0144] Specifically, when any parsed initial tag does not meet the first condition and the second condition, determine whether this parsed initial tag meets the third condition; or, when any first target tag does not meet the second condition, determine whether this first target tag meets the third condition; or, when any parsed initial tag or first target tag meets the second condition, and after obtaining the second target tag through step S305, determine whether this second target tag meets the third condition.

[0145] Specifically, based on the second regular expression corresponding to each tag name, each tag value of the tag input in step S306 under this tag name is matched to determine whether any tag value matches the second regular expression corresponding to its tag name.

[0146] If each tag value matches the second regular expression corresponding to its tag name, it is determined that the tag input in step S306 does not meet the third condition, and step S308 is entered; if any tag value does not match the second regular expression corresponding to its tag name, it is determined that the tag input in step S306 meets the third condition, and step S307 is entered.

[0147] Step S307: Perform mapping conversion on the tag value;

[0148] Specifically, when the tag input in step S306 meets the third condition, the tag value in the tag is mapped and converted to obtain a third target tag. This step is similar to step S234 and will not be elaborated here.

[0149] Step S308: Determine whether each parsed initial tag meets the fourth condition;

[0150] Specifically, based on the third regular expression, match the search instruction to determine whether there is a preset tag value in the search instruction that matches the third regular expression. If there is no tag value in the search instruction that matches the third regular expression, it is determined that no extraction omission problem has occurred, and at least one parsed initial tag does not meet the fourth condition, and step S310 is entered.

[0151] If there is a preset tag value in the search instruction that matches the third regular expression, determine whether each parsed initial tag does not contain the preset tag value. If any parsed initial tag contains the preset tag value, at least one parsed initial tag does not meet the fourth condition, and step S310 is entered.

[0152] If each parsed initial tag does not contain the preset tag value, it is determined that each parsed initial tag meets the fourth condition, and step S309 is entered.

[0153] Step S309: Extract the tag value matching the third regular expression from the search instruction, and combine the tag value with the corresponding tag name;

[0154] Specifically, if each parsed initial tag meets the fourth condition, extract the tag value matching the third regular expression from the search instruction, and combine the tag value with the corresponding tag name into a fourth target tag. This step is similar to step S235 and will not be elaborated here.

[0155] Step S310: Output the target tag.

[0156] Specifically, when any parsed initial tag does not meet the first condition, the second condition, and the third condition, use the parsed initial tag as the target tag; or, when any first target tag does not meet the second condition and the third condition, use the first target tag as the target tag; or, when any second target tag does not meet the third condition, use the second target tag as the target tag.

[0157] Alternatively, when any of the parsed initial tags, first target tags, or second target tags meets the third condition, and after obtaining the third target tag through step S307, the third target tag is used as the target tag; or, when each of the parsed initial tags meets the fourth condition, and after obtaining the fourth target tag through step S309, the fourth target tag is used as the target tag. Among them, the number of target tags can be multiple.

[0158] It can be understood that when any of the parsed initial tags does not meet the fourth condition, the output target tags do not include the fourth target tag.

[0159] In the embodiments of the present application, by performing calibration processing on the parsed initial tags, the present application can solve the problems of incorrect content extraction and missed extraction by the large language model LLM, improve the accuracy and comprehensiveness of tag extraction, and further improve the accuracy of image retrieval when performing image retrieval based on the target tags subsequently.

[0160] In some embodiments, the target tag includes a target tag name and a target tag value. The target tag name is the tag name in the target tag, and the target tag value is the tag value in the target tag. One target tag name corresponds to at least one target tag value. After parsing and calibrating the initial tag to obtain the corresponding target tag, the method further includes: displaying the target information corresponding to each target tag value on the display interface, and performing tag clarification in response to the user's selection operation on the target information.

[0161] Among them, tag clarification refers to updating the target tag to obtain the updated target tag. For example: when there is tag ambiguity in the search instruction input by the user, that is, when the same target tag value corresponds to at least two images, it is necessary to display the target information corresponding to each target tag value on the display interface so that the user can select and confirm the correct target tag value, thereby achieving tag clarification.

[0162] For example: in the image set, there are relevant images of "Wang Xiaoer" and relevant images of "Li Xiaoer". When the search instruction input by the user is "Photos of Xiaoer studying seriously", when the target tag name obtained in step S203 is "Name" and the target tag value is "Xiaoer", since "Wang Xiaoer" and "Li Xiaoer" have the same name, that is, the tag is ambiguous, directly retrieving the image set based on this target tag will return relevant photos of both people, which does not match the user's search target. Therefore, in this case, it is necessary to interact with the user through the display interface for the user to select and confirm the correct target tag value.

[0163] Please refer to Figure 4 , Figure 4 which is a schematic flowchart of a process for performing tag clarification provided by the embodiments of the present application;

[0164] As shown Figure 4 in the figure, the process for label clarification includes:

[0165] Step S401: In the label information library corresponding to any target label name, query the quantity of the target label value corresponding to the target label name and the target information;

[0166] Wherein, each target label name corresponds to a label information library, and the label information library is used to store the target information of the target label value corresponding to the target label name, and the target information is the information corresponding to the target label value in the label information library. The label information library and the target information can be set by those skilled in the art according to the picture set, and no limitation is made here.

[0167] Specifically, in the label information library corresponding to any target label name, query the quantity of the target label value corresponding to the target label name and the target information corresponding to the target label value.

[0168] For example: The target label name is "Name", the target label value is "Xiao Er". In the label information library corresponding to "Name", query the quantity of names containing "Xiao Er", and the target information corresponding to each name containing "Xiao Er". Two names containing "Xiao Er" are queried from the label information library: "Wang Xiao Er" and "Li Xiao Er", then the quantity of the target label value is two, and the target information corresponding to the two target label values are "Wang Xiao Er" and "Li Xiao Er" respectively. In some embodiments, the target information may further include the school roll information, age information, etc. of "Wang Xiao Er" or "Li Xiao Er", and no limitation is made here.

[0169] Step S402: Determine whether the quantity of the target label value is at least two;

[0170] Specifically, if the quantity of the target label values corresponding to the same target label name is at least two, then enter Step S403; if the quantity of the target label values corresponding to the same target label name is one, then end the process, without the user needing to make a selection and confirmation, and directly enter Step S204: Retrieve the picture set based on the target label to obtain at least one target picture.

[0171] Step S403: Display each target information;

[0172] Specifically, when the quantity of the target label values corresponding to the same target label name is at least two, display the target information corresponding to each target label value on the display interface.

[0173] For example: Display the name, school roll information, age information, etc. of "Wang Xiao Er" and "Li Xiao Er" on the display interface.

[0174] Step S404: In response to the user's selection operation on the target information, update the target label to obtain the updated target label.

[0175] Specifically, when at least two pieces of target information are displayed on the display interface, in response to the user's selection operation on one of the target information, update the target label based on the selected target information to obtain the updated target label.

[0176] For example: After the user selects the target information of "Wang Xiaoer" on the display interface, update the original target label, that is, the target label name is "Name" and the target label value is "Xiaoer", to the target label name is "Name" and the target label value is "Wang Xiaoer".

[0177] After label clarification, the method further includes: retrieving the picture set based on the updated target label to obtain several target pictures that match the updated target label; determining the arrangement order of the several target pictures based on the updated target label, and displaying the target pictures in the arrangement order.

[0178] Specifically, execute Step S204 and Step S205 with the updated target label.

[0179] In the embodiments of the present application, when the number of target label values corresponding to the same target label name is at least two, display the target information corresponding to each target label value on the display interface, and in response to the user's selection operation on the target information, update the target label, and retrieve the picture set based on the updated target label to obtain the matching target pictures, and display the target pictures in the arrangement order. The present application can achieve label clarification through interaction with the user, so as to improve the accuracy of picture retrieval when performing picture retrieval based on the updated target label subsequently.

[0180] Step S204: Retrieve the picture set based on the target label to obtain at least one target picture;

[0181] Among them, the picture set is used to store several pictures, each picture has picture information and picture labels. The picture information is the relevant information of the picture, and the picture information includes but is not limited to information such as the resolution, format, size, date and time, and storage location of the picture. The picture labels include the label name and label value of the picture, and the picture labels can be marked by those skilled in the art according to the true information of the picture.

[0182] In the embodiments of the present application, the picture information and picture tags of pictures are stored in a relational database. The relational database includes a picture information table and a picture tag table. Among them, the picture information table is used to store picture information, and the picture tag table is used to store picture tags. The relational database organizes and stores data through a table structure, follows the relational model, and uses Structured Query Language (SQL) for data management. The relational database includes, but is not limited to, MySQL relational databases, etc. Preferably, the embodiments of the present application adopt a MySQL relational database.

[0183] In the embodiments of the present application, step S204 specifically includes steps S241 - S243:

[0184] Step S241: Combine a number of target tags into a first SQL query statement based on the AND logic;

[0185] Among them, the first SQL query statement is an SQL query statement obtained by combining target tags through the AND logic. The first SQL query statement is used to screen out picture tags that simultaneously contain all target tags from the MySQL relational database to determine the picture information corresponding to the picture tag.

[0186] Specifically, the combination logic of tags is the AND logic. The AND logic is used to combine a number of target tags into an SQL query statement, that is, the first SQL query statement.

[0187] Step S242: Retrieve the matching first picture information in the relational database based on the first SQL query statement;

[0188] Among them, the first picture information is the picture information of the target picture stored in the picture information table.

[0189] Specifically, the picture tag table is retrieved through the first SQL query statement to screen out picture tags that simultaneously contain all target tags, and the picture information corresponding to the picture tag in the picture information table is determined, that is, the first picture information.

[0190] Step S243: Retrieve the matching target picture in the picture set based on the first picture information.

[0191] Among them, the target picture is a picture that meets all the target tags.

[0192] Specifically, retrieve the target picture that matches the first picture information from the picture set.

[0193] In the embodiment of the present application, by using AND logic, several target tags are combined into a first SQL query statement, and based on the first SQL query statement, matching first picture information is retrieved from a relational database, and target pictures matching the first picture information are retrieved from a picture set. The present application can make the target pictures meet all the target tags, improving the matching degree and accuracy between the target pictures and the search instruction input by the user.

[0194] Step S205: Determine the arrangement order of the target pictures based on the target tags, and display the target pictures in the arrangement order.

[0195] Specifically, when retrieving the picture set based on the target tags and at least two matching target pictures are retrieved, determine the arrangement order of these target pictures based on the target tags, and display these target pictures on the display interface in the arrangement order.

[0196] It can be understood that when retrieving the picture set based on the target tags and only one matching target picture is retrieved, there is no need to determine the arrangement order, and directly display the target picture on the display interface.

[0197] In the embodiment of the present application, the target tags are multi-value tags or single-value tags. The number of target tag values in the multi-value tags is at least two, and the number of target tag values in the single-value tags is one. Each target picture corresponds to a date and time, and the date and time includes the shooting date (year, month, day) and shooting time (hour, minute, second) of the target picture.

[0198] In the embodiment of the present application, the step of determining the arrangement order of the target pictures based on the target tags includes: when any one of the target tags is a multi-value tag, determine the number of target tag values matched by each target picture in the multi-value tag, and use the order from largest to smallest of the numbers as the arrangement order of the corresponding target pictures.

[0199] Specifically, sort the target pictures according to the number of target tag values hit by the target pictures in the multi-value tag. When any one of the target tags is a multi-value tag, the picture tags of each target picture retrieved in step S204 include at least one target tag value of the target tag. Determine the number of target tag values matched by each target picture in the target tag, and use the order from largest to smallest of the numbers as the arrangement order of the corresponding target pictures.

[0200] For example: The target label name of a target label is "topic label", and the target label values are "highlight moment" or "birthday party". This target label is a multi-value label. In step S204, two target pictures are retrieved. Target picture A has picture labels with label name "topic label" and label values "highlight moment" and "birthday party", and target picture B has a picture label with label name "topic label" and label value "highlight moment". That is, the number of target label values matched by target picture A in this target label is 2, and the number of target label values matched by target picture B in this target label is 1. Compared with target picture B, target picture A has a higher matching degree with the search instruction, and target picture A is arranged in front of target picture B.

[0201] In some embodiments, when the quantities corresponding to at least two target pictures are the same, the descending order of date and time is used as the arrangement order of the corresponding target pictures.

[0202] Among them, the descending order of date and time means the order of date and time from the latest to the oldest, that is, the latest date is ranked in the front, and when the dates are the same, the latest time is ranked in the front.

[0203] Specifically, when the number of target label values matched by at least two target pictures in a multi-value label is the same, these target pictures are sorted according to the date and time of the target pictures, and the target pictures with newer date and time are ranked in the front.

[0204] In some embodiments, the step of determining the arrangement order of target pictures based on target labels includes: when each target label is a single-value label, using the descending order of date and time as the arrangement order of the corresponding target pictures.

[0205] Specifically, when each target label is a single-value label, these target pictures are sorted according to the date and time of the target pictures, and the target pictures with newer date and time are ranked in the front.

[0206] Further, after determining the arrangement order of target pictures based on target labels, the target pictures are displayed on the display interface according to this arrangement order.

[0207] In the embodiments of the present application, by determining the arrangement order of target pictures based on target labels and displaying the target pictures according to the arrangement order, the present application can rearrange the retrieved target pictures according to the matching degree between the target pictures and the search instruction, so that the target pictures are displayed on the display interface in descending order of the matching degree with the search instruction, making the target pictures that best match the search instruction be preferentially displayed in the front, so as to improve the logic and orderliness of picture display and optimize the user's search experience.

[0208] In the embodiments of the present application, when no matching target image is retrieved from the image set based on the target tag, the method further includes: correcting the search instruction to generate a recommended search instruction. The recommended search instruction is an instruction obtained by correcting the search instruction based on the target tag, and the recommended search instruction is used to more accurately match and retrieve the images stored in the image set.

[0209] For example: The user enters a correct search instruction, but there are no images that meet the requirements in the image set, resulting in no target image being retrieved. It is necessary to correct the search instruction to generate a recommended search instruction.

[0210] Or, the user enters an incorrect search instruction, which does not match the user's original search target and no matching target image can be retrieved from the image set. The search instruction is corrected to generate a recommended search instruction. For example: The correct search instruction is "Zhang Xiaosi in the second small class", but the user enters it incorrectly, and the actual search instruction entered is "Zhang Xiaosi in the first small class". At this time, no matching target image can be retrieved from the image set.

[0211] Among them, the steps of correcting the search instruction to generate a recommended search instruction specifically include steps S1 - S7:

[0212] Step S1: Combine a number of target tags into a second SQL query statement based on the OR logic;

[0213] The second SQL query statement is an SQL query statement obtained by combining the target tags through the OR logic. The second SQL query statement is used to screen out the image tags in the MySQL relational database that contain at least one target tag to determine the image information corresponding to each image tag.

[0214] Specifically, the combination logic of the tags is the OR logic. A number of target tags are combined into an SQL query statement, that is, the second SQL query statement, using the OR logic.

[0215] Step S2: Retrieve the matching second image information in the relational database based on the second SQL query statement;

[0216] The second image information is the image information of the recommended images stored in the image information table.

[0217] Specifically, the image tag table is retrieved through the second SQL query statement to screen out the image tags that contain at least one target tag, and the image information corresponding to each image tag in the image information table is determined, that is, the second image information.

[0218] Step S3: Retrieve the recommended images in the image set;

[0219] Among them, the recommended image is the image in the image set that matches the second image information, that is, the recommended image is the image that satisfies at least one target label.

[0220] Specifically, retrieve the recommended images that match the second image information from the image set.

[0221] Step S4: Calculate the image score for each recommended image based on the weight value of the target label corresponding to each recommended image;

[0222] Among them, each target label corresponds to a weight value, and the image score is the sum of the weight values of the target labels corresponding to the recommended image. The weight value of the target label can be set by those skilled in the art according to the priority of the target label, and no limitation is made here. For example: the priority of the target label with the target label name "Name" is higher than that of the target label with the target label name "Season", the weight value of the target label with the target label name "Name" is set to 2, and the weight value of the target label with the target label name "Season" is set to 1.

[0223] Specifically, for each recommended image, calculate the sum of the weight values of each target label included in the image label of the recommended image, and use the obtained result as the image score of the recommended image.

[0224] For example: Step S203 obtains four target labels, and the target label names of these four target labels are "Name", "Class", "Season", and "Theme Label" in sequence. The image label of the recommended image A only satisfies the target label names and target label values of three of these target labels, that is, the target label with the target label name "Name", the target label with the target label name "Class", and the target label with the target label name "Season". Among them, the weight value of the target label with the target label name "Name" is 2, the weight value of the target label with the target label name "Class" is 1.5, and the weight value of the target label with the target label name "Season" is 1. Then the image score of the recommended image A is 2 + 1.5 + 1 = 4.5.

[0225] Step S5: Select a preset number of recommended images from several recommended images in descending order of the image score;

[0226] Specifically, select the preset number of recommended images with the highest scores from several recommended images in descending order of the image score. The preset number can be set by those skilled in the art according to the image label, and no limitation is made here. Exemplarily, the preset number is 3.

[0227] Step S6: Generate a recommended search instruction based on the preset number of recommended images and the search instruction;

[0228] Specifically, the recommended search instruction is generated by direct replacement or by template filling.

[0229] In some embodiments, when the recommended search instruction is generated by direct replacement, the preset number is set to 1, that is, only one recommended picture with the highest score is selected in step S5. Step S6 specifically includes: determining the recommended picture with the highest score, and replacing the third tag value in the search instruction with the second tag value of the recommended picture to generate the recommended search instruction.

[0230] Wherein, the second tag value is different from the third tag value, the second tag value and the third tag value correspond to the same target tag name, the third tag value is the target tag value corresponding to any target tag in the search instruction, and the second tag value is the tag value that is different from the third tag value but corresponds to the same target tag name as the third tag value in the picture tags of the recommended picture with the highest score.

[0231] Specifically, the second tag value in the recommended picture with the highest score is used to replace the third tag value in the search instruction, and the replaced instruction is the recommended search instruction.

[0232] For example: the search instruction is "Zhang Xiaosi in Class One", there are two target tags, one is the target tag with the target tag name "Name" and the target tag value "Zhang Xiaosi", and the other is the target tag with the target tag name "Class" and the target tag value "Class One". For the recommended picture with the highest score, its picture tags include the tag with the tag name "Name" and the tag value "Zhang Xiaosi", and the tag with the tag name "Class" and the tag value "Class Two". The second tag value is "Class Two", the third tag value is "Class One", and the recommended search instruction is "Zhang Xiaosi in Class Two".

[0233] In the embodiments of the present application, by replacing the third tag value in the search instruction with the second tag value in the recommended picture with the highest score to generate the recommended search instruction, the present application can make the generated recommended search instruction have the same format as the search instruction input by the user, and can enhance the accuracy of the recommended search instruction.

[0234] In some embodiments, when the recommended search instruction is generated by template filling, step S6 specifically includes: selecting any preset instruction template, and substituting the tag values of the preset number of recommended pictures into the preset instruction template respectively to generate the preset number of recommended search instructions. Wherein, the preset instruction template is the template format of the search instruction, and the number and specific format of the preset instruction template can be set by those skilled in the art according to the picture tags, and are not limited herein.

[0235] Specifically, a preset instruction template is randomly selected from a number of preset instruction templates. For a preset number of recommended images, the tag values under the image tags of each recommended image are respectively substituted into the positions where the corresponding tag names are located in the preset instruction template, so as to generate corresponding recommended search instructions.

[0236] For example: the template format of the preset instruction template is "the ZZ image of class XX's student YY". The tag values under the image tags of the top 3 recommended images with the highest scores are respectively filled into the corresponding positions in the preset instruction template, so as to obtain 3 complete recommended search instructions.

[0237] In some embodiments, by substituting the tag values of a preset number of recommended images into the preset instruction template respectively to generate a preset number of recommended search instructions, the present application can make the recommended search instructions contain the tag values extracted and summarized.

[0238] For example: the template format of the preset instruction template is "the ZZ image of class XX's student YY". The search instruction input by the user is "student Zhang Xiaosi in class Xiaoban". At this time, the search instruction input by the user does not contain the specific value of "ZZ", and "ZZ" is the tag value that needs to be extracted and summarized. By substituting the tag values of a preset number of recommended images into the preset instruction template respectively, the specific value of "ZZ" can be obtained. Or, the search instruction input by the user contains the specific value of "ZZ", but this value is incorrect and no matching target image can be retrieved from the image set. By substituting the tag values of a preset number of recommended images into the preset instruction template respectively, the real specific value of "ZZ" can be obtained.

[0239] Step S7: Display the recommended search instructions.

[0240] Specifically, the recommended search instructions are displayed on the display interface so that the user can select a suitable recommended retrieval instruction and further retrieve the target image according to the recommended search instruction.

[0241] Compared with the existing solution, when the search instruction input by the user is incorrect, for example, there is a spelling mistake in the search instruction, resulting in the inability to retrieve a matching target image from the image set, the process is directly ended and the search instruction is not supported for correction. The present application corrects the search instruction, generates and displays the recommended search instructions. On the one hand, it can reduce the time for the user to adjust and input a new search instruction, making the user experience more friendly and smooth. On the other hand, it can improve the matching degree between the recommended search instructions and the images in the image set, facilitating the retrieval of target images that meet the user's search objectives based on the recommended search instructions in the subsequent process and improving the retrieval success rate.

[0242] Further, after displaying the recommended search instructions, the method further includes: in response to a user's selection operation on any one of the recommended search instructions, using the selected recommended search instruction as a new search instruction, and re-executing steps S201 to S205 to display target pictures retrieved based on the new search instruction on the display interface.

[0243] In an embodiment of the present application, by providing a picture retrieval method, the picture retrieval method includes: obtaining a search instruction; extracting information from the search instruction based on a large language model to obtain an initial label; performing parsing and correction processing on the initial label to obtain a corresponding target label; retrieving a picture set based on the target label to obtain at least one target picture; determining the arrangement order of the target pictures based on the target label, and displaying the target pictures in the arrangement order.

[0244] On the one hand, by obtaining a search instruction, extracting information from the search instruction based on a large language model to obtain an initial label, performing parsing and correction processing on the initial label to obtain a corresponding target label, and retrieving a picture set based on the target label to obtain at least one target picture, the present application can directly extract the initial label from the search instruction through the large language model and perform correction processing on it, improving the accuracy and comprehensiveness of label extraction, and further improving the accuracy of picture retrieval; on the other hand, by determining the arrangement order of the target pictures based on the target label and displaying the target pictures in the arrangement order, the present application can optimize the arrangement order of the target pictures according to the target label, improve the logic and coherence of picture display, and optimize the user's search experience.

[0245] Please refer to Figure 5 , Figure 5 which is a schematic structural diagram of a picture retrieval device provided by an embodiment of the present application;

[0246] Among them, the picture retrieval device is applied to an electronic device, and the electronic device includes a terminal or a server. Specifically, the picture retrieval device is applied to one or at least two processors of the electronic device.

[0247] As Figure 5 shown, the picture retrieval device 500 includes:

[0248] A text acquisition module 501, configured to obtain a search instruction.

[0249] An information extraction module 502, configured to extract information from the search instruction based on a large language model to obtain an initial label.

[0250] A label parsing module 503, configured to perform parsing and correction processing on the initial label to obtain a corresponding target label.

[0251] The image retrieval module 504 is used to retrieve an image set based on a target tag to obtain at least one target image.

[0252] The post-processing module 505 is used to determine the arrangement order of the target images based on the target tag and display the target images in the arrangement order.

[0253] In some embodiments, the image retrieval device 500 further includes a search instruction recommendation module. The search instruction recommendation module is configured to: when retrieving the image set based on the target tag and no matching target image is retrieved, combine several target tags into a second SQL query statement based on the OR logic; retrieve matching second image information in the relational database based on the second SQL query statement; retrieve recommended images in the image set, where the recommended images are images in the image set that match the second image information; calculate the image score of each recommended image based on the weight value of the target tag corresponding to each recommended image; select a preset number of recommended images from several recommended images in descending order of the image score; generate a recommended search instruction based on the preset number of recommended images and the search instruction; display the recommended search instruction.

[0254] In some embodiments, the target tag includes a target tag name and a target tag value. The image retrieval device 500 further includes an information confirmation module. The information confirmation module is configured to: after parsing and correcting the initial tag to obtain the corresponding target tag, query the number and target information of the target tag value corresponding to the target tag name in any tag information library corresponding to the target tag name, where the target information is the information corresponding to the target tag value in the tag information library; when the number is at least two, display each target information; in response to the user's selection operation on the target information, update the target tag to obtain the updated target tag; retrieve an image set based on the updated target tag to obtain several target images that match the updated target tag; determine the arrangement order of the several target images based on the updated target tag and display the target images in the arrangement order.

[0255] In the embodiments of the present application, the image retrieval device can also be built by hardware devices. For example, the image retrieval device can be built by one or more than two chips, and each chip can work in coordination with each other to complete the image retrieval method described in the above embodiments. For another example, the image retrieval device can also be built by various logic devices, such as being built by a general-purpose processor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), a single-chip microcomputer, an ARM processor (Advanced RISC Machines, ARM), or other programmable logic devices, discrete gate or transistor logic, discrete hardware components, or any combination of these components.

[0256] The image retrieval device in the embodiments of the present application can be a device, or a component, an integrated circuit, or a chip in a terminal. The device can be a mobile electronic device or a non-mobile electronic device. Exemplarily, the mobile electronic device can be a mobile phone, a tablet computer, a laptop computer, a handheld computer, a vehicle-mounted electronic device, a wearable device, an ultra-mobile personal computer (UMPC), a netbook, or a personal digital assistant (PDA), etc., and the non-mobile electronic device can be a server, a network attached storage (NAS), a personal computer (PC), a television (TV), a teller machine, or a self-service machine, etc. The embodiments of the present application do not make specific limitations.

[0257] The image retrieval device in the embodiments of the present application can be a device with an operating system. The operating system can be the Android operating system, the iOS operating system, or other possible operating systems. The embodiments of the present application do not make specific limitations.

[0258] The image retrieval device provided by the embodiments of the present application can implement Figure 2 the various processes implemented. To avoid repetition, they will not be elaborated here.

[0259] It should be noted that the above image retrieval device can execute the image retrieval method provided in the above embodiments, and has the corresponding functional modules and beneficial effects for executing the method. For the technical details not described in detail in the embodiments of the image retrieval device, reference can be made to the image retrieval method provided in the above embodiments.

[0260] In an embodiment of the present application, by providing an image retrieval device, the image retrieval device includes: a text acquisition module for acquiring a search instruction; an information extraction module for extracting information from the search instruction based on a large language model to obtain an initial label; a label parsing module for parsing and correcting the initial label to obtain a corresponding target label; an image retrieval module for retrieving an image set based on the target label to obtain at least one target image; and a post-processing module for determining the arrangement order of the target images based on the target label and displaying the target images in the arrangement order.

[0261] On the one hand, through the information extraction module extracting information from the search instruction based on the large language model, the label parsing module parsing and correcting the initial label, and the image retrieval module retrieving the image set based on the target label, the present application can directly extract the initial label from the search instruction through the large language model and perform correction processing on it, improving the accuracy and comprehensiveness of label extraction, and further improving the accuracy of image retrieval. On the other hand, through the post-processing module determining the arrangement order of the target images based on the target label and displaying the target images in the arrangement order, the present application can optimize the arrangement order of the target images according to the target label, improve the logic and coherence of image display, and optimize the user's search experience.

[0262] Please refer to Figure 6 , Figure 6 which is a schematic structural diagram of an electronic device provided by an embodiment of the present application;

[0263] As Figure 6 shown, the electronic device 600 includes one or more processors 601 and a memory 602. Among them, Figure 6 one processor 601 is taken as an example herein.

[0264] The processor 601 and the memory 602 can be connected through a bus or other means, Figure 6 and taking connection through a bus as an example herein.

[0265] The processor 601 is used to provide computing and control capabilities to control the electronic device 600 to execute corresponding tasks. For example, controlling the electronic device 600 to execute the image retrieval method in any of the above method embodiments.

[0266] On the one hand, the present application can directly extract the initial label from the search instruction through the large language model and perform correction processing on it, improving the accuracy and comprehensiveness of label extraction, and further improving the accuracy of image retrieval. On the other hand, the present application can optimize the arrangement order of the target images according to the target label, improve the logic and coherence of image display, and optimize the user's search experience.

[0267] The processor 601 may be a general-purpose processor, including a Central Processing Unit (CPU), a Network Processor (NP), a hardware chip, or any combination thereof; it may also be a Digital Signal Processing (DSP), an Application Specific Integrated Circuit (ASIC), a Programmable Logic Device (PLD), or a combination thereof. The above PLD may be a Complex Programmable Logic Device (CPLD), a Field-Programmable Gate Array (FPGA), a Generic Array Logic (GAL), or any combination thereof.

[0268] The memory 602, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs, non-transitory computer-executable programs, and modules, such as the program instructions / modules corresponding to the picture retrieval method in the embodiments of the present application. By running the non-transitory software programs, instructions, and modules stored in the memory 602, the processor 601 can implement the picture retrieval method in any of the above method embodiments. Specifically, the memory 602 may include a Volatile Memory (VM), such as a Random Access Memory (RAM); the memory 602 may also include a Non-Volatile Memory (NVM), such as a Read-Only Memory (ROM), a Flash memory, a Hard Disk Drive (HDD), or a Solid-State Drive (SSD), or other non-transitory solid-state storage devices; the memory 602 may also include a combination of the above types of memories.

[0269] The memory 602 may include high-speed random access memory and may also include non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other non-volatile solid-state storage devices. In some embodiments, the memory 602 optionally includes a memory remotely set relative to the processor 601, and these remote memories can be connected to the processor 601 through a network. Examples of the above network include but are not limited to the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.

[0270] One or more modules are stored in the memory 602 and, when executed by one or more processors 601, perform the picture retrieval method in any of the above method embodiments. For example, execute each of the steps described above. Figure 2 shown steps.

[0271] In the embodiments of the present application, the electronic device 600 may also have components such as a wired or wireless network interface and an input / output interface for input / output. The electronic device 600 may also include other components for implementing the functions of the device, which will not be elaborated here.

[0272] The embodiments of the present application also provide a non-volatile computer-readable storage medium, such as a memory including program code. The above program code can be executed by a processor to complete the picture retrieval method in the above embodiments. For example, the non-volatile computer-readable storage medium may be a read-only memory (ROM), a random access memory (RAM), a compact disc read-only memory (CDROM), a magnetic tape, a floppy disk, and an optical data storage device, etc.

[0273] The embodiments of the present application also provide a computer program product. The computer program product includes one or more pieces of program code, and the program code is stored in a computer-readable storage medium. The processor of the electronic device reads the program code from the computer-readable storage medium, and the processor executes the program code to complete the method steps of the picture retrieval method provided in the above embodiments.

[0274] Those of ordinary skill in the art can understand that all or part of the steps of implementing the above embodiments can be completed by hardware, or can be completed by hardware related to program code. The program can be stored in a computer-readable storage medium, and the above-mentioned storage medium can be a read-only memory, a magnetic disk, or an optical disc, etc.

[0275] Through the description of the above embodiments, those of ordinary skill in the art can clearly understand that each embodiment can be implemented by means of software plus a general hardware platform, and of course, it can also be implemented by hardware. Those of ordinary skill in the art can understand that all or part of the processes of implementing the methods in the above embodiments can be completed by a computer program instructing related hardware. The program can be stored in a computer-readable storage medium. When the program is executed, it may include the processes of the embodiments of the above methods. Among them, the storage medium may be a magnetic disk, an optical disc, a read-only memory (ROM), or a random access memory (RAM), etc.

[0276] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than limiting them; under the idea of the present application, the technical features in the above embodiments or different embodiments can also be combined, and the steps can be implemented in any order, and there are many other variations in different aspects of the present application as described above. For the sake of brevity, they are not provided in detail; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present application.

Claims

1. A method for image retrieval, characterized in that, Including: Obtain a search instruction; Perform information extraction on the search instruction based on a large language model to obtain an initial label; Perform parsing and correction processing on the initial label to obtain a corresponding target label; Retrieve an image set based on the target label to obtain at least one target image; Determine the arrangement order of the target images based on the target label and display the target images in the arrangement order.

2. The method according to claim 1, characterized in that, The large language model includes an information extraction function, and each initial label includes a label name and a label value. Each input parameter of the information extraction function corresponds to a label name; The performing information extraction on the search instruction based on the large language model to obtain an initial label includes: Parse the search instruction based on the information extraction function, extract the label value corresponding to the input parameter from the search instruction, and combine each label name with the corresponding label value into an initial label with a preset structure; The method further includes: Define the configuration information of the information extraction function in an extensible markup language file, where the configuration information includes input parameters; When adding a new label name to any image in the image set, add the input parameter corresponding to the new label name to the file.

3. The method according to claim 1, wherein The performing parsing and correction processing on the initial label to obtain a corresponding target label includes: Based on the format of each initial label, perform parsing on the initial label according to the preset parsing method corresponding to the format to obtain a parsed initial label, where the parsed initial label includes a label name and a label value, and each label name corresponds to at least one label value; When any parsed initial label meets the first condition, delete the first label value in the parsed initial label to obtain a corresponding target label; Wherein, the first condition includes: any label value corresponding to any label name is not within the preset range value corresponding to the label name, and the first label value is the label value not within the preset range value corresponding to any label name; When any parsed initial label meets the second condition, perform position replacement on the label value to generate a target label, where each label value included in the target label is located under the label name corresponding to the first regular expression matched by the label value; Wherein, the second condition includes: when matching each label value based on the first regular expression corresponding to each label name, the label name corresponding to the first regular expression matched by any label value is different from the label name corresponding to the label value; 4. The method according to claim 3, characterized in that The performing parsing and correction processing on the initial label to obtain a corresponding target label further includes: When any parsed initial label meets the third condition, perform mapping conversion on the label value to generate a target label, where each label value included in the target label matches the second regular expression corresponding to the label name; Wherein, the third condition includes: when matching the label value corresponding to each label name based on the second regular expression corresponding to each label name, any label value does not match the second regular expression; When each of the parsed initial tags satisfies the fourth condition, extract the tag value matched by the third regular expression from the search instruction, and combine the tag value with the corresponding tag name to form a target tag; Among them, the fourth condition includes: matching the search instruction based on the third regular expression, there is a preset tag value in the search instruction that matches the third regular expression, and each parsed initial tag does not contain the preset tag value.

5. The method according to claim 1, wherein The picture information and picture tags of pictures are stored in a relational database, and the relational database includes a picture information table and a picture tag table. Among them, the picture information table is used to store picture information, and the picture tag table is used to store picture tags; The retrieving the picture set based on the target tag to obtain at least one target picture includes: Combining a plurality of the target tags into a first SQL query statement based on the AND logic; Retrieving matching first picture information in the relational database based on the first SQL query statement; Retrieving matching target pictures in the picture set based on the first picture information.

6. The method according to claim 5, wherein When no matching target picture is retrieved based on the target tag for the picture set, the method further includes: Combining a plurality of the target tags into a second SQL query statement based on the OR logic; Retrieving matching second picture information in the relational database based on the second SQL query statement; Retrieving recommended pictures in the picture set, where the recommended pictures are pictures in the picture set that match the second picture information; Calculating the picture score of each recommended picture based on the weight value of the target tag corresponding to each recommended picture; Selecting a preset number of recommended pictures from the plurality of recommended pictures in descending order of the picture score; Generating a recommended search instruction based on the preset number of recommended pictures and the search instruction; Displaying the recommended search instruction.

7. The method according to claim 6, characterized in that The generating a recommended search instruction based on the preset number of recommended pictures and the search instruction includes: Determining the recommended picture with the highest score, and replacing the third tag value in the search instruction with the second tag value of the recommended picture to generate a recommended search instruction, where the second tag value is different from the third tag value, and the second tag value and the third tag value correspond to the same target tag name; Alternatively, select any preset instruction template, and substitute the tag values of the preset number of recommended pictures into the preset instruction template to generate a preset number of recommended search instructions.

8. The method according to claim 1, wherein The target tag is a multi-value tag or a single-value tag, the target tag includes a target tag value, the number of target tag values in the multi-value tag is at least two, and each target picture corresponds to a date and time; The determining the arrangement order of the target pictures based on the target tag and displaying the target pictures in the arrangement order includes: When any one of the target tags is a multi-value tag, determining the number of target tag values matched by each target picture in the multi-value tag, and using the order from large to small of the number as the arrangement order of the corresponding target pictures; When the quantities corresponding to at least two target pictures are the same, use the descending order of the date and time as the arrangement order of the corresponding target pictures; Alternatively, when each of the target tags is a single-value tag, use the descending order of the date and time as the arrangement order of the corresponding target pictures; Display the target pictures according to the arrangement order.

9. The method according to any one of claims 1-8, characterized in that, The target tag includes a target tag name and a target tag value; After parsing and correcting the initial tag to obtain the corresponding target tag, the method further includes: In the tag information library corresponding to any of the target tag names, query the quantity and target information of the target tag value corresponding to the target tag name, where the target information is the information corresponding to the target tag value in the tag information library; When the quantity is at least two, display each of the target information; In response to the user's selection operation on the target information, update the target tag to obtain an updated target tag; Retrieve the picture set based on the updated target tag to obtain a number of target pictures that match the updated target tag; Determine the arrangement order of the number of target pictures based on the updated target tag, and display the target pictures according to the arrangement order.

10. An electronic device, characterized in that, Includes: At least one processor, and A memory communicatively connected to the at least one processor, where The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method according to any one of claims 1-9.