Data search method and device, computer equipment and storage medium

By using structured templates to extract structured tags and generate structured text in multimodal data search, the problem of inaccurate search results in the prior art is solved, and higher search accuracy and efficiency are achieved.

CN119938880APending Publication Date: 2025-05-06SHANGHAI LIANYING ZHIYUAN MEDICAL TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411999116.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-31
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

Existing multimodal algorithms cannot accurately reflect the correlation between multimodal data, resulting in inaccurate data search results.

Method used

By determining the first feature vector corresponding to the input data and matching it with the second feature vector corresponding to the candidate data, the structured template is used to extract structured labels and generate structured text, and superimposing the encoding results of text data and structured text to obtain accurate feature vectors.

Benefits of technology

The coded information can accurately reflect the correlation between multimodal data and improve the accuracy and efficiency of data search.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119938880A_ABST
    Figure CN119938880A_ABST
Patent Text Reader

Abstract

The invention relates to a data search method and device, computer equipment and a storage medium, and the method comprises the steps: determining a first feature vector corresponding to input data; matching the first feature vector with a second feature vector corresponding to each piece of candidate data, and determining the candidate data corresponding to the second feature vector matched with the first feature vector as target data; if the input data is a text, the first feature vector is associated with a feature vector corresponding to the text data and a feature vector of a structured text corresponding to the text data; if the candidate data is the text, the second feature vector is associated with the feature vector corresponding to the text data and the feature vector of the structured text corresponding to the text data. Through the data search method and device, the problem that the data search result is inaccurate due to the fact that the encoding information of the current multi-modal algorithm cannot accurately reflect the association between the multi-modal data is solved, the encoding information can accurately reflect the association between the multi-modal data, and the accuracy of data search is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of data processing technology, and in particular to a data search method, apparatus, computer equipment and storage medium. Background Art

[0002] Multimodal algorithms can process multiple types of data at the same time, such as images, text, and audio. With the development of artificial intelligence and neural network technology, multimodal algorithms are gradually applied to various fields. For example, in medical information systems, multimodal information association is performed on medical imaging information and medical text reports to facilitate mutual search of multimodal data, which is beneficial for doctors to query and refer to similar cases. However, the encoding information of current multimodal algorithms cannot accurately reflect the association between multimodal data, especially the association between relevant features that are clinically valuable, resulting in inaccurate data search results.

[0003] There is currently no effective solution to the problem that the encoding information of the current multimodal algorithm in the relevant technology cannot accurately reflect the relationship between the multimodal data, resulting in inaccurate data search results. Summary of the invention

[0004] In this embodiment, a data search method, apparatus, computer device and storage medium are provided to solve the problem in the related art that the encoding information of the current multimodal algorithm cannot accurately reflect the relationship between multimodal data, resulting in inaccurate data search results.

[0005] In a first aspect, a data search method is provided in this embodiment, the method comprising:

[0006] determining a first eigenvector corresponding to the input data;

[0007] Matching the first feature vector with a second feature vector corresponding to each candidate data, and determining the candidate data corresponding to the second feature vector that matches the first feature vector as the target data;

[0008] Wherein, if the input data is text, the first feature vector is associated with a feature vector corresponding to the text data and a feature vector of a structured text corresponding to the text data;

[0009] If the candidate data is text, the second feature vector is associated with a feature vector corresponding to the text data and a feature vector of a structured text corresponding to the text data.

[0010] In some embodiments, the input data is text, and determining the first feature vector includes:

[0011] Extracting multiple structured tags from the text data based on a preset structured template;

[0012] Based on each of the structured tags, generating the structured text corresponding to the text data;

[0013] The encoding result of the text data and the encoding result of the structured text are superimposed to obtain the first feature vector.

[0014] In some of the embodiments, the candidate data is text, and determining the second feature vector includes:

[0015] Extracting multiple structured tags from the text data based on a preset structured template;

[0016] Based on each of the structured tags, generating the structured text corresponding to the text data;

[0017] The encoding result of the text data and the encoding result of the structured text are superimposed to obtain the second feature vector.

[0018] In some embodiments, the input data is at least one structured label, and determining a first feature vector corresponding to the input data includes:

[0019] Based on the at least one structured tag, generating a structured text corresponding to the at least one structured tag;

[0020] The feature vector corresponding to the structured text is taken as the first feature vector.

[0021] In some embodiments, both the input data and the candidate data are image data, and the method further comprises:

[0022] Determining a first similarity between the first feature vector and an encoding of a preset label;

[0023] Determining a second similarity between each of the second feature vectors and the encoding of the preset label;

[0024] Based on the first similarity and each of the second similarities, candidate data corresponding to the second feature vector matching the first feature vector is determined as target data.

[0025] In some embodiments, there are multiple preset tags, and the method further includes:

[0026] Traversing each of the preset tags, and determining a third similarity between the first feature vector and each of the second feature vectors based on the first similarity associated with the encoding of the preset tag and each of the second similarities;

[0027] Based on a preset weight corresponding to the code of each of the preset tags, weighted averaging the third similarities to obtain a fourth similarity between the first feature vector and each of the second feature vectors;

[0028] Based on each of the fourth similarities, candidate data corresponding to the second feature vector that matches the first feature vector is determined as the target data.

[0029] In some of the embodiments, it also includes:

[0030] Obtaining a second feature vector corresponding to each of the candidate data;

[0031] storing each second eigenvector in a vector database;

[0032] The matching of the first feature vector with the second feature vector corresponding to each candidate data refers to matching the first feature vector with each second feature vector in the vector database.

[0033] In a second aspect, a data search device is provided in this embodiment, the device comprising:

[0034] An encoding module, configured to determine a first eigenvector corresponding to the input data;

[0035] A search module, configured to match the first feature vector with a second feature vector corresponding to each candidate data, and determine that the candidate data corresponding to the second feature vector matching the first feature vector is the target data;

[0036] Wherein, if the input data is text, the first feature vector is associated with a feature vector corresponding to the text data and a feature vector of a structured text corresponding to the text data;

[0037] If the candidate data is text, the second feature vector is associated with a feature vector corresponding to the text data and a feature vector of a structured text corresponding to the text data.

[0038] In a third aspect, a computer device is provided in this embodiment, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the data search method described in the first aspect when executing the computer program.

[0039] In a fourth aspect, in this embodiment, a storage medium is provided, on which a computer program is stored, and when the program is executed by a processor, the data search method described in the first aspect is implemented.

[0040] Compared with the related art, the data search method, apparatus, computer equipment and storage medium provided in this embodiment determine the first feature vector corresponding to the input data; match the first feature vector with the second feature vector corresponding to each candidate data, and determine that the candidate data corresponding to the second feature vector matching the first feature vector is the target data; wherein, if the input data is text, the first feature vector is associated with the feature vector corresponding to the text data and the feature vector of the structured text corresponding to the text data; if the candidate data is text, the second feature vector is associated with the feature vector corresponding to the text data and the feature vector of the structured text corresponding to the text data, which solves the problem that the encoding information of the current multimodal algorithm cannot accurately reflect the association between multimodal data, resulting in inaccurate data search results, and realizes that the encoding information can accurately reflect the association between multimodal data, thereby improving the accuracy of data search.

[0041] Details of one or more embodiments of the present application are set forth in the following drawings and description to make other features, objects, and advantages of the present application more readily apparent. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:

[0043] Figure 1 It is a hardware structure block diagram of a terminal device of a data search method provided by an embodiment of the present application;

[0044] Figure 2 is a flow chart of a data search method provided by an embodiment of the present application;

[0045] Figure 3 is a schematic diagram of a structured text generation method provided in an embodiment of the present application;

[0046] Figure 4 is a flow chart of a data search method based on structured tags provided in one embodiment of the present application;

[0047] Figure 5 is a flowchart of a method for searching images by image provided by an embodiment of the present application;

[0048] Figure 6 It is a flowchart of a data search method provided by an embodiment of the present application;

[0049] Figure 7 is a flowchart of a data search method provided by another embodiment of the present application;

[0050] Figure 8is a flow chart of a data search method provided by a preferred embodiment of the present application;

[0051] Fig. 9 It is a structural block diagram of a data search device provided in one embodiment of the present application.

[0052] In the figure: 102, processor; 104, memory; 106, transmission device; 108, input and output device; 10, encoding module; 20, search module. DETAILED DESCRIPTION

[0053] In order to more clearly understand the purpose, technical solutions and advantages of the present application, the present application is described and illustrated below in conjunction with the accompanying drawings and embodiments.

[0054] Unless otherwise defined, the technical terms or scientific terms involved in this application shall have the general meaning understood by people with ordinary skills in the technical field to which this application belongs. The words "one", "a", "a", "the", "these" and the like in this application do not represent quantitative restrictions, and they can be singular or plural. The terms "include", "comprise", "have" and any variants thereof involved in this application are intended to cover non-exclusive inclusions; for example, a process, method and system, product or device comprising a series of steps or modules (units) is not limited to the listed steps or modules (units), but may include unlisted steps or modules (units), or may include other steps or modules (units) inherent to these processes, methods, products or devices. The words "connect", "connected", "coupled" and the like involved in this application are not limited to physical or mechanical connections, but may include electrical connections, whether directly or indirectly. The "multiple" involved in this application refers to two or more. "And / or" describes the association relationship of associated objects, indicating that there can be three relationships. For example, "A and / or B" can mean: A exists alone, A and B exist at the same time, and B exists alone. Usually, the character " / " indicates that the objects associated with each other are in an "or" relationship. The terms "first", "second", "third", etc. involved in this application are only used to distinguish similar objects and do not represent a specific ordering of the objects.

[0055] The method embodiment provided in this embodiment can be executed in a terminal, a computer or a similar computing device. For example, running on a terminal, Figure 1 FIG. 1 is a hardware structure diagram of a terminal of the data search method of this embodiment. Figure 1 As shown, the terminal may include one or more ( Figure 1Only one is shown in the figure) processor 102 and memory 104 for storing data, wherein processor 102 may include but is not limited to a processing device such as a microprocessor MCU or a programmable logic device FPGA. The above terminal may also include a transmission device 106 and an input and output device 108 for communication functions. It can be understood by those skilled in the art that Figure 1 The structure shown is only for illustration and does not limit the structure of the above terminal. Figure 1 More or fewer components as shown, or with Figure 1 Different configurations shown.

[0056] The memory 104 can be used to store computer programs, for example, software programs and modules of application software, such as the computer program corresponding to the data search method in this embodiment. The processor 102 executes various functional applications and data processing by running the computer program stored in the memory 104, that is, to implement the above method. The memory 104 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some examples, the memory 104 may further include a memory remotely arranged relative to the processor 102, and these remote memories may be connected to the terminal via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0057] The transmission device 106 is used to receive or send data via a network. The above network includes a wireless network provided by the communication provider of the terminal. In one example, the transmission device 106 includes a network adapter (Network Interface Controller, referred to as NIC), which can be connected to other network devices through a base station so as to communicate with the Internet. In one example, the transmission device 106 can be a radio frequency (Radio Frequency, referred to as RF) module, which is used to communicate with the Internet wirelessly.

[0058] In this embodiment, a data search method is provided. Figure 2 is a flow chart of the data search method of this embodiment. Figure 2 As shown, the process includes the following steps:

[0059] Step S210, determining a first eigenvector corresponding to the input data;

[0060] Step S220, matching the first feature vector with the second feature vector corresponding to each candidate data, and determining that the candidate data corresponding to the second feature vector matching the first feature vector is the target data; wherein, if the input data is text, the first feature vector is associated with the feature vector corresponding to the text data and the feature vector of the structured text corresponding to the text data; if the candidate data is text, the second feature vector is associated with the feature vector corresponding to the text data and the feature vector of the structured text corresponding to the text data.

[0061] Specifically, the input data is encoded to obtain a first feature vector corresponding to the input data, and each candidate data is encoded to obtain a second feature vector corresponding to the candidate data. The input data and the candidate data can be images, texts, and structured tags, etc., such as searching for images by text, searching for texts by images, searching for images by structured tags, etc., which are not limited here. Among them, structured tags refer to standardized representations formed based on the induction of key information.

[0062] If the input data is text, multiple structured tags are extracted from the text data according to a preset structured template, and structured text corresponding to the text data is generated based on each structured tag, and the text data and the structured text are encoded respectively to obtain a feature vector corresponding to the text data and a feature vector of the structured text, and the feature vector corresponding to the text data and the feature vector of the structured text are added to obtain a first feature vector. Similarly, if the candidate data is text, multiple structured tags are extracted from the text data according to a preset structured template, and structured text corresponding to the text data is generated based on each structured tag, and the text data and the structured text are encoded respectively to obtain a feature vector corresponding to the text data and a feature vector of the structured text, and the feature vector corresponding to the text data and the feature vector of the structured text are added to obtain a second feature vector.

[0063] Among them, the preset structured template is used to indicate the structured label category extracted from the text data, including multiple labels, such as lesion category, lesion location, lesion manifestation, etc. The lesion manifestation includes the numerical characteristics of organs or tumors and other parts and the levels corresponding to the numerical characteristics. For example, the size of a high-density shadow is about 15*15mm. The structured template can provide multiple labels in single-select, multiple-select or judgment form, which is not limited here. The structured text generated based on each structured label is a standardized text after the original text data is accurately refined. The corresponding structured text can be generated according to each group of different structured labels to obtain multiple structured texts. Each structured text contains key features in the text data that are associated with the field or scene to which the data belongs, such as the key features used to describe the results of brain magnetic resonance imaging, including brain tissue morphology, brain lesion location and its manifestation, etc., and the key features used to describe the results of abdominal ultrasound scans, including organ morphology and size, echo characteristics, etc. For example, Figure 3 As shown, the preset structured template includes lesion category, lesion location and lesion manifestation. In the search scenario of brain scan related data, according to the label category indicated by the structured template, corresponding multiple structured labels are extracted from the text data. The extracted structured labels include brain stem, high-density shadow, display size 15*15, etc. Based on each structured label, the corresponding examination description is generated as structured text, for example, "high-density shadow can be seen in the brain stem, the size of which is about 15*15mm".

[0064] Furthermore, the first feature vector is matched with the second feature vector corresponding to each candidate data to obtain the similarity between the first feature vector and each second feature vector, and the candidate data corresponding to the second feature vector matching the first feature vector is determined as the target data according to the calculated similarity. For example, the calculated similarities are sorted, and the first n candidate data with the largest similarity are selected as the target data, or when the calculated similarity exceeds a preset similarity threshold, the current candidate data is determined as the target data, etc.

[0065] It should be noted that the above encoding processing for input data and candidate data can be implemented by using an encoding model, and the encoding model includes an image encoding structure and a text encoding structure, and the image encoding structure and the text encoding structure can be customized. For example, the encoding part of STU-Net-B is used as the image encoding structure, and the image encoding process includes encoding the image data to obtain a feature map of 256 channels, compressing the dimension of each channel through a pooling operation, thereby converting the 256 channel feature map into a 256*1 feature vector, and using a text embedding model as a text encoding structure, the text encoding process includes encoding to obtain a feature vector corresponding to the text data and a feature vector of the structured text, mapping the 1536*1 feature vector to a 256*1 feature vector through a fully connected layer, and adding the feature vector corresponding to the text data processed by the fully connected layer and the feature vector of the structured text, so as to achieve alignment of multimodal data in the same dimension and make them comparable in the same dimension. The encoding model can be obtained by training a pre-trained model based on a constructed domain dataset. For example, in the medical field, the domain dataset can be data for a single imaging site or a single disease, ensuring that a high-quality model is trained with a smaller amount of data, so that the trained model can better express the characteristics of the corresponding field or scenario, which helps to improve the accuracy of subsequent data searches.

[0066] Taking the brain scanning scenario as an example, a training data set is constructed. The training data set contains multiple groups of sample brain CT data and structured templates for brain CT. Each group of brain CT data includes brain CT images and brain CT image reports that are multimodal. That is, each group of brain CT data is usually a brain CT image and a brain CT image report obtained from the same examination of the target individual. The sample data in the training data set is enhanced, such as randomly rotating the brain CT image, randomly removing some text content in the image report, and generating structured text corresponding to the image report. Expanding the training data set helps to improve the generalization of the model and effectively constrain the model training direction through structured text. Based on the loss function constructed by the encoding results of the brain CT image and the encoding results of the brain CT image report, the pre-trained model is trained with the data-enhanced training data set to train and adjust the image encoding structure parameters and the fully connected layer parameters of the text encoding in the model. During the model training process, the feature vector of the text is obtained by adding the feature vector corresponding to the original text data and the feature vector of the structured text, thereby creating a contrastive learning paradigm with dual-channel input.

[0067] In addition, the model training parameters can be customized. For example, the preset training batch size is 1000 and the number of training rounds is 10. The pre-trained model is trained based on the training data set. The training process uses the Adaptive Moment Estimation (Adam) optimizer, and the initial learning rate of the optimizer is preferably 10 -6 .

[0068] Multimodal algorithms can process multiple types of data at the same time, such as images, text, and audio. With the development of artificial intelligence and neural network technology, multimodal algorithms are gradually applied to various fields. For example, in medical information systems, multimodal information association is performed on medical imaging information and text reports to facilitate mutual search of data, which is beneficial for doctors to query and refer to similar cases. However, the current multimodal algorithms cannot accurately encode text information, resulting in inaccurate data search results.

[0069] Compared with the prior art, the present application determines the first feature vector corresponding to the input data; matches the first feature vector with the second feature vector corresponding to each candidate data, and determines the candidate data corresponding to the second feature vector matching the first feature vector as the target data; wherein, if the input data is text, the first feature vector is associated with the feature vector corresponding to the text data and the feature vector of the structured text corresponding to the text data; if the candidate data is text, the second feature vector is associated with the feature vector corresponding to the text data and the feature vector of the structured text corresponding to the text data. Based on this, by embedding the key features of a specific field or scene into the encoding for representation, the text encoding can better capture the semantic information related to the field or scene, which helps to align the multimodal features, thereby accurately and efficiently searching for data based on the feature vector obtained by encoding, solving the problem that the encoding information of the current multimodal algorithm cannot accurately reflect the association between multimodal data, resulting in inaccurate data search results, and achieving adaptation to specific fields or scenes. The encoding information can accurately reflect the association between multimodal data, improve the accuracy and efficiency of data search, and is suitable for all encodable data types for mutual search, with strong scalability, so that it can be applied to similar case query and reference, scientific research data grouping and cleaning and other scenarios.

[0070] In some embodiments, the input data is text, and determining the first feature vector comprises the following steps:

[0071] Extract multiple structured tags from text data based on preset structured templates;

[0072] Based on each structured tag, generate structured text corresponding to the text data;

[0073] The encoding result of the text data and the encoding result of the structured text are superimposed to obtain a first feature vector.

[0074] In this embodiment, a structured template is predefined according to the actual field or specific scenario. The structured template includes multiple tags, for example, the preset structured template includes lesion category, lesion location, lesion manifestation, etc., which are used to indicate the structured tag category extracted from the text data.

[0075] When the input data is text, multiple structured tags are extracted from the text data according to the tag category indicated by the structured template, and structured text corresponding to the text data is generated based on each structured tag. The structured text is a standardized text after the original text data is accurately refined, and it contains key features in the text data that are associated with the field or scene to which the data belongs. Then, the text data and the structured text are encoded respectively to obtain the feature vector corresponding to the text data and the feature vector of the structured text, and the feature vector corresponding to the text data and the feature vector of the structured text are added to obtain the first feature vector.

[0076] Through this embodiment, based on a preset structured template, multiple structured tags are extracted from text data, and based on each structured tag, a structured text corresponding to the text data is generated, and then the encoding result of the text data and the encoding result of the structured text are superimposed to obtain a first feature vector, so as to accurately capture the semantic information related to the field or scene in the text data, realize accurate encoding of the text data, and help improve the accuracy of data search. At the same time, the text encoding includes the feature vector corresponding to the text data and the feature vector of the structured text, so that it is not limited to strict label expression and has wider adaptability.

[0077] In some embodiments, the candidate data is text, and determining the second feature vector comprises the following steps:

[0078] Extract multiple structured tags from text data based on preset structured templates;

[0079] Based on each structured tag, generate structured text corresponding to the text data;

[0080] The encoding result of the text data and the encoding result of the structured text are superimposed to obtain a second feature vector.

[0081] In this embodiment, a structured template is predefined according to the actual field or specific scenario. The structured template includes multiple tags, for example, the preset structured template includes lesion category, lesion location, lesion manifestation, etc., which are used to indicate the structured tag category extracted from the text data.

[0082] When the input data is text, multiple structured tags are extracted from the text data according to the tag category indicated by the structured template, and structured text corresponding to the text data is generated based on each structured tag. The structured text is a standardized text after the original text data is accurately refined, and it contains key features in the text data that are associated with the field or scene to which the data belongs. Then, the text data and the structured text are encoded respectively to obtain the feature vector corresponding to the text data and the feature vector of the structured text, and the feature vector corresponding to the text data and the feature vector of the structured text are added to obtain the second feature vector.

[0083] Through this embodiment, based on a preset structured template, multiple structured tags are extracted from text data, and based on each structured tag, a structured text corresponding to the text data is generated, and then the encoding result of the text data and the encoding result of the structured text are superimposed to obtain a second feature vector, so as to accurately capture the semantic information related to the field or scene in the text data, realize accurate encoding of text data, and help improve the accuracy of data search.

[0084] In some embodiments, the input data is at least one structured label, and determining a first feature vector corresponding to the input data comprises the following steps:

[0085] Based on at least one structured tag, generate a structured text corresponding to the at least one structured tag;

[0086] The feature vector corresponding to the structured text is taken as the first feature vector.

[0087] Specifically, Figure 4 As shown, when the input data is a single or multiple structured tags, the corresponding structured text is generated based on the single or multiple structured tags, and the feature vector obtained by encoding the structured text is used as the first feature vector. Among them, the structured tag refers to a standardized representation formed based on the induction of key information, such as the location and manifestation of the lesion.

[0088] Taking the structured tag search image as an example, the data search process is explained. Generate a structured text corresponding to at least one structured tag, encode the structured text, use the encoded feature vector as the first feature vector T1, and compare the first feature vector T1 with the second feature vector I corresponding to each candidate image in the candidate image library. k Matching is performed to obtain the similarity I between the first feature vector and each second feature vector k T1, select the image data with greater similarity as the target data.

[0089] It should be noted that, in this embodiment, structured text may also be directly input to perform data search, so as to search for images or texts that match the structured text.

[0090] Through this embodiment, if the input data is at least one structured tag, a structured text corresponding to it is generated based on at least one structured tag, and the feature vector corresponding to the structured text is used as the first feature vector. In this way, accurate data search is achieved based on the input structured tag, and mutual search between multimodal data is supported, which helps to expand application scenarios.

[0091] In some embodiments, both the input data and the candidate data are image data, and the data search method further includes the following steps:

[0092] Determining a first similarity between the first feature vector and an encoding of a preset label;

[0093] Determining a second similarity between each second feature vector and an encoding of a preset label;

[0094] Based on the first similarity and each second similarity, candidate data corresponding to the second feature vector matching the first feature vector is determined as target data.

[0095] Specifically, when the input data and the candidate data are both image data, the input data is encoded to obtain a first feature vector corresponding to the input data, a first similarity between the first feature vector and the encoding of a preset label is calculated, a second similarity between each second feature vector and the encoding of the preset label is calculated, and then based on the first similarity and each second similarity, the similarity between the first feature vector and different second feature vectors is calculated, and image data with greater similarity is selected as target data.

[0096] It should be noted that the preset labels are usually determined by the structured template in a specific field or scenario, that is, the preset labels are specific instances of the labels defined by the structured template, and the preset labels can be flexibly selected according to the actual image search requirements. For example, the preset structured template includes a lesion manifestation label, and the corresponding preset label can be the display size of the high-density shadow, etc., so as to search for images with similar high-density shadow display sizes.

[0097] Through this embodiment, a first similarity between a first feature vector and an encoding of a preset label is determined, a second similarity between each second feature vector and the encoding of the preset label is determined, and based on the first similarity and each second similarity, candidate data corresponding to a second feature vector matching the first feature vector is determined as target data, thereby accurately matching images with the encoding of the preset label as a medium, thereby achieving fast and accurate image search.

[0098] In some embodiments, there are multiple preset tags, and the data search method further includes the following steps:

[0099] Traversing each preset tag, and determining a third similarity between the first feature vector and each second feature vector based on the first similarity and each second similarity associated with the encoding of the preset tag;

[0100] Based on a preset weight corresponding to the code of each preset tag, weighted averaging the third similarities to obtain a fourth similarity between the first feature vector and each second feature vector;

[0101] Based on each fourth similarity, candidate data corresponding to the second feature vector matching the first feature vector is determined as target data.

[0102] Specifically, when there are multiple preset tags, each preset tag is encoded to obtain the encoding of the preset tag, and the first similarity between the first feature vector and the encoding of the preset tag, and the second similarity between each second feature vector and the encoding of each preset tag are calculated.

[0103] Traverse each preset tag, and calculate the third similarity between the first feature vector and each second feature vector based on the first similarity and each second similarity associated with the encoding of the same preset tag. The third similarity refers to the similarity between the first feature vector and each second feature vector calculated with the encoding of a preset tag as the medium. Then, based on the preset weight corresponding to the encoding of each preset tag, perform weighted averaging on each third similarity to obtain the fourth similarity between the first feature vector and each second feature vector. The fourth similarity refers to the similarity between the first feature vector and each second feature vector calculated with the encoding of all preset tags as the medium, and select the image data with the larger fourth similarity as the target data.

[0104] It should be noted that the preset tags in this embodiment are usually determined by a structured template in a specific field or scenario, that is, the preset tags are specific instances of tags defined by the structured template, and the preset tags can be flexibly selected according to actual image search requirements.

[0105] For example, Figure 5 As shown in FIG. 1 , in a brain scanning scenario, according to a preset structured template, multiple structured tags are extracted from an image report associated with an input image as preset tags, a structured text corresponding to each preset tag is generated, and the image report and the structured text are encoded respectively to obtain a feature vector corresponding to the image report and a feature vector of the structured text, and the feature vector corresponding to the image report and the feature vector of each structured text are added to obtain the encoding T of each preset tag. L , the encoding of the preset label is usually in vector form. Obtain the first feature vector I0 corresponding to the input image, and the second feature vector I corresponding to each candidate image in the candidate image library k , calculate the first similarity I0T between the first feature vector and the encoding of the preset label L , calculate the second similarity between each second feature vector and the code of the preset tag, calculate the third similarity between the first feature vector and each second feature vector based on the first similarity associated with the code of the same preset tag and each second similarity, and then perform weighted averaging on each third similarity based on the preset weight corresponding to the code of each preset tag to obtain a fourth similarity S between the first feature vector and each second feature vector ik , select the image data with the fourth larger similarity as the target data.

[0106] Through this embodiment, each preset tag is traversed, and the third similarity between the first feature vector and each second feature vector is determined based on the first similarity and each second similarity associated with the encoding of the preset tag. Based on the preset weight corresponding to the encoding of each preset tag, each third similarity is weighted averaged to obtain the fourth similarity between the first feature vector and each second feature vector. Then, based on each fourth similarity, the candidate data corresponding to the second feature vector matching the first feature vector is determined as the target data, so that accurate matching between images is performed with the encoding of multiple preset tags as the medium, and fast and accurate image search is achieved.

[0107] In some of the embodiments, the following steps are also included:

[0108] Obtain the second eigenvector corresponding to each candidate data;

[0109] storing each second eigenvector in a vector database;

[0110] Matching the first feature vector with the second feature vector corresponding to each candidate data refers to matching the first feature vector with each second feature vector in the vector database.

[0111] Specifically, each candidate data is encoded to obtain a second feature vector corresponding to the candidate data, and each second feature vector is stored in a vector database, so that the candidate data is pre-encoded and pre-stored to avoid real-time encoding processing during data search.

[0112] Furthermore, during data search, the first feature vector is matched with each second feature vector in the vector database, and candidate data corresponding to the second feature vector matching the first feature vector is determined as the target data.

[0113] Through this embodiment, the second feature vector corresponding to each candidate data is obtained, and each second feature vector is stored in a vector database. During the data search process, the first feature vector is matched with each second feature vector in the vector database, so as to avoid a large amount of real-time calculation during data search and effectively improve the data search efficiency.

[0114] The present embodiment is described and illustrated by means of specific examples below.

[0115] Figure 6 is a flow chart of the data search method of this embodiment, such as Figure 6 As shown, when searching for images with text, the data search method specifically includes the following steps:

[0116] When searching for images with text, multiple structured tags are extracted from the input text data according to a preset structured template, and structured text corresponding to the text data is generated based on each structured tag. The text data and the structured text are encoded respectively to obtain a feature vector corresponding to the text data and a feature vector of the structured text, and the feature vector corresponding to the text data and the feature vector of the structured text are added to obtain a first feature vector T1.

[0117] Further, obtain the second feature vector I corresponding to each candidate image in the candidate image library k , match the first feature vector with each second feature vector, and calculate the similarity I between the first feature vector and each second feature vector k T1, from each candidate image, select at least one image data with a large similarity as the target data matching the input text data.

[0118] Figure 7 is a flow chart of the data search method of this embodiment, such as Figure 7 As shown, when searching for text with an image, the data search method specifically includes the following steps:

[0119] When searching for text with an image, the input image data is encoded to obtain a first feature vector I1 corresponding to the input image data. For each candidate text data, multiple structured tags are extracted from the text data according to a preset structured template, and a structured text corresponding to the text data is generated based on each structured tag. The text data and the structured text are encoded to obtain a feature vector corresponding to the text data and a feature vector of the structured text, and the feature vector corresponding to the text data and the feature vector of the structured text are added to obtain a second feature vector T N .

[0120] Furthermore, the first feature vector is matched with each second feature vector, and the similarity I1T between the first feature vector and each second feature vector is calculated. N , from each candidate text data, at least one text data with a large similarity is selected as the target data that matches the input image data.

[0121] The present embodiment is described and illustrated below through preferred embodiments.

[0122] Figure 8 is a flow chart of the data search method of this preferred embodiment, such as Figure 8 As shown, the data search method includes the following steps:

[0123] Step S810, when the input data is text data, extracting a plurality of structured tags from the text data based on a preset structured template;

[0124] Step S820, generating a structured text corresponding to the text data based on each structured tag, and superimposing the encoding result of the text data and the encoding result of the structured text to obtain a first feature vector;

[0125] Step S830, encoding each candidate data to obtain a second feature vector corresponding to each candidate data;

[0126] Step S840, matching the first feature vector with the second feature vector corresponding to each candidate data to calculate the similarity between the first feature vector and each second feature vector;

[0127] Step S850: According to the similarity calculation result, at least one candidate data with a greater similarity is selected as the target data.

[0128] Through this embodiment, when the input data is text data, multiple structured tags are extracted from the text data based on a preset structured template, and based on each structured tag, a structured text corresponding to the text data is generated, and the encoding result of the text data and the encoding result of the structured text are superimposed to obtain a first feature vector.

[0129] Furthermore, each candidate data is encoded to obtain a second eigenvector corresponding to each candidate data, and the first eigenvector is matched with the second eigenvector corresponding to each candidate data to calculate the similarity between the first eigenvector and each second eigenvector, and then based on the similarity calculation result, at least one candidate data with a larger similarity is selected as the target data. This solves the problem that the encoding information of the current multimodal algorithm cannot accurately reflect the association between multimodal data, resulting in inaccurate data search results, and achieves that the encoding information can accurately reflect the association between multimodal data, thereby improving the accuracy of data search.

[0130] It should be noted that the steps shown in the above process or the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.

[0131] In this embodiment, a data search device is also provided, which is used to implement the above-mentioned embodiments and preferred implementation modes, and the descriptions that have been made will not be repeated. The terms "module", "unit", "subunit" and the like used below can implement a combination of software and / or hardware of a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, the implementation of hardware, or a combination of software and hardware is also possible and conceivable.

[0132] Fig. 9 is a structural block diagram of the data search device of this embodiment, such as Fig. 9 As shown, the device comprises:

[0133] The encoding module 10 is used to determine a first feature vector corresponding to the input data;

[0134] A search module 20, for matching the first feature vector with a second feature vector corresponding to each candidate data, and determining the candidate data corresponding to the second feature vector matching the first feature vector as the target data;

[0135] Wherein, if the input data is text, the first feature vector is associated with a feature vector corresponding to the text data and a feature vector of a structured text corresponding to the text data;

[0136] If the candidate data is text, the second feature vector is associated with the feature vector corresponding to the text data and the feature vector of the structured text corresponding to the text data.

[0137] Through the device provided in this embodiment, the first feature vector corresponding to the input data is determined; the first feature vector is matched with the second feature vector corresponding to each candidate data, and the candidate data corresponding to the second feature vector matching the first feature vector is determined as the target data; wherein, if the input data is text, the first feature vector is associated with the feature vector corresponding to the text data and the feature vector of the structured text corresponding to the text data; if the candidate data is text, the second feature vector is associated with the feature vector corresponding to the text data and the feature vector of the structured text corresponding to the text data, which solves the problem that the encoding information of the current multimodal algorithm cannot accurately reflect the association between multimodal data, resulting in inaccurate data search results, and realizes that the encoding information can accurately reflect the association between multimodal data, thereby improving the accuracy of data search.

[0138] In some of the embodiments, the encoding module 10 is also used to extract multiple structured tags from the text data based on a preset structured template; generate structured text corresponding to the text data based on each structured tag; and superimpose the encoding result of the text data and the encoding result of the structured text to obtain a first feature vector.

[0139] In some of the embodiments, the encoding module 10 is also used to extract multiple structured tags from the text data based on a preset structured template; generate structured text corresponding to the text data based on each structured tag; and superimpose the encoding result of the text data and the encoding result of the structured text to obtain a second feature vector.

[0140] In some of the embodiments, the encoding module 10 is further used to generate a structured text corresponding to at least one structured tag based on the at least one structured tag; and the feature vector corresponding to the structured text is used as the first feature vector.

[0141] In some of the embodiments, the search module 20 is also used to determine a first similarity between the first feature vector and the encoding of a preset label; determine a second similarity between each second feature vector and the encoding of the preset label; and based on the first similarity and each second similarity, determine that the candidate data corresponding to the second feature vector matching the first feature vector is the target data.

[0142] In some of the embodiments, the search module 20 is further used to traverse each preset tag, and determine the third similarity between the first feature vector and each second feature vector based on the first similarity and each second similarity associated with the encoding of the preset tag; based on the preset weight corresponding to the encoding of each preset tag, perform weighted averaging on each third similarity to obtain the fourth similarity between the first feature vector and each second feature vector; based on each fourth similarity, determine that the candidate data corresponding to the second feature vector matching the first feature vector is the target data.

[0143] In some of these embodiments, Fig. 9 On the basis of, the device also includes a pre-storage module for obtaining the second feature vector corresponding to each candidate data; storing each second feature vector in a vector database; matching the first feature vector with the second feature vector corresponding to each candidate data means matching the first feature vector with each second feature vector in the vector database.

[0144] It should be noted that the above modules can be functional modules or program modules, and can be implemented by software or hardware. For modules implemented by hardware, the above modules can be located in the same processor; or the above modules can be located in different processors in any combination.

[0145] In this embodiment, a computer device is further provided, including a memory and a processor, wherein a computer program is stored in the memory, and the processor is configured to run the computer program to execute the steps in any one of the above method embodiments.

[0146] Optionally, the computer device may further include a transmission device and an input / output device, wherein the transmission device is connected to the processor, and the input / output device is connected to the processor.

[0147] Optionally, in this embodiment, the processor may be configured to perform the following steps through a computer program:

[0148] S1, determine the first eigenvector corresponding to the input data;

[0149] S2, matching the first feature vector with the second feature vector corresponding to each candidate data, and determining that the candidate data corresponding to the second feature vector matching the first feature vector is the target data; wherein, if the input data is text, the first feature vector is associated with the feature vector corresponding to the text data and the feature vector of the structured text corresponding to the text data; if the candidate data is text, the second feature vector is associated with the feature vector corresponding to the text data and the feature vector of the structured text corresponding to the text data.

[0150] It should be noted that the specific examples in this embodiment can refer to the examples described in the above embodiments and optional implementation modes, and will not be repeated in this embodiment.

[0151] In addition, in combination with the data search method provided in the above embodiments, a storage medium may be provided in this embodiment to implement the data search method. The storage medium stores a computer program; when the computer program is executed by a processor, any one of the data search methods in the above embodiments is implemented.

[0152] It should be understood that the specific embodiments described herein are only used to explain the application, rather than to limit it. Based on the embodiments provided in this application, all other embodiments obtained by ordinary technicians in this field without creative work are within the protection scope of this application.

[0153] Obviously, the drawings are only some examples or embodiments of the present application. For ordinary technicians in the field, the present application can also be applied to other similar situations based on these drawings without creative work. In addition, it is understandable that although the work done in this development process may be complicated and lengthy, for ordinary technicians in the field, certain changes in design, manufacturing or production based on the technical content disclosed in this application are only conventional technical means and should not be regarded as insufficient content disclosed in this application.

[0154] The term "embodiment" in this application refers to a specific feature, structure or characteristic described in conjunction with the embodiment that can be included in at least one embodiment of the present application. The appearance of this phrase in various locations in the specification does not necessarily mean the same embodiment, nor does it mean that it is mutually exclusive with other embodiments and is independent or optional. It is clearly or implicitly understood by those of ordinary skill in the art that the embodiments described in this application can be combined with other embodiments without conflict.

[0155] The above-mentioned embodiments only express several implementation methods of the present application, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the scope of patent protection. It should be pointed out that, for a person of ordinary skill in the art, several variations and improvements can be made without departing from the concept of the present application, and these all belong to the scope of protection of the present application. Therefore, the scope of protection of the present application shall be subject to the attached claims.

Claims

1. A data search method, characterized in that: The method comprises: determining a first eigenvector corresponding to the input data; Matching the first feature vector with a second feature vector corresponding to each candidate data, and determining the candidate data corresponding to the second feature vector that matches the first feature vector as the target data; Wherein, if the input data is text, the first feature vector is associated with a feature vector corresponding to the text data and a feature vector of a structured text corresponding to the text data; If the candidate data is text, the second feature vector is associated with a feature vector corresponding to the text data and a feature vector of a structured text corresponding to the text data.

2. The data search method according to claim 1, characterized in that: The input data is text, and determining the first eigenvector includes: Extracting multiple structured tags from the text data based on a preset structured template; Based on each of the structured tags, generating the structured text corresponding to the text data; The encoding result of the text data and the encoding result of the structured text are superimposed to obtain the first feature vector.

3. The data search method according to claim 1, characterized in that: The candidate data is text, and determining the second feature vector includes: Extracting multiple structured tags from the text data based on a preset structured template; Based on each of the structured tags, generating the structured text corresponding to the text data; The encoding result of the text data and the encoding result of the structured text are superimposed to obtain the second feature vector.

4. The data search method according to claim 1, characterized in that: The input data is at least one structured label, and determining a first feature vector corresponding to the input data includes: Based on the at least one structured tag, generating a structured text corresponding to the at least one structured tag; The feature vector corresponding to the structured text is taken as the first feature vector.

5. The data search method according to claim 1, characterized in that: The input data and the candidate data are both image data, and the method further includes: Determining a first similarity between the first feature vector and an encoding of a preset label; Determining a second similarity between each of the second feature vectors and the encoding of the preset label; Based on the first similarity and each of the second similarities, candidate data corresponding to the second feature vector matching the first feature vector is determined as target data.

6. The data search method according to claim 5, characterized in that: There are multiple preset tags, and the method further includes: Traversing each of the preset tags, and determining a third similarity between the first feature vector and each of the second feature vectors based on the first similarity associated with the encoding of the preset tag and each of the second similarities; Based on a preset weight corresponding to the code of each of the preset tags, weighted averaging the third similarities to obtain a fourth similarity between the first feature vector and each of the second feature vectors; Based on each of the fourth similarities, candidate data corresponding to the second feature vector that matches the first feature vector is determined as the target data.

7. The data search method according to claim 1, characterized in that: Also includes: Obtaining a second feature vector corresponding to each of the candidate data; storing each second eigenvector in a vector database; The matching of the first feature vector with the second feature vector corresponding to each candidate data refers to matching the first feature vector with each second feature vector in the vector database.

8. A data search device, characterized in that: The device comprises: An encoding module, configured to determine a first eigenvector corresponding to the input data; A search module, configured to match the first feature vector with a second feature vector corresponding to each candidate data, and determine that the candidate data corresponding to the second feature vector matching the first feature vector is the target data; Wherein, if the input data is text, the first feature vector is associated with a feature vector corresponding to the text data and a feature vector of a structured text corresponding to the text data; If the candidate data is text, the second feature vector is associated with a feature vector corresponding to the text data and a feature vector of a structured text corresponding to the text data.

9. A computer device comprising a memory and a processor, characterized in that: A computer program is stored in the memory, and the processor is configured to run the computer program to perform the steps of the data search method according to any one of claims 1 to 7.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the data search method according to any one of claims 1 to 7 are implemented.