Information retrieval method and system
By separating the database data into multiple types and extracting features, and sorting similarity with user search information, the problem that the single-modal information retrieval model cannot meet the diverse information needs is solved, and the comprehensiveness and accuracy of multi-modal information retrieval is achieved.
Patent Information
- Application Number
- CN202510578639.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-07
- Publication Date
- 2025-06-03
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The single-modal information retrieval mode in the prior art cannot meet users' needs for diversified information, and the search results are not comprehensive enough.
By separating the data in the database into text data, picture data and video data, and extracting global and local features respectively, sorting similarity with user search information, multimodal search results are generated.
It breaks through the limitations of traditional single-modal retrieval, provides multi-modal target data, meets users' needs for diversified information, and improves the comprehensiveness and accuracy of the search structure.
Smart Images

Figure CN120086395A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data processing, and particularly relates to an information retrieval method and system. Background Art
[0002] With the advent of the digital age, data information has changed from the traditional single-modal form to the multi-modal form. Multi-modal refers to the way of expressing, communicating and understanding by using information of multiple different forms or perception channels, such as text form, picture form and video form.
[0003] Information retrieval is an important way for users to obtain information. Specifically, it refers to the process of accurately finding the target information corresponding to the text information provided by the user in a huge and complex database.
[0004] The existing information retrieval mode is a single-modal mode, that is, according to the text information provided by the user, only the text target information of the same modality can be indexed, and the picture target information or video target information cannot be obtained. This retrieval method is difficult to meet the user's needs for diversified information, and the retrieval results are not comprehensive enough. Summary of the Invention
[0005] Aiming at the deficiencies of the prior art, the purpose of the present invention is to provide an information retrieval method and system, aiming to solve the technical problem that the single-modal information retrieval mode in the prior art cannot meet the user's needs for diversified information and the retrieval results are not comprehensive enough.
[0006] To achieve the above purpose, in a first aspect, an embodiment of the present application provides an information retrieval method, including the following steps: Partition a number of data in the database into a number of text data, a number of picture data and a number of video data, perform word segmentation processing on the text data to obtain a number of word data, and obtain the global text feature and the local text feature set of the text data based on the number of word data; Adjust the picture data to a fixed size, and obtain the global picture feature of the picture data through a first feature extraction network. Partition the picture data into a number of image blocks to obtain a local picture feature set. Cut the video data into a number of frame pictures, and obtain the global video feature and the local video feature set corresponding to the video data based on the number of frame pictures; Obtain the retrieval information of the user, convert the retrieval information into a global retrieval feature and a local retrieval feature set, and perform similarity sorting on a number of the text data through the global retrieval feature, the local retrieval feature set, the global text feature and the local text feature set to generate a text retrieval result; Performing similarity sorting on several pieces of the picture data by means of the global retrieval feature, the local retrieval feature set, the global picture feature, and the local picture feature set to generate a picture retrieval result, and performing similarity sorting on several pieces of the video data by means of the global retrieval feature, the local retrieval feature set, the global video feature, and the local video feature set to generate a video retrieval result.
[0007] Compared with the prior art, the beneficial effects of the present invention are as follows: By partitioning data into the text data, the picture data, and the video data, the retrieval limitation of the traditional single modality is broken through, multi-modal target data is provided, the needs of users for diversified information are met, and the comprehensiveness of the retrieval structure is improved; By obtaining the global text feature, the local text feature set, the global picture feature, the local picture feature set, the global video feature, and the local video feature set, and obtaining the similarity between the retrieval information and the former, an implementation approach for multi-modal data retrieval is provided; By distinguishing between the overall and local features, combining the overall meaning and detailed information, the representation ability of the content is enhanced, and the accuracy of the retrieval is improved.
[0008] Further, the step of obtaining the global text feature and the local text feature set of the text data based on several pieces of the word data includes: Converting the word data into word vectors, and performing sorting processing on several of the word vectors from the head end to the tail end direction based on the word order of the text data; From the head end to the tail end direction, assigning a time step to each of the word vectors to convert several of the word vectors into several forward vectors; From the tail end to the head end direction, assigning a time step to each of the word vectors to convert several of the word vectors into several backward vectors; Concatenating the forward vectors and the backward vectors at the same time step into text feature vectors, combining several of the text feature vectors into a local text feature set, and selecting the text feature vector at the last time step as the global text feature.
[0009] Furthermore, the acquisition formula for the forward vector is: , wherein, represents the forward vector at the t-th time step from the head end to the tail end direction, represents the update gate information at the t-th time step from the head end to the tail end direction, represents the forward vector at the (t - 1)-th time step from the head end to the tail end direction, represents the candidate forward vector at the t-th time step from the head end to the tail end direction, denotes element-wise multiplication, where, , where, denotes the update gate weight, denotes the word vector corresponding to the t-th time step in the direction from the head end to the tail end, denotes the sigmoid activation function; , where, denotes the candidate weight, denotes the reset gate information of the t-th time step in the direction from the head end to the tail end, denotes the hyperbolic tangent function, where, , where, denotes the reset gate weight.
[0010] Furthermore, the step of partitioning the picture data into a plurality of image blocks to obtain a local picture feature set includes: Generating a plurality of region bounding boxes in the picture data through a region localization network, and partitioning the picture data based on the region bounding boxes to obtain a plurality of image blocks; Obtaining local image features of the image blocks through a second feature extraction network, and combining the plurality of local image features into a local picture feature set.
[0011] Furthermore, the step of obtaining the global video feature and the local video feature set corresponding to the video data based on the plurality of frame pictures includes: Obtaining the picture features of each frame picture through a third feature extraction network, and performing an averaging process on all the picture features to obtain the global video feature; Partitioning all the picture features into a plurality of feature groups based on the time frames, obtaining the averaged features of each feature group, and combining the plurality of averaged features into a local video feature set.
[0012] Furthermore, the local retrieval feature set includes a plurality of retrieval sub-features, and the step of sorting the plurality of text data by similarity through the global retrieval feature, the local retrieval feature set, the global text feature, and the local text feature set includes: Obtaining the global similarity between the global retrieval feature and the global text feature; Obtaining the attention weights between the retrieval sub-features and the text feature vectors, and updating the text feature vectors to aligned feature vectors based on the attention weights; Obtain the local sub - similarity between the retrieved sub - feature and the alignment feature vector, and perform an averaging process on several of the local sub - similarities to obtain the local similarity; Obtain the final similarity through the global similarity and the local similarity, and perform a similarity ranking on several of the text data based on the final similarity.
[0013] Furthermore, the formula for obtaining the global similarity is: , where, represents the global similarity between the global retrieval feature and the global text feature corresponding to the m - th text data, represents the global retrieval feature, represents the global text feature corresponding to the m - th text data, represents the dot product, represents the norm of the vector; The formula for obtaining the attention weight is: , where, represents the attention weight between the i - th retrieved sub - feature and the j - th text feature vector in the local text feature set corresponding to the m - th text data, represents the i - th retrieved sub - feature, represents the j - th text feature vector in the local text feature set corresponding to the m - th text data, represents the Query projection matrix, represents the Key projection matrix, represents the scaling factor, represents the transpose symbol; The formula for obtaining the alignment feature vector is: , where, represents the alignment feature vector corresponding to the i - th retrieved sub - feature, represents the number of text feature vectors in the local text feature set corresponding to the m - th text data.
[0014] In a second aspect, an embodiment of the present application provides an information retrieval system, which is applied to the information retrieval method described in the first aspect above. The system includes: The formula for obtaining the global similarity is: , where, represents the global similarity between the global retrieval feature and the global text feature corresponding to the m - th text data, Represents the global retrieval feature, Represents the global text feature corresponding to the m-th text data, Represents the dot product (inner product of vectors), Represents the magnitude of a vector; The formula for obtaining the attention weight is: , where, Represents the attention weight between the i-th retrieval sub-feature and the j-th text feature vector in the local text feature set corresponding to the m-th text data, Represents the i-th retrieval sub-feature, Represents the j-th text feature vector in the local text feature set corresponding to the m-th text data, Represents the Query projection matrix, Represents the Key projection matrix, Represents the scaling factor, Represents the transpose symbol; The formula for obtaining the alignment feature vector is: , where, Represents the alignment feature vector corresponding to the i-th retrieval sub-feature, Represents the number of text feature vectors in the local text feature set corresponding to the m-th text data.
[0015] In a third aspect, an embodiment of the present application provides a computer, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the information retrieval method described in the first aspect above is implemented.
[0016] In a fourth aspect, an embodiment of the present application provides a storage medium, on which a computer program is stored. When the computer program is executed by a processor, the information retrieval method described in the first aspect above is implemented. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] Figure 1 Is the flowchart of the information retrieval method in the first embodiment of the present invention; Figure 2 Is the structural block diagram of the information retrieval system in the second embodiment of the present invention; The following specific embodiments will further illustrate the present invention in conjunction with the above-mentioned drawings. SPECIFIC EMBODIMENTS
[0018] To facilitate the understanding of the present invention, the present invention will be described more comprehensively below with reference to the relevant accompanying drawings. Several embodiments of the present invention are shown in the drawings. However, the present invention can be implemented in many different forms and is not limited to the embodiments described herein. On the contrary, these embodiments are provided to make the disclosure of the present invention more thorough and comprehensive.
[0019] It should be noted that when an element is referred to as being "fixed on" another element, it can be directly on the other element or there may also be an intermediate element. When an element is considered to be "connected" to another element, it can be directly connected to the other element or there may be an intermediate element at the same time. The terms "vertical", "horizontal", "left", "right" and similar expressions used herein are for illustrative purposes only.
[0020] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present invention belongs. The terms used in the description of the present invention herein are only for the purpose of describing specific embodiments and are not intended to limit the present invention. The term "and / or" used herein includes any and all combinations of one or more of the related listed items.
[0021] Please refer to Figure 1 , the information retrieval method provided by the first embodiment of the present invention includes the following steps: S10: Partition a number of data in the database into a number of text data, a number of picture data, and a number of video data, perform word segmentation processing on the text data to obtain a number of word data, and obtain the global text feature and local text feature set of the text data based on the number of word data; The step S10 includes: S110: Convert the word data into word vectors, and sort a number of the word vectors from the head end to the tail end direction based on the word order of the text data; Assume that a certain text data is a red balloon under the blue sky. After word segmentation and sorting, the word data is blue sky, red, and balloon. After conversion into vectors, they are vector 1, vector 2, and vector 3.
[0022] S120: Assign a time step to each word vector from the head end to the tail end direction to convert a number of the word vectors into a number of forward vectors; After assigning the time step, vector 1 corresponds to the first time step, vector 2 corresponds to the second time step, and vector 3 corresponds to the third time step.
[0023] The acquisition formula of the forward vector is: , Among them, represents the forward vector at the t-th time step from the head end to the tail end, represents the update gate information at the t-th time step from the head end to the tail end, represents the forward vector at the (t - 1)-th time step from the head end to the tail end, represents the candidate forward vector at the t-th time step from the head end to the tail end, represents element-wise multiplication, where , Among them, represents the update gate weight, represents the word vector corresponding to the t-th time step from the head end to the tail end, represents the sigmoid activation function; , Among them, represents the candidate weight, represents the reset gate information at the t-th time step from the head end to the tail end, represents the hyperbolic tangent function, where , Among them, represents the reset gate weight. It should be noted that when converting vector 1 into a forward vector, there is no forward vector at the (t - 1)-th time step. At this time, takes a value of 0. By setting the time step, different vectors can be gradually superimposed. The forward vector formed by converting the vector 3 corresponding to the 3rd time step includes the meanings of all vectors.
[0024] S130: From the tail end to the head end direction, assign a time step to each of the word vectors to convert a plurality of the word vectors into a plurality of reverse vectors; The obtaining method of the reverse vector is the same as that of the forward vector, and no further description will be given here. It can be understood that still taking blue sky, red, balloon, which respectively correspond to vector 1, vector 2, and vector 3. At this time, vector 1 corresponds to the third time step, vector 2 corresponds to the second time step, and vector 3 corresponds to the first time step.
[0025] S140: Concatenate the forward vector and the reverse vector at the same time step into a text feature vector, combine a plurality of the text feature vectors into a local text feature set, and select the text feature vector at the last time step as the global text feature.
[0026] By means of two-way processing, the semantic dependencies before and after the words are captured, and then through vector splicing, a more comprehensive text feature vector of the text is formed, avoiding the occurrence of local semantic loss and ensuring a comprehensive understanding of the meaning of the text data.
[0027] S20: Adjust the picture data to a fixed size, and obtain the global picture features of the picture data through a first feature extraction network. Divide the picture data into several image blocks to obtain a local picture feature set. Cut the video data into several frame pictures, and obtain the global video features and local video feature sets corresponding to the video data based on the several frame pictures; In this embodiment, the first feature extraction network is a convolutional neural network (CNN) ResNet-50, and the global picture features are output through its fully connected layer.
[0028] The step S20 includes: S210: Generate several region bounding boxes in the picture data through a region localization network, and divide the picture data based on the region bounding boxes to obtain several image blocks; In this embodiment, the region localization network is a target detection model FasterR-CNN. Assume that one of the picture data is a balloon floating in the sky, then the region bounding boxes are the sky bounding box and the balloon bounding box.
[0029] S220: Obtain the local image features of the image blocks through a second feature extraction network, and combine the several local image features into a local picture feature set; In this embodiment, the second feature extraction network is also a convolutional neural network (CNN) ResNet-50 to ensure the consistency between features.
[0030] S230: Obtain the picture features of each frame picture through a third feature extraction network, and perform averaging processing on all the picture features to obtain the global video features; In this embodiment, the third feature extraction network is still a convolutional neural network (CNN) ResNet-50.
[0031] S240: Divide all the picture features into several feature groups based on the time frames, obtain the averaged features of each feature group, and combine the several averaged features into a local video feature set; By obtaining the local picture feature set and the local video feature set, the omission of details caused by only obtaining the global picture features and the global video features is avoided, ensuring the accurate representation of the picture data and the video data.
[0032] S30: Obtain the retrieval information of the user, convert the retrieval information into a global retrieval feature and a set of local retrieval features, and perform similarity ranking on a number of the text data through the global retrieval feature, the set of local retrieval features, the global text feature, and the set of local text features to generate a text retrieval result; The set of local retrieval features includes a number of retrieval sub-features. The retrieval information is a piece of retrieval text, which has the same essential meaning as the text data, and the retrieval sub-feature has the same essential meaning as the text feature vector. Therefore, the way of obtaining the global retrieval feature is the same as that of the global text feature, and the way of obtaining the set of local retrieval features is the same as that of the set of local text features, which will not be elaborated here.
[0033] The step S30 includes: S310: Obtain the global similarity between the global retrieval feature and the global text feature; The formula for obtaining the global similarity is: , where, represents the global similarity between the global retrieval feature and the global text feature corresponding to the m-th text data, represents the global retrieval feature, represents the global text feature corresponding to the m-th text data, represents dot product (inner product of vectors), represents the norm of the vector.
[0034] S320: Obtain the attention weight between the retrieval sub-feature and the text feature vector, and update the text feature vector to an aligned feature vector based on the attention weight; The formula for obtaining the attention weight is: , where, represents the attention weight between the i-th retrieval sub-feature and the j-th text feature vector in the set of local text features corresponding to the m-th text data, represents the i-th retrieval sub-feature, represents the j-th text feature vector in the set of local text features corresponding to the m-th text data, represents the Query projection matrix, represents the Key projection matrix, represents the scaling factor, represents the transpose symbol; The formula for obtaining the aligned feature vector is: , where, represents the alignment feature vector corresponding to the i-th retrieval sub-feature, represents the number of text feature vectors in the local text feature set corresponding to the m-th text data.
[0035] S330: Obtain the local sub-similarity between the retrieval sub-feature and the alignment feature vector, and perform averaging processing on a number of the local sub-similarities to obtain the local similarity; The obtaining method of the local sub-similarity is the same as that of the global similarity, and will not be elaborated here.
[0036] S340: Obtain the final similarity through the global similarity and the local similarity, and perform similarity ranking on a number of the text data based on the final similarity; It can be understood that by respectively assigning corresponding fusion weights to the global similarity and the local similarity to fuse them into the final similarity, after obtaining the final similarity between each text data and the retrieval information, the text data can be sorted from large to small according to the value of the final similarity, and the text retrieval result is formed.
[0037] S40: Perform similarity ranking on a number of the picture data through the global retrieval feature, the local retrieval feature set, the global picture feature and the local picture feature set to generate a picture retrieval result, and perform similarity ranking on a number of the video data through the global retrieval feature, the local retrieval feature set, the global video feature and the local video feature set to generate a video retrieval result; The similarity ranking method between the retrieval information and the picture data and the similarity ranking method between the retrieval information and the video data are the same as the similarity ranking method between the retrieval information and the text data, which are all to obtain the similarity between vectors, and will not be elaborated here. After obtaining the text retrieval result, the picture retrieval result and the video retrieval result, the information retrieval can be completed.
[0038] By partitioning the data into the text data, the picture data and the video data, the traditional single-modal retrieval limitation is broken through, multi-modal target data is provided, the user's demand for diversified information is met, and the comprehensiveness of the retrieval structure is improved; by obtaining the global text feature, the local text feature set, the global picture feature, the local picture feature set, the global video feature and the local video feature set, and obtaining the similarity between the retrieval information and the former, an implementation approach for multi-modal data retrieval is provided; by distinguishing the overall and local features, combining the overall meaning and detailed information, the representation ability of the content is enhanced, and the accuracy of the retrieval is improved.
[0039] Please refer to Figure 2 For the second embodiment of the present invention, an information retrieval system is provided. This system is applied to the information retrieval method in the above embodiment, and the parts that have been described will not be elaborated again. As used hereinafter, terms such as "module", "unit", "sub-unit", etc. can be a combination of software and / or hardware that can achieve a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, implementation in hardware, or a combination of software and hardware is also possible and contemplated.
[0040] The system includes: A classification module 10, configured to partition a number of data in a database into a number of text data, a number of picture data, and a number of video data, perform word segmentation processing on the text data to obtain a number of word data, and obtain the global text feature and the local text feature set of the text data based on the number of word data; The classification module 10 includes: A first unit, configured to convert the word data into word vectors, and perform sorting processing on the number of word vectors from the head end to the tail end direction based on the word order of the text data; A second unit, configured to assign a time step to each word vector from the head end to the tail end direction to convert the number of word vectors into a number of forward vectors; A third unit, configured to assign a time step to each word vector from the tail end to the head end direction to convert the number of word vectors into a number of reverse vectors; A fourth unit, configured to splice the forward vectors and the reverse vectors at the same time step into a text feature vector, combine the number of text feature vectors into a local text feature set, and select the text feature vector at the last time step as the global text feature; An extraction module 20, configured to adjust the picture data to a fixed size, obtain the global picture feature of the picture data through a first feature extraction network, partition the picture data into a number of image blocks to obtain a local picture feature set, cut the video data into a number of frame pictures, and obtain the global video feature and the local video feature set corresponding to the video data based on the number of frame pictures; The extraction module 20 includes: A fifth unit, configured to generate a number of regional bounding boxes in the picture data through a regional localization network, and partition the picture data based on the regional bounding boxes to obtain a number of image blocks; A sixth unit, configured to obtain the local image feature of the image block through a second feature extraction network, and combine the number of local image features into a local picture feature set; The seventh unit is configured to obtain the picture features of each of the frame pictures through a third feature extraction network, and perform an averaging process on all the picture features to obtain global video features; The eighth unit is configured to partition all the picture features into several feature groups based on time frames, obtain the averaged features of each feature group, and combine the several averaged features into a local video feature set; The first analysis module 30 is configured to obtain the retrieval information of a user, convert the retrieval information into a global retrieval feature and a local retrieval feature set, and perform a similarity ranking on several pieces of the text data through the global retrieval feature, the local retrieval feature set, the global text feature, and the local text feature set to generate a text retrieval result; The first analysis module 30 includes: The ninth unit is configured to obtain the global similarity between the global retrieval feature and the global text feature; The tenth unit is configured to obtain the attention weight between the retrieval sub-feature and the text feature vector, and update the text feature vector to an aligned feature vector based on the attention weight; The eleventh unit is configured to obtain the local sub-similarity between the retrieval sub-feature and the aligned feature vector, and perform an averaging process on several local sub-similarities to obtain a local similarity; The twelfth unit is configured to obtain a final similarity through the global similarity and the local similarity, and perform a similarity ranking on several pieces of the text data based on the final similarity; The second analysis module 40 is configured to perform a similarity ranking on several pieces of the picture data through the global retrieval feature, the local retrieval feature set, the global picture feature, and the local picture feature set to generate a picture retrieval result, and perform a similarity ranking on several pieces of the video data through the global retrieval feature, the local retrieval feature set, the global video feature, and the local video feature set to generate a video retrieval result.
[0041] The present invention further provides a computer, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the information retrieval method described in the above technical solution is implemented.
[0042] The present invention further provides a storage medium, on which a computer program is stored. When the computer program is executed by a processor, the information retrieval method described in the above technical solution is implemented.
[0043] In the description of this specification, the descriptions referring to terms such as "one embodiment", "some embodiments", "examples", "specific examples", or "some examples", etc., mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples.
[0044] The above-described embodiments merely represent several implementation manners of the present invention. The descriptions thereof are relatively specific and detailed, but should not be construed as a limitation on the scope of the patent for the present invention. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several modifications and improvements can still be made, and these all belong to the protection scope of the present invention. Therefore, the protection scope of the patent for the present invention shall be subject to the appended claims.
Claims
1. An information retrieval method, characterized in that: The following steps are involved: Segmenting a number of data in a database into a number of text data, a number of image data, and a number of video data, performing word segmentation processing on the text data to obtain a number of word data, and obtaining a global text feature and a local text feature set of the text data based on the number of word data; The image data is adjusted to a fixed size, and global image features of the image data are obtained through a first feature extraction network, the image data is segmented into a plurality of image blocks to obtain a local image feature set, the video data is segmented into a plurality of frame images, and global video features and local video feature sets corresponding to the video data are obtained based on the plurality of frame images; Acquire the user's search information, convert the search information into a global search feature and a local search feature set, and sort the similarity of a plurality of the text data by the global search feature, the local search feature set, the global text feature, and the local text feature set to generate a text search result; The image data are sorted by similarity using the global retrieval feature, the local retrieval feature set, the global image feature and the local image feature set to generate an image retrieval result; the video data are sorted by similarity using the global retrieval feature, the local retrieval feature set, the global video feature and the local video feature set to generate a video retrieval result.
2. The information retrieval method according to claim 1, characterized in that: The step of obtaining the global text feature and the local text feature set of the text data based on the plurality of word data comprises: Convert the word data into word vectors, and sort the word vectors from the beginning to the end based on the word order of the text data; Assigning a time step to each of the word vectors from the head end to the tail end to convert a plurality of the word vectors into a plurality of forward vectors; Assigning a time step to each of the word vectors from the tail end to the head end to convert a plurality of the word vectors into a plurality of reverse vectors; The forward vector and the reverse vector at the same time step are concatenated into a text feature vector, a plurality of the text feature vectors are combined into a local text feature set, and the text feature vector at the last time step is selected as a global text feature.
3. The information retrieval method according to claim 2, characterized in that: The formula for obtaining the forward vector is: , in, represents the forward vector from the head to the tail at the tth time step, Represents the update gate information of the tth time step from the head end to the tail end, represents the forward vector from the head to the tail at the t-1th time step, represents the candidate forward vector at the tth time step from the head to the tail, represents element-wise multiplication, where , in, represents the update gate weight, represents the word vector corresponding to the t-th time step from the beginning to the end, Represents the sigmoid activation function; , in, represents the candidate weight, Represents the reset gate information of the tth time step from the head end to the tail end, represents the hyperbolic tangent function, where , in, Reset gate weights.
4. The information retrieval method according to claim 1, characterized in that: The step of dividing the image data into a plurality of image blocks to obtain a local image feature set comprises: Generate a plurality of region boundary boxes in the image data through a region positioning network, and segment the image data based on the region boundary boxes to obtain a plurality of image blocks; The local image features of the image block are obtained through a second feature extraction network, and a plurality of the local image features are combined into a local image feature set.
5. The information retrieval method according to claim 1, characterized in that: The step of acquiring the global video features and the local video feature set corresponding to the video data based on the plurality of frame images comprises: Obtaining the image features of each frame of the image through a third feature extraction network, and performing averaging processing on all the image features to obtain global video features; Based on the time frame, all the image features are divided into a number of feature groups, the averaged features of each feature group are obtained, and a number of the averaged features are combined into a local video feature set.
6. The information retrieval method according to claim 2, characterized in that: The local search feature set includes a plurality of search sub-features, and the step of sorting the plurality of text data by similarity using the global search feature, the local search feature set, the global text feature, and the local text feature set includes: Obtaining a global similarity between the global search feature and the global text feature; Acquire an attention weight between the retrieval sub-feature and the text feature vector, and update the text feature vector to an alignment feature vector based on the attention weight; Obtaining a local sub-similarity between the retrieval sub-feature and the alignment feature vector, and performing averaging processing on a plurality of the local sub-similarity to obtain a local similarity; A final similarity is obtained through the global similarity and the local similarity, and a plurality of the text data are sorted by similarity based on the final similarity.
7. The information retrieval method according to claim 6, characterized in that: The formula for obtaining the global similarity is: , in, represents the global similarity between the global retrieval feature and the global text feature corresponding to the mth text data, represents the global retrieval feature, represents the global text feature corresponding to the mth text data, represents dot product, Represents the magnitude of a vector; The formula for obtaining the attention weight is: , in, represents the attention weight between the i-th retrieval sub-feature and the j-th text feature vector in the local text feature set corresponding to the m-th text data, represents the i-th retrieval sub-feature, represents the jth text feature vector in the local text feature set corresponding to the mth text data, represents the Query projection matrix, represents the Key projection matrix, represents the scaling factor, Represents the transposition character; The formula for obtaining the alignment feature vector is: , in, represents the aligned feature vector corresponding to the i-th retrieval sub-feature, Represents the number of text feature vectors in the local text feature set corresponding to the mth text data.
8. An information retrieval system, applied to the information retrieval method according to any one of claims 1 to 7, characterized in that: The system comprises: A classification module is used to separate a number of data in the database into a number of text data, a number of image data and a number of video data, perform word segmentation processing on the text data to obtain a number of word data, and obtain a global text feature and a local text feature set of the text data based on the number of word data; An extraction module, used for adjusting the image data to a fixed size, obtaining global image features of the image data through a first feature extraction network, segmenting the image data into a plurality of image blocks to obtain a local image feature set, dividing the video data into a plurality of frame images, and obtaining global video features and a local video feature set corresponding to the video data based on the plurality of frame images; A first analysis module is used to obtain the user's search information, convert the search information into a global search feature and a local search feature set, and perform similarity sorting on a plurality of the text data using the global search feature, the local search feature set, the global text feature, and the local text feature set to generate a text search result; The second analysis module is used to sort the image data by similarity using the global retrieval feature, the local retrieval feature set, the global image feature and the local image feature set to generate an image retrieval result, and to sort the video data by similarity using the global retrieval feature, the local retrieval feature set, the global video feature and the local video feature set to generate a video retrieval result.
9. A computer comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the information retrieval method according to any one of claims 1 to 7 is implemented.
10. A storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the information retrieval method as described in any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Strong-correlation unsupervised cross-modal retrieval method guided by information amount
CN116756363A
Text similarity calculation method and device, equipment and storage medium
CN116956867A
Image-text retrieval method based on information enhancement and multi-modal global local feature alignment
CN119646272A
Cross-modal retrieval system and method based on pre-training model and recall and ranking
WO2023065617A1