Cross-modal data retrieval method and related equipment

By using cross-modal data retrieval methods in multimodal data retrieval of power grid equipment, the hash code of semantic features is solved, and the problem of difficult to capture semantic correlation features in the prior art is achieved, achieving higher retrieval accuracy and comprehensive retrieval of power grid data information.

CN120030174APending Publication Date: 2025-05-23STATE GRID INFORMATION & TELECOMM GRP CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411916044.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-24
Publication Date
2025-05-23

AI Technical Summary

Technical Problem

The prior art is difficult to effectively capture the semantic correlation characteristics in multimodal data of power grid equipment, resulting in the inability to support the accurate counting needs of power grid users.

Method used

A cross-modal data retrieval method is proposed. By receiving the user's text and image retrieval data, the basic features are extracted and hash codes are generated, and the training semantic features are used as supervision information is used to calculate the Hamming distance between the search data and the database data, and the corresponding database data is displayed in a preset manner.

Benefits of technology

It improves the accuracy of cross-modal search results, enhances the semantic hierarchy correlation between grid equipment text and images, and helps grid users to retrieve more comprehensive data information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120030174A_ABST
    Figure CN120030174A_ABST
Patent Text Reader

Abstract

One or more embodiments of the invention provide a cross-modal data retrieval method and related equipment. The method comprises the following steps: receiving retrieval data of a user, wherein the types of the retrieval data comprise texts and images; inputting the retrieval data into a basic feature extraction model to obtain basic features of the retrieval data; inputting the basic features into a hash code generation model to obtain a hash code of the retrieval data; the hash code generation model is obtained by training semantic features corresponding to the basic features for training as supervision information; and calculating a Hamming distance between the Hash code of the retrieval data and the Hash code of database data, and displaying the database data according to the Hamming distance and a preset mode. Through the method provided by the invention, the accuracy of a cross-modal data retrieval result can be effectively improved, and the cross-modal retrieval effect is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] One or more embodiments of the present application relate to the field of computer technology, and in particular, to a cross-modal data retrieval method and related devices. Background Art

[0002] With the rapid development of computer and big data technology in the field of power grids, and the improvement of electronic storage standards for power grid equipment data, multimodal power grid equipment data has exploded. For example, multimodal data such as project equipment manuals, equipment photos, work tickets, on-site maintenance photos, etc. They often coexist and complement each other, forming a complex feature that is semantically similar and interrelated. For example, the appearance of equipment can be represented by equipment photos, while the specific parameters and operation steps of equipment need to be more intuitively expressed in text.

[0003] Multimodal retrieval solutions in related technologies are mostly implemented based on content tag similarity, which makes it difficult to capture the semantic association characteristics of domain words, and thus cannot support the precise number query needs of power grid users. Summary of the invention

[0004] In view of this, an object of one or more embodiments of the present application is to propose a cross-modal data retrieval method and related devices to solve the problems raised by the background technology.

[0005] Based on the above objectives, one or more embodiments of the present application provide a cross-modal data retrieval method, including:

[0006] Receiving search data from a user, wherein the types of the search data include text and image;

[0007] Inputting the search data into a basic feature extraction model to obtain basic features of the search data;

[0008] Inputting the basic features into a hash code generation model to obtain a hash code of the retrieval data; the hash code generation model is trained using semantic features corresponding to the training basic features as supervision information;

[0009] The Hamming distance between the hash code of the retrieved data and the hash code of the database data is calculated, and the database data is displayed in a preset manner according to the Hamming distance.

[0010] Optionally, the training step of the hash code generation model includes:

[0011] Obtain training data and obtain corresponding basic features for training;

[0012] According to the basic features for training and the neural network model, the hash code generation model is obtained by taking the category label as supervision information;

[0013] The category labels indicate semantic features between the training data.

[0014] Optionally, the step of acquiring the category label includes:

[0015] Inputting the training data into a word vector embedding model to generate a word vector, wherein the types of the training data include text and words in an image;

[0016] The word vectors are clustered into multiple label categories through a hierarchical clustering algorithm.

[0017] Optionally, use the Skip-gram model in the Word2vec tool as the word vector embedding model.

[0018] Optionally, the step of acquiring training data includes:

[0019] Acquire equipment data of the power system, where the types of the equipment data include text and image;

[0020] The natural language processing tool is used to extract the word segments in the device data.

[0021] Optionally, the search data is input into a basic feature extraction model to obtain basic features of the search data, including:

[0022] In response to receiving the text data, inputting the text data into a Sent2vec model to obtain basic text features;

[0023] In response to receiving the image data, the image data is input into the EfficientNetV2 model to obtain basic image features.

[0024] Optionally, the MBConv module and the Fused-MBConv module of the EfficientNetV2 model are implemented using a convolution kernel of size 3*3.

[0025] Based on the same inventive concept, one or more embodiments of the present application further provide a cross-modal data retrieval device, including:

[0026] A receiving module, configured to receive search data of a user, wherein the types of the search data include text and image;

[0027] A basic feature extraction module is configured to input the search data into a basic feature extraction model to obtain basic features of the search data;

[0028] A generation module is configured to input the basic features into a hash code generation model to obtain a hash code of the retrieval data; the hash code generation model is trained using the semantic features corresponding to the training basic features as supervision information;

[0029] The retrieval module is configured to calculate the Hamming distance between the hash code of the retrieval data and the hash code of the database data, and display the database data in a preset manner according to the Hamming distance.

[0030] Based on the same inventive concept, one or more embodiments of the present application also provide an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the cross-modal data retrieval method as described in any one of the above is implemented.

[0031] Based on the same inventive concept, one or more embodiments of the present application further provide a non-transitory computer-readable storage medium, wherein the non-transitory computer-readable storage medium stores computer instructions, and the computer instructions are used to enable the computer to execute any of the above-mentioned cross-modal data retrieval methods.

[0032] From the above, it can be seen that the cross-modal data retrieval method provided by one or more embodiments of the present application receives the user's retrieval data, and the types of the retrieval data include text and image; inputs the retrieval data into a basic feature extraction model to obtain the basic features of the retrieval data; inputs the basic features into a hash code generation model to obtain a hash code of the retrieval data; the hash code generation model is trained using the semantic features corresponding to the training basic features as supervision information; calculates the Hamming distance between the hash code of the retrieval data and the hash code of the database data, and displays the database data in a preset manner according to the Hamming distance.

[0033] This application is based on the basic features of the retrieval data and uses the semantic features of the data as an aid to effectively improve the accuracy of cross-modal retrieval results and improve the cross-modal retrieval effect.

[0034] The cross-modal data retrieval device, electronic device and computer-readable storage medium provided in the present application can all implement the steps of the above-mentioned cross-modal data retrieval method, and therefore also have the beneficial effects of the above-mentioned cross-modal data retrieval method. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] In order to more clearly illustrate one or more embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only one or more embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0036] Figure 1 A schematic diagram of a flow chart of a cross-modal data retrieval method according to one or more embodiments of the present application;

[0037] Figure 2 A schematic diagram of the structure of a cross-modal data retrieval device according to one or more embodiments of the present application;

[0038] Figure 3 A schematic diagram of the hardware structure of an electronic device according to one or more embodiments of the present application. DETAILED DESCRIPTION

[0039] In order to make the objectives, technical solutions and advantages of the present disclosure more clearly understood, the present disclosure is further described in detail below in combination with specific embodiments and with reference to the accompanying drawings.

[0040] It should be noted that, unless otherwise defined, the technical terms or scientific terms used in one or more embodiments of the present application should be understood by people with ordinary skills in the field to which the present disclosure belongs. The "first", "second" and similar words used in one or more embodiments of the present application do not indicate any order, quantity or importance, but are only used to distinguish different components. "Include" or "comprise" and similar words mean that the elements or objects appearing before the word cover the elements or objects listed after the word and their equivalents, without excluding other elements or objects. "Connect" or "connected" and similar words are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. "Up", "down", "left", "right" and the like are only used to indicate relative positional relationships. When the absolute position of the described object changes, the relative positional relationship may also change accordingly.

[0041] As described in the background technology section, related technologies often use content tag similarity as the basis for cross-modal retrieval. However, in the context of power networks, the multimodal data of equipment in the power grid data center is numerous and scattered, and data association and information fusion are insufficient. When faced with massive multimodal data of equipment in the power grid center, related technologies find it difficult to capture the semantic association features of domain words, and thus cannot support the precise number query needs of power grid users. Therefore, there is an urgent need to study the association enhancement technology of equipment text description and image multimodal knowledge, design a cross-modal retrieval method for power grid equipment images and texts that integrates semantic features, and mine the semantic hierarchical association between power grid equipment text and images, which can help power grid users retrieve more comprehensive data information.

[0042] refer to Figure 1 The cross-modal data retrieval method of one or more embodiments of the present application comprises the following steps:

[0043] Step S101: receiving search data from a user, where the types of the search data include text and image.

[0044] Step S102: input the above retrieval data into a basic feature extraction model to obtain the basic features of the above retrieval data.

[0045] Step S103: input the basic features into a hash code generation model to obtain a hash code for the retrieval data; the hash code generation model is trained using the semantic features corresponding to the training basic features as supervision information.

[0046] Step S104: Calculate the Hamming distance between the hash code of the search data and the hash code of the database data, and display the database data in a preset manner according to the Hamming distance.

[0047] The retrieved data in step S101 includes text and images, such as project equipment manuals, equipment photos, work tickets, on-site maintenance photos, etc.

[0048] Step S102 requires first selecting a basic feature extraction model according to the type of the retrieved data. Specifically, for text data, the Sent2vec model can be used to extract basic features; for image data, the EfficientNetV2 model can be used to extract basic features.

[0049] Sent2vec is an unsupervised learning text embedding model used to generate sentence vectors with semantic information. It is an extension of Word2vec, capturing the overall semantics of a sentence by directly training the context at the sentence level. Sent2vec uses pre-trained word vectors combined with weighted average or feature enhancement technology, and can handle sentences of varying lengths while retaining important semantic relationships. It is suitable for tasks such as text classification, retrieval, and question-answering systems, and is particularly outstanding in semantic similarity calculations. Considering that power texts have the characteristics of less noise and strong semantics, the Sent2vec model is used.

[0050] In an embodiment of the present application, in order to improve the retrieval effect, the Sent2vec model can be used to embed the entire sentence rather than being limited to a single word, so that the semantic structure information in the sentence is retained and the semantic features are more complete and comprehensive.

[0051] EfficientNetV2 is an efficient convolutional neural network (CNN) designed to improve the model's training speed and inference efficiency while maintaining high accuracy. It is an upgraded version of EfficientNet, using new architecture design and optimization strategies, such as improved convolution modules (Fused-MBConv) and adaptive optimization network scaling methods, to achieve higher performance with fewer computing resources. Compared with traditional networks, EfficientNetV2 can better balance the depth, width and resolution of the model, and is particularly suitable for large-scale image classification, feature extraction and application scenarios with limited computing resources.

[0052] In the embodiments of the present application, the MBConv module and the Fused-MBConv module of the EfficientNetV2 model are implemented using a convolution kernel of size 3*3.

[0053] Afterwards, the above basic features are input into the hash code generation model to obtain the hash code of the retrieved data. In order to improve the retrieval effect, in the embodiment of the present application, the semantic features of the data are integrated in the process of generating the hash code.

[0054] Specifically, the above-mentioned hash code generation model uses data semantic features as supervision information during the training process. In an embodiment of the present application, the training steps of the hash code generation model may include: obtaining training data and obtaining corresponding basic features for training; according to the basic features for training and the neural network model, using the above-mentioned category labels as supervision information, obtaining the above-mentioned hash code generation model; wherein the above-mentioned category labels indicate the semantic features between the above-mentioned training data.

[0055] The step of obtaining the above-mentioned category labels may include: inputting the above-mentioned training data into a word vector embedding model to generate a word vector, and the types of the above-mentioned training data include text and text in an image; clustering the above-mentioned word vector into multiple label categories through a hierarchical clustering algorithm.

[0056] Specifically, the above process can be explained as follows: inputting the training data including text data and text in the picture into the word vector embedding model, such as Skip-gram in the Word2vec tool, to obtain the word vector; then constructing the word vector matrix based on the above word vector; using the hierarchical clustering algorithm to divide the word vectors in the word vector matrix into different label categories, forming the potential semantic hierarchical structure association features between the device text and the image label. The above hierarchical clustering algorithm can use the Hierarchical K-means method.

[0057] Hash code is an efficient representation method in cross-modal retrieval technology, which is used to map high-dimensional heterogeneous data (such as text and images) into a common low-dimensional hash space. It generates binary hash codes with similar semantics to achieve fast matching between different modal data. It performs well in processing large-scale data and is suitable for tasks such as image and text retrieval and multimodal recommendation. By comparing the hash code obtained in step S103 with the hash code of the data in the database, the database data similar to the retrieval data can be determined.

[0058] In an embodiment of the present application, the Hamming distance between the hash code of the search data and the hash code of the database data may be calculated, and the corresponding database data may be displayed from small to large according to the Hamming distance.

[0059] It should be noted that the method of one or more embodiments of the present application can be performed by a single device, such as a computer or server. The method of this embodiment can also be applied to a distributed scenario and completed by multiple devices cooperating with each other. In the case of such a distributed scenario, one of the multiple devices can only perform one or more steps in the method of one or more embodiments of the present application, and the multiple devices will interact with each other to complete the described method.

[0060] It should be noted that the above describes specific embodiments of the present application. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recorded in the claims can be performed in an order different from that in the embodiments and still achieve the desired results. In addition, the processes depicted in the accompanying drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0061] Based on the same inventive concept, corresponding to any of the above-mentioned embodiments and methods, the present application also provides a cross-modal data retrieval device. Figure 2 As shown, the device comprises:

[0062] The receiving module 11 is configured to receive the user's search data, where the types of the search data include text and image;

[0063] A basic feature extraction module 12 is configured to input the search data into a basic feature extraction model to obtain basic features of the search data;

[0064] The generation module 13 is configured to input the basic features into a hash code generation model to obtain a hash code of the retrieval data; the hash code generation model is trained using the semantic features corresponding to the training basic features as supervision information;

[0065] The search module 14 is configured to calculate the Hamming distance between the hash code of the search data and the hash code of the database data, and display the database data in a preset manner according to the Hamming distance.

[0066] For the convenience of description, the above devices are described in terms of functions and modules. Of course, when implementing one or more embodiments of the present application, the functions of each module can be implemented in the same or multiple software and / or hardware.

[0067] The apparatus of the above-mentioned embodiment is used to implement the corresponding method in the above-mentioned embodiment, and has the beneficial effects of the corresponding method embodiment, which will not be described in detail here.

[0068] Figure 3 A more specific schematic diagram of the hardware structure of an electronic device provided in this embodiment is shown, and the device may include: a processor 1010, a memory 1020, an input / output interface 1030, a communication interface 1040, and a bus 1050. The processor 1010, the memory 1020, the input / output interface 1030, and the communication interface 1040 are connected to each other through the bus 1050 in the device.

[0069] The processor 1010 can be implemented by a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of the present application.

[0070] The memory 1020 may be implemented in the form of ROM (Read Only Memory), RAM (Random Access Memory), static storage device, dynamic storage device, etc. The memory 1020 may store an operating system and other application programs. When the technical solution provided in the embodiment of the present application is implemented by software or firmware, the relevant program code is stored in the memory 1020 and is called and executed by the processor 1010.

[0071] The input / output interface 1030 is used to connect the input / output module to realize information input and output. The input / output module can be configured in the device as a component (not shown in the figure), or it can be externally connected to the device to provide corresponding functions. The input device may include a keyboard, a mouse, a touch screen, a microphone, various sensors, etc., and the output device may include a display, a speaker, a vibrator, an indicator light, etc.

[0072] The communication interface 1040 is used to connect a communication module (not shown) to realize communication interaction between the device and other devices. The communication module can realize communication through a wired mode (such as USB, network cable, etc.) or a wireless mode (such as mobile network, WIFI, Bluetooth, etc.).

[0073] The bus 1050 includes a path that transmits information between the various components of the device (eg, the processor 1010, the memory 1020, the input / output interface 1030, and the communication interface 1040).

[0074] It should be noted that, although the above device only shows the processor 1010, the memory 1020, the input / output interface 1030, the communication interface 1040 and the bus 1050, in the specific implementation process, the device may also include other components necessary for normal operation. In addition, it can be understood by those skilled in the art that the above device may also only include the components necessary for implementing the embodiment of the present application, and does not necessarily include all the components shown in the figure.

[0075] The electronic device of the above embodiment is used to implement the corresponding method in the above embodiment, and has the beneficial effects of the corresponding method embodiment, which will not be described in detail here.

[0076] The computer-readable medium of this embodiment includes permanent and non-permanent, removable and non-removable media, and information storage can be implemented by any method or technology. Information can be computer-readable instructions, data structures, modules of programs, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, read-only compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device.

[0077] Those skilled in the art should understand that the discussion of any of the above embodiments is merely illustrative and is not intended to imply that the scope of the present disclosure (including the claims) is limited to these examples. Based on the concept of the present disclosure, the technical features in the above embodiments or different embodiments may be combined, the steps may be implemented in any order, and there are many other variations of different aspects of one or more embodiments of the present application as described above, which are not provided in detail for the sake of simplicity.

[0078] In addition, to simplify the description and discussion, and in order not to make one or more embodiments of the present application difficult to understand, the known power / ground connections to the integrated circuit (IC) chip and other components may or may not be shown in the provided drawings. In addition, the device can be shown in the form of a block diagram to avoid making one or more embodiments of the present application difficult to understand, and this also takes into account the fact that the details of the implementation of these block diagram devices are highly dependent on the platform on which one or more embodiments of the present application will be implemented (that is, these details should be fully within the scope of understanding of those skilled in the art). In the case of elaborating specific details (e.g., circuits) to describe exemplary embodiments of the present disclosure, it is obvious to those skilled in the art that one or more embodiments of the present application can be implemented without these specific details or when these specific details are changed. Therefore, these descriptions should be considered illustrative rather than restrictive.

[0079] Although the present disclosure has been described in conjunction with specific embodiments of the present disclosure, many replacements, modifications and variations of these embodiments will be apparent to those skilled in the art from the foregoing description. For example, other memory architectures (e.g., dynamic RAM (DRAM)) may use the embodiments discussed.

[0080] One or more embodiments of the present application are intended to cover all such substitutions, modifications and variations that fall within the broad scope of the appended claims. Therefore, any omissions, modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of one or more embodiments of the present application should be included in the scope of protection of this disclosure.

Claims

1. A cross-modal data retrieval method, characterized in that: include: Receiving search data from a user, wherein the types of the search data include text and image; Inputting the search data into a basic feature extraction model to obtain basic features of the search data; Inputting the basic features into a hash code generation model to obtain a hash code of the retrieval data; the hash code generation model is trained using the semantic features corresponding to the training basic features as supervision information; The Hamming distance between the hash code of the retrieved data and the hash code of the database data is calculated, and the database data is displayed in a preset manner according to the Hamming distance.

2. The method according to claim 1, characterized in that The training steps of the hash code generation model include: Obtain training data and obtain corresponding basic features for training; According to the basic features for training and the neural network model, the hash code generation model is obtained by taking the category label as supervision information; The category labels indicate semantic features between the training data.

3. The method according to claim 2, characterized in that The step of obtaining the category label includes: Inputting the training data into a word vector embedding model to generate a word vector, wherein the types of the training data include text and words in an image; The word vectors are clustered into multiple label categories through a hierarchical clustering algorithm.

4. The method according to claim 3, characterized in that: The Skip-gram model in the Word2vec tool is used as the word vector embedding model.

5. The method according to claim 3, characterized in that: The step of obtaining the training data comprises: Acquire equipment data of the power system, where the types of the equipment data include text and image; The natural language processing tool is used to extract the word segments in the device data.

6. The method according to claim 1, characterized in that Inputting the search data into a basic feature extraction model to obtain basic features of the search data includes: In response to receiving the text data, inputting the text data into a Sent2vec model to obtain basic text features; In response to receiving the image data, the image data is input into the EfficientNetV2 model to obtain basic image features.

7. The method according to claim 6, characterized in that The MBConv module and Fused-MBConv module of the EfficientNetV2 model are implemented using a convolution kernel of size 3*3.

8. A cross-modal data retrieval device, characterized in that: include: A receiving module, configured to receive search data of a user, wherein the types of the search data include text and image; A basic feature extraction module is configured to input the search data into a basic feature extraction model to obtain basic features of the search data; A generation module is configured to input the basic features into a hash code generation model to obtain a hash code of the retrieval data; the hash code generation model is trained using the semantic features corresponding to the training basic features as supervision information; The retrieval module is configured to calculate the Hamming distance between the hash code of the retrieval data and the hash code of the database data, and display the database data in a preset manner according to the Hamming distance.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, the method according to any one of claims 1 to 7 is implemented.

10. A non-transitory computer-readable storage medium, characterized in that: The non-transitory computer-readable storage medium stores computer instructions, and the computer instructions are used to cause the computer to execute any one of claims 1 to 7.