Multimodal Data Storage Method, Device, Electronic Device, Storage Medium and Program Product

Through the multimodal data storage method, the storage space and associated information of multimodal data are determined, which solves the problem of difficult fusion and high resource consumption of different modal data and realizes efficient data fusion and labeling.

CN119513366BActive Publication Date: 2025-07-22BEIJING GUODIANTONG NETWORK TECH CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411438210.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-10-15
Publication Date
2025-07-22
Estimated Expiration
2044-10-15

AI Technical Summary

Technical Problem

In the prior art, data of different modes are difficult to effectively integrate, and the annotation and collection of multimodal data requires a lot of manpower and material resources.

Method used

By determining the labeling data of multimodal data, a first storage space is obtained based on the labeling data, the first data association information is determined, the second storage space is obtained, and the third storage space is obtained through the first label association information, and a fifth label information is obtained based on the unlabeled data and the association information of the annotated data, and the unlabeled data is stored in multiple storage spaces.

Benefits of technology

It realizes efficient fusion and labeling of multimodal data, reduces resource consumption, and improves data processing efficiency and accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119513366B_ABST
    Figure CN119513366B_ABST
Patent Text Reader

Abstract

The present disclosure provides a multimodal data storage method, apparatus, electronic device, storage medium, and program product, including: determining annotation data of multimodal data, and obtaining a first storage space based on the annotation data; determining first data association information of the first storage space, and obtaining a second storage space based on the first data association information; determining first label association information of the first storage space, and obtaining a third storage space based on the first label association information; determining unannotated data of the multimodal data, and obtaining second data association information based on the unannotated data and the annotation data; annotating the unannotated data based on the second data association information to obtain fifth label information, and storing the unannotated data in the first storage space, the second storage space, and the third storage space based on the fifth label information. The annotation and collection of the multimodal data in the present disclosure do not consume excessive resources and can effectively fuse the data together.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of artificial intelligence technology, and in particular, to a multi-modal data storage method, apparatus, electronic device, storage medium, and program product. Background Art

[0002] This section aims to provide background or context for the embodiments of the present disclosure stated in the claims. The description herein is not admitted to be prior art merely because it is included in this section.

[0003] Data often exists in multiple forms, such as text, images, audio, etc., which are called multi-modal data; the storage of multi-modal data is an important research direction in the field of deep learning, which aims to effectively integrate data of different modalities to improve the perception and understanding ability of the model.

[0004] However, in the related art, data of different modalities have different feature representations and semantic structures, so it is impossible to effectively fuse the data together, and the annotation and collection of multi-modal data require a large amount of human and material resources. Summary of the Invention

[0005] In view of this, the purpose of the present disclosure is to propose a multi-modal data storage method, apparatus, electronic device, storage medium, and program product, which at least solves one of the technical problems in the related art to a certain extent.

[0006] Based on the above purpose, in the first aspect of an exemplary embodiment of the present disclosure, a multi-modal data storage method is provided, which is applied to a server. The method includes:

[0007] Determine the annotation data of the multi-modal data, and obtain a first storage space based on the annotation data;

[0008] Determine the first data association information of the first storage space, and obtain a second storage space based on the first data association information;

[0009] Determine the first label association information of the first storage space, and obtain a third storage space based on the first label association information;

[0010] Determine the unannotated data of the multi-modal data, and obtain second data association information based on the unannotated data and the annotation data;

[0011] Annotate the unannotated data based on the second data association information to obtain fifth label information, and store the unannotated data in the first storage space, the second storage space, and the third storage space based on the fifth label information.

[0012] Based on the same inventive concept, a second aspect of the exemplary embodiments of the present disclosure provides a multimodal data storage device, including:

[0013] A first space determination module, configured to determine the labeled data of the multimodal data, and obtain a first storage space based on the labeled data;

[0014] A second space determination module, configured to determine first data association information of the first storage space, and obtain a second storage space based on the first data association information;

[0015] A third space determination module, configured to determine first label association information of the first storage space, and obtain a third storage space based on the first label association information;

[0016] A second association determination module, configured to determine the unlabeled data of the multimodal data, and obtain second data association information based on the unlabeled data and the labeled data;

[0017] A data storage determination module, configured to label the unlabeled data based on the second data association information to obtain fifth label information, and store the unlabeled data in the first storage space, the second storage space, and the third storage space based on the fifth label information.

[0018] Based on the same inventive concept, a third aspect of the exemplary embodiments of the present disclosure provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, where when the processor executes the program, the method described in the first aspect is implemented.

[0019] Based on the same inventive concept, a fourth aspect of the exemplary embodiments of the present disclosure provides a non-transitory computer-readable storage medium, where the non-transitory computer-readable storage medium stores computer instructions, and the computer instructions are used to cause a computer to execute the method described in the first aspect.

[0020] Based on the same inventive concept, a fifth aspect of the exemplary embodiments of the present disclosure provides a computer program product, including computer program instructions, where when the computer program instructions run on a computer, the computer is caused to execute the method described in the first aspect.

[0021] As can be seen from the above, the multimodal data storage method, device, electronic device, storage medium, and program product provided by the embodiments of the present disclosure, the method includes:

[0022] Determine the labeled data of the multimodal data, obtain a first storage space based on the labeled data, determine the first data association information of the first storage space, obtain a second storage space based on the first data association information, determine the first label association information of the first storage space, obtain a third storage space based on the first label association information, determine the unlabeled data of the multimodal data, obtain second data association information based on the unlabeled data and the labeled data, label the unlabeled data based on the second data association information to obtain fifth label information, and store the unlabeled data in the first storage space, the second storage space, and the third storage space based on the fifth label information. The labeling and collection of the multimodal data in this disclosure do not consume excessive resources and can effectively fuse the data together. Description of the Drawings

[0023] In order to more clearly illustrate the technical solutions in this disclosure or related technologies, the following will briefly introduce the drawings required for use in the embodiments or related technology descriptions. Obviously, the drawings in the following descriptions are only the embodiments of this disclosure. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0024] Figure 1 A schematic diagram of an application scenario of the multimodal data storage method provided by an exemplary embodiment of this disclosure;

[0025] Figure 2 A schematic flowchart of the multimodal data storage method provided by an exemplary embodiment of this disclosure;

[0026] Figure 3 A schematic structural diagram of the multimodal data storage device provided by an exemplary embodiment of this disclosure;

[0027] Figure 4 A schematic diagram of the hardware structure of an electronic device for multimodal data storage provided by an exemplary embodiment of this disclosure. Detailed Embodiments

[0028] It can be understood that before using the technical solutions disclosed in the embodiments of this application, the types, usage scopes, usage scenarios, etc. of the personal information involved in this application should be informed to users and user authorization should be obtained in an appropriate manner in accordance with relevant laws and regulations.

[0029] For example, when responding to an active request from a user, a prompt message is sent to the user to clearly prompt the user that the operation requested by the user will require obtaining and using the user's personal information. Thus, the user can autonomously choose whether to provide personal information to software or hardware such as an electronic device, an application program, a server, or a storage medium that performs the operations of the technical solution of this application based on the prompt message.

[0030] As an optional but non-limiting implementation manner, the manner of sending a prompt message to the user in response to receiving an active request from the user can be, for example, in the form of a pop-up window, and the prompt message can be presented in text in the pop-up window. In addition, the pop-up window can also carry a selection control for the user to choose "agree" or "disagree" to provide personal information to the electronic device.

[0031] It can be understood that the above notification and user authorization process is only illustrative and does not limit the implementation manner of this application. Other manners that comply with relevant laws and regulations can also be applied to the implementation manner of this application.

[0032] It can be understood that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) should comply with the requirements of corresponding laws, regulations, and related regulations.

[0033] To make the purpose, technical solution, and advantages of this disclosure clearer and more understandable, the principles and spirit of this disclosure will be described below with reference to several exemplary embodiments. It should be understood that these embodiments are only provided to enable those skilled in the art to better understand and then implement this disclosure, rather than limiting the scope of this disclosure in any way. On the contrary, these embodiments are provided to make this disclosure more thorough and complete, and to be able to convey the scope of this disclosure completely to those skilled in the art.

[0034] In this article, it should be understood that any number of elements in the drawings is for illustration rather than limitation, and any naming is only for distinction and does not have any limiting meaning.

[0035] It should be noted that, unless otherwise defined, the technical terms or scientific terms used in the embodiments of the present disclosure should have the ordinary meanings understood by those of ordinary skill in the field to which the present disclosure belongs. The "first", "second" and similar terms used in the embodiments of the present disclosure do not denote any order, quantity or importance, but are only used to distinguish different components. Words such as "including" or "comprising" mean that the elements or objects appearing before this word cover the elements or objects listed after this word and their equivalents, without excluding other elements or objects. Words such as "connected" or "linked" are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. "Up", "down", "left", "right", etc. are only used to represent relative position relationships, and when the absolute position of the object being described changes, the relative position relationship may also change accordingly. The article "a" or "an" before an element does not exclude the existence of multiple such elements.

[0036] Next, with reference to several representative embodiments of the present disclosure, the principles and spirit of the present disclosure will be elaborated in detail.

[0037] As described in the background art, in the related art, data of different modalities have different feature representations and semantic structures. Therefore, it is impossible to effectively fuse the data together, and the annotation and collection of multi-modal data require a large amount of human and material resources. Specifically:

[0038] First of all, different modalities of data have unique feature representation methods. Text data is carried by natural language and contains rich information such as vocabulary, grammar and semantics; image data exists in the form of a pixel matrix and expresses information through visual features such as color, shape and texture; audio data transmits information through acoustic features such as the frequency, amplitude and duration of sound waves. The differences between these features make it difficult to directly fuse different modalities of data, and appropriate feature transformation and matching methods need to be found to effectively integrate them.

[0039] Secondly, there are also differences in the semantic structures of different modalities of data. Text data usually has a clear semantic structure, such as sentences, paragraphs and texts, etc., while image and audio data lack such a clear hierarchical structure and need to be understood and analyzed through technologies such as image recognition and speech recognition. This difference in semantic structure makes it difficult to fuse different modalities of data at the semantic level, and technologies such as natural language processing and computer vision need to be used to achieve semantic-level integration.

[0040] In addition, the annotation and collection of multimodal data also require a large amount of human and material resources. The annotation of multimodal data requires specialized annotation for different types of modal data, such as text annotation, image annotation, and audio annotation, etc. This requires a large number of professional annotators to work for a long time and is prone to annotation errors. The collection of multimodal data needs to obtain different types of modal data from different sources, such as the network, databases, and sensors, etc. This requires the establishment of a complex data acquisition system and is prone to data quality problems.

[0041] In summary, multimodal data fusion and annotation face huge challenges and new technologies and methods need to be developed to solve these problems in order to improve the efficiency and accuracy of multimodal data processing.

[0042] To solve the above problems, the present disclosure provides a multimodal data storage method, apparatus, electronic device, storage medium, and program product solution, specifically including:

[0043] Determine the annotation data of the multimodal data, obtain a first storage space based on the annotation data, determine the first data association information of the first storage space, obtain a second storage space based on the first data association information, determine the first label association information of the first storage space, obtain a third storage space based on the first label association information, determine the unannotated data of the multimodal data, obtain second data association information based on the unannotated data and the annotation data, annotate the unannotated data based on the second data association information to obtain fifth label information, and store the unannotated data in the first storage space, the second storage space, and the third storage space based on the fifth label information. In the present invention, for the annotated multimodal data, after knowledge fusion in different dimensions, a data interaction path is established according to the relationship between different dimensions, and the multimodal data after knowledge fusion is stored in different storage spaces. For the unannotated multimodal data, the present invention makes annotations according to the connection between the multimodal data after knowledge fusion and the unannotated multimodal data, and stores them according to the annotation results.

[0044] After introducing the basic principles of the present disclosure, the following specifically introduces various non-limiting implementation manners of the present disclosure.

[0045] Refer to Figure 1 , which is a schematic diagram of an application scenario of the multimodal data storage method provided by an exemplary embodiment of the present disclosure.

[0046] In this application scenario, it includes a terminal device 101 and a server 102. Among them, both the terminal device 101 and the server 102 can be connected through a wired or wireless communication network to achieve data interaction.

[0047] The terminal device 101 can be an electronic device near the user side with data transmission and multimedia input / output functions, including but not limited to desktop computers, mobile phones, mobile computers, tablet computers, media players, smart wearable devices, personal digital assistants (PDAs), or other electronic devices capable of implementing the above functions. The electronic device may include a processor and a display screen with touch input function, the display screen is used to present a graphical user interface, the graphical user interface can display an application interface, and the processor is used to process application data, generate a graphical user interface, and control the display of the graphical user interface on the display screen.

[0048] The server 102 can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms.

[0049] In some exemplary embodiments, the multimodal data storage method can run on the terminal device 101 or the server 102.

[0050] When the multimodal data storage method runs on the server 102, the server 102 is used to provide multimodal data storage services to the users of the terminal device 101.

[0051] The server 102 determines the annotation data of the multimodal data, and the server 102 obtains a first storage space based on the annotation data;

[0052] The server 102 determines the first data association information of the first storage space, and the server 102 obtains a second storage space based on the first data association information;

[0053] The server 102 determines the first label association information of the first storage space, and the server 102 obtains a third storage space based on the first label association information;

[0054] The server 102 determines the unannotated data of the multimodal data, and the server 102 obtains second data association information based on the unannotated data and the annotation data;

[0055] The server 102 annotates the unannotated data based on the second data association information to obtain fifth label information, and the server 102 stores the unannotated data in the first storage space, the second storage space, and the third storage space based on the fifth label information.

[0056] It should be noted that the above application scenarios are only shown for the convenience of understanding the spirit and principles of the present disclosure, and the embodiments of the present disclosure are not limited in this regard. On the contrary, the embodiments of the present disclosure can be applied to any applicable scenario.

[0057] Reference Figure 2 , a multi-modal data storage method, applied to a server, the method comprising the following steps:

[0058] Step S210, determine the annotation data of the multi-modal data, and obtain a first storage space based on the annotation data.

[0059] Specifically, when implemented, the ways to determine the annotation data of the multi-modal data include, but are not limited to, at least one of the following:

[0060] Obtain annotated multi-modal data such as text, images, videos, audio, etc. through publicly available open-source databases, obtain annotated multi-modal data such as text, images, videos, audio, etc. through social media platforms, or obtain annotated multi-modal data such as text, images, videos, audio, etc. through specific data collection methods.

[0061] In the above exemplary embodiments, the ways to obtain the annotation data of the multi-modal data are introduced. Next, the ways to obtain the first storage space will be introduced:

[0062] Specifically, when implemented, the ways to obtain a first storage space based on the annotation data include:

[0063] First, perform data cleaning on the original annotation data collected by the server, delete duplicate records in the dataset, and ensure that each data sample is unique;

[0064] Secondly, if the obtained annotation data is text data, remove irrelevant characters, HTML tags, or stop words in the text; if the obtained annotation data is image data, use filters (such as Gaussian filtering, median filtering) to reduce noise in the image; if the obtained annotation data is audio data, apply signal processing techniques (such as spectral subtraction, wavelet transform) to remove background noise;

[0065] Then, convert the data from different sources into a unified format for easy processing and analysis;

[0066] Finally, design a reasonable data storage structure for easy data management and fast access.

[0067] Step S220, determine the first data association information of the first storage space, and obtain a second storage space based on the first data association information.

[0068] In specific implementation, the first data association information refers to:

[0069] In the first storage space, the first association relationship between different modality data; this relationship can be understood as that due to one or more identical tags existing between different data, a necessary connection is generated.

[0070] As a specific embodiment, for example, there is multimodal data about "beach" and multimodal data about "vacation" in the storage space. The first association relationship can be the connection between "beach" and "vacation" because the beach is usually associated with the vacation experience. In this case, if we see a picture of a beach, we may automatically associate with the concept of vacation, even if the word "vacation" is not directly mentioned in the text description or user comments.

[0071] In the above exemplary embodiment, the first data association information is introduced. Next, the method for obtaining the second tag information is introduced:

[0072] In this exemplary embodiment, based on the first data association information, tag fusion is performed on the labeled data in the first storage space to obtain the second tag information, including:

[0073] Feature extraction is performed on the labeled data to obtain the first tag information, and based on the first data association information, tag fusion is performed on the first tag information to obtain the second tag information.

[0074] In specific implementation, the method for performing feature extraction on the labeled data to obtain the first tag information includes but is not limited to at least one of the following:

[0075] If the obtained labeled data is text data, then convert the text into a set of words and use the frequency of word occurrence as a feature;

[0076] If the obtained labeled data is image data, then count the number of pixels of different colors in the image as a feature;

[0077] If the obtained labeled data is video data, decompose the video into frames and perform image feature extraction on each frame;

[0078] If the obtained labeled data is audio data, then extract the Mel-frequency cepstral coefficients of the audio signal as a feature;

[0079] Label the data according to predefined rules, for example, extract keywords according to the text content as the first tag information; in this solution, the first tag information includes at least two different tags so that this solution can capture the diversity and complexity of multimodal data.

[0080] In specific implementation, the method for performing label fusion on the first label information based on the first data association information to obtain the second label information is as follows:

[0081] Generate the second label information according to the first data association information among the first label information. For example, if two texts have the same first label, they may have the same second label, indicating that they belong to the same theme. If two texts do not have the same first label, we perform knowledge fusion on the multimodal data according to the first data association information. By fusing the labels, new second label information is generated. In this solution, the second label information includes at least two different labels, so that this solution can capture the diversity and complexity of the multimodal data.

[0082] In the above exemplary embodiment, the method for obtaining the second label information is introduced. Next, the method for obtaining the second storage space is introduced:

[0083] In this exemplary embodiment, obtaining the second storage space based on the first data association information includes:

[0084] Perform label fusion on the labeled data in the first storage space based on the first data association information to obtain the second label information, and construct the second storage space based on the second label information.

[0085] In specific implementation, constructing the second storage space based on the second label information includes but is not limited to at least one of the following:

[0086] Construct the second storage space based on the storage structure of the file system according to the second label information, construct the second storage space based on the storage structure of the database according to the second label information, construct the second storage space based on the storage structure of distributed storage according to the second label information, or construct the second storage space based on the storage structure of cloud storage according to the second label information.

[0087] Step S230: Determine the first label association information of the first storage space, and obtain the third storage space based on the first label association information.

[0088] In specific implementation, the first label association information refers to:

[0089] The possible relationships between different tags in the labeled data. This relationship can be understood as the relationship between tags generated when setting the labeling rules, or it can also be the relationship between tags obtained through big data. For example: the inevitable relationship between a person's name and a company (such as an employment relationship), the inevitable relationship between a person's name and another person's name (such as a spousal relationship, a kinship, etc.), the inevitable relationship between companies (such as a subordinate relationship, etc.), or it can be understood that in a certain dimension of data, there are different tags, resulting in an inevitable connection between the tags. For example: someone likes playing basketball, someone interacts with someone else, and that someone else also likes basketball, that is, the objective connection (having a common hobby) between a person's name + a sport, and including having a common occupation, etc., or it can be exemplified as all being cats, just the difference between black cats, white cats, and calico cats.

[0090] In the above exemplary embodiment, the method of the first tag association information is introduced. Next, the method of obtaining the third storage space is introduced:

[0091] Determine the first tag association information of the first tag information, perform tag fusion on the first tag information based on the first tag association information to obtain the third tag information, and construct the third storage space based on the third tag information.

[0092] Specifically, when implementing, the advantage of performing tag fusion on the first tag information based on the first tag association information is to fully fuse the data of other dimensions corresponding to the tags according to the inevitable connection between the tags of the data in the same dimension, so that the fused data exhibits multi-dimensional characteristics.

[0093] Specifically, when implementing, the third tag information includes at least two different tags, so that this solution can capture the diversity and complexity of multi-modal data.

[0094] Specifically, when implementing, the methods of constructing the third storage space based on the third tag information include but are not limited to at least one of the following:

[0095] Construct the third storage space based on the storage structure of the file system for the third tag information, construct the third storage space based on the storage structure of the database for the third tag information, construct the third storage space based on the distributed storage structure for the third tag information, or construct the third storage space based on the cloud storage structure for the third tag information.

[0096] In the above exemplary embodiments, the methods for constructing the first storage space, the second storage space, and the third storage space based on the obtained labeled multimodal data are introduced respectively. When establishing different dimensions and different new contents, it is also necessary to consider what necessary connections exist among these newly formed fusion knowledges, so as to establish the necessary connections among different storage spaces. It should be noted here that storing different labels in different storage spaces aims to ensure data security and reduce the storage pressure of the storage space when calling data. Since the label fusion methods include content-dominated and label-dominated methods, as the amount of data continues to increase, the order of magnitude of their fusion will inevitably increase exponentially. Then, in the first storage space, that is, the storage of (the first label information), the storage pressure will gradually increase, and there are also risks such as data loss when expanding the first storage space. Therefore, it is necessary to release the storage pressure through the multi-storage space method. And by obtaining the necessary connections among the fusions, a data interaction channel between the fusion knowledge and the original knowledge is well established, and the second storage space and the third storage space can be wirelessly expanded according to this channel.

[0097] Step S240: Determine the unlabeled data of the multimodal data, and obtain second data association information based on the unlabeled data and the labeled data.

[0098] Specifically in implementation, the methods for determining the unlabeled data of the multimodal data include, but are not limited to, at least one of the following:

[0099] Obtain unlabeled multimodal data such as text, images, videos, and audios through publicly available open-source databases, obtain unlabeled multimodal data such as text, images, videos, and audios through social media platforms, or obtain unlabeled multimodal data such as text, images, videos, and audios through specific data collection methods.

[0100] Specifically in implementation, the second data association information refers to:

[0101] It refers to the correlation relationship between the unlabeled multimodal data and the labeled multimodal data. It is used to guide the annotation process of the unlabeled data and determine its storage location.

[0102] Step S250: Label the unlabeled data based on the second data association information to obtain fifth label information, and store the unlabeled data in the first storage space, the second storage space, and the third storage space based on the fifth label information.

[0103] In this exemplary embodiment, labeling the unlabeled data based on the second data association information to obtain fifth label information includes:

[0104] Extract features from the unlabeled data to obtain unlabeled data feature information, label the unlabeled data based on the unlabeled data feature information to obtain fourth label information, and label the fourth label information based on the second data association information to obtain fifth label information.

[0105] In specific implementation, the methods for extracting features from the unlabeled data to obtain unlabeled data feature information include, but are not limited to, at least one of the following:

[0106] If the obtained unlabeled data is image data, count the number of pixels of different colors in the image as features;

[0107] If the obtained unlabeled data is video data, decompose the video into frames and extract image features for each frame;

[0108] If the obtained unlabeled data is audio data, extract the Mel-frequency cepstral coefficients of the audio signal as features;

[0109] Obtain unlabeled data feature information based on the above methods.

[0110] In specific implementation, the methods for labeling the unlabeled data based on the unlabeled data feature information to obtain fourth label information include:

[0111] Label the data according to predefined rules. For example, extract keywords from the text content as the fourth label information; in this solution, the fourth label information contains at least two different labels to enable this solution to capture the diversity and complexity of multimodal data.

[0112] In specific implementation, label the fourth label information based on the second data association information to obtain fifth label information:

[0113] For unlabeled multimodal data, when labeling the third label, it is also possible to obtain whether there is "second data association information" between the unlabeled data and the labeled data. Here, the second data association information can be understood as the possible existence of relevance based on the labeled content for the unlabeled content corresponding to the fourth label information; and then directly obtain the third label of the labeled content according to this relevance and assign it to the unlabeled data. That is to say, when labeling the third label for this part of the content, it is not necessary to label the fourth label information, but directly label the fifth label information for this part, thus saving the labeling of duplicate content (or approximate content, such as the same pictures that may exist between images, the same audio that may exist between sound and images, the same text content that may exist between text and images, etc.).

[0114] In the above exemplary embodiment, the method for obtaining the fifth tag information is introduced. Next, the method for storing the unlabeled data in the first storage space, the second storage space, and the third storage space is introduced:

[0115] In this exemplary embodiment, storing the unlabeled data in the first storage space, the second storage space, and the third storage space based on the fifth tag information includes:

[0116] Determine the second tag association information of the fourth tag information, perform tag fusion on the fourth tag information based on the second tag association information to obtain the sixth tag information, and store the unlabeled data in the first storage space, the second storage space, and the third storage space based on the fourth tag information, the fifth tag information, and the sixth tag information.

[0117] Specifically, the second tag association information refers to:

[0118] The possible relationships between different tags in the unlabeled data. This relationship can be understood as the relationship between tags generated when setting the annotation rules, or the relationship between tags obtained through big data.

[0119] Specifically, the method for storing the unlabeled data in the first storage space, the second storage space, and the third storage space based on the fourth tag information, the fifth tag information, and the sixth tag information includes:

[0120] Store the unlabeled data with the fourth tag information and / or the fifth tag information in the first storage space; here, the storage of the unlabeled data is based on the update of the data without duplicate content in the storage space.

[0121] Store the unlabeled data with the fifth tag information in the first storage space and the second storage space. Among them, the second multimodal data with the third tag information stored in the first storage space establishes a first mapping relationship with the unlabeled data with the fourth tag information and / or the fifth tag information, and according to the first mapping relationship, when the data is called according to the fourth tag information in the first storage space, the unlabeled data corresponding to the fifth tag information is presented; here, the storage of the unlabeled data is based on the update operation of the data with duplicate content corresponding to the fifth tag in different storage spaces.

[0122] Store the unlabeled data with the sixth label in the first storage space, the second storage space, and the third storage space. Among them, a second mapping relationship is established between the unlabeled data with the sixth label stored in the first storage space and the unlabeled data with the fourth label information and / or the fifth label information. According to the second mapping relationship, when data is called according to the first label or the second label in the first storage space, the unlabeled data with the sixth label is presented;

[0123] A third mapping relationship is established between the second multimodal data with the sixth label stored in the second storage space and the second unlabeled data with the fifth label. According to the third mapping relationship, when data is called according to the fifth label in the second storage space, the unlabeled data corresponding to the sixth label is presented; Here, the storage of the unlabeled data is based on the update operation of the data with repeated content corresponding to the sixth label in different storage spaces.

[0124] It should be noted that the method of the embodiments of the present disclosure can be executed by a single device, such as a computer or a server. The method of this embodiment can also be applied to a distributed scenario and completed by multiple devices cooperating with each other. In this case of a distributed scenario, one of the multiple devices can only execute one or more steps of the method of the embodiments of the present disclosure, and these multiple devices will interact with each other to complete the described method.

[0125] It should be noted that some embodiments of the present disclosure have been described above. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in a different order than in the above embodiments and still achieve the desired result. Additionally, the processes depicted in the figures do not necessarily require the particular order or sequential order shown to achieve the desired result. In certain implementations, multitasking and parallel processing are also possible or may be advantageous.

[0126] Based on the same inventive concept, corresponding to the method of any of the above embodiments, the present disclosure also provides a multimodal data storage device.

[0127] Refer to Figure 3 , the multimodal data storage device includes:

[0128] The first space determination module 310 is configured to determine the labeled data of the multimodal data and obtain the first storage space based on the labeled data;

[0129] The second space determination module 320 is configured to determine the first data association information of the first storage space and obtain the second storage space based on the first data association information;

[0130] A third space determination module 330, configured to determine first tag association information of the first storage space, and obtain a third storage space based on the first tag association information;

[0131] A second association determination module 340, configured to determine unlabeled data of the multimodal data, and obtain second data association information based on the unlabeled data and the labeled data;

[0132] A data storage determination module 350, configured to label the unlabeled data based on the second data association information to obtain fifth tag information, and store the unlabeled data in the first storage space, the second storage space, and the third storage space based on the fifth tag information.

[0133] In this exemplary embodiment, the first space determination module 310 is specifically configured to:

[0134] Determine labeled data of the multimodal data, and obtain a first storage space based on the labeled data.

[0135] In this exemplary embodiment, the second space determination module 320 is specifically configured to:

[0136] Determine first data association information of the first storage space, extract features from the labeled data to obtain first tag information, perform tag fusion on the first tag information based on the first data association information to obtain the second tag information, and construct the second storage space based on the second tag information.

[0137] In this exemplary embodiment, the third space determination module 330 is specifically configured to:

[0138] Determine the first tag association information of the first tag information, perform tag fusion on the first tag information based on the first tag association information to obtain third tag information, and construct the third storage space based on the third tag information.

[0139] In this exemplary embodiment, the second association determination module 340 is specifically configured to:

[0140] Determine unlabeled data of the multimodal data, and obtain second data association information based on the unlabeled data and the labeled data.

[0141] In this exemplary embodiment, the data storage determination module 350 is specifically configured to:

[0142] Feature extraction is performed on the unlabeled data to obtain unlabeled data feature information. Based on the unlabeled data feature information, the unlabeled data is labeled to obtain fourth label information. Based on the second data association information, the fourth label information is labeled to obtain fifth label information. The second label association information of the fourth label information is determined. Based on the second label association information, label fusion is performed on the fourth label information to obtain sixth label information. Based on the fourth label information, the fifth label information, and the sixth label information, the unlabeled data is stored in the first storage space, the second storage space, and the third storage space.

[0143] For the convenience of description, when describing the above device, various modules are described separately according to their functions. Of course, when implementing the present disclosure, the functions of each module can be implemented in the same or multiple software and / or hardware.

[0144] The device of the above embodiment is used to implement the corresponding multi-modal data storage method in any of the foregoing embodiments, and has the beneficial effects of the corresponding method embodiments, which will not be elaborated here.

[0145] Based on the same inventive concept, corresponding to the method of any of the above embodiments, the present disclosure further provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the multi-modal data storage method described in any of the above embodiments.

[0146] Figure 4 FIG. shows a more specific schematic diagram of the hardware structure of the electronic device provided in this embodiment. The device may include: a processor 1010, a memory 1020, an input / output interface 1030, a communication interface 1040, and a bus 1050. Among them, the processor 1010, the memory 1020, the input / output interface 1030, and the communication interface 1040 are communicatively connected to each other inside the device through the bus 1050.

[0147] The processor 1010 may be implemented in a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, etc., and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this specification.

[0148] The memory 1020 may be implemented in the form of a ROM (Read Only Memory), a RAM (Random Access Memory), a static storage device, a dynamic storage device, etc. The memory 1020 may store an operating system and other application programs. When implementing the technical solutions provided in the embodiments of this specification through software or firmware, the relevant program codes are stored in the memory 1020 and called and executed by the processor 1010.

[0149] The input / output interface 1030 is used to connect to an input / output module to achieve information input and output. The input / output module may be configured as a component in the device (not shown in the figure) or externally connected to the device to provide corresponding functions. The input devices may include a keyboard, a mouse, a touch screen, a microphone, various sensors, etc., and the output devices may include a display, a speaker, a vibrator, an indicator light, etc.

[0150] The communication interface 1040 is used to connect to a communication module (not shown in the figure) to achieve communication interaction between this device and other devices. The communication module may achieve communication through a wired manner (such as USB, network cable, etc.) or through a wireless manner (such as a mobile network, WIFI, Bluetooth, etc.).

[0151] The bus 1050 includes a path for transmitting information between various components of the device (such as the processor 1010, the memory 1020, the input / output interface 1030, and the communication interface 1040).

[0152] It should be noted that although the above device only shows the processor 1010, the memory 1020, the input / output interface 1030, the communication interface 1040, and the bus 1050, in the specific implementation process, the device may also include other components necessary for normal operation. In addition, those skilled in the art can understand that the above device may also only include the components necessary for implementing the solutions of the embodiments of this specification, and do not necessarily include all the components shown in the figure.

[0153] The electronic device in the above embodiment is used to implement the corresponding multi-modal data storage method in any of the foregoing embodiments, and has the beneficial effects of the corresponding method embodiments, which will not be elaborated here.

[0154] Based on the same inventive concept, corresponding to the method in any of the above embodiments, the present disclosure also provides a non-transitory computer-readable storage medium storing computer instructions for causing the computer to execute the multi-modal data storage method as described in any of the foregoing embodiments.

[0155] The computer-readable medium of this embodiment includes permanent and non-permanent, removable and non-removable media, and information storage can be implemented by any method or technology. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory, or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD), or other optical storage, magnetic cassette tapes, magnetic tape magnetic disk storage, or other magnetic storage devices, or any other non-transmission medium that can be used to store information accessible by a computing device.

[0156] The above non-transitory computer-readable storage medium can be any available medium or data storage device accessible by a computer, including but not limited to magnetic memory (such as floppy disks, hard disks, magnetic tapes, magneto-optical discs (MO), etc.), optical memory (such as CDs, DVDs, BDs, HVDs, etc.), and semiconductor memory (such as ROM, EPROM, EEPROM, non-volatile memory (NAND FLASH), solid-state drives (SSD)), etc.

[0157] The computer instructions stored in the storage medium of the above embodiment are used to cause the computer to execute the multi-modal data storage method described in any one of the above exemplary method embodiments, and have the beneficial effects of the corresponding method embodiments, which will not be elaborated here.

[0158] Based on the same inventive concept, corresponding to the multi-modal data storage method described in any of the above embodiments, the present disclosure also provides a computer program product, which includes computer program instructions. In some embodiments, the computer program instructions can be executed by one or more processors of the computer to cause the computer and / or the processor to execute the multi-modal data storage method. Corresponding to the execution subjects corresponding to the steps in each embodiment of the multi-modal data storage method, the processor that executes the corresponding steps can belong to the corresponding execution subject.

[0159] The computer program product of the above embodiment is used to cause the computer and / or the processor to execute the multi-modal data storage method described in any of the above embodiments, and has the beneficial effects of the corresponding method embodiments, which will not be elaborated here.

[0160] As is known to those skilled in the art, the embodiments of the present disclosure can be implemented as a system, a method, or a computer program product. Therefore, the present disclosure can be specifically implemented in the following forms, namely: all hardware, all software (including firmware, resident software, microcode, etc.), or a combination of hardware and software, which is generally referred to as "circuit", "module", or "system" in this article. In addition, in some embodiments, the present disclosure can also be implemented in the form of a computer program product in one or more computer-readable media, which contains computer-readable program code.

[0161] Any combination of one or more computer-readable media can be adopted. The computer-readable media can be a computer-readable signal medium or a computer-readable storage medium. The computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (non-exhaustive examples) of the computer-readable storage medium can include, for example: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In this document, the computer-readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or in combination with an instruction execution system, apparatus, or device.

[0162] The computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries the computer-readable program code. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium can also be any computer-readable medium other than the computer-readable storage medium, and this computer-readable medium can send, propagate, or transmit a program for use by or in combination with an instruction execution system, apparatus, or device.

[0163] The program code contained on the computer-readable medium can be transmitted by any appropriate medium, including but not limited to wireless, wire, optical cable, RF, etc., or any suitable combination of the above.

[0164] Computer program code for performing the operations of this disclosure can be written in one or more programming languages or combinations thereof, including object-oriented programming languages such as Java, Smalltalk, C++, and also including conventional procedural programming languages such as the "C" language or similar programming languages. The program code can execute entirely on the user's computer, partially on the user's computer, execute as a stand-alone software package, execute partially on the user's computer and partially on a remote computer, or execute entirely on the remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or alternatively, can be connected to an external computer (e.g., through the Internet using an Internet service provider).

[0165] It should be understood that each block of the flowchart and / or block diagram, and combinations of blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine, such that the computer program instructions, when executed by the computer or other programmable data processing apparatus, create means for implementing the functions / operations specified in the blocks of the flowchart and / or block diagram.

[0166] These computer program instructions can also be stored in a computer-readable medium that can direct a computer or other programmable data processing apparatus to function in a particular manner, such that the instructions stored in the computer-readable medium produce an article of manufacture including instructions means for implementing the functions / operations specified in the blocks of the flowchart and / or block diagram.

[0167] The computer program instructions can also be loaded onto a computer, other programmable data processing apparatus, or other device, causing a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, such that the instructions executed on the computer or other programmable apparatus provide a process for implementing the functions / operations specified in the blocks of the flowchart and / or block diagram.

[0168] In addition, although the operations of the method of this disclosure are depicted in the figures in a particular order, this is not required or implied to perform the operations in that particular order, or to perform all of the illustrated operations to achieve the desired result. Instead, the steps depicted in the flowchart can be reordered. Additionally or alternatively, certain steps can be omitted, multiple steps can be combined into one step for execution, and / or one step can be decomposed into multiple steps for execution.

[0169] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present application. Among them, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code, and the above-mentioned module, program segment, or part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than that marked in the accompanying drawings. For example, two consecutively represented blocks may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram or flowchart, as well as the combination of blocks in the block diagram or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.

[0170] It should be noted that although several modules or units of the device for action execution are mentioned in the above detailed description, such a division is not mandatory. In fact, according to the embodiments of the present application, the features and functions of the two or more modules or units described above can be embodied in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided and embodied by multiple modules or units.

[0171] Those of ordinary skill in the art should understand that: the discussion of any of the above embodiments is only exemplary and is not intended to imply that the scope of the present application (including the claims) is limited to these examples; under the idea of the present application, the technical features between the above embodiments or different embodiments can also be combined, the steps can be implemented in any order, and there are many other variations of different aspects of the embodiments of the present application as described above, and they are not provided in detail for the sake of brevity.

[0172] In addition, for the sake of simplicity of description and discussion, and in order not to make the embodiments of the present application difficult to understand, the well-known power / ground connections to integrated circuit (IC) chips and other components may or may not be shown in the provided drawings. In addition, the device may be shown in the form of a block diagram in order to avoid making the embodiments of the present application difficult to understand, and this also takes into account the fact that the details of the implementation of these block diagram devices are highly dependent on the platform on which the embodiments of the present application will be implemented (that is, these details should be completely within the understanding of those skilled in the art). In the case where specific details (such as circuits) are set forth to describe the exemplary embodiments of the present application, it will be apparent to those skilled in the art that the embodiments of the present application can be implemented without these specific details or with variations of these specific details. Therefore, these descriptions should be considered illustrative rather than restrictive.

[0173] Although the present application has been described in connection with specific embodiments thereof, many alternatives, modifications, and variations of these embodiments will be apparent to those of ordinary skill in the art in light of the foregoing description. For example, other memory architectures (e.g., dynamic RAM (DRAM)) may be used with the embodiments discussed.

[0174] Embodiments of the present application are intended to cover all such alternatives, modifications, and variations that fall within the broad scope of the appended claims. Accordingly, any omissions, modifications, equivalent substitutions, improvements, etc. made within the spirit and principle of the embodiments of the present application shall be included within the protection scope of the present application.

[0175] Although the spirit and principles of the present disclosure have been described with reference to several specific embodiments, it should be understood that the present disclosure is not limited to the specific embodiments disclosed, and the division of each aspect does not mean that the features in these aspects cannot be combined for benefit. This division is only for convenience of expression. The present disclosure is intended to cover various modifications and equivalent arrangements included within the spirit and scope of the appended claims. The scope of the appended claims is to be accorded the broadest interpretation so as to encompass all such modifications and equivalent structures and functions.

Claims

1. A multimodal data storage method, characterized in that, Including: Determine the labeled data of the multimodal data, and obtain a first storage space based on the labeled data; Determine the first data association information of the first storage space, extract features from the labeled data, and obtain first label information; Perform label fusion on the first label information based on the first data association information to obtain second label information, and construct a second storage space based on the second label information; wherein, the first data association information refers to the first association relationship between different modality data in the first storage space, and a necessary connection is generated due to one or more identical labels existing between the different modality data; Determine the first label association information of the first label information in the first storage space, perform label fusion on the first label information based on the first label association information to obtain third label information, and construct a third storage space based on the third label information; wherein, the first label association information refers to the possible relationship between different labels in the labeled data; Determine the unlabeled data of the multimodal data, and obtain second data association information based on the unlabeled data and the labeled data; Label the unlabeled data based on the second data association information to obtain fifth label information, and store the unlabeled data in the first storage space, the second storage space, and the third storage space based on the fifth label information.

2. The method according to claim 1, characterized in that, The labeling the unlabeled data based on the second data association information to obtain fifth label information includes: Extract features from the unlabeled data to obtain unlabeled data feature information; Label the unlabeled data based on the unlabeled data feature information to obtain fourth label information; Label the fourth label information based on the second data association information to obtain fifth label information.

3. The method according to claim 2, wherein The storing the unlabeled data in the first storage space, the second storage space, and the third storage space based on the fifth label information includes: Determine the second label association information of the fourth label information, perform label fusion on the fourth label information based on the second label association information to obtain sixth label information; Store the unlabeled data in the first storage space, the second storage space, and the third storage space based on the fourth label information, the fifth label information, and the sixth label information.

4. A multimodal data storage device, characterized in that, Including: A first space determination module, configured to determine the labeled data of the multimodal data, and obtain a first storage space based on the labeled data; A second space determination module, configured to determine the first data association information of the first storage space, extract features from the labeled data, and obtain first label information; Perform label fusion on the first label information based on the first data association information to obtain second label information, and construct a second storage space based on the second label information; wherein, the first data association information refers to the first association relationship between different modality data in the first storage space, and a necessary connection is generated due to one or more identical labels existing between the different modality data. A third space determination module, configured to determine the first label association information of the first label information in the first storage space, perform label fusion on the first label information based on the first label association information to obtain third label information, and construct a third storage space based on the third label information; wherein, the first label association information refers to the possible relationship between different labels in the labeled data. A second association determination module, configured to determine the unlabeled data of the multimodal data, and obtain second data association information based on the unlabeled data and the labeled data. A data storage determination module, configured to label the unlabeled data based on the second data association information to obtain fifth label information, and store the unlabeled data in the first storage space, the second storage space, and the third storage space based on the fifth label information.

5. An electronic device, characterized in that, It includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the method according to any one of claims 1 to 3.

6. A non-transitory computer-readable storage medium, characterized in that, The non-transitory computer-readable storage medium stores computer instructions for causing a computer to execute the method according to any one of claims 1 to 3.

7. A computer program product, characterized in that, It includes computer program instructions that, when running on a computer, cause the computer to execute the method according to any one of claims 1 to 3.

Citation Information

Patent Citations

  • Association identification method and device of multi-modal data, equipment and storage medium

    CN117611845A

  • Method for processing multimodal data using neural network, device, and medium

    EP4109347A2