Training data deduplication method, device and equipment
The method improves training data deduplication by extracting and fusing multi-dimensional features from multi-modal data, addressing inefficiencies and imprecision in existing methods to enhance the accuracy and efficiency of computer vision algorithms.
Patent Information
- Application Number
- CN202510219738.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-26
- Publication Date
- 2025-07-15
AI Technical Summary
In the prior art, the training data deduplication efficiency is low and the accuracy is poor, which affects the model effect.
By extracting multi-dimensional feature information of multi-modal training data and fusion, deduplication is used to use similarity calculations, and feature information is stored in combination with object storage and vector databases.
Improve the efficiency and accuracy of training data deduplication and optimize the model training process.
Smart Images

Figure CN120316408A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data processing, and in particular, to a method, device, and equipment for deduplicating training data. Background Art
[0002] At present, there is a vast amount of image data and related text and audio data on the Internet, which provides great convenience for the training of computer vision algorithms. However, due to the large volume of these data, there are many duplicate and similar data. During the model training process, due to the appearance of a large number of similar data, it is easy to produce the problem of sample imbalance, resulting in bias in the trained model and affecting the use effect of the model.
[0003] The existing technology mainly deduplicates training data through manual methods or primary retrieval methods. The efficiency and reliability of manual deduplication are poor, and the deduplication efficiency of the primary retrieval method is relatively high, but the accuracy is poor. Summary of the Invention
[0004] The present invention provides a method, device, and equipment for deduplicating training data to solve the defects that the efficiency and reliability of manual deduplication in the existing technology are poor, and the accuracy of deduplication by the primary retrieval method is poor.
[0005] In a first aspect, the present invention provides a method for deduplicating training data, including: Obtaining multiple new multimodal training data, and obtaining multiple historical multimodal training data from a sample database; Extracting first multi-dimensional feature information corresponding to each new multimodal training data, and extracting second multi-dimensional feature information corresponding to each historical multimodal training data; Fusing the first multi-dimensional feature information corresponding to each new multimodal training data to obtain first fusion feature information corresponding to each new multimodal training data, and fusing the second multi-dimensional feature information corresponding to each historical multimodal training data to obtain second fusion feature information corresponding to each historical multimodal training data; Deduplicating the multiple new multimodal training data based on the first fusion feature information corresponding to each new multimodal training data and the second fusion feature information corresponding to each historical multimodal training data.
[0006] In some embodiments, each piece of new multimodal training data includes each image sample data, and text sample data and / or audio sample data associated with each image sample data; each piece of historical multimodal training data includes each historical image sample data, and historical text sample data and / or historical audio sample data associated with each historical image sample data; the first multi-dimensional feature information includes first image features, first text features, and / or first audio features; the second multi-dimensional feature information includes second image features, second text features, and / or second audio features.
[0007] In some embodiments, extracting the first multi-dimensional feature information corresponding to each piece of new multimodal training data includes: Performing feature extraction on each image sample data to obtain first image features corresponding to each image sample data; Performing feature extraction on the text sample data associated with each image sample data to obtain first text features corresponding to the text sample data associated with each image sample data, and / or performing feature extraction on the audio sample data associated with each image sample data to obtain first audio features corresponding to the audio sample data associated with each image sample data; Extracting the second multi-dimensional feature information corresponding to each piece of historical multimodal training data includes: Performing feature extraction on each historical image sample data to obtain second image features corresponding to each historical image sample data; Performing feature extraction on the historical text sample data associated with each historical image sample data to obtain second text features corresponding to the historical text sample data associated with each historical image sample data, and / or performing feature extraction on the historical audio sample data associated with each historical image sample data to obtain second audio features corresponding to the historical audio sample data associated with each historical image sample data.
[0008] In some embodiments, fusing the first multi-dimensional feature information corresponding to each piece of new multimodal training data to obtain first fused feature information corresponding to each piece of new multimodal training data includes: Determining the weights of the first image features, the first text features, and / or the first audio features corresponding to each piece of new multimodal training data; Based on the weights of the first image features, the first text features, and / or the first audio features, fusing the first image features, the first text features, and / or the first audio features corresponding to each piece of new multimodal training data to obtain first fused feature information corresponding to each piece of new multimodal training data.
[0009] In some embodiments, fusing the second multi-dimensional feature information corresponding to each piece of historical multi-modal training data to obtain the second fused feature information corresponding to each piece of historical multi-modal training data includes: Determine the weights of the second image features, the weights of the second text features, and / or the weights of the second audio features corresponding to each piece of historical multi-modal training data; Based on the weights of the second image features, the weights of the second text features, and / or the weights of the second audio features, fuse the second image features, the second text features, and / or the second audio features corresponding to each piece of historical multi-modal training data to obtain the second fused feature information corresponding to each piece of historical multi-modal training data.
[0010] In some embodiments, de-duplicating the multiple pieces of new multi-modal training data based on the first fused feature information corresponding to each piece of new multi-modal training data and the second fused feature information corresponding to each piece of historical multi-modal training data includes: Calculate the similarity between each piece of new multi-modal training data and each piece of historical multi-modal training data based on the first fused feature information corresponding to each piece of new multi-modal training data and the second fused feature information corresponding to each piece of historical multi-modal training data; De-duplicate the multiple pieces of new multi-modal training data according to the similarity between each piece of new multi-modal training data and each piece of historical multi-modal training data and a preset similarity threshold.
[0011] In some embodiments, the method further includes: Store the multiple pieces of new multi-modal training data and multiple pieces of historical multi-modal training data in an object storage database; Store the first multi-dimensional feature information corresponding to each piece of new multi-modal training data and the second multi-dimensional feature information corresponding to each piece of historical multi-modal training data in a first vector database; Store the first fused feature information corresponding to each piece of new multi-modal training data and the second fused feature information corresponding to each piece of historical multi-modal training data in a second vector database.
[0012] In some embodiments, the method further includes: Obtain a retrieval request of a user, determine a target retrieval method based on the retrieval request of the user, and retrieve each piece of new multi-modal training data according to the target retrieval method to obtain a retrieval result; The target retrieval method includes at least one of the following: Retrieve historical multi-modal training data similar to each piece of new multi-modal training data from the object storage database; Retrieve second multi-dimensional feature information similar to the first multi-dimensional feature information from the first vector database; Retrieve second fusion feature information similar to the first fusion feature information from the second vector database.
[0013] In a second aspect, the present invention further provides a training data deduplication device, including: An acquisition unit, configured to acquire multiple new multi-modal training data, and acquire multiple historical multi-modal training data from a sample database; A feature extraction unit, configured to extract first multi-dimensional feature information corresponding to each new multi-modal training data, and extract second multi-dimensional feature information corresponding to each historical multi-modal training data; A feature fusion unit, configured to fuse the first multi-dimensional feature information corresponding to each new multi-modal training data to obtain first fusion feature information corresponding to each new multi-modal training data, and fuse the second multi-dimensional feature information corresponding to each historical multi-modal training data to obtain second fusion feature information corresponding to each historical multi-modal training data; A data deduplication unit, configured to deduplicate the multiple new multi-modal training data based on the first fusion feature information corresponding to each new multi-modal training data and the second fusion feature information corresponding to each historical multi-modal training data.
[0014] In a third aspect, the present invention further provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, where when the processor executes the computer program, it implements the training data deduplication method as described in any one of the above.
[0015] The training data deduplication method, device, and equipment provided by the present invention improve the efficiency and accuracy of training data deduplication by acquiring multiple new multi-modal training data and multiple historical multi-modal training data; extracting first multi-dimensional feature information corresponding to each new multi-modal training data and second multi-dimensional feature information corresponding to each historical multi-modal training data; fusing the first multi-dimensional feature information corresponding to each new multi-modal training data to obtain corresponding first fusion feature information, and fusing the second multi-dimensional feature information corresponding to each historical multi-modal training data to obtain corresponding second fusion feature information; and deduplicating the multiple new multi-modal training data based on the first fusion feature information corresponding to each new multi-modal training data and the second fusion feature information corresponding to each historical multi-modal training data. Description of the Drawings
[0016] To more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the accompanying drawings required for the description of the embodiments or the prior art. Obviously, the accompanying drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can be obtained based on these drawings.
[0017] Figure 1 is the architecture diagram of the training data deduplication system provided by the embodiments of the present invention.
[0018] Figure 2 is one of the flow schematic diagrams of the training data deduplication method provided by the embodiments of the present invention.
[0019] Figure 3 is the second flow schematic diagram of the training data deduplication method provided by the embodiments of the present invention.
[0020] Figure 4 is the structural schematic diagram of the training data deduplication device provided by the embodiments of the present invention.
[0021] Figure 5 is the structural schematic diagram of the electronic device provided by the embodiments of the present invention. Detailed Embodiments
[0022] To make the objectives, technical solutions and advantages of the present invention clearer, the following will clearly and completely describe the technical solutions in the present invention with reference to the accompanying drawings in the present invention. Obviously, the described embodiments are some but not all of the embodiments of the present invention. Based on the embodiments in the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the protection scope of the present invention.
[0023] The terms "first", "second", etc. in the present invention are used to distinguish similar objects, rather than to describe a specific order or sequence. It should be understood that such terms can be interchanged under appropriate circumstances so that the embodiments of the present invention can be implemented in an order other than those illustrated or described herein, and the objects distinguished by "first" and "second" are generally of the same category, and the number of objects is not limited. For example, the first object can be one or multiple. In addition, the "and / or" in the present invention represents at least one of the connected objects, and the character " / " generally represents an "or" relationship between the associated objects before and after.
[0024] Figure 1 is the architecture diagram of the training data deduplication system provided by the embodiments of the present invention. As Figure 1As shown in the figure, a training data deduplication system is provided. The system includes an application service layer, an algorithm model layer, and a data storage layer. The application service layer is used to deduplicate multiple new multi-modal training data of computer vision algorithms, and remove duplicate or similar new multi-modal training data. The algorithm model layer is used to perform feature processing on multiple new multi-modal training data and multiple historical multi-modal training data to obtain corresponding feature data. The data storage layer includes an object storage database and a vector database. The object database is used to store multiple new multi-modal training data and multiple historical multi-modal training data, and the vector database is used to store the corresponding feature data.
[0025] Optionally, the algorithm model layer at least includes: Message-Digest Algorithm 5 (MD5), Secure Hash Algorithm (SHA), algorithm models such as Residual Network (ResNet) and Visual Geometry Group (VGG).
[0026] Figure 2 One of the flow diagrams of the training data deduplication method provided by the embodiments of the present invention. As Figure 2 shown, a training data deduplication method is provided. Taking the application of this training data deduplication method to the Figure 1 system in the figure as an example, it includes the following steps: Step 210, Step 220, Step 230, and Step 240. The steps of this method flow are only a possible implementation manner of the present invention.
[0027] Step 210: Obtain multiple new multi-modal training data, and obtain multiple historical multi-modal training data from the sample database; Step 220: Extract the first multi-dimensional feature information corresponding to each new multi-modal training data, and extract the second multi-dimensional feature information corresponding to each historical multi-modal training data; Step 230: Fuse the first multi-dimensional feature information corresponding to each new multi-modal training data to obtain the first fusion feature information corresponding to each new multi-modal training data, and fuse the second multi-dimensional feature information corresponding to each historical multi-modal training data to obtain the second fusion feature information corresponding to each historical multi-modal training data; Step 240: Deduplicate multiple new multi-modal training data based on the first fusion feature information corresponding to each new multi-modal training data and the second fusion feature information corresponding to each historical multi-modal training data.
[0028] It should be noted that multiple new multimodal training data and multiple historical multimodal training data are both used to train the computer vision algorithm model.
[0029] Optionally, a one-way hash function (such as MD5, SHA, etc.) or a convolutional neural network (ResNet, VGG, etc.) is used to extract the first multi-dimensional feature information corresponding to each new multimodal training data and the second multi-dimensional feature information corresponding to each historical multimodal training data.
[0030] Optionally, if there is historical multimodal training data that is the same as or similar to each new multimodal training data in the sample database, then this new multimodal training data is deleted; if there is no historical multimodal training data that is the same as or similar to each new multimodal training data in the sample database, then this new multimodal training data is added to the sample database.
[0031] In some embodiments, each new multimodal training data includes each image sample data, as well as the text sample data and / or audio sample data associated with each image sample data; each historical multimodal training data includes each historical image sample data, as well as the historical text sample data and / or historical audio sample data associated with each historical image sample data; the first multi-dimensional feature information includes the first image feature, the first text feature, and / or the first audio feature; the second multi-dimensional feature information includes the second image feature, the second text feature, and / or the second audio feature.
[0032] Optionally, the text sample data associated with each image sample data includes data such as an image title, an image description, image annotation information, and the text content in the image.
[0033] Optionally, the historical text sample data associated with each historical image sample data includes data such as a historical image title, a historical image description, historical image annotation information, and the text content in the historical image.
[0034] Optionally, the first image feature includes, but is not limited to: color feature, texture feature, shape feature, and spatial relationship feature.
[0035] Optionally, the second image feature includes, but is not limited to: color feature, texture feature, shape feature, and spatial relationship feature.
[0036] In an embodiment of the present invention, by obtaining multiple pieces of new multimodal training data and multiple pieces of historical multimodal training data; extracting first multi-dimensional feature information corresponding to each piece of new multimodal training data, and extracting second multi-dimensional feature information corresponding to each piece of historical multimodal training data; fusing the first multi-dimensional feature information corresponding to each piece of new multimodal training data to obtain corresponding first fused feature information, and fusing the second multi-dimensional feature information corresponding to each piece of historical multimodal training data to obtain corresponding second fused feature information; based on the first fused feature information corresponding to each piece of new multimodal training data and the second fused feature information corresponding to each piece of historical multimodal training data, duplicate data in the multiple pieces of new multimodal training data is removed, improving the efficiency and accuracy of duplicate data removal for training data.
[0037] It should be noted that each embodiment of the present invention can be freely combined, the order can be swapped, or each can be executed independently, and does not need to rely on or depend on a fixed execution order.
[0038] In some embodiments, extracting the first multi-dimensional feature information corresponding to each piece of new multimodal training data includes: Performing feature extraction on each image sample data to obtain a first image feature corresponding to each image sample data; Performing feature extraction on the text sample data associated with each image sample data to obtain a first text feature corresponding to the text sample data associated with each image sample data, and / or performing feature extraction on the audio sample data associated with each image sample data to obtain a first audio feature corresponding to the audio sample data associated with each image sample data; Extracting the second multi-dimensional feature information corresponding to each piece of historical multimodal training data includes: Performing feature extraction on each historical image sample data to obtain a second image feature corresponding to each historical image sample data; Performing feature extraction on the historical text sample data associated with each historical image sample data to obtain a second text feature corresponding to the historical text sample data associated with each historical image sample data, and / or performing feature extraction on the historical audio sample data associated with each historical image sample data to obtain a second audio feature corresponding to the historical audio sample data associated with each historical image sample data.
[0039] Optionally, a convolutional neural network is used to perform feature extraction on each image sample data and each historical image sample data.
[0040] Optionally, for images with fewer pixels, MD5 is used to calculate the digital fingerprint information and 128-bit data is output. For images with more pixels, SHA-256 is used to calculate the digital fingerprint information and 256-bit data is output. For animated images and short videos, SHA-512 is used to calculate the digital fingerprint information and 512-bit digital fingerprint information is output.
[0041] Optionally, natural language processing technology is adopted to extract features from the text sample data associated with each image sample data and the historical text sample data associated with each historical image sample data.
[0042] Optionally, acoustic analysis technology is adopted to extract features from the audio sample data associated with each image sample data and the historical audio sample data associated with each historical image sample data.
[0043] In some embodiments, fusing the first multi-dimensional feature information corresponding to each new multi-modal training data to obtain the first fused feature information corresponding to each new multi-modal training data includes: Determining the weights of the first image feature, the first text feature, and / or the first audio feature corresponding to each new multi-modal training data; Based on the weights of the first image feature, the first text feature, and / or the first audio feature, fusing the first image feature, the first text feature, and / or the first audio feature corresponding to each new multi-modal training data to obtain the first fused feature information corresponding to each new multi-modal training data.
[0044] Optionally, determining the first association relationship of the first image feature, the first text feature, and / or the first audio feature corresponding to each new multi-modal training data.
[0045] Optionally, according to the first association relationship, determining the weights of the first image feature, the first text feature, and / or the first audio feature corresponding to each new multi-modal training data.
[0046] For example, the weight of the first image feature is 0.6, the weight of the first text feature is 0.3, and the weight of the first audio feature is 0.1.
[0047] In some embodiments, fusing the second multi-dimensional feature information corresponding to each historical multi-modal training data to obtain the second fused feature information corresponding to each historical multi-modal training data includes: Determining the weights of the second image feature, the second text feature, and / or the second audio feature corresponding to each historical multi-modal training data; Based on the weights of the second image features, the weights of the second text features, and / or the weights of the second audio features, fuse the second image features, the second text features, and / or the second audio features corresponding to each piece of historical multimodal training data to obtain the second fused feature information corresponding to each piece of historical multimodal training data.
[0048] Optionally, determine the second association relationships of the second image features, the second text features, and / or the second audio features corresponding to each piece of historical multimodal training data.
[0049] Optionally, based on the second association relationships, determine the weights of the second image features, the weights of the second text features, and / or the weights of the second audio features corresponding to each piece of historical multimodal training data.
[0050] For example, the weight of the second image feature is 0.6, the weight of the second text feature is 0.3, and the weight of the second audio feature is 0.1.
[0051] In some embodiments, based on the first fused feature information corresponding to each piece of new multimodal training data and the second fused feature information corresponding to each piece of historical multimodal training data, deduplicate the multiple pieces of new multimodal training data, including: Based on the first fused feature information corresponding to each piece of new multimodal training data and the second fused feature information corresponding to each piece of historical multimodal training data, calculate the similarity between each piece of new multimodal training data and each piece of historical multimodal training data; According to the similarity between each piece of new multimodal training data and each piece of historical multimodal training data and a preset similarity threshold, deduplicate the multiple pieces of new multimodal training data.
[0052] Optionally, based on the first fused feature information corresponding to each piece of new multimodal training data, calculate the similarity between different pieces of new multimodal training data pairwise.
[0053] Optionally, according to the similarity between different pieces of new multimodal training data pairwise and a pre-review similarity threshold, deduplicate the multiple pieces of new multimodal training data.
[0054] In some embodiments, the above method further includes: Store the multiple pieces of new multimodal training data and the multiple pieces of historical multimodal training data in an object storage database; Store the first multi-dimensional feature information corresponding to each piece of new multimodal training data and the second multi-dimensional feature information corresponding to each piece of historical multimodal training data in a first vector database; Store the first fused feature information corresponding to each piece of new multimodal training data and the second fused feature information corresponding to each piece of historical multimodal training data in a second vector database.
[0055] Optionally, the object storage database is a database such as Ceph or MinIO, and the first vector database and the second vector database are databases such as Milvus or Chroma.
[0056] It can be understood that by classifying and storing the original training data, the corresponding feature information, and the fused feature information, a fast response can be made to the high-concurrency retrieval of multiple users, which helps users to perform hierarchical retrieval on the training data and improves the user experience.
[0057] Figure 3 This is the second flowchart of the training data deduplication method provided by the embodiment of the present invention. As Figure 3 shown, in some embodiments, the above method further includes: Obtain the retrieval request of the user, determine the target retrieval method based on the retrieval request of the user, and retrieve each new multimodal training data according to the target retrieval method to obtain a retrieval result; The target retrieval method includes at least one of the following: Retrieve historical multimodal training data similar to each new multimodal training data from the object storage database; Retrieve second multi-dimensional feature information similar to the first multi-dimensional feature information from the first vector database; Retrieve second fused feature information similar to the first fused feature information from the second vector database.
[0058] Optionally, determine the mapping relationship between each new multimodal training data, the first multi-dimensional feature information, and the first fused feature information.
[0059] Optionally, determine the mapping relationship between each historical multimodal training data, the second multi-dimensional feature information, and the second fused feature information.
[0060] Optionally, retrieve each new multimodal training data according to the ID of each new multimodal training data.
[0061] Optionally, obtain the retrieval requests of multiple users, and determine the target retrieval method of each user based on the retrieval request of each user.
[0062] Optionally, perform parallel retrieval on each new multimodal training data corresponding to each user according to the target retrieval method of each user to obtain a retrieval result.
[0063] Next, the training data deduplication device provided by the embodiment of the present invention will be described. The training data deduplication device described below can be correspondingly referred to the training data deduplication method described above.
[0064] Figure 4The structural schematic diagram of the training data deduplication device provided by the embodiment of the present invention is as follows Figure 4 As shown, the training data deduplication device 400 includes: An acquisition unit 410, configured to acquire multiple new multimodal training data, and acquire multiple historical multimodal training data from a sample database; A feature extraction unit 420, configured to extract first multi-dimensional feature information corresponding to each new multimodal training data, and extract second multi-dimensional feature information corresponding to each historical multimodal training data; A feature fusion unit 430, configured to fuse the first multi-dimensional feature information corresponding to each new multimodal training data to obtain first fusion feature information corresponding to each new multimodal training data, and fuse the second multi-dimensional feature information corresponding to each historical multimodal training data to obtain second fusion feature information corresponding to each historical multimodal training data; A data deduplication unit 440, configured to deduplicate multiple new multimodal training data based on the first fusion feature information corresponding to each new multimodal training data and the second fusion feature information corresponding to each historical multimodal training data.
[0065] Optionally, each new multimodal training data includes each image sample data, and text sample data and / or audio sample data associated with each image sample data; each historical multimodal training data includes each historical image sample data, and historical text sample data and / or historical audio sample data associated with each historical image sample data; the first multi-dimensional feature information includes first image features, first text features, and / or first audio features; the second multi-dimensional feature information includes second image features, second text features, and / or second audio features.
[0066] Optionally, extracting the first multi-dimensional feature information corresponding to each new multimodal training data includes: Performing feature extraction on each image sample data to obtain first image features corresponding to each image sample data; Performing feature extraction on text sample data associated with each image sample data to obtain first text features corresponding to the text sample data associated with each image sample data, and / or performing feature extraction on audio sample data associated with each image sample data to obtain first audio features corresponding to the audio sample data associated with each image sample data; Extracting the second multi-dimensional feature information corresponding to each historical multimodal training data includes: Performing feature extraction on each historical image sample data to obtain second image features corresponding to each historical image sample data; Extract features from the historical text sample data associated with each historical image sample data to obtain the second text features corresponding to the historical text sample data associated with each historical image sample data, and / or extract features from the historical audio sample data associated with each historical image sample data to obtain the second audio features corresponding to the historical audio sample data associated with each historical image sample data.
[0067] Optionally, fuse the first multi-dimensional feature information corresponding to each new multi-modal training data to obtain the first fused feature information corresponding to each new multi-modal training data, including: Determine the weights of the first image features, the first text features, and / or the first audio features corresponding to each new multi-modal training data; Based on the weights of the first image features, the first text features, and / or the first audio features, fuse the first image features, the first text features, and / or the first audio features corresponding to each new multi-modal training data to obtain the first fused feature information corresponding to each new multi-modal training data.
[0068] Optionally, fuse the second multi-dimensional feature information corresponding to each historical multi-modal training data to obtain the second fused feature information corresponding to each historical multi-modal training data, including: Determine the weights of the second image features, the second text features, and / or the second audio features corresponding to each historical multi-modal training data; Based on the weights of the second image features, the second text features, and / or the second audio features, fuse the second image features, the second text features, and / or the second audio features corresponding to each historical multi-modal training data to obtain the second fused feature information corresponding to each historical multi-modal training data.
[0069] Optionally, based on the first fused feature information corresponding to each new multi-modal training data and the second fused feature information corresponding to each historical multi-modal training data, deduplicate multiple new multi-modal training data, including: Calculate the similarity between each new multi-modal training data and each historical multi-modal training data based on the first fused feature information corresponding to each new multi-modal training data and the second fused feature information corresponding to each historical multi-modal training data; Deduplicate multiple new multi-modal training data according to the similarity between each new multi-modal training data and each historical multi-modal training data and a preset similarity threshold.
[0070] Optionally, the training data deduplication device 400 further includes: An object storage unit for storing multiple new multimodal training data and multiple historical multimodal training data into an object storage database; A first vector storage unit for storing the first multi-dimensional feature information corresponding to each new multimodal training data and the second multi-dimensional feature information corresponding to each historical multimodal training data into a first vector database; A second vector storage unit for storing the first fusion feature information corresponding to each new multimodal training data and the second fusion feature information corresponding to each historical multimodal training data into a second vector database.
[0071] Optionally, the training data deduplication device 400 further includes: A retrieval unit for obtaining a user's retrieval request, determining a target retrieval method based on the user's retrieval request, and retrieving each new multimodal training data according to the target retrieval method to obtain a retrieval result; The target retrieval method includes at least one of the following: Retrieving historical multimodal training data similar to each new multimodal training data from the object storage database; Retrieving second multi-dimensional feature information similar to the first multi-dimensional feature information from the first vector database; Retrieving second fusion feature information similar to the first fusion feature information from the second vector database.
[0072] It should be noted here that the training data deduplication device provided in the embodiment of the present invention can implement all the method steps implemented in the above-mentioned training data deduplication method embodiment and can achieve the same technical effect. The same parts and beneficial effects as those in the method embodiment are not specifically described in this embodiment.
[0073] Figure 5 It is a schematic structural diagram of an electronic device provided in an embodiment of the present invention, as Figure 5As shown, the electronic device may include: a processor 510, a communications interface 520, a memory 530, and a communication bus 540. Among them, the processor 510, the communications interface 520, and the memory 530 complete mutual communication through the communication bus 540. The processor 510 may call logic instructions in the memory 530 to execute a method for deduplicating training data, and the method includes: obtaining multiple pieces of new multimodal training data, and obtaining multiple pieces of historical multimodal training data from a sample database; extracting first multi-dimensional feature information corresponding to each piece of new multimodal training data, and extracting second multi-dimensional feature information corresponding to each piece of historical multimodal training data; fusing the first multi-dimensional feature information corresponding to each piece of new multimodal training data to obtain first fused feature information corresponding to each piece of new multimodal training data, and fusing the second multi-dimensional feature information corresponding to each piece of historical multimodal training data to obtain second fused feature information corresponding to each piece of historical multimodal training data; based on the first fused feature information corresponding to each piece of new multimodal training data and the second fused feature information corresponding to each piece of historical multimodal training data, deduplicate the multiple pieces of new multimodal training data.
[0074] In addition, when the logic instructions in the above-mentioned memory 530 are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: various media such as a USB flash drive, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk, or an optical disc that can store program codes.
[0075] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. A person of ordinary skill in the art can understand and implement it without creative labor.
[0076] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on such an understanding, the essence of the above technical solution, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.
[0077] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for deduplicating training data, characterized in that, Including: Obtain multiple new multimodal training data, and obtain multiple historical multimodal training data from the sample database; Extract the first multi-dimensional feature information corresponding to each new multimodal training data, and extract the second multi-dimensional feature information corresponding to each historical multimodal training data; Fuse the first multi-dimensional feature information corresponding to each new multimodal training data to obtain the first fusion feature information corresponding to each new multimodal training data, and fuse the second multi-dimensional feature information corresponding to each historical multimodal training data to obtain the second fusion feature information corresponding to each historical multimodal training data; Deduplicate the multiple new multimodal training data based on the first fusion feature information corresponding to each new multimodal training data and the second fusion feature information corresponding to each historical multimodal training data.
2. The training data deduplication method according to claim 1, wherein Each new multimodal training data includes each image sample data, and the text sample data and / or audio sample data associated with each image sample data; each historical multimodal training data includes each historical image sample data, and the historical text sample data and / or historical audio sample data associated with each historical image sample data; the first multi-dimensional feature information includes first image features, first text features, and / or first audio features; the second multi-dimensional feature information includes second image features, second text features, and / or second audio features.
3. The training data deduplication method according to claim 2, wherein The extraction of the first multi-dimensional feature information corresponding to each new multimodal training data includes: Perform feature extraction on each image sample data to obtain the first image features corresponding to each image sample data; Perform feature extraction on the text sample data associated with each image sample data to obtain the first text features corresponding to the text sample data associated with each image sample data, and / or perform feature extraction on the audio sample data associated with each image sample data to obtain the first audio features corresponding to the audio sample data associated with each image sample data; The extraction of the second multi-dimensional feature information corresponding to each historical multimodal training data includes: Perform feature extraction on each historical image sample data to obtain the second image features corresponding to each historical image sample data; Perform feature extraction on the historical text sample data associated with each historical image sample data to obtain the second text features corresponding to the historical text sample data associated with each historical image sample data, and / or perform feature extraction on the historical audio sample data associated with each historical image sample data to obtain the second audio features corresponding to the historical audio sample data associated with each historical image sample data.
4. The training data deduplication method according to claim 2, wherein The fusion of the first multi-dimensional feature information corresponding to each new multimodal training data to obtain the first fusion feature information corresponding to each new multimodal training data includes: Determine the weights of the first image features, the weights of the first text features, and / or the weights of the first audio features corresponding to each new multimodal training data; Based on the weights of the first image features, the weights of the first text features, and / or the weights of the first audio features, fuse the first image features, the first text features, and / or the first audio features corresponding to each piece of the new multimodal training data to obtain the first fusion feature information corresponding to each piece of the new multimodal training data.
5. The training data deduplication method according to claim 2, wherein The fusing the second multi-dimensional feature information corresponding to each piece of the historical multimodal training data to obtain the second fusion feature information corresponding to each piece of the historical multimodal training data includes: Determine the weights of the second image features, the weights of the second text features, and / or the weights of the second audio features corresponding to each piece of the historical multimodal training data; Based on the weights of the second image features, the weights of the second text features, and / or the weights of the second audio features, fuse the second image features, the second text features, and / or the second audio features corresponding to each piece of the historical multimodal training data to obtain the second fusion feature information corresponding to each piece of the historical multimodal training data.
6. The training data deduplication method according to claim 2, wherein The deduplicating the multiple pieces of new multimodal training data based on the first fusion feature information corresponding to each piece of the new multimodal training data and the second fusion feature information corresponding to each piece of the historical multimodal training data includes: Based on the first fusion feature information corresponding to each piece of the new multimodal training data and the second fusion feature information corresponding to each piece of the historical multimodal training data, calculate the similarity between each piece of the new multimodal training data and each piece of the historical multimodal training data; According to the similarity between each piece of the new multimodal training data and each piece of the historical multimodal training data and a preset similarity threshold, deduplicate the multiple pieces of new multimodal training data.
7. The training data deduplication method according to claim 1, wherein The method further includes: Store the multiple pieces of new multimodal training data and multiple pieces of historical multimodal training data into an object storage database; Store the first multi-dimensional feature information corresponding to each piece of the new multimodal training data and the second multi-dimensional feature information corresponding to each piece of the historical multimodal training data into a first vector database; Store the first fusion feature information corresponding to each piece of the new multimodal training data and the second fusion feature information corresponding to each piece of the historical multimodal training data into a second vector database.
8. The training data deduplication method according to claim 7, wherein The method further includes: Obtain a retrieval request of a user, determine a target retrieval method based on the retrieval request of the user, and according to the target retrieval method, retrieve each piece of the new multimodal training data to obtain a retrieval result; The target retrieval method includes at least one of the following: Retrieve historical multimodal training data similar to each piece of the new multimodal training data from the object storage database; Retrieve second multi-dimensional feature information similar to the first multi-dimensional feature information from the first vector database; Retrieve second fusion feature information similar to the first fusion feature information from the second vector database.
9. A training data deduplication device, characterized in that, Includes: An obtaining unit, configured to obtain multiple pieces of new multimodal training data and obtain multiple pieces of historical multimodal training data from a sample database; A feature extraction unit, configured to extract first multi-dimensional feature information corresponding to each new multi-modal training data, and extract second multi-dimensional feature information corresponding to each historical multi-modal training data; A feature fusion unit, configured to fuse the first multi-dimensional feature information corresponding to each new multi-modal training data to obtain first fused feature information corresponding to each new multi-modal training data, and fuse the second multi-dimensional feature information corresponding to each historical multi-modal training data to obtain second fused feature information corresponding to each historical multi-modal training data; A data deduplication unit, configured to deduplicate the multiple new multi-modal training data based on the first fused feature information corresponding to each new multi-modal training data and the second fused feature information corresponding to each historical multi-modal training data.
10. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and running on the processor, characterized in that, When the processor executes the computer program, it implements the training data deduplication method according to any one of claims 1 to 8.