Semantic recognition-based training data redundancy elimination method and system for large model
By building a domain knowledge graph and combining a multi-level coding mechanism, identifying and removing redundant information from the training data, the problem of data redundancy and quality in the existing technology is solved, and the data set quality improvement and model training efficiency are improved.
Patent Information
- Application Number
- CN202510472445.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-16
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-04-16
AI Technical Summary
The prior art has shortcomings in semantic recognition and data cleaning, and it is difficult to effectively identify and remove redundant information from training data, affecting the generalization ability of the model, and the traditional manual labeling and auditing mechanism is inefficient and easy to introduce human errors.
By building a domain knowledge graph, combining semantic association mining, data preprocessing and multi-level coding mechanisms, redundant data can be identified and removed, and data governance processes can be optimized through data quality evaluation.
Effectively optimize the data governance process, improve the quality and representativeness of data sets, improve the training efficiency and accuracy of deep learning models, and ensure that the model achieves its maximum potential when processing complex data.
Smart Images

Figure CN119988842A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of natural language processing technology, and in particular to a method and system for eliminating redundant training data of a large model based on semantic recognition. Background Technology
[0002] Against the backdrop of the rapid development of information technology today, artificial intelligence (AI) and machine learning (ML) have become important forces driving innovation in all walks of life. Especially in the field of natural language processing (NLP), large models based on semantic recognition have gradually emerged as a powerful tool for understanding human language, extracting semantic information, and generating natural language text. In recent years, with the advancement of deep learning technology, researchers have made significant progress in the construction and application of knowledge graphs. Knowledge graphs can effectively capture and organize domain knowledge by representing entities and relationships, providing strong semantic support for downstream tasks. At the same time, in the training process of large models, the quality and diversity of data have an increasingly obvious impact on model performance. Therefore, how to effectively eliminate redundant data and improve data quality has become a research hotspot.
[0003] Although current technology has achieved certain achievements in semantic recognition and data cleaning, there are still many shortcomings. Existing data preprocessing methods mostly focus on simple deduplication or cleaning based on specific rules, and lack a deep understanding of the intrinsic semantic structure of the data. As a result, in the process of data redundancy elimination, some redundant information due to semantic similarity may be missed, which in turn affects the generalization ability of the model. In addition, when faced with large-scale and diversified data sets, traditional manual annotation and review mechanisms are inefficient, prone to human errors, and reduce data quality. Therefore, there is an urgent need for a more efficient redundant data elimination method that can combine advanced semantic recognition and knowledge graph technology to comprehensively analyze and process large-scale data. SUMMARY OF THE INVENTION
[0004] In view of the above existing problems, the present invention is proposed.
[0005] Therefore, the present invention provides a method for eliminating redundant training data of a large model based on semantic recognition, which can have significant advantages in enriching semantic understanding, improving data quality and optimizing model training efficiency, and effectively solves the problems of data redundancy and substandard quality in traditional methods.
[0006] In order to solve the above technical problems, the present invention provides the following technical solutions: a method for eliminating redundant training data based on a large model of semantic recognition, comprising: constructing a domain knowledge graph, performing semantic association mining based on the domain knowledge graph and identifying the redundant parts of the data in the domain knowledge graph; removing irrelevant data in combination with semantic recognition technology during the data preprocessing stage of constructing the domain knowledge graph; after identifying the redundant parts in the data, removing the redundant data through a data cleaning algorithm; and optimizing the data governance process in combination with data quality assessment.
[0007] As a preferred solution of the method for eliminating redundant training data of a large model based on semantic recognition described in the present invention, wherein: the construction of a domain knowledge graph includes collecting domain professional literature and large-scale data sets; combining a domain vocabulary to vectorize the content of domain professional literature; based on content recognition, using a graph attention network model to extract the relationship between the contents of the quantitative representation;
[0008] Run the GNN-based entity alignment algorithm, review and adjust the alignment results, use the meta-path-based similarity measurement method to verify and optimize the accuracy of entity alignment, implement multimodal information fusion technology, integrate and analyze relationship information from different sources, integrate the aligned entities and fused relationships into a fine-grained knowledge graph, and perform periodic updates and maintenance.
[0009] As a preferred solution of the method for eliminating redundant training data of a large model based on semantic recognition described in the present invention, wherein: the semantic association mining includes analyzing multi-hop relationships between entities using a graph neural network; extracting entities and relationships from the data of the domain knowledge graph to generate a graph G structure including nodes V and edges E,
[0010] Select the graph neural network architecture and initialize the feature vector of each node , get the initial feature matrix of all nodes , the node feature vector is updated through multiple rounds of propagation: ;
[0011] Among them, is the feature vector of node i after the kth round of transmission, is the feature vector of node i after the k-1th round of transmission, is the set of neighbor nodes of node i, is an aggregation function that combines the features of node i's neighbor node j together, Used to update the characteristic function of node i, is the feature vector of node i’s neighbor node j after the k-1th round of transmission;
[0012] Through multiple rounds of propagation updates, we finally get the high-dimensional feature vector of the node. The high-dimensional feature vector of each node contains the information of multi-hop neighbors.
[0013] As a preferred solution of the method for eliminating redundant training data of a large model based on semantic recognition described in the present invention, wherein: the semantic association mining also includes identifying redundant instances in the data by calculating semantic similarity, for any two entities and , calculate their similarity, similarity is defined as: ;
[0014] Among them, Is Entity and The similarity score between , Set the similarity threshold for the feature vector of node i's neighbor node j after the kth round of transmission , when , think and is redundant.
[0015] As a preferred solution of the method for eliminating redundant training data of a large model based on semantic recognition described in the present invention, wherein: the elimination of irrelevant data includes unifying the format of the original data obtained from different data sources, removing obvious errors and missing values, and storing them in a database;
[0016] Using a multi-level coding mechanism, the original data is divided into three levels for coding:
[0017] Extract the basic features of the original data to generate basic vectors; use the pre-trained deep learning model to perform contextual semantic encoding on the text or image in the original data and convert it into a high-dimensional semantic vector; calculate the similarity between high-dimensional semantic vectors based on the knowledge graph in the target field, generate a weighted semantic feature vector, and map each data point in the original data to a feature space with higher explanatory power;
[0018] When the original data enters the encoding stage, it will be matched with the knowledge graph to generate high-level semantic features;
[0019] Perform a neighborhood search for each data point. If the number of points in the neighborhood of the current data point, including itself, is greater than the density threshold, it is considered a core point; otherwise, it is considered an edge point; the core point and its neighboring points are grouped into the same cluster. If a core point can reach another core point, the points adjacent to the two core points will also be merged into the same cluster;
[0020] Mark all edge points and unclassified points as outliers, that is, data points that cannot be classified into any cluster formed by core points, and are marked outside the specified density and size group;
[0021] Randomly sample the automatically marked data points according to the proportion or review them according to the number of outliers. After the review is completed, enter the modification results and update the marking database.
[0022] As a preferred solution of the method for eliminating redundant training data of a large model based on semantic recognition described in the present invention, wherein: the redundant data removal includes generating an identifier for each data point, marking the data point as "normal" or "redundant", creating a tag set M, ;
[0023] Among them, is the marked data point, represents a redundant data set. According to the tag set M, an operation to remove redundant data is defined: ;
[0024] Here, is to mark redundant data points, is the data set after removing redundant data, represents the difference operation of the set, which means removing all data points marked as redundant from D.
[0025] As a preferred solution of the method for eliminating redundant training data of a large model based on semantic recognition described in the present invention, wherein: the combined data quality assessment includes: evaluating the data set after eliminating redundant data Perform quality assessment and define the quality assessment function Q to measure the quality of the retained data in the dataset: ;
[0026] Among them, is a marker data point The weights of are dynamically adjusted by selecting relevance based on the features: ;
[0027] Among them, is the mean of the features corresponding to the current data set, which is used to reflect the overall distribution of all data points, so that the feature weights can be adjusted in real time according to the actual data distribution. is the eigenvector.
[0028] As a preferred solution of the training data redundancy elimination system based on the semantic recognition large model described in the present invention, it includes: a knowledge graph construction module, a semantic association mining module, a data preprocessing module, a redundant data identification and elimination module, and a data quality assessment and optimization module;
[0029] The knowledge graph construction module is responsible for constructing a domain knowledge graph to provide basic data and semantic references for subsequent semantic analysis and association mining;
[0030] The semantic association mining module uses the constructed knowledge graph to analyze the relationship between entities in the data and mine semantic information;
[0031] The data preprocessing module cleans and formats the raw data to ensure the quality and consistency of the data and provide reliable input for subsequent analysis;
[0032] The redundant data identification and elimination module identifies and eliminates redundant data to improve the simplicity and effectiveness of the data set;
[0033] The data quality assessment and optimization module, after removing redundant data, conducts quality assessment on the data set to ensure the integrity and accuracy of the data, while optimizing the data governance process.
[0034] Beneficial effects of the present invention: This method effectively optimizes the data governance process by constructing a domain knowledge graph, combining semantic association mining, and implementing data cleaning and quality assessment. This technical solution can improve the quality and representativeness of the data set while identifying and eliminating redundant data, thereby improving the training efficiency and accuracy of the deep learning model. In addition, the present invention effectively integrates information from multiple data sources and uses graph neural network technology to strengthen the mining of semantic relationships, ensuring that the model can maximize its potential when processing complex data. Brief Description of the Figures
[0035] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for describing the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work. Among them:
[0036] Figure 1 A flow chart of a method for eliminating redundant training data of a large model based on semantic recognition provided by an embodiment of the present invention.
[0037] Figure 2 A schematic diagram of the working modules of a system for eliminating redundant training data based on a large model of semantic recognition provided by an embodiment of the present invention. Specific implementation method
[0038] In order to make the above-mentioned purposes, features and advantages of the present invention more obvious and easy to understand, the specific implementation methods of the present invention are described in detail below in conjunction with the drawings of the specification. Obviously, the examples described are part of the embodiments of the present invention, but not all of them. Based on the embodiments of the present invention, all other embodiments obtained by ordinary persons in the art without creative work should fall within the scope of protection of the present invention.
[0039] In the following description, many specific details are set forth to facilitate a full understanding of the present invention, but the present invention can also be implemented in other ways different from the description, and those skilled in the art can make similar generalizations without violating the connotation of the present invention, so the present invention is not limited to the specific embodiments disclosed below.
[0040] Secondly, the term "one embodiment" or "embodiment" as used herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the present invention. The term "in one embodiment" that appears in different places in this specification does not necessarily refer to the same embodiment, nor is it a separate or selective embodiment that is mutually exclusive with other embodiments.
[0041] The present invention is described in detail with reference to the schematic diagram. When describing the embodiments of the present invention, for the sake of convenience, the cross-sectional diagram showing the device structure will not be partially enlarged according to the general scale, and the schematic diagram is only an example, which should not limit the scope of protection of the present invention. In addition, in actual production, the three-dimensional space dimensions of length, width and depth should be included.
[0042] At the same time, in the description of the present invention, it should be noted that the directions or positional relationships indicated by the terms "upper, lower, inner and outer" are based on the directions or positional relationships shown in the drawings, and are only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific direction, be constructed and operated in a specific direction, and therefore cannot be understood as limiting the present invention. In addition, the terms "first, second or third" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance.
[0043] Unless otherwise clearly specified and limited, the terms "install, connect, connect" in the present invention should be understood in a broad sense, for example: it can be a fixed connection, a detachable connection or an integral connection; it can also be a mechanical connection, an electrical connection or a direct connection, or it can be an indirect connection through an intermediate medium, or it can be the internal connection of two components. For ordinary technicians in this field, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.
[0044] Example 1, refer to Figure 1, which is the first embodiment of the present invention, and provides a method for eliminating redundant training data of a large model based on semantic recognition, including:
[0045] S1: Construct a domain knowledge graph, perform semantic association mining based on the domain knowledge graph, and identify redundant parts of the data in the domain knowledge graph.
[0046] Furthermore, the construction of the domain knowledge graph includes collecting domain professional literature and large-scale data sets; combining the domain vocabulary to vectorize the content of the domain professional literature; based on content recognition, using the graph attention network model to extract the relationship between the quantitatively represented content;
[0047] Run the GNN-based entity alignment algorithm, review and adjust the alignment results, use the meta-path-based similarity measurement method to verify and optimize the accuracy of entity alignment, implement multimodal information fusion technology, integrate and analyze relationship information from different sources, integrate the aligned entities and fused relationships into a fine-grained knowledge graph, and perform periodic updates and maintenance.
[0048] Furthermore, we use graph neural networks to analyze multi-hop relationships between entities; extract entities and relationships from the data of the domain knowledge graph to generate the structure of the graph G, which includes nodes V and edges E.
[0049] Select the graph neural network architecture and initialize the feature vector of each node , get the initial feature matrix of all nodes , the node feature vector is updated through multiple rounds of propagation: ;
[0050] Among them, is the feature vector of node i after the kth round of transmission, is the feature vector of node i after the k-1th round of transmission, is the set of neighbor nodes of node i, is an aggregation function that combines the features of node i's neighbor node j together, Used to update the characteristic function of node i, is the feature vector of node i’s neighbor node j after the k-1th round of transmission;
[0051] Through multiple rounds of propagation updates, we finally get the high-dimensional feature vector of the node. The high-dimensional feature vector of each node contains the information of multi-hop neighbors.
[0052] Through semantic similarity calculation, redundant instances in the data are identified. For any two entities and , calculate their similarity, the similarity is defined as: ;
[0053] Among them, Is Entity and The similarity score between , Set the similarity threshold for the feature vector of node i's neighbor node j after the kth round of transmission , when , think and is redundant.
[0054] It should be noted that the similarity score here is cosine similarity. When When =1, it means that the two vectors have the same direction, which means they are very similar.
[0055] When When =0, it means the angle between the two vectors is 90, indicating that there is no similarity between them.
[0056] When When =-1, it means the directions of the two vectors are completely opposite.
[0057] It should be noted that by building a knowledge graph, a clear domain knowledge framework is formed, which helps to understand the complex relationships in the field. With the help of the Graph Attention Network (GAT), the relationship between entities can be more accurately identified and extracted, improving data integration capabilities. By periodically updating and maintaining the knowledge graph, the identification of redundant data is ensured to be real-time and accurate.
[0058] S2: In the data preprocessing stage of constructing the domain knowledge graph, semantic recognition technology is used to eliminate irrelevant data.
[0059] Furthermore, the raw data obtained from different data sources are unified in format, obvious errors and missing values are removed, and stored in the database;
[0060] Using a multi-level coding mechanism, the original data is divided into three levels for coding:
[0061] Extract the basic features of the original data to generate basic vectors; use the pre-trained deep learning model to perform contextual semantic encoding on the text or image in the original data and convert it into a high-dimensional semantic vector; calculate the similarity between high-dimensional semantic vectors based on the knowledge graph in the target field, generate a weighted semantic feature vector, and map each data point in the original data to a feature space with higher explanatory power;
[0062] When the original data enters the encoding stage, it will be matched with the knowledge graph to generate high-level semantic features;
[0063] Perform a neighborhood search for each data point. If the number of points in the neighborhood of the current data point, including itself, is greater than the density threshold, it is considered a core point; otherwise, it is considered an edge point; the core point and its neighboring points are grouped into the same cluster. If a core point can reach another core point, the points adjacent to the two core points will also be merged into the same cluster;
[0064] Mark all edge points and unclassified points as outliers, that is, data points that cannot be classified into any cluster formed by core points, and mark data points outside the specified density and size group;
[0065] Randomly sample the automatically marked data points according to the proportion or review them according to the number of outliers. After the review is completed, enter the modification results and update the marking database.
[0066] It should be noted that by unifying the format and removing errors, the quality of the data is significantly improved, which is convenient for subsequent analysis. A multi-level encoding mechanism is adopted to convert the original data into higher-dimensional semantic features, which improves the performance of subsequent models. Through the density clustering method, core points and outliers are accurately identified, which is convenient for further data processing and cleaning.
[0067] S3: After identifying the redundant parts in the data, remove the redundant data through the data cleaning algorithm.
[0068] Furthermore, generate an identifier for each data point, mark the data point as "normal" or "redundant", and create a label set M, ;
[0069] Among them, is the marked data point, represents a redundant data set. According to the tag set M, an operation to remove redundant data is defined: ;
[0070] Here, is to mark redundant data points, is the data set after removing redundant data, represents the difference operation of the set, which means removing all data points marked as redundant from D.
[0071] It should be noted that an identifier is generated for each data point to facilitate management and tracking of data status, ensuring the validity and reliability of the data. Set operations are used to quickly eliminate redundant data, reduce the size of the data set, and improve subsequent processing efficiency.
[0072] S4: Combine data quality assessment to optimize data governance processes.
[0073] Furthermore, the combined data quality assessment includes: Perform quality assessment and define the quality assessment function Q to measure the quality of the retained data in the dataset: ;
[0074] Among them, is a marker data point The weights of are dynamically adjusted by selecting relevance based on the features: ;
[0075] Among them, is the mean of the features corresponding to the current data set, which is used to reflect the overall distribution of all data points, so that the feature weights can be adjusted in real time according to the actual data distribution. is the eigenvector.
[0076] Furthermore, establish a user feedback channel, collect application feedback on data usage, and adjust data governance strategies in a timely manner. Classify and mark user feedback for subsequent analysis. Feedback content can be divided into "data accuracy issues", "data missing feedback", "user experience suggestions", etc., to facilitate the team to identify priority issues. Regularly generate reports on data governance effectiveness and quality assessment for management decision-making reference, and promote continuous improvement of data governance processes.
[0077] It should be noted that the integrity and validity of the final data set are ensured through quality assessment, providing guarantees for the use of data. By updating feature weights in real time, the data governance process is made flexible and adaptable, and can respond to changes in the environment and data in a timely manner.
[0078] Example 2, the second embodiment of the present invention, is different from the previous embodiment in that:
[0079] If the functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for a computer device (which can be a personal computer, server, or network device, etc.) to perform all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage media include: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk, and other media that can store program codes.
[0080] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as an ordered list of executable instructions for implementing logical functions, and can be specifically implemented in any computer-readable medium for use by an instruction execution system, device or apparatus (such as a computer-based system, a system including a processor, or other system that can fetch instructions from an instruction execution system, device or apparatus and execute instructions), or in combination with these instruction execution systems, devices or apparatuses. For the purpose of this specification, "computer-readable medium" can be any device that can contain, store, communicate, propagate or transmit a program for use by an instruction execution system, device or apparatus, or in combination with these instruction execution systems, devices or apparatuses.
[0081] More specific examples of computer-readable media (a non-exhaustive list) include the following: an electrical connection with one or more wirings (electronic devices), a portable computer disk case (magnetic devices), a random access memory (RAM), a read-only memory (ROM), an erasable and programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disk read-only memory (CDROM). In addition, the computer-readable medium may even be a paper or other suitable medium on which the program is printed, since the program may be obtained electronically, for example, by optically scanning the paper or other medium, followed by editing, deciphering or, if necessary, processing in another suitable manner, and then stored in a computer memory.
[0082] It should be understood that various parts of the present invention can be implemented in hardware, software, firmware or a combination thereof. In the above embodiments, multiple steps or methods can be implemented in software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, it can be implemented by any one of the following technologies known in the art or their combination: a discrete logic circuit with a logic gate circuit for implementing a logic function on a data signal, a dedicated integrated circuit with a suitable combination of logic gate circuits, a programmable gate array (PGA), a field programmable gate array (FPGA), etc.
[0083] Example 3, refer to Figure 2 , which is an embodiment of the present invention, provides a training data redundancy elimination system based on a large model of semantic recognition, including a knowledge graph construction module, a semantic association mining module, a data preprocessing module, a redundant data identification and elimination module, and a data quality assessment and optimization module;
[0084] The knowledge graph construction module is responsible for constructing a domain knowledge graph to provide basic data and semantic reference for subsequent semantic analysis and association mining;
[0085] The semantic association mining module uses the constructed knowledge graph to analyze the relationship between entities in the data and mine semantic information;
[0086] The data preprocessing module cleans and formats the raw data to ensure the quality and consistency of the data and provide reliable input for subsequent analysis;
[0087] The redundant data identification and elimination module identifies and eliminates redundant data to improve the simplicity and effectiveness of the data set;
[0088] The data quality assessment and optimization module, after removing redundant data, conducts quality assessment on the data set to ensure the integrity and accuracy of the data, while optimizing the data governance process.
[0089] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention rather than to limit it. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present invention may be modified or replaced by equivalents without departing from the spirit and scope of the technical solutions of the present invention, which should all be included in the scope of the claims of the present invention.
Claims
1. A method for eliminating redundant training data of a large model based on semantic recognition, characterized in that: include, Construct a domain knowledge graph, perform semantic association mining based on the domain knowledge graph, and identify redundant parts of data in the domain knowledge graph; In the data preprocessing stage of constructing the domain knowledge graph, semantic recognition technology is used to eliminate irrelevant data; After identifying the redundant parts in the data, the redundant data is removed through the data cleaning algorithm; Combine data quality assessment to optimize data governance processes.
2. The method for eliminating redundant training data of a large model based on semantic recognition as claimed in claim 1, characterized in that: The construction of the domain knowledge graph includes collecting domain professional literature and large-scale data sets; Combined with the domain vocabulary, the content of domain professional literature is vectorized; Based on content recognition, the graph attention network model is used to extract and quantify the relationship between the contents. Run the GNN-based entity alignment algorithm, review and adjust the alignment results, use the meta-path-based similarity measurement method to verify and optimize the accuracy of entity alignment, implement multimodal information fusion technology, integrate and analyze relationship information from different sources, integrate the aligned entities and fused relationships into a fine-grained knowledge graph, and perform periodic updates and maintenance.
3. The method for eliminating redundant training data of a large model based on semantic recognition as claimed in claim 2, characterized in that: The semantic association mining includes analyzing the multi-hop relationship between entities using a graph neural network; extracting entities and relationships from the data of the domain knowledge graph to generate a graph G structure including nodes V and edges E. Select the graph neural network architecture and initialize the feature vector of each node , get the initial feature matrix of all nodes , the feature vector of the node is updated through multiple rounds of propagation: ; in, is the feature vector of node i after the kth round of transmission, is the feature vector of node i after the k-1th round of transmission, is the set of neighbor nodes of node i, is an aggregation function that integrates the features of node i’s neighbor node j together. The characteristic function used to update node i, is the feature vector of node i’s neighbor node j after the k-1th round of transmission; Through multiple rounds of propagation updates, we finally get the high-dimensional feature vector of the node. The high-dimensional feature vector of each node contains the information of multi-hop neighbors.
4. The method for eliminating redundant training data of a large model based on semantic recognition as claimed in claim 3, characterized in that: The semantic association mining also includes identifying redundant instances in the data by calculating semantic similarity. and , calculate their similarity, the similarity is defined as: ; in, For Entity and The similarity score between Set the similarity threshold for the feature vector of node i’s neighbor node j after the kth round of transmission ,when When and is redundant.
5. The method for eliminating redundant training data of a large model based on semantic recognition as claimed in claim 4, characterized in that: Eliminating irrelevant data includes unifying the format of raw data obtained from different data sources, removing obvious errors and missing values, and storing them in a database; A multi-level coding mechanism is used to divide the original data into three levels for coding: Extract the basic features of the original data to generate basic vectors; use the pre-trained deep learning model to perform contextual semantic encoding on the text or image in the original data and convert it into a high-dimensional semantic vector; calculate the similarity between high-dimensional semantic vectors based on the knowledge graph in the target domain, generate a weighted semantic feature vector, and map each data point in the original data to a feature space with higher explanatory power; When the raw data enters the encoding stage, it will be matched with the knowledge graph to generate high-level semantic features; Perform a neighborhood search for each data point. If the number of points in the neighborhood of the current data point, including itself, is greater than the density threshold, it is considered a core point; otherwise, it is considered an edge point. The core point and its neighboring points are grouped into the same cluster. If a core point reaches another core point, the points adjacent to the two core points will also be merged into the same cluster. Mark all edge points and unclassified points as outliers, that is, data points that cannot be classified into any cluster formed by core points, and mark data points outside the specified density and size group; The automatically marked data points are randomly sampled in proportion or reviewed according to the number of outliers. After the review is completed, the modification results are entered and the marking database is updated.
6. The method for eliminating redundant training data of a large model based on semantic recognition as claimed in claim 5, characterized in that: The redundant data removal includes generating an identifier for each data point, marking the data point as "normal" or "redundant", creating a marker set M, ; in, are the marked data points, Represents a set of redundant data. According to the tag set M, an operation to remove redundant data is defined: ; here, To mark redundant data points, This is the dataset after removing redundant data. represents the difference operation of the set, which means removing all data points marked as redundant from D.
7. The method for eliminating redundant training data of a large model based on semantic recognition according to claim 6, characterized in that: The combined data quality assessment includes: Perform quality assessment and define the quality assessment function Q to measure the quality of the retained data in the dataset: ; in, is a labeled data point The weights are adjusted dynamically by selecting relevance based on the features: ; in, It is the mean of the features corresponding to the current data set, which is used to reflect the overall distribution of all data points, so that the feature weights can be adjusted in real time according to the actual data distribution. is the eigenvector.
8. A system for eliminating redundant training data of a large model based on semantic recognition, applied to a method for eliminating redundant training data of a large model based on semantic recognition as claimed in any one of claims 1 to 7, characterized in that: It includes knowledge graph construction module, semantic association mining module, data preprocessing module, redundant data identification and elimination module, and data quality assessment and optimization module; The knowledge graph construction module is responsible for constructing a domain knowledge graph to provide basic data and semantic references for subsequent semantic analysis and association mining; The semantic association mining module uses the constructed knowledge graph to analyze the relationship between entities in the data and mine semantic information; The data preprocessing module cleans and formats the raw data to ensure the quality and consistency of the data and provide reliable input for subsequent analysis; The redundant data identification and elimination module identifies and eliminates redundant data to improve the simplicity and effectiveness of the data set; The data quality assessment and optimization module performs quality assessment on the data set after removing redundant data to ensure the integrity and accuracy of the data, while optimizing the data governance process.
Citation Information
Patent Citations
Detection method for RDF data redundancy semantics
CN114692646A
Method for removing redundant information of public opinion information
CN117764074A
Intelligent medical inquiry method and system based on graph Transform
CN117954081A
Recommendation method and device based on knowledge graph, equipment and storage medium
CN119149697A
Ingredient recommendation method based on knowledge graph, and device and storage medium
WO2024140432A1
Cited By
Natural language understanding model training method and device, electronic equipment and storage medium
CN120748378A
Multi-modal large model incremental training data screening method
CN121145961A
Multimodal large model incremental training data screening method
CN121145961B