Adaptive filtering method for noise data in high-quality dataset construction

CN122527481APending Publication Date: 2026-08-07NANJING XIANWEI INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
NANJING XIANWEI INFORMATION TECH CO LTD
Filing Date
2026-07-07
Publication Date
2026-08-07

AI Technical Summary

Technical Problem

[0005]为解决现有技术中存在的边缘语义骨架易塌缩、跨场景关联结构断裂、AI模型推理精度降低的问题,本发明的目的在于解决上述缺陷,进而提出高质量数据集构建中的噪声数据自适应过滤方法

Benefits of technology

(1)本发明通过获取增量样本数据集合并进行场景分类划分与边缘关联分析,构建场景边缘关联网络;基于场景边缘关联网络将对应的增量样本数据集合进行串联,生成场景继承关联路径并识别骨架支撑区域,汇总得到边缘语义骨架集合;对边缘语义骨架集合进行稳定性分析与场景关联跨度检测,并通过冻结保护机制筛选得到长期保留映射结果;基于长期保留映射结果对噪声过滤边界进行重构,生成噪声过滤分层边界;采用差异化过滤策略根据噪声过滤分层边界对增量样本数据集合进行协同过滤调度,生成高质量数据集,从而在长周期处理过程中精准区分低频继承片段与无效噪声数据,完整保留边缘语义骨架并高效剔除无效噪声数据。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122527481A_ABST
    Figure CN122527481A_ABST
Patent Text Reader

Abstract

The application provides a noise data adaptive filtering method in high-quality data set construction, and relates to the technical field of artificial intelligence, which comprises the following steps: acquiring an incremental sample data set and performing scene classification and edge correlation analysis to construct a scene edge correlation network; based on the scene edge correlation network and the incremental sample data set, a scene inheritance correlation path is generated, a skeleton support area is identified, and an edge semantic skeleton set is obtained by summarizing; the edge semantic skeleton set is subjected to stability analysis and scene correlation span detection, and a long-term retention mapping result is screened; based on the long-term retention mapping result, a noise filtering boundary is reconstructed, and a noise filtering layered boundary is generated; and according to the noise filtering layered boundary, the incremental sample data set is subjected to cooperative filtering scheduling, and a high-quality data set is generated, so that low-frequency inheritance fragments and invalid noise data can be accurately distinguished in a long-period processing process, the edge semantic skeleton can be completely retained, and the invalid noise data can be efficiently removed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of artificial intelligence technology, and more specifically, to an adaptive filtering method for noisy data in the construction of high-quality datasets. Background Technology

[0002] In the intelligent transportation industry, various business scenarios, such as traffic construction, accident linkage, special dispatching, and extreme operations, are diverse and complex. In recent years, on-site business data based on incremental training has been widely used in dataset construction and AI model iteration. Through multiple rounds of data cleaning and noise removal, the quality of the dataset and the model training effect have been improved to some extent. However, in application scenarios with long-term incremental training and continuous expansion of industry knowledge bases, this model is limited by the traditional single noise judgment logic. Low-frequency data is uniformly treated as noise, and a large amount of low-frequency edge scene data is continuously isolated and removed in multiple rounds of filtering, resulting in the gradual damage to the semantic structure of the dataset. The reasoning ability of AI models in complex traffic scenarios cannot be guaranteed in the long term.

[0003] Existing technologies mainly rely on the frequency of data occurrence to set fixed filtering rules, resulting in a single dimension of judgment. This makes it difficult to accurately distinguish between effective noise and marginal effective data in long-term incremental data processing scenarios. Furthermore, during the filtering process, low-frequency key data with cross-scenario correlation characteristics are easily deleted, leading to problems such as semantic skeleton collapse and structural degradation of the dataset. This results in a decrease in the cross-scenario inference accuracy of the backend AI model, failing to meet the requirements of constructing high-quality training datasets for complex business scenarios in intelligent transportation.

[0004] Therefore, it is necessary to provide an adaptive filtering method for noisy data in the construction of high-quality datasets to solve the above-mentioned technical problems. Summary of the Invention

[0005] To address the problems of easy collapse of edge semantic skeletons, broken cross-scene association structures, and reduced inference accuracy of AI models in existing technologies, the present invention aims to solve the above defects and propose an adaptive filtering method for noisy data in the construction of high-quality datasets.

[0006] The present invention adopts the following technical solution.

[0007] This invention discloses an adaptive filtering method for noisy data in the construction of high-quality datasets, the method comprising: The incremental sample dataset is acquired and merged to perform scene classification and edge correlation analysis, and a scene edge correlation network is constructed. Based on the scene edge association network, the corresponding incremental sample data sets are concatenated to generate scene inheritance association paths and identify skeleton support regions, and the edge semantic skeleton set is obtained by summarizing them. The edge semantic skeleton set is subjected to stability analysis and scene association span detection, and long-term retained mapping results are obtained by filtering through the freeze protection mechanism; Based on the long-term retention mapping results, the noise filtering boundary is reconstructed to generate a layered noise filtering boundary. A differentiated filtering strategy is adopted to perform collaborative filtering scheduling on the incremental sample data set according to the noise filtering hierarchical boundary, thereby generating a high-quality dataset.

[0008] Preferably, the incremental sample dataset is acquired, merged, and used for scene classification to generate a scene-layered sample set, specifically including: The incremental sample dataset is semantically parsed using a scene classification algorithm. Multiple basic data fragments are extracted and input into the scene classification model. The scene affiliation probability of each of the multiple basic data fragments corresponding to different scene categories is calculated. Based on the scene affiliation probability, the scene categories corresponding to the multiple basic data segments are determined, and the multiple basic data segments are stratified and classified according to the scene categories, and the scene stratification sample set is generated.

[0009] Preferably, edge correlation analysis is performed on the scene hierarchical sample set to construct the scene edge correlation network, specifically including: A low-frequency data aggregation algorithm is used to perform frequency statistics and cross-scene reuse detection on the scene-layered samples in the scene-layered sample set, and the low-frequency inheritance strength of the scene-layered samples is calculated. A preset inheritance threshold is used to filter and integrate the scene-layered samples whose low-frequency inheritance strength is greater than or equal to the inheritance threshold, thereby generating a set of low-frequency inheritance fragments. A cross-scene relationship mining algorithm is used to analyze and identify the causal relationships, treatment continuity relationships, risk transmission relationships, and result inheritance relationships among multiple low-frequency inheritance fragments in the set of low-frequency inheritance fragments, and to calculate the edge association strength among multiple low-frequency inheritance fragments. A preset association threshold is set, and low-frequency inherited segments with edge association strength greater than or equal to the association threshold are retained. The retained low-frequency inherited segments are used as association nodes and the corresponding edge association strengths are used as association edges to construct the scene edge association network.

[0010] Preferably, the step of concatenating the corresponding incremental sample data sets based on the scene edge association network to generate scene inheritance association paths and identify skeleton support regions, and summarizing them to obtain an edge semantic skeleton set, specifically includes: The scene edge association network is traversed in depth-first order using a path extraction algorithm. Starting from the associated node, multiple associated nodes are sequentially connected along the associated edge in the order of the causal relationship, the treatment continuation relationship, the risk transmission relationship, and the result inheritance relationship to generate multiple scene inheritance association paths. The cross-scene connectivity capability of each scene inheritance association path is evaluated by the skeleton support evaluation model. The skeleton support strength of each scene inheritance association path is calculated. A skeleton support strength threshold is preset and the skeleton support strength is filtered. Multiple scene inheritance association paths with skeleton support strength greater than or equal to the skeleton support strength threshold are retained. The retained scene inheritance association paths are marked as the skeleton support region. Extract and integrate the skeleton support region, the associated nodes corresponding to the skeleton support region, the associated edges, the skeleton support strength, and the number of cross-scene categories to generate an edge semantic skeleton. Summarize the edge semantic skeletons of multiple skeleton support regions to generate the edge semantic skeleton set.

[0011] Preferably, the step of performing stability analysis and scene association span detection on the edge semantic skeleton set, and filtering out long-term retained mapping results through a freeze protection mechanism, specifically includes: The stability analysis of each edge semantic skeleton in the edge semantic skeleton set is performed by the skeleton stability evaluation model, and the periodic inheritance strength corresponding to each edge semantic skeleton is calculated. Based on the periodic inheritance strength, a scene association span detection algorithm is used to detect the scene range connected by each edge semantic skeleton, and the scene association span value corresponding to each edge semantic skeleton is calculated. A preset retention threshold is set, and the freezing protection mechanism is used to filter and protect the edge semantic skeleton set based on the periodic inheritance strength and the scene association span value. If the scene association span value is greater than or equal to the retention threshold, the corresponding edge semantic skeleton is marked as the long-term retention mapping result.

[0012] Preferably, the step of reconstructing the noise filtering boundary based on the long-term retention mapping result to generate a hierarchical noise filtering boundary specifically includes: Based on the long-term retention mapping result, a backtracking check is performed on the current noise filtering boundary to identify the overlap range between the current noise filtering boundary and the long-term retention mapping result, and the overlap ratio and conflict ratio are calculated using the overlap range. An adaptive shrinkage algorithm is used to reconstruct the noise filtering boundary based on the overlap ratio and the conflict ratio to generate the noise filtering layer boundary.

[0013] Preferably, the step of reconstructing the noise filtering boundary using an adaptive shrinkage algorithm based on the overlap ratio and the conflict ratio to generate the noise filtering layer boundary specifically includes: Preset protection boundary threshold and buffer boundary threshold; Obtain the noise filtering threshold corresponding to the noise filtering boundary; The overlap ratio and the conflict ratio are weighted and fused to calculate the boundary reconstruction strength; If the boundary reconstruction intensity is greater than or equal to the protection boundary threshold, then the noise filtering threshold corresponding to the noise filtering boundary is expanded to the boundary reconstruction intensity, and the protection range of the reconstructed noise filtering boundary is marked as the skeleton protection area. If the boundary reconstruction strength is less than the protection boundary threshold, and the boundary reconstruction strength is greater than or equal to the buffer boundary threshold, then the noise filtering threshold corresponding to the current noise filtering boundary is maintained, and the protection range of the noise filtering boundary is marked as the observation retention area.

[0014] If the boundary reconstruction intensity is less than the buffer boundary threshold, the noise filtering threshold corresponding to the noise filtering boundary is reduced to the boundary reconstruction intensity, and the protection range of the reconstructed noise filtering boundary is marked as a candidate filtering region.

[0015] Preferably, the step of employing a differentiated filtering strategy to collaboratively filter and schedule the incremental sample data set according to the noise filtering hierarchical boundary to generate a high-quality dataset specifically includes: Based on the noise filtering layer boundary, calculate the adaptive filtering intensity of multiple basic data segments in the incremental sample data set; The adaptive filtering intensity is matched with the noise filtering layer boundary to determine the region type to which the basic data segment corresponding to the adaptive filtering intensity belongs. A differentiated filtering strategy is used for collaborative filtering scheduling to integrate the retained basic data segments and generate the high-quality dataset.

[0016] An adaptive filtering system for noisy data in the construction of high-quality datasets, the system comprising: The scene classification and edge analysis module is used to acquire incremental sample datasets, merge them, perform scene classification and edge correlation analysis, and construct a scene edge correlation network. The skeleton support recognition and extraction module is used to concatenate the corresponding incremental sample data set based on the scene edge association network, generate scene inheritance association path and identify skeleton support region, and summarize to obtain edge semantic skeleton set. The stability assessment and freeze protection module is used to perform stability analysis and scene association span detection on the edge semantic skeleton set, and to filter the long-term retention mapping results through the freeze protection mechanism. The noise filtering boundary reconstruction module is used to reconstruct the noise filtering boundary based on the long-term retained mapping result, and generate a layered noise filtering boundary. The collaborative filtering scheduling module is used to perform collaborative filtering scheduling on the incremental sample data set according to the noise filtering hierarchical boundary using a differentiated filtering strategy, thereby generating a high-quality dataset.

[0017] The beneficial effects of the present invention are as follows: Compared with the prior art, the present invention has the following advantages: (1) This invention constructs a scene edge association network by merging incremental sample datasets to classify and analyze the scene and perform edge association analysis. Based on the scene edge association network, the corresponding incremental sample datasets are concatenated to generate scene inheritance association paths and identify skeleton support areas, and the edge semantic skeleton set is obtained. The edge semantic skeleton set is subjected to stability analysis and scene association span detection, and long-term retention mapping results are obtained by filtering through the freeze protection mechanism. Based on the long-term retention mapping results, the noise filtering boundary is reconstructed to generate a noise filtering hierarchical boundary. A differentiated filtering strategy is adopted to perform collaborative filtering scheduling on the incremental sample dataset according to the noise filtering hierarchical boundary to generate a high-quality dataset, thereby accurately distinguishing low-frequency inheritance fragments and invalid noise data in the long-term processing, completely preserving the edge semantic skeleton and efficiently removing invalid noise data.

[0018] (2) This invention identifies low-frequency inheritance fragments that recur across scenes through a low-frequency data aggregation algorithm, and constructs a scene edge association network by combining it with a cross-scene relationship mining algorithm. This accurately mines effective edge data that is difficult to identify using traditional frequency statistics methods, solving the problem of low-frequency inheritance fragments being misjudged as noise during multiple rounds of data cleaning. This invention identifies the set of edge semantic skeletons supporting multi-scene association transmission by extracting scene inheritance association paths and evaluating the skeleton support strength. It also uses a freeze protection mechanism to continuously retain edge semantic skeletons that have been stable for a long time and have a large cross-scene association span, avoiding the phenomenon of semantic skeleton collapse and knowledge chain breakage in the dataset during incremental training. At the same time, this invention dynamically reconstructs the noise filtering boundary based on the long-term retained mapping results, adaptively adjusts the filtering range according to the overlap ratio and conflict ratio, and adopts a hierarchical filtering strategy to differentiate the basic data fragments in different regions. This achieves the synergistic optimization of noise data removal and key semantic retention, solving the problem of insufficient adaptability of traditional fixed threshold filtering mechanisms. This improves the structural integrity of the incremental sample dataset, scene association ability, cross-scene inference accuracy of the backend AI model, and long-term iterative stability. Attached Figure Description

[0019] Figure 1 A flowchart of an adaptive filtering method for noisy data in the construction of a high-quality dataset provided in this embodiment of the invention; Figure 2 A system block diagram of an adaptive filtering system for noisy data in the construction of a high-quality dataset provided in an embodiment of the present invention. Detailed Implementation

[0020] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0021] like Figure 1 The diagram shown is a flowchart of the noisy data adaptive filtering method in the construction of a high-quality dataset provided in an embodiment of the present invention. Figure 1 The execution subject of the method shown can be a software and / or hardware device. The execution subject of this invention can include, but is not limited to, at least one of the following: user equipment, network equipment, etc. User equipment can include, but is not limited to, computers, smartphones, personal digital assistants (PDAs), and the aforementioned electronic devices. Network equipment can include, but is not limited to, a single network server, a server group consisting of multiple network servers, or a cloud based on cloud computing consisting of a large number of computers or network servers. Cloud computing is a type of distributed computing, consisting of a super virtual computer composed of a group of loosely coupled computers. This embodiment does not limit this. Steps S1 to S5 are detailed below: Step S1: Obtain incremental sample datasets, merge them, perform scene classification and edge association analysis, and construct scene edge association network.

[0022] The incremental sample datasets are acquired, merged, and used for scene classification to generate a scene-stratified sample set, specifically including: The incremental sample dataset is semantically parsed using a scene classification algorithm. Multiple basic data fragments are extracted and input into the scene classification model. The scene affiliation probability of each of the multiple basic data fragments corresponding to different scene categories is calculated. Based on the scene affiliation probability, the scene categories corresponding to the multiple basic data segments are determined, and the multiple basic data segments are stratified and classified according to the scene categories, and the scene stratification sample set is generated.

[0023] The incremental sample dataset refers to the set of data samples newly collected, labeled, or accessed during the continuous iteration of the high-quality dataset. This can include text data, image data, video data, audio data, log data, or multimodal fusion data, specifically including traffic construction records, accident handling logs, special dispatch instructions, and extreme operating condition reports. In scenarios involving long-term incremental training and continuous expansion of the industry knowledge base, the incremental sample dataset enters the high-quality dataset construction system in multiple rounds and from multiple sources. This results in high data heterogeneity and blurred scene boundaries. If a unified filter is directly applied to the incremental sample dataset, low-frequency inherited fragments that occur frequently and are related across scenes are easily misjudged as noise and removed, leading to a gradual damage to the semantic skeleton of the high-quality dataset.

[0024] Therefore, the incremental sample dataset is first subjected to scenario-level parsing to identify the scenario type to which each basic data segment belongs. A scenario classification algorithm is then used to perform semantic parsing on the incremental sample dataset, decomposing each incremental sample data into four basic data segments: business object, event type, handling action, and result description. These extracted basic data segments are input into a pre-trained scenario classification model to calculate the scenario affiliation probability of each basic data segment for four scenarios: traffic construction, accident linkage, special scheduling, and extreme operation. The basic data segments are assigned to the scenario category with the highest scenario affiliation probability. If a basic data segment's scenario affiliation probability exceeds a preset scenario affiliation probability threshold in two or more scenarios, it is marked as a cross-scenario stratified sample. Cross-scenario stratified samples with high matching degrees across multiple scenarios are retained to avoid missing low-frequency inherited segments with marginal associations. Finally, all basic data segments are stratified and summarized according to scenario category to generate a scenario-stratified sample set. The scenario affiliation probability threshold can be set to 0.8.

[0025] Edge correlation analysis is performed on the scene hierarchical sample set to construct the scene edge correlation network, specifically including: A low-frequency data aggregation algorithm is used to perform frequency statistics and cross-scene reuse detection on the scene-layered samples in the scene-layered sample set, and the low-frequency inheritance strength of the scene-layered samples is calculated. A preset inheritance threshold is used to filter and integrate the scene-layered samples whose low-frequency inheritance strength is greater than or equal to the inheritance threshold, thereby generating a set of low-frequency inheritance fragments. A cross-scene relationship mining algorithm is used to analyze and identify the causal relationships, treatment continuity relationships, risk transmission relationships, and result inheritance relationships among multiple low-frequency inheritance fragments in the set of low-frequency inheritance fragments, and to calculate the edge association strength among multiple low-frequency inheritance fragments. A preset association threshold is set, and low-frequency inherited segments with edge association strength greater than or equal to the association threshold are retained. The retained low-frequency inherited segments are used as association nodes and the corresponding edge association strengths are used as association edges to construct the scene edge association network.

[0026] For scenario-layered samples that appear infrequently but are repeatedly referenced and connect multiple scenarios, such as traffic construction, accident linkage, special dispatching, or extreme operation scenarios, these samples should not be treated as ordinary low-frequency noise. The more times a scenario-layered sample is referenced by different scenarios and the more scenario categories it involves, the stronger its cross-scenario support effect. Instead, the low-frequency inheritance strength should be calculated, and the corresponding calculation formula is as follows: ; In the formula, L represents the low-frequency inheritance strength; M represents the number of times the layered sample of this scene is referenced in different scenes; This indicates the number of scene categories involved in the scene stratification sample; N represents the total number of times the scene stratification sample appears in the incremental period count.

[0027] Subsequently, an inheritance threshold is preset, and low-frequency inheritance intensity is filtered out. Scene-level samples with low-frequency inheritance intensity greater than or equal to the inheritance threshold are retained and integrated to form a low-frequency inheritance fragment set. The inheritance threshold is set to a value of 2.50–3.00 based on the distribution of low-frequency inheritance intensity.

[0028] Furthermore, a cross-scenario relationship mining algorithm is employed to analyze whether there are causal relationships, continuation relationships, risk transmission relationships, or outcome inheritance relationships among different low-frequency inheritance fragments. Causal relationships characterize the causal association between events in a previous scenario and events in a subsequent scenario. Continuation relationships characterize the continuity of processing flows in different scenarios. Risk transmission relationships characterize the propagation process of risk factors between different scenarios; outcome inheritance relationships characterize the inherited impact of historical results on subsequent scenarios. For example, construction lockdowns leading to vehicle detours, vehicle detours triggering accident linkage, and accident linkage triggering extreme scheduling—this low-frequency inheritance fragment, although not frequent in a single scenario, constitutes a complex scenario reasoning chain as a whole. By calculating the edge association strength between multiple low-frequency inheritance fragments, the edge association structure between two low-frequency inheritance fragments is identified. The corresponding calculation formula is as follows: ; In the formula, G represents the edge association strength; U represents the semantic similarity ratio between low-frequency inherited segments; V represents the number of cross-scene references between low-frequency inherited segments; W represents the number of disposal acceptances between low-frequency inherited segments; and K represents the number of conflicts between low-frequency inherited segments.

[0029] A preset association threshold is used to filter edge association strength, retaining low-frequency inherited fragments with edge association strength greater than or equal to the threshold. Finally, using the retained low-frequency inherited fragments as associated nodes and their corresponding edge association strengths as associated edges, association relationships are established between these nodes, constructing a scene edge association network. This scene edge association network describes the inheritance path, propagation path, and association structure of low-frequency inherited fragments between different scenes. The association threshold ranges from 3.6 to 5.0.

[0030] Step S2: Based on the scene edge association network, the corresponding incremental sample data sets are concatenated to generate scene inheritance association paths and identify skeleton support regions, and the edge semantic skeleton set is obtained by summarizing.

[0031] The step of concatenating the corresponding incremental sample data sets based on the scene edge association network to generate scene inheritance association paths and identify skeleton support regions, and summarizing them to obtain an edge semantic skeleton set, specifically including: The scene edge association network is traversed in depth-first order using a path extraction algorithm. Starting from the associated node, multiple associated nodes are sequentially connected along the associated edge in the order of the causal relationship, the treatment continuation relationship, the risk transmission relationship, and the result inheritance relationship to generate multiple scene inheritance association paths. The cross-scene connectivity capability of each scene inheritance association path is evaluated by the skeleton support evaluation model. The skeleton support strength of each scene inheritance association path is calculated. A skeleton support strength threshold is preset and the skeleton support strength is filtered. Multiple scene inheritance association paths with skeleton support strength greater than or equal to the skeleton support strength threshold are retained. The retained scene inheritance association paths are marked as the skeleton support region. Extract and integrate the skeleton support region, the associated nodes corresponding to the skeleton support region, the associated edges, the skeleton support strength, and the number of cross-scene categories to generate an edge semantic skeleton. Summarize the edge semantic skeletons of multiple skeleton support regions to generate the edge semantic skeleton set.

[0032] In this context, the scene inheritance association path refers to a directed path in the scene edge association network formed by sequentially connecting multiple associated nodes along causal relationships, handling continuation relationships, risk transmission relationships, and result inheritance relationships. This corresponds to a complete reasoning chain for complex business scenarios, such as construction closures, lane compression, accident occurrences, emergency detours, and extreme scheduling. By using sequential concatenation constraints, the problem of meaningless jump paths generated during the traversal of the scene edge association network can be solved, thereby improving the semantic continuity and logical integrity of the scene inheritance association path.

[0033] The skeleton support area refers to the scene inheritance association path that connects multiple scenes simultaneously and plays a key supporting role in cross-scene reasoning. It has cross-scene connection capabilities. The absence of the skeleton support area will lead to the break of logical association between multiple scenes.

[0034] An edge semantic skeleton refers to a semantic unit obtained by structurally encapsulating the skeleton's supporting region. It includes at least a set of associated nodes, a set of associated edges, a skeleton support strength value, and the number of cross-scene categories. The edge semantic skeleton set is a collection of all edge semantic skeletons, used to represent all supporting structures in the incremental sample data set that have long-term inheritance value, cross-scene propagation value, and structural support value.

[0035] The skeleton support evaluation model assesses the edge association strength, path integrity, cross-cycle inheritance capability, and fracture status of the inherited paths in each scenario, and calculates the corresponding skeleton support strength. In the formula, B represents the skeleton support strength; E represents the average edge association strength in the scene inheritance association path; H represents the completeness ratio of the continuous inheritance path, obtained by dividing the number of fully inherited association nodes by the total number of association nodes in the scene inheritance association path; J represents the number of incremental cycles crossed; and d represents the number of broken nodes in the scene inheritance association path. A higher completeness ratio of the continuous inheritance path indicates a more complete semantic inheritance relationship in the scene inheritance association path; a larger number of incremental cycles crossed indicates that the scene inheritance association path remains stable over a longer period; and a smaller number of broken nodes in the scene inheritance association path indicates a more stable structure of the scene inheritance association path.

[0036] A preset skeleton support strength threshold is used to filter the skeleton support strength corresponding to each scene inheritance association path. Scene inheritance association paths with skeleton support strength greater than or equal to the skeleton support strength threshold are retained, and these retained scene inheritance association paths are marked as skeleton support regions. The skeleton support region is not a physical region in the traditional sense, but a high-value semantic association region composed of multiple associated nodes and their associated edges, reflecting the long-term stable cross-scene semantic inheritance structure within the incremental sample data set. The skeleton support strength threshold can be set to 8~10.

[0037] After identifying the skeleton support region, the associated nodes, associated edges, skeleton support strength, and number of cross-scene categories corresponding to the skeleton support region are further extracted to generate the corresponding edge semantic skeleton. The edge semantic skeletons generated from multiple skeleton support regions are then aggregated to form an edge semantic skeleton set.

[0038] Step S3: Perform stability analysis and scene association span detection on the edge semantic skeleton set, and obtain long-term retained mapping results through the freeze protection mechanism.

[0039] The process of performing stability analysis and scene association span detection on the edge semantic skeleton set, and obtaining long-term retained mapping results through a freeze protection mechanism, specifically includes: The stability analysis of each edge semantic skeleton in the edge semantic skeleton set is performed by the skeleton stability evaluation model, and the periodic inheritance strength corresponding to each edge semantic skeleton is calculated. Based on the periodic inheritance strength, a scene association span detection algorithm is used to detect the scene range connected by each edge semantic skeleton, and the scene association span value corresponding to each edge semantic skeleton is calculated. A preset retention threshold is established, and the freezing protection mechanism is used to filter and protect the edge semantic skeleton set based on the periodic inheritance strength and the scene association span value. If the scene association span value is greater than or equal to the retention threshold, the corresponding edge semantic skeleton is marked as the long-term retention mapping result. In this invention, the retention threshold is generally set to 5.

[0040] Although the edge semantic skeleton set has identified the skeleton support regions with cross-scene connectivity, some edge structures may lack continuous stability due to occasional factors during long-term incremental training. If all edge semantic skeletons are retained for a long time, it may cause the retention range to expand excessively and the noise filtering ability to be weakened.

[0041] The long-term stability and cross-scenario support span of each edge semantic skeleton are evaluated by the skeleton stability assessment model. Edge semantic skeletons with long-term retention value are selected and protected by the freezing protection mechanism to generate long-term retention mapping results.

[0042] Periodic inheritance strength characterizes the ability of a current marginal semantic skeleton to be retained and inherited across multiple incremental update cycles. A higher periodic inheritance strength indicates that a marginal semantic skeleton consistently appears across multiple incremental update cycles and maintains the structural integrity of the inference path; conversely, a lower periodic inheritance strength indicates that the marginal semantic skeleton is a short-term, sporadic structure. The formula for calculating periodic inheritance strength is as follows: ; In the formula, C represents the periodic inheritance strength; P represents the number of consecutive incremental periods of the edge semantic skeleton; R represents the number of times the edge semantic skeleton is reused in the historical incremental periods; A represents the scene inheritance continuity ratio of the edge semantic skeleton, which is obtained by dividing the number of complete inheritance chains by the number of required inheritance chains; and F represents the number of times the edge semantic skeleton breaks between incremental periods.

[0043] Based on the periodic inheritance strength, a scene association span detection algorithm is used to analyze the scene range connected to each edge semantic skeleton, and evaluate the cross-scene generalization ability of the edge semantic skeleton. The corresponding calculation formula for the scene association span value is as follows: ; In the formula, D represents the scene association span value; C represents the periodic inheritance strength; and g represents the number of scene categories connected by the edge semantic skeleton. O represents the scene hierarchy span connected by the edge semantic skeleton; O represents the number of isolated fragments generated after removing the edge semantic skeleton.

[0044] A preset retention threshold is set. If the scene association span value is greater than or equal to the retention threshold, it is determined that the current edge semantic skeleton has both high cross-scene connectivity and strong long-term inheritance capability. A freeze protection mechanism is then used to mark the current edge semantic skeleton as a long-term retention mapping result and protect it. This solves the problem that traditional noise filtering methods only filter based on frequency of occurrence or local statistical features, which leads to the accidental deletion of edge semantic skeletons.

[0045] Step S4: Reconstruct the noise filtering boundary based on the long-term retention mapping result to generate a noise filtering layer boundary; The step of reconstructing the noise filtering boundary based on the long-term retained mapping result to generate a layered noise filtering boundary specifically includes: Based on the long-term retention mapping result, a backtracking check is performed on the current noise filtering boundary to identify the overlap range between the current noise filtering boundary and the long-term retention mapping result, and the overlap ratio and conflict ratio are calculated using the overlap range. An adaptive shrinkage algorithm is used to reconstruct the noise filtering boundary based on the overlap ratio and the conflict ratio to generate the noise filtering layer boundary.

[0046] Among them, the noise filtering boundary refers to the noise judgment boundary established based on the statistical results of historical data, which is used to distinguish between retained data and filtered data.

[0047] During the backtracking process, the edge semantic skeletons corresponding to the long-term preserved mapping results are mapped to the feature space corresponding to the noise filtering boundary. The positional relationship between the long-term preserved mapping results and the noise filtering boundary is detected, and the overlap range between them is identified. The overlap ratio is obtained by dividing the number of basic data segments in the overlapping region between the long-term preserved mapping results and the noise filtering boundary by the number of basic data segments in the overall region corresponding to the long-term preserved mapping results. The conflict ratio is obtained by identifying and calculating the number of edge semantic skeletons located inside the noise filtering boundary and belonging to the long-term preserved mapping results, divided by the total number of edge semantic skeletons in the long-term preserved mapping results. This conflict ratio is used to assess the degree of conflict between the noise filtering boundary and the long-term preserved mapping results. A higher conflict ratio indicates a more significant false filtering effect of the noise filtering boundary on the long-term preserved mapping results.

[0048] The step of reconstructing the noise filtering boundary using an adaptive shrinkage algorithm based on the overlap ratio and the conflict ratio to generate the noise filtering layer boundary specifically includes: Preset protection boundary threshold and buffer boundary threshold; Obtain the noise filtering threshold corresponding to the noise filtering boundary; The overlap ratio and the conflict ratio are weighted and fused to calculate the boundary reconstruction strength; If the boundary reconstruction intensity is greater than or equal to the protection boundary threshold, then the noise filtering threshold corresponding to the noise filtering boundary is expanded to the boundary reconstruction intensity, and the protection range of the reconstructed noise filtering boundary is marked as the skeleton protection area. If the boundary reconstruction strength is less than the protection boundary threshold, and the boundary reconstruction strength is greater than or equal to the buffer boundary threshold, then the noise filtering threshold corresponding to the current noise filtering boundary is maintained, and the protection range of the noise filtering boundary is marked as the observation retention area. If the boundary reconstruction intensity is less than the buffer boundary threshold, the noise filtering threshold corresponding to the noise filtering boundary is reduced to the boundary reconstruction intensity, and the protection range of the reconstructed noise filtering boundary is marked as a candidate filtering region.

[0049] Traditional, single noise filtering boundaries cannot accurately remove noise and protect the skeleton, and are prone to misclassifying edge semantic skeletons in long-term retained mapping results as noise. By reconstructing the noise filtering boundaries, different filtering rules can be set for data regions of different importance.

[0050] Pre-set protection boundary thresholds and buffer boundary thresholds. The protection boundary threshold determines whether a large-scale expansion of the current noise filtering boundary is needed; the buffer boundary threshold determines whether the current noise filtering boundary needs to be kept stable, and can be set to 0.5. Then, obtain the noise filtering threshold corresponding to the noise filtering boundary. This threshold is a specific value used to define whether to retain or filter data, and in traditional technologies, it is a single fixed value, which can be set to 0.5. Through multiple rounds of incremental training, the importance of long-term retention mapping results is calculated to dynamically adjust the overlap ratio weight coefficient and the conflict ratio weight coefficient. The overlap ratio is multiplied by the overlap ratio weight coefficient, and the conflict ratio is multiplied by the conflict ratio weight coefficient. The results are then fused to generate the boundary reconstruction strength. The overlap ratio weight coefficient and the conflict ratio weight coefficient both range from 0.4 to 0.6, set by the long-term retention importance and the misfiltering risk level, respectively.

[0051] When the boundary reconstruction strength is greater than or equal to the protection boundary threshold, it indicates that the current noise filtering boundary has already covered a large number of long-term retained mapping results, posing a high risk of false filtering. In this case, the noise filtering threshold corresponding to the current noise filtering boundary is extended outward to the adjustment range corresponding to the boundary reconstruction strength, thereby expanding the protection range for the edge semantic skeleton. The adjusted protection range is marked as the skeleton protection area. The skeleton protection area is used to store the edge semantic skeletons that need to be retained, preventing them from being mistakenly deleted during the noise filtering process. The protection boundary threshold can be set to 0.7.

[0052] When the boundary reconstruction strength is less than the protective boundary threshold but greater than or equal to the buffer boundary threshold, it indicates that there is a certain degree of cross-influence between the current noise filtering boundary and the long-term retention mapping result, but it has not yet reached the level requiring large-scale adjustment. In this case, the noise filtering threshold corresponding to the current noise filtering boundary remains unchanged, and the corresponding region is marked as the observation retention region. Forced filtering is not applied to the sample data in the observation retention region; instead, changes in the observation retention region are continuously observed during subsequent rounds of incremental updates.

[0053] When the boundary reconstruction intensity is less than the buffer boundary threshold, it indicates that there is essentially no significant conflict between the current noise filtering boundary and the long-term retained mapping result. In this case, an adaptive shrinkage strategy is used to shrink the current noise filtering boundary, adjusting the noise filtering threshold to the range corresponding to the boundary reconstruction intensity, and marking the adjusted region as a candidate filtering region. Sample data in the candidate filtering region are preferentially used as noise filtering targets.

[0054] By reconstructing the noise filtering boundary, the traditional single noise filtering boundary is transformed into a hierarchical noise filtering boundary that includes a skeleton protection region, an observation retention region, and a candidate filtering region. This hierarchical noise filtering boundary can dynamically adjust the filtering range based on long-term retention mapping results, prioritizing the protection of the marginal semantic skeleton while improving the filtering efficiency of low-value noisy sample data.

[0055] Step S5: Using a differentiated filtering strategy, the incremental sample data set is collaboratively filtered and scheduled according to the noise filtering hierarchical boundary to generate a high-quality dataset.

[0056] The step of employing a differentiated filtering strategy to collaboratively filter and schedule the incremental sample data set according to the noise filtering hierarchical boundary to generate a high-quality dataset specifically includes: Based on the noise filtering layer boundary, calculate the adaptive filtering intensity of multiple basic data segments in the incremental sample data set; The adaptive filtering intensity is matched with the noise filtering layer boundary to determine the region type to which the basic data segment corresponding to the adaptive filtering intensity belongs. A differentiated filtering strategy is used for collaborative filtering scheduling to integrate the retained basic data segments and generate the high-quality dataset.

[0057] The contribution of basic data fragments in different region types to improving the quality of high-quality datasets varies. Using a uniform filtering intensity might lead to the accidental deletion of basic data fragments in skeleton-protected regions or the incorrect retention of noisy sample data in candidate filtering regions. Therefore, a differentiated filtering strategy is adopted, which involves collaboratively scheduling the incremental sample data set based on the noise filtering hierarchical boundaries to achieve refined filtering of sample data from different regions.

[0058] Based on the noise filtering layer boundary, multiple basic data segments in the incremental sample dataset are analyzed to calculate the corresponding adaptive filtering intensity. In the formula, This represents the basic filtering strength, which is set according to the quality target of the high-quality dataset and can be set to 0.5. The region protection determination value for the basic data segment is represented by _{Q}; Q represents the number of ordinary noise hits for the basic data segment in the current incremental period; and C represents the period inheritance strength corresponding to the basic data segment. The adaptive filtering strength characterizes the probability that the basic data segment is identified as noise sample data and the priority of performing filtering operations.

[0059] The adaptive filtering intensity is matched with the noise filtering layer boundary. Based on the region type where the basic data fragment is located, it is divided into the corresponding skeleton protection region, observation retention region, or candidate filtering region. If the basic data fragment is located in the skeleton protection region, it indicates that the basic data fragment has a strong correlation with the long-term retention mapping result, is an important part of the marginal semantic skeleton, and is protected by a strong protection strategy. If the basic data fragment is located in the observation retention region, it indicates that the basic data fragment has some retention value, but its long-term contribution ability is not yet fully determined. The observation filtering strategy is used to temporarily retain the basic data fragment, and it is continuously tracked in subsequent incremental update cycles. If the basic data fragment is located in the candidate filtering region, it indicates that the basic data fragment has a weak correlation with the long-term retention mapping result and has a low necessity for protection. The enhanced filtering strategy is used to filter the basic data fragment. Finally, all retained basic data fragments are integrated to generate a high-quality dataset.

[0060] like Figure 2 The diagram shown is a system block diagram of an adaptive filtering system for noisy data in the construction of a high-quality dataset provided in an embodiment of the present invention. The system includes: The scene classification and edge analysis module is used to acquire incremental sample datasets, merge them, perform scene classification and edge correlation analysis, and construct a scene edge correlation network. The skeleton support recognition and extraction module is used to concatenate the corresponding incremental sample data set based on the scene edge association network, generate scene inheritance association path and identify skeleton support region, and summarize to obtain edge semantic skeleton set. The stability assessment and freeze protection module is used to perform stability analysis and scene association span detection on the edge semantic skeleton set, and to filter the long-term retention mapping results through the freeze protection mechanism. The noise filtering boundary reconstruction module is used to reconstruct the noise filtering boundary based on the long-term retained mapping result, and generate a layered noise filtering boundary. The collaborative filtering scheduling module is used to perform collaborative filtering scheduling on the incremental sample data set according to the noise filtering hierarchical boundary using a differentiated filtering strategy, thereby generating a high-quality dataset.

[0061] Figure 2 The apparatus of the illustrated embodiment can be used to perform corresponding actions. Figure 1 The steps in the example of the adaptive filtering method for noisy data in the construction of the high-quality dataset shown are similar in principle and technical effect, and will not be repeated here.

[0062] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. An adaptive filtering method for noisy data in the construction of high-quality datasets, characterized in that, The method includes: The incremental sample dataset is acquired and merged to perform scene classification and edge correlation analysis, and a scene edge correlation network is constructed. Based on the scene edge association network, the corresponding incremental sample data sets are concatenated to generate scene inheritance association paths and identify skeleton support regions, and the edge semantic skeleton set is obtained by summarizing them. The edge semantic skeleton set is subjected to stability analysis and scene association span detection, and long-term retained mapping results are obtained by filtering through the freeze protection mechanism; Based on the long-term retention mapping results, the noise filtering boundary is reconstructed to generate a layered noise filtering boundary. A differentiated filtering strategy is adopted to perform collaborative filtering scheduling on the incremental sample data set according to the noise filtering hierarchical boundary, thereby generating a high-quality dataset.

2. The method for adaptive filtering of noisy data in the construction of high-quality datasets according to claim 1, characterized in that, The incremental sample datasets are acquired, merged, and used for scene classification to generate a scene-stratified sample set, specifically including: The incremental sample dataset is semantically parsed using a scene classification algorithm. Multiple basic data fragments are extracted and input into the scene classification model. The scene affiliation probability of each of the multiple basic data fragments corresponding to different scene categories is calculated. Based on the scene affiliation probability, the scene categories corresponding to the multiple basic data segments are determined, and the multiple basic data segments are stratified and classified according to the scene categories, and the scene stratification sample set is generated.

3. The method for adaptive filtering of noisy data in the construction of high-quality datasets according to claim 2, characterized in that, Edge correlation analysis is performed on the scene hierarchical sample set to construct the scene edge correlation network, specifically including: A low-frequency data aggregation algorithm is used to perform frequency statistics and cross-scene reuse detection on the scene-layered samples in the scene-layered sample set, and the low-frequency inheritance strength of the scene-layered samples is calculated. A preset inheritance threshold is used to filter and integrate the scene-layered samples whose low-frequency inheritance strength is greater than or equal to the inheritance threshold, thereby generating a set of low-frequency inheritance fragments. A cross-scene relationship mining algorithm is used to analyze and identify the causal relationships, treatment continuity relationships, risk transmission relationships, and result inheritance relationships among multiple low-frequency inheritance fragments in the set of low-frequency inheritance fragments, and to calculate the edge association strength among multiple low-frequency inheritance fragments. A preset association threshold is set, and low-frequency inherited segments with edge association strength greater than or equal to the association threshold are retained. The retained low-frequency inherited segments are used as association nodes and the corresponding edge association strengths are used as association edges to construct the scene edge association network.

4. The method for adaptive filtering of noisy data in the construction of high-quality datasets according to claim 3, characterized in that, The step of concatenating the corresponding incremental sample data sets based on the scene edge association network to generate scene inheritance association paths and identify skeleton support regions, and summarizing them to obtain an edge semantic skeleton set, specifically including: The scene edge association network is traversed in depth-first order using a path extraction algorithm. Starting from the associated node, multiple associated nodes are sequentially connected along the associated edge in the order of the causal relationship, the treatment continuation relationship, the risk transmission relationship, and the result inheritance relationship to generate multiple scene inheritance association paths. The cross-scene connectivity capability of each scene inheritance association path is evaluated by the skeleton support evaluation model. The skeleton support strength of each scene inheritance association path is calculated. A skeleton support strength threshold is preset and the skeleton support strength is filtered. Multiple scene inheritance association paths with skeleton support strength greater than or equal to the skeleton support strength threshold are retained. The retained scene inheritance association paths are marked as the skeleton support region. Extract and integrate the skeleton support region, the associated nodes corresponding to the skeleton support region, the associated edges, the skeleton support strength, and the number of cross-scene categories to generate an edge semantic skeleton. Summarize the edge semantic skeletons of multiple skeleton support regions to generate the edge semantic skeleton set.

5. The method for adaptive filtering of noisy data in the construction of high-quality datasets according to claim 1, characterized in that, The process of performing stability analysis and scene association span detection on the edge semantic skeleton set, and obtaining long-term retained mapping results through a freeze protection mechanism, specifically includes: The stability analysis of each edge semantic skeleton in the edge semantic skeleton set is performed by the skeleton stability evaluation model, and the periodic inheritance strength corresponding to each edge semantic skeleton is calculated. Based on the periodic inheritance strength, a scene association span detection algorithm is used to detect the scene range connected by each edge semantic skeleton, and the scene association span value corresponding to each edge semantic skeleton is calculated. A preset retention threshold is set, and the freezing protection mechanism is used to filter and protect the edge semantic skeleton set based on the periodic inheritance strength and the scene association span value. If the scene association span value is greater than or equal to the retention threshold, the corresponding edge semantic skeleton is marked as the long-term retention mapping result.

6. The method for adaptive filtering of noisy data in the construction of high-quality datasets according to claim 1, characterized in that, The process of reconstructing the noise filtering boundary based on the long-term retained mapping result to generate a hierarchical noise filtering boundary specifically includes: Based on the long-term retention mapping result, a backtracking check is performed on the current noise filtering boundary to identify the overlap range between the current noise filtering boundary and the long-term retention mapping result, and the overlap ratio and conflict ratio are calculated using the overlap range. An adaptive shrinkage algorithm is used to reconstruct the noise filtering boundary based on the overlap ratio and the conflict ratio to generate the noise filtering layer boundary.

7. The method for adaptive filtering of noisy data in the construction of high-quality datasets according to claim 6, characterized in that, The step of reconstructing the noise filtering boundary using an adaptive shrinkage algorithm based on the overlap ratio and the conflict ratio to generate the noise filtering layer boundary specifically includes: Preset protection boundary threshold and buffer boundary threshold; Obtain the noise filtering threshold corresponding to the noise filtering boundary; The overlap ratio and the conflict ratio are weighted and fused to calculate the boundary reconstruction strength; If the boundary reconstruction intensity is greater than or equal to the protection boundary threshold, then the noise filtering threshold corresponding to the noise filtering boundary is expanded to the boundary reconstruction intensity, and the protection range of the reconstructed noise filtering boundary is marked as the skeleton protection area. If the boundary reconstruction strength is less than the protection boundary threshold, and the boundary reconstruction strength is greater than or equal to the buffer boundary threshold, then the noise filtering threshold corresponding to the current noise filtering boundary is maintained, and the protection range of the noise filtering boundary is marked as the observation retention area. If the boundary reconstruction intensity is less than the buffer boundary threshold, the noise filtering threshold corresponding to the noise filtering boundary is reduced to the boundary reconstruction intensity, and the protection range of the reconstructed noise filtering boundary is marked as a candidate filtering region.

8. The method for adaptive filtering of noisy data in the construction of high-quality datasets according to claim 2, characterized in that, The step of employing a differentiated filtering strategy to collaboratively filter and schedule the incremental sample data set according to the noise filtering hierarchical boundary to generate a high-quality dataset specifically includes: Based on the noise filtering layer boundary, calculate the adaptive filtering intensity of multiple basic data segments in the incremental sample data set; The adaptive filtering intensity is matched with the noise filtering layer boundary to determine the region type to which the basic data segment corresponding to the adaptive filtering intensity belongs. A differentiated filtering strategy is used for collaborative filtering scheduling to integrate the retained basic data segments and generate the high-quality dataset.

9. An adaptive filtering system for noisy data in the construction of high-quality datasets, characterized in that: The system, which is applied to the adaptive filtering method for noisy data in the construction of high-quality datasets as described in any one of claims 1-8, comprises: The scene classification and edge analysis module is used to acquire incremental sample datasets, merge them, perform scene classification and edge correlation analysis, and construct a scene edge correlation network. The skeleton support recognition and extraction module is used to concatenate the corresponding incremental sample data set based on the scene edge association network, generate scene inheritance association path and identify skeleton support region, and summarize to obtain edge semantic skeleton set. The stability assessment and freeze protection module is used to perform stability analysis and scene association span detection on the edge semantic skeleton set, and to filter the long-term retention mapping results through the freeze protection mechanism. The noise filtering boundary reconstruction module is used to reconstruct the noise filtering boundary based on the long-term retained mapping result, and generate a layered noise filtering boundary. The collaborative filtering scheduling module is used to perform collaborative filtering scheduling on the incremental sample data set according to the noise filtering hierarchical boundary using a differentiated filtering strategy, thereby generating a high-quality dataset.