A material data cleaning and grading screening method and system
By acquiring document structure integrity and functional description similarity analysis, deviations and compliance issues in material data are dynamically identified, solving the insufficient accuracy and compliance of data cleaning and hierarchical screening in existing technologies, and achieving efficient resource allocation and task execution monitoring.
Patent Information
- Application Number
- CN202511351255.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-22
- Publication Date
- 2026-01-23
- Estimated Expiration
- 2045-09-22
AI Technical Summary
Existing technologies lack in-depth analysis of document structure logic and key field distribution in material data cleaning and classification screening, leading to the misremoval or incorrect classification of valid data, difficulty in identifying functionally repetitive content, inaccurate compliance verification, and an imbalance between resource allocation and task execution, thus affecting project efficiency and quality.
By acquiring document structure integrity information and key field missing distribution patterns, document integrity deviation labels are generated. Deviation correction is performed by combining functional description similarity distribution intervals, identifying technically infeasible items and compliance deviations, dynamically identifying the matching status of resource configuration and task execution, and forming a data structure identifier for full-process tracking.
It significantly improves the accuracy of information screening and the sensitivity of anomaly identification, accurately detects unreasonable resource allocation and imbalance in task execution rhythm, and enhances the pertinence of progress monitoring and the accuracy of project execution.
Smart Images

Figure CN120849534B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of materials information processing technology, and in particular to a method and system for cleaning and classifying materials data. Background Technology
[0002] The field of materials information processing technology involves the collection, integration, preprocessing, analysis, and utilization of various types of materials data. Its core aspects include the standardized processing of basic materials data, the classification and analysis of materials performance parameters, and the extraction and management of data information for materials design and application. This technology aims to support the research and development of new materials and the optimization of materials performance. Through the standardized processing and in-depth analysis of massive amounts of materials information, it provides data support and decision-making basis for materials engineering. Traditional materials data cleaning and classification screening methods refer to the technical means of preliminary cleaning and screening of raw materials data during the data processing process. These methods primarily address issues such as missing, abnormal, redundant, or inconsistent data. Non-standard data is removed, corrected, or supplemented through rule-based judgment or manual annotation. The data is then hierarchically classified and screened according to manually set standards to eliminate data items that do not meet the criteria and to effectively classify and organize the data. Data cleaning and classification are performed using methods such as static threshold setting, field logical relationship judgment, or manual comparison and screening based on specific field content.
[0003] Existing technologies primarily rely on rule-based judgment or manual annotation for data cleaning and hierarchical screening, lacking in-depth analysis of document structure logic and key field distribution. In scenarios with incomplete documents or numerous missing fields, this can easily lead to the misremoval of valid data or incorrect classification. Furthermore, when using static thresholds and field comparisons to process functional content, it is difficult to identify semantic repetitions and technical overlaps between descriptions, making it difficult to identify and filter out repetitive content. In the compliance verification stage, the reliance on manually set screening standards lacks linkage with actual operational norms, resulting in the omission of some non-compliant information. Especially when there are dynamic differences between resource allocation and task execution, it cannot reflect the trend of process deviation, which can easily cause task execution to deviate from the target plan, thereby affecting the overall efficiency and quality of project progress. Summary of the Invention
[0004] The purpose of this invention is to address the shortcomings of existing technologies by proposing a method and system for cleaning and classifying material data.
[0005] To achieve the above objectives, the present invention adopts the following technical solution: a method for cleaning and classifying material data, comprising the following steps:
[0006] S1: Obtain information on the document structure integrity, key field missing distribution patterns, and logical vulnerability marker list from the digital project initiation materials; extract the document paragraph missing rate and logical connection strength; and generate document integrity deviation tags.
[0007] S2: Based on the document integrity deviation label, select a hierarchical screening mode that is suitable for the current document characteristics, extract the functional description similarity distribution range from the original project database, and perform deviation correction and classification processing with the benchmark similarity to obtain the functional duplication screening label.
[0008] S3: Call the function repetition screening label, extract the function description similarity segment number, identify the technical feasibility assessment data within the segment, compare the technically infeasible items with the function description segment, record the number of overlapping time periods, and generate a list of technically infeasible items affecting segments;
[0009] S4: Based on the list of technically infeasible items affecting the affected areas, filter out unmarked document compliance verification nodes, extract compliance management specifications and operational data, map the distribution of construction content and compliance requirements, determine whether there is a compliance deviation, and obtain the compliance deviation fluctuation group.
[0010] As a further aspect of the present invention, the document integrity deviation label includes deviation number, missing status label, logical correlation deviation amount, and pattern classification; the functional duplication screening label includes similarity level, fluctuation type, benchmark comparison result, and task association number; the list of technically infeasible item impact segments includes device type, impact time segment, number of overlapping time periods, and affected task number; and the compliance deviation fluctuation group includes task distribution unevenness number, operation duration record, task completion deviation amount, and operation matching degree.
[0011] As a further aspect of the present invention, the step of obtaining the document integrity deviation tag specifically includes:
[0012] S111: Obtain document structure integrity information, key field missing distribution patterns, and logical vulnerability marker list from the digital project initiation materials; extract document paragraph missing rate and logical correlation strength; match document integrity requirements with real-time distribution; compare time range with integrity requirements; and generate partition node document record time periods.
[0013] S112: Based on the document record time period of the partition node, extract the overlapping time period between the document integrity status and the required time interval, calculate the ratio of overlapping time to the total required duration, filter nodes with an overlap ratio lower than the benchmark value, and obtain the partition node integrity coverage deviation rate based on the number of integrity status annotations.
[0014] S113: Based on the partition node integrity coverage deviation rate, determine the deviation status of the node number, identify the node number whose deviation rate exceeds the node synchronization threshold, integrate the node number, integrity coverage information and deviation rate value, and generate a document integrity deviation label.
[0015] As a further aspect of the present invention, the step of obtaining the functional repeatability screening label specifically includes:
[0016] S211: Based on the document integrity deviation label, identify the hierarchical screening mode and functional description similarity distribution characteristics that are adapted to the current document characteristics, extract the start and end fluctuation range of real-time functional description similarity, calculate the start and end fluctuation difference of functional description similarity segment, compare it with the benchmark similarity fluctuation range, and obtain the functional description similarity fluctuation deviation value.
[0017] S212: Call the function description similarity fluctuation deviation value, combine the segment distribution, fluctuation trend and adjustment frequency, uniformly collect the graded screening mode deviation data, identify and calculate the adaptive deviation degree according to the segment number, determine the fluctuation direction according to the adjustment frequency, and obtain the functional repeatability screening label.
[0018] As a further aspect of the present invention, the step of obtaining the list of technically infeasible items affecting the region specifically comprises:
[0019] S311: Call the function repetition screening label, filter the function description similarity deviation task segment number, extract the function description similarity time period according to the node association table, process the segment function description time period according to the time dimension, identify the function description similarity time period index table, and obtain the function description similarity task time period set.
[0020] S312: Based on the set of time periods for the function description similarity task, collect the evaluation data of technical infeasibility items in the same period, identify the table of technical infeasibility item impact, determine the daily technical infeasibility item interference based on the interference threshold, match it with the time period for the function description similarity task, determine whether there is an abnormal association of function description similarity, and generate a list of technical infeasibility item impacted sections.
[0021] As a further aspect of the present invention, the step of obtaining the compliance deviation fluctuation group specifically includes:
[0022] S411: Based on the list of affected sections by the aforementioned technical infeasibility items, filter out unmarked document compliance verification nodes, extract the task list, and obtain the task set of nodes not affected by technical infeasibility items;
[0023] S412: Call the task set of nodes not affected by technical infeasibility, match compliance management specifications and operation data, extract the planned task volume by task number, count the number of operators and real-time operation time period, and generate a compliance execution status matching dataset.
[0024] S413: Based on the compliance execution status matching dataset, evaluate the degree of matching between task execution efficiency and operation distribution, identify efficiency fluctuation nodes, calculate fluctuation identification difference value, mark tasks with fluctuations exceeding the benchmark value as abnormal nodes, identify compliance task execution matching deviations, and obtain compliance deviation fluctuation groups.
[0025] As a further aspect of the present invention, the method further includes step S5:
[0026] S5: Call the compliance deviation fluctuation group, extract the abnormal fluctuation task group, calculate the ratio of resource input to task in the project establishment model, record the difference distribution between resource allocation cycle and task execution cycle, and generate project progress monitoring structure indicators.
[0027] The project initiation progress monitoring structure indicators include resource allocation ratio, task intensity level, execution cycle difference, and project initiation efficiency indicators.
[0028] As a further aspect of the present invention, the steps for obtaining the project initiation progress monitoring structural indicators are specifically as follows:
[0029] S511: Call the compliance deviation fluctuation group, filter nodes that exceed the threshold, record the time interval and the magnitude of task volume change, and obtain the progress fluctuation abnormal identifier set;
[0030] S512: Based on the project establishment model resources corresponding to the nodes in the progress fluctuation anomaly identifier set, identify the node resource input ratio sequence, extract the abnormal distribution interval and compare the critical coefficient, record the ratio offset direction and node number, and form a project establishment resource matching offset index group.
[0031] S513: Based on the project initiation resource matching offset index group, extract the time period for project initiation model resource deployment and task execution, identify the difference between the resource deployment cycle and the operation cycle, and sort and label them according to the progress benchmark to generate project progress monitoring structure indicators.
[0032] The material data cleaning and grading screening system is used to perform the above-mentioned material data cleaning and grading screening method. The system includes:
[0033] The integrity deviation extraction module obtains document structure integrity information, key field missing distribution patterns and logical vulnerability marker lists from digital project initiation materials, extracts document paragraph missing rate and logical connection strength, compares document integrity requirements with real-time distribution, and generates document integrity deviation tags.
[0034] The hierarchical screening mode classification module locates the hierarchical screening mode that is suitable for the current document characteristics based on the document integrity deviation label, extracts the document node identifier, process node and benchmark similarity, calculates the difference between the on-site similarity and the benchmark similarity, classifies and labels the difference type, and generates node mode adaptability identifier.
[0035] The infeasibility identification module, based on the node pattern adaptive identifier, filters task components with lagging functional description similarity, locates the corresponding segment, extracts the interference period of technical infeasibility evaluation, determines the overlap with the operation time period, filters frequently interfered segments, and generates a set of interference mappings for infeasibility in project progress.
[0036] The compliance diagnosis module, based on the infeasibility interference mapping set of the project initiation progress, eliminates interference segment tasks, extracts compliance management specifications and operation data, matches task assignments and operation time periods, calculates the ratio of operation volume to task, identifies task-intensive and inefficient component tasks, and obtains the node operation execution deviation set.
[0037] Based on the node operation execution deviation set, the resource allocation analysis module locates the resource input records of the task in the project initiation model, extracts the ratio of task quantity to resource quantity allocation, compares the difference between the operation cycle and the input cycle, maps the task resource usage and progress status, and forms a project initiation progress monitoring structure index.
[0038] Compared with the prior art, the advantages and positive effects of the present invention are as follows:
[0039] In this invention, by extracting information on document structural integrity and logical vulnerabilities, a quantitative expression of document deviation characteristics is achieved. Combined with interval analysis and deviation correction of functional description similarity, redundant content and functional duplication items are effectively identified. Furthermore, by comparing technical feasibility assessment data with functional sections, the interference range of infeasible items in different description sections is clarified. Through the mapping between compliance verification nodes and management specifications, dynamic identification of compliance deviations is achieved. On this basis, the matching of resource input and task volume is correlated to accurately capture phenomena such as unreasonable resource allocation and imbalance in task execution rhythm. Ultimately, a data structure identifier for full-process tracking can be formed, significantly improving the accuracy of information screening, the sensitivity of anomaly identification, and the targeting of progress monitoring. Attached Figure Description
[0040] Figure 1 This is a schematic diagram of the workflow of the present invention;
[0041] Figure 2 This is a flowchart illustrating the process of obtaining document integrity deviation tags in this invention;
[0042] Figure 3 This is a flowchart illustrating the process of obtaining functional repeatability screening tags in this invention.
[0043] Figure 4 This is a flowchart illustrating the process of obtaining the list of technically infeasible sections in this invention.
[0044] Figure 5 This is a flowchart illustrating the process of obtaining the compliance deviation fluctuation group in this invention.
[0045] Figure 6 This is a flowchart illustrating the process of obtaining structural indicators for project initiation progress monitoring in this invention. Detailed Implementation
[0046] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention.
[0047] In the description of this invention, it should be understood that the terms "length," "width," "upper," "lower," "front," "rear," "left," "right," "vertical," "horizontal," "top," "bottom," "inner," and "outer," etc., indicating orientation or positional relationships, are based on the orientation or positional relationships shown in the accompanying drawings and are only for the convenience of describing the invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation, and therefore should not be construed as a limitation of the invention. Furthermore, in the description of this invention, "a plurality of" means two or more, unless otherwise explicitly specified.
[0048] Example 1
[0049] Please see Figure 1 This invention provides a technical solution: a method for cleaning and classifying material data, comprising the following steps:
[0050] S1: Obtain information on the document structure integrity, key field missing distribution patterns, and logical vulnerability marker list from the digital project initiation materials; extract the document paragraph missing rate and logical connection strength; and generate document integrity deviation tags.
[0051] S2: Based on the document integrity deviation label, select the hierarchical screening mode that is suitable for the current document characteristics, extract the functional description similarity distribution range from the original project database, and perform deviation correction and classification processing with the benchmark similarity to obtain the functional duplication screening label.
[0052] S3: Call the functional duplication screening label, extract the functional description similarity segment number, identify the technical feasibility assessment data within the segment, compare the technically infeasible items with the functional description segment, record the number of overlapping time periods, and generate a list of segments affected by technically infeasible items.
[0053] S4: Based on the list of technically infeasible items affecting the area, filter out unmarked document compliance verification nodes, extract compliance management specifications and operational data, map the distribution of construction content and compliance requirements, determine whether there is a compliance deviation, and obtain the compliance deviation fluctuation group;
[0054] S5: Call the compliance deviation fluctuation group, extract the abnormal fluctuation task group, calculate the ratio of resource input to task volume in the project establishment model, record the difference distribution between resource allocation cycle and task execution cycle, and generate project progress monitoring structure indicators.
[0055] Document integrity deviation labels include deviation number, missing status label, logical correlation deviation amount, and pattern classification; functional duplication screening labels include similarity level, fluctuation type, benchmark comparison result, and task association number; the list of technically infeasible items affecting the affected area includes equipment type, affected time period, number of overlapping time periods, and affected task number; compliance deviation fluctuation group includes task distribution unevenness number, operation duration record, task completion deviation amount, and operation matching degree; and project progress monitoring structure indicators include resource allocation ratio, task intensity level, execution cycle difference, and project efficiency indicators.
[0056] Please see Figure 2 The specific steps for obtaining document integrity deviation tags are as follows:
[0057] S111: Obtain document structure integrity information, key field missing distribution patterns, and logical vulnerability marker list from the digital project initiation materials; extract document paragraph missing rate and logical correlation strength; match document integrity requirements with real-time distribution; compare time range with integrity requirements; and generate partition node document record time periods.
[0058] This process involves extracting document structure integrity information, key field missing distribution patterns, and logical vulnerability marker lists from digital project initiation materials. For a software development project's initiation documents, the process begins by retrieving the project's initiation documents from project management, such as the initiation documents for the "Intelligent Manufacturing System V1.0 Development Project." These documents include project requirements specifications, system design documents, and test plans. From these documents, structural integrity information is extracted, such as the completeness of chapters in the requirements specifications. The process checks for missing or out-of-order chapters and stores this information as a "Structural Integrity Table," as shown in Table 1. Simultaneously, the process identifies the distribution patterns of missing key fields, such as the "functionality" field in the requirements specifications. Key fields such as "Module Name," "Interface Definition," and "Performance Metrics" were analyzed, and the percentage of missing fields across all requirement documents was statistically analyzed. For example, if 15 out of 100 requirement documents were found to be missing the "Interface Definition" field (a 15% missing rate), this information was recorded as a "Key Field Missing Distribution Table." Documents were also marked with logical vulnerabilities. For instance, if a mismatch was found between the input of functional module A and the output of functional module B in the system design document, it was marked as a logical vulnerability, creating a "Logical Vulnerability Marking List." Similarly, in the "Intelligent Manufacturing System V1.0 Development Project," an incomplete "User Management" section was detected in the requirement document, lacking a "User Permissions" subsection. Furthermore, a missing "Data..." section was also detected in the design document. According to the transmission protocol, key fields are missing in 50% of the documents, and a logical flaw was found in the relationship between "test cases" and "requirements specification" in the test plan. Based on this information, the document paragraph missing rate and logical connection strength are calculated. For the requirements specification of the "Intelligent Manufacturing System V1.0 Development Project," which has a total of 100 paragraphs, 5 paragraphs are actually missing, so the paragraph missing rate is 5 / 100 = 0.05, or 5%. The logical connection strength is quantified by analyzing cross-references and data flow matching in the documents, expressed as a value between 0 and 1, where 1 represents complete connection and 0 represents no connection. For example, the average logical connection strength of the project documents is 0.85. The document's paragraph missing rate of 0.05 and logical association strength of 0.85 are matched against the preset document integrity requirements. Assuming the preset requirements stipulate a paragraph missing rate below 0.02 and a logical association strength above 0.9, the current missing rate of 0.05 is higher than 0.02, and the logical association strength of 0.85 is lower than 0.9, resulting in a mismatch. The system further compares the current project document's time frame with the integrity requirements. For example, if the project plan document's completion date is June 30th, and the integrity requirement mandates 80% document structure integrity by June 25th, but only 60% was actually completed by June 25th, the system combines the above matching and comparison results. For example, regarding "Intelligent Manufacturing System V1..."The project initiation documents for the "0 Development Project" are divided into partition nodes such as "Requirements Analysis Phase Documents," "System Design Phase Documents," and "Testing Phase Documents." The integrity information of each partition node is recorded within a specific time period. For example, the requirements analysis phase document records data from June 1st to June 20th, with a structural integrity deviation, a key field missing rate of 0.15, and a logical association strength of 0.75. This data is used to generate the partition node document recording period.
[0059] Table 1: Structural Integrity Table
[0060]
[0061] As shown in Table 1, this table records the structural integrity information of a software development project initiation document, including the document name, chapter name, and corresponding structural integrity status.
[0062] S112: Based on the document record time period of the partition node, extract the overlapping time period between the document integrity status and the required time interval, calculate the ratio of overlapping time to the total required duration, filter nodes with an overlap ratio lower than the benchmark value, and obtain the partition node integrity coverage deviation rate based on the number of integrity status annotations.
[0063] Based on the document recording period of the partitioned nodes, for example, the recording period of "Requirements Analysis Phase Documents" is from June 1st to June 20th, the overlapping period between the document integrity status and the requirement time interval is extracted. For example, if the document integrity status update period is from June 10th to June 22nd, and the requirement time interval is from June 1st to June 20th, then the overlapping period is from June 10th to June 20th, a total of 11 days. The ratio of the overlapping time to the total requirement duration is calculated. For example, if the total requirement duration is 20 days (June 1st to June 20th), and the overlapping period is 11 days, then the ratio is 11 / 20 = 0.55, or 55%. Nodes with an overlap ratio lower than the benchmark value are filtered out. The benchmark value is set empirically based on project type, project cycle, and document importance. For example, for requirement documents of core functional modules, the benchmark... A value of 0.8 indicates that at least 80% overlap is required to guarantee document validity. For non-core documents, the baseline value can be set to 0.6. For example, for the requirements analysis phase document of the "Intelligent Manufacturing System V1.0 Development Project", the baseline value is set to 0.70. Since the calculated overlap ratio of 0.55 is lower than the baseline value of 0.70, the "Requirements Analysis Phase Document" node is selected. Based on the number of integrity status annotations, for example, in the selected "Requirements Analysis Phase Document" node, the number of annotations for "Incomplete Document Structure" is 5, the number of annotations for "Missing Key Fields" is 3, and the number of annotations for "Logical Vulnerabilities" is 2, for a total of 10 integrity status annotations. The number of annotations is used to calculate the partition node integrity coverage deviation rate to obtain the partition node integrity coverage deviation rate.
[0064] S113: Based on the partition node integrity coverage deviation rate, determine the deviation status of the node number, identify the node number whose deviation rate exceeds the node synchronization threshold, integrate the node number, integrity coverage information and deviation rate value, and generate a document integrity deviation label.
[0065] Based on the integrity coverage deviation rate of partitioned nodes, for example, the integrity coverage deviation rate of "Requirements Analysis Phase Documents" is 0.25, the node number is judged for deviation status. The deviation rate is compared with the preset "Node Synchronization Threshold". The node synchronization threshold is set based on project management experience and risk control requirements. For example, for high-priority core document nodes, the threshold can be set to 0.2, meaning that a deviation rate exceeding 20% is considered abnormal. For low-priority non-core document nodes, the threshold can be set to 0.35, meaning that a deviation rate exceeding 35% is considered abnormal. For example, for the "Requirements Analysis Phase Documents" node, its node synchronization threshold is set to 0.20 because this document belongs to the core... The document identification process identifies node numbers whose deviation rate exceeds the node synchronization threshold. Since 0.25 is greater than 0.20, the node number of "Requirements Analysis Phase Document" (e.g., node number "DOC-REQ-001") is identified as a node with a deviation rate exceeding the threshold. The process integrates the node number, integrity coverage information, and deviation rate value. For example, integrating the node number "DOC-REQ-001", its integrity coverage information includes 5 structural incompletenesses, 3 missing key fields, and 2 logical vulnerabilities, with a deviation rate value of 0.25. This generates a document integrity deviation label, for example, "DOC-REQ-001: Integrity Deviation (Structure Incompleteness: 5, Key Field Missing: 3, Logical Vulnerability: 2, Deviation Rate: 0.25)".
[0066] Please see Figure 3 The specific steps for obtaining functional repeatability screening labels are as follows:
[0067] S211: Based on document integrity deviation tags, identify the hierarchical screening mode and functional description similarity distribution characteristics that are adapted to the current document characteristics, extract the start and end fluctuation ranges of real-time functional description similarity, and use the formula:
[0068] ;
[0069] Calculate the start and end fluctuation difference of the functional description similarity segment, compare it with the benchmark similarity fluctuation range, and obtain the functional description similarity fluctuation deviation value;
[0070] in, The difference in the start and end fluctuations of the segment representing the similarity in functional description. This represents the number of sampling points contained in the similarity segment. Representing the The weighting coefficients for each sampling point Representing the The functional description similarity value of each sampling point This represents the average functional description similarity of all sampled points within this similarity segment. This represents the initial similarity value of the current similarity segment. This represents the termination similarity value of the current similarity segment. A constant used to prevent the denominator from being zero, representing a positive number that approaches zero;
[0071] Based on document integrity deviation labels, such as "DOC-REQ-001: Integrity Deviation (Incomplete Structure: 5, Missing Key Fields: 3, Logical Vulnerabilities: 2, Deviation Rate: 0.25)," we identify a tiered screening mode and functional description similarity distribution characteristics that are suitable for the current document's features. For the requirement document "DOC-REQ-001," its characteristic is identified as a "requirement specification document," and based on this characteristic, the corresponding tiered screening mode is selected. For example, the "Key Functional Module Priority Screening Mode" is adopted. Under this mode, functional descriptions related to core business logic are prioritized for screening. Simultaneously, the functional description similarity distribution characteristics of the document are extracted. For example, through analysis of historical requirement documents, it is found that the functional description similarity of requirement specification documents follows a normal distribution, with an average similarity between 0.7 and 0.9. The real-time fluctuation range of the initial and final functional description similarity is monitored. The changes in the similarity of functional descriptions in the "DOC-REQ-001" document are monitored in real time. For example, within a certain time window, the functional description similarity of the document changes from the initial value... Fluctuation to the termination value Using formula The calculation function describes the start and end fluctuation difference of similarity segments, where, The difference between the start and end points of the functional description similarity segment quantifies the overall fluctuation of functional description similarity within a specific time period. This represents the number of sampling points contained in the similarity segment, indicating the number of functional description similarity data collected within this time period. Representing the The weight coefficients of each sampling point are used to reflect the importance of different sampling points in the fluctuation calculation. For example, the sampling point closer to the current time has a higher weight. Representing the The functional description similarity value for each sampling point is the specifically measured numerical value of functional description similarity. This represents the average functional description similarity of all sampled points within the similarity segment, used to measure the average similarity level of that segment. This represents the initial similarity value of the current similarity segment; it is the similarity value at the beginning of this time period. The termination similarity value represents the current similarity segment; it is the similarity value at the end of this time period. A positive number approaching zero is used as a constant to prevent the denominator from being zero; for example, taking... The left half of the formula The weighted dispersion of similarity sampling points relative to the mean is calculated, reflecting the fluctuation range of similarity within a segment. (Right half) The relative changes in the start and end values of similarity segments were calculated to reflect the trend of similarity. The two values were added together and their absolute values were taken to comprehensively assess the fluctuation range and trend of similarity, providing a comprehensive fluctuation difference index. Specifically, the calculation assumed... Each sampling point has a similarity value for its functional description. They are respectively , , , , Corresponding weight coefficient They are respectively , , , , Initial similarity value Termination similarity value , ;
[0072] First, calculate the average. : ;
[0073] Then calculate the summation term:
[0074] ;
[0075] Next, calculate the left half: ;
[0076] Calculate the right half: ;
[0077] Final calculation : ;
[0078] The advantage of this formula lies in its ability to comprehensively quantify the fluctuations in functional description similarity by combining the weighted dispersion of similarity sampling points relative to the average value with the relative changes in the start and end values of similarity segments. This helps to more accurately identify potential problems in functional repeatability screening, and the calculated difference in the start and end fluctuations of functional description similarity segments can be used to further quantify these fluctuations. Compared to the baseline similarity fluctuation range, the "baseline similarity fluctuation range" is determined based on historical project data and domain expert experience. For example, for stable and mature systems, the baseline similarity fluctuation range is set to [0.01, 0.05], indicating a smaller allowable fluctuation range. For highly innovative projects or projects with frequently changing requirements, the baseline range is set to [0.05, 0.15]. For example, the baseline similarity fluctuation range for this project is [0.03, 0.08]. If the similarity fluctuation exceeds the baseline range [0.03, 0.08], the functional description similarity fluctuation deviation value is obtained. For example, the functional description similarity fluctuation deviation value is... This indicates that the fluctuation exceeds the normal range.
[0079] S212: Call the function description similarity fluctuation deviation value, combine the segment distribution, fluctuation trend and adjustment frequency, uniformly collect the graded screening mode deviation data, identify and calculate the adaptive deviation degree according to the segment number, determine the fluctuation direction according to the adjustment frequency, and obtain the function repeatability screening label.
[0080] The similarity fluctuation deviation value of the function descriptions is used. For example, if the function description similarity fluctuation deviation value is 0.01886, the distribution of the corresponding function description segments in the entire document is analyzed, along with the segment distribution, fluctuation trend, and adjustment frequency. For example, this deviation is mainly concentrated in the "User Permission Management" and "Data Report Generation" function modules. The fluctuation trend is also analyzed; for example, the deviation shows a continuous upward trend, along with the adjustment frequency of related function descriptions. For instance, the "User Permission Management" function description has been adjusted 5 times in the past week, and the "Data Report Generation" function description has been adjusted 4 times. The functional descriptions were adjusted three times. To standardize and collect the deviation data for the tiered screening model, the following information was integrated: fluctuation deviation value of 0.01886, segment distribution (user access management, data report generation), fluctuation trend (continuously rising), and adjustment frequency (user access management 5 times, data report generation 3 times). This formed a tiered screening model deviation dataset. Adaptability deviation was identified and calculated according to segment numbers. Based on the number of each functional segment (e.g., "FUNC-UM-001" represents user access management, "FUNC-DR-002" represents data report generation...), the data was processed accordingly. The adaptive deviation is calculated by comparing the fluctuation deviation value of the functional description similarity of the segment with the historical average fluctuation deviation value of the segment. For example, the current fluctuation deviation value of the "FUNC-UM-001" segment is 0.01886, and its historical average fluctuation deviation value is 0.01, so the adaptive deviation is 0.01886 - 0.01 = 0.00886. The fluctuation direction is determined based on the adjustment frequency. For example, the adjustment frequency of the "User Permission Management" functional description is 5 times, which is higher than the average adjustment frequency of 2 times, so its fluctuation direction is determined. For "positive fluctuation" (i.e., increased fluctuation), the "Data Report Generation" function description adjusts 3 times, which is close to the average adjustment frequency. The fluctuation direction is judged to be "stable fluctuation", and a functional repeatability screening label is obtained. For example, the functional repeatability screening label generated for "FUNC-UM-001" is "high repeatability risk (fluctuation direction: positive fluctuation, fitness deviation: 0.00886)", and the functional repeatability screening label generated for "FUNC-DR-002" is "medium repeatability risk (fluctuation direction: stable fluctuation, fitness deviation: 0.005)".
[0081] Please see Figure 4 The specific steps for obtaining the list of technically infeasible affected sections are as follows:
[0082] S311: Call the function repetition screening label, filter the function description similarity deviation task segment number, extract the function description similarity time period according to the node association table, process the segment function description time period by time dimension, identify the function description similarity time period index table, and obtain the function description similarity task time period set.
[0083] The system invokes the functional repetition screening label, for example, "FUNC-UM-001: High repetition risk (fluctuation direction: positive fluctuation, adaptability bias: 0.00886)". It then filters the functional description similarity deviation task segment numbers, selecting those with "high repetition risk" or "medium repetition risk" from the labels. For example, it filters out the segment numbers "FUNC-UM-001" and "FUNC-DR-002". Based on the node association table, it extracts the functional description similarity time period and queries the pre-established "node association table," which records the mapping relationship between functional segment numbers and corresponding functional description similarity time periods. For example, the query reveals the functional description similarity time period for "FUNC-UM-001". For the period "June 10th - June 15th", the functional description time periods are processed according to the time dimension. The extracted functional description similarity time periods are organized and categorized in chronological order. For example, if multiple functional segments have fluctuating similarity within the same time period, they will be grouped into the same processing batch. An index table of functional description similarity time periods is identified. An index table is generated based on the processed functional description time periods. For example, the index table contains the start and end times of the time periods and the associated functional segment numbers. The set of functional description similarity task time periods is obtained. For example, the set obtained is {"June 10th - June 15th" associated with "FUNC-UM-001", "June 12th - June 17th" associated with "FUNC-DR-002"}.
[0084] S312: Based on the set of task time periods for functional description similarity, collect the evaluation data of technical infeasibility items in the same period, identify the impact table of technical infeasibility items, judge the daily technical infeasibility item interference based on the interference threshold, match it with the task time periods for functional description similarity, determine whether there is an abnormal association of functional description similarity, and generate a list of technical infeasibility item impact sections.
[0085] Based on the set of task time periods with similarity in functional descriptions, for example, {"June 10th - June 15th" associated with "FUNC-UM-001", "June 12th - June 17th" associated with "FUNC-DR-002"}, technical infeasibility assessment data for the same period is collected. For each task time period, technical infeasibility assessment data for that time period is collected from the project technical assessment database. For example, for the "June 10th - June 15th" time period, technical infeasibility assessment data related to "FUNC-UM-001" is collected. For instance, the assessment report indicates a technical compatibility issue when integrating third-party authentication services into the user permissions module, highlighting the technical infeasibility... The Impact Table for Technical Infeasibility: Based on the collected assessment data of technical infeasibility items, an impact table for technical infeasibility items is identified. This table lists specific technical infeasibility items and their scope and severity of impact. For example, it is identified that "third-party authentication service compatibility issues" will affect the development progress of the "user management module." Daily technical infeasibility interference is determined based on an interference threshold. The interference threshold is set according to the project's risk tolerance and technical complexity. For example, for tasks on the critical path, the threshold is set to 0.1 (i.e., a task delay exceeding 10% due to a technical infeasibility item is considered interference), while for non-critical tasks, the threshold can be set to 0.25. Whether feasible items interfere with the task is determined. For example, if the "third-party authentication service compatibility issue" caused a 2-day delay to the development of "FUNC-UM-001" on June 12th, and the ratio of this delay to the original planned task time exceeds a preset interference threshold of 0.1, then interference exists. For example, if the original plan was for the task to be completed in 10 days, a 2-day delay results in a delay ratio of 2 / 10 = 0.2, which is greater than 0.1, thus indicating interference. The interference of technically infeasible items is matched with the similarity of the functional description to the task time period. For example, the similarity between "the interference of the third-party authentication service compatibility issue on June 12th" and "FUNC-UM-001" is compared. During the task period (June 10th-June 15th), matching was performed to determine if there were any abnormal associations in the similarity of function descriptions. For example, the matching revealed that during the period of interference from the "third-party certification service compatibility issue" (June 12th), the similarity fluctuation deviation value of the function description "FUNC-UM-001" increased sharply from 0.01886 to 0.05. This indicates that technical infeasibility caused instability in the function description. A list of affected segments due to technical infeasibility was generated. For example, the generated list entries were {Segment Number: FUNC-UM-001, Technical Infeasibility: Third-Party Certification Service Compatibility Issue, Interference Date: June 12th, Abnormal Association in Similarity: Exists}.
[0086] Please see Figure 5 The specific steps for obtaining the compliance deviation volatility group are as follows:
[0087] S411: Based on the list of technically infeasible items affecting the segment, filter out unmarked document compliance verification nodes, extract the task list, and obtain the task set of nodes not affected by technically infeasible items;
[0088] Based on the list of technically infeasible impact segments, for example, {Segment Number: FUNC-UM-001, Technically Infeasible Item: Third-Party Certification Service Compatibility Issue, Interference Date: June 12th, Abnormal Similarity Association: Exists}, filter unmarked document compliance verification nodes. From all document compliance verification nodes to be verified, filter out those nodes that do not appear in the list of technically infeasible impact segments. For example, assuming that in addition to FUNC-UM-001, there are also nodes FUNC-DR-002 (data report generation) and FUNC-RE-003 (report export), if the list of technically infeasible impact segments only contains FUNC-UM-001, then FUNC-DR-002 and... FUNC-RE-003 is an unmarked node. Extract the task list. For the unmarked document compliance verification node, extract the task list associated with it. For example, for the FUNC-DR-002 node, the extracted task list includes "Data report format standardization check" and "Report data consistency check". Obtain the task set of the node that is not affected by technical infeasibility. For example, the obtained task set is {Task number: TASK-DR-001, Task name: Data report format standardization check, Associated node: FUNC-DR-002} and {Task number: TASK-DR-002, Task name: Report data consistency check, Associated node: FUNC-DR-002}.
[0089] S412: Call the task set of nodes not affected by technical infeasibility, match compliance management specifications and operational data, extract the planned task volume by task number, count the number of operators and real-time operation time period, and generate a compliance execution status matching dataset.
[0090] The system invokes task sets unaffected by technical infeasibility issues, such as {Task ID: TASK-DR-001, Task Name: Data Report Format Standardization Check, Associated Node: FUNC-DR-002} and {Task ID: TASK-DR-002, Task Name: Report Data Consistency Verification, Associated Node: FUNC-DR-002}. It matches compliance management specifications with operational data, matching the execution status of each task with preset "compliance management specifications." For example, the compliance management specification for the "Data Report Format Standardization Check" task requires completion within 2 hours and execution by 2 quality inspectors with intermediate-level authority. The system then matches actual operational data with this specification; for instance, the actual operational data for task TASK-DR-001 shows that the task took 2.5 hours and was executed by 1 quality inspector with senior authority. This task is limited to quality inspectors and one junior-level quality inspector. The planned task quantity is extracted by task number. For example, the planned task quantity for TASK-DR-00101 is to inspect 100 reports, and the planned task quantity for TASK-DR-002 is to verify 200 data entries. The number of operators and the real-time operation period are also considered. For example, the number of operators for task TASK-DR-001 is 2, and the real-time operation period is June 16th, 9:00-11:30. A compliance execution matching dataset is generated. For example, the generated dataset contains {Task Number: TASK-DR-001, Planned Task Quantity: 100 reports, Actual Number of Operators: 2, Real-time Operation Period: June 16th, 9:00-11:30, Compliance Matching Result: Partially Mismatched}.
[0091] S413: Match the dataset based on compliance performance, assess the degree of matching between task execution efficiency and operation distribution, identify efficiency fluctuation nodes, and use the following formula:
[0092] ;
[0093] Calculate the fluctuation identification difference value, mark tasks with fluctuations exceeding the benchmark value as abnormal nodes, identify the compliance task execution matching deviation, and obtain the compliance deviation fluctuation group;
[0094] in, Represents the difference value for identifying fluctuations. This represents the actual execution weight of the task at time t. This represents the average operation time of the task node at time t. This represents the resource consumption value of the task at time t. This represents the available resource capacity of a task node at time t. This represents the total number of time intervals during task execution. This represents the historical average fluctuation value under the baseline task set;
[0095] Based on the compliance execution status, a dataset is matched. For example, {Task ID: TASK-DR-001, Planned task quantity: 100 reports, Actual number of operators: 2, Real-time operation time: June 16th, 9:00-11:30, Compliance specification matching result: Partially inconsistent}. The degree of matching between task execution efficiency and operation distribution is evaluated. The execution efficiency of task TASK-DR-001 (100 reports / 2.5 hours) is assessed to see if it matches the expected efficiency. The distribution of the number of operators (2) and the operation time (2.5 hours) is also evaluated to see if it is reasonable. For example, if the plan is for 2 people to complete 100 reports in 2 hours, and the actual completion time is 2 people in 2.5 hours, the efficiency is slightly lower than expected. Efficiency fluctuation points are identified by comparing the actual execution efficiency with the planned execution efficiency. For example, task TASK-DR-001 is identified as an efficiency fluctuation point because it takes longer than planned. A formula is used to identify these fluctuation points. Calculate the fluctuation identification difference value, where, This represents the variance in fluctuation identification, quantifying the deviation between task execution efficiency and the degree of matching between the operation distribution. Representing the task in time The actual execution weight value reflects the importance of the task execution or the proportion of resources invested at that point in time. Represents the task node in time The average operation time represents the average time spent executing the task at that point in time. Representing the task in time The resource consumption value represents the actual amount of resources (such as CPU time, memory, and network bandwidth) consumed during task execution at that point in time. Represents the task node in time The available resource capacity value represents the maximum amount of resources available to the task at that point in time. This represents the total number of time intervals during task execution, indicating the total number of time periods the task lasts from start to finish. Representing the historical mean volatility under the baseline task set, it is used to provide a standard reference point for assessing current volatility differences. The formula is expressed through the numerator... The weighted total time consumed during the actual execution of the task was calculated, reflecting the actual efficiency. The denominator... The combined impact of resource consumption and available capacity was quantified, reflecting the rationality of operational distribution, and ultimately compared with historical fluctuation averages. By comparing the values, the fluctuation identification difference value is obtained.
[0096] The innovation lies in the fact that this formula comprehensively considers multiple dimensions of parameters, such as actual execution weight, average operation time, operation resource consumption, and available resource capacity, to fully assess the matching degree between task execution efficiency and operation distribution. This allows for a more accurate identification of compliance deviations and fluctuations. Specifically, the calculation assumes the total number of execution time intervals for task TASK-DR-001 is [not specified]. (For example, divided into morning and afternoon time periods), the parameter values are shown in Table 2. ;
[0097] Table 2: TASK-DR-001 Task Execution Parameters
[0098]
[0099] As shown in Table 2, this table records the actual execution weight value, average operation time, operation resource consumption value and available resource capacity value of TASK-DR-001 task in two time intervals.
[0100] Calculate the numerator : ;
[0101] Calculate the denominator :
[0102] ; ;
[0103] calculate : ;
[0104] Tasks with fluctuations exceeding a baseline value are marked as anomalous nodes. The baseline value is set based on project compliance requirements and task importance. For example, for core compliance tasks, the baseline value can be set to 0.1, meaning a fluctuation difference exceeding 10% is considered anomalous. For non-core compliance tasks, the baseline value can be set to 0.2. For instance, the baseline value for compliance tasks is 0.10, based on the calculated fluctuation difference value. The value exceeded the baseline of 0.10, so the TASK-DR-001 task was marked as an abnormal node. The compliance task execution matching deviation was identified. This result indicates that the execution efficiency and operation distribution of the TASK-DR-001 task deviate significantly from the compliance requirements, and there is a potential compliance risk. A compliance deviation fluctuation group was obtained. For example, the TASK-DR-001 task and its related deviation information (e.g., execution efficiency is lower than expected, and the number of operators does not conform to the specifications) are classified into the compliance deviation fluctuation group.
[0105] Please see Figure 6 The specific steps for obtaining the structural indicators for project progress monitoring are as follows:
[0106] S511: Call the compliance deviation fluctuation group, filter nodes that exceed the threshold, record the time interval and the magnitude of task volume change, and obtain the progress fluctuation anomaly identifier set;
[0107] The compliance deviation fluctuation group is invoked, for example, which includes task TASK-DR-001 marked as an anomalous node. Nodes exceeding a threshold are filtered out. This threshold is set based on project schedule management requirements and risk levels. For example, for nodes affecting project milestones, the threshold is set to "fluctuation difference greater than 0.1," and for minor nodes, it is set to "fluctuation difference greater than 0.2." The fluctuation identification difference value of each task in the compliance deviation fluctuation group is compared with the preset threshold. For example, if the fluctuation identification difference value of task TASK-DR-001 is 0.14277, then... If the threshold is set to 0.12, the node exceeds the threshold. Record the time interval and the change in task volume. For nodes that exceed the threshold, such as TASK-DR-001, record the time interval of the anomaly, such as June 16, and the change in task volume, such as the original plan to check 100 reports, but due to efficiency issues, only 80 reports were completed that day, and the change in task volume was -20%. Obtain the progress fluctuation anomaly identifier set, such as the identifier set containing {task number: TASK-DR-001, anomaly time interval: June 16, task volume change: -20%}.
[0108] S512: Based on the project model resources corresponding to the nodes in the schedule fluctuation anomaly identifier set, identify the node resource input ratio sequence, extract the abnormal distribution interval and compare the critical coefficient, record the ratio offset direction and node number, and form a project resource matching offset index group.
[0109] Based on the schedule fluctuation anomaly identifier set, for example, {Task Number: TASK-DR-001, Anomaly Time Interval: June 16th, Task Volume Change Range: -20%}, identify the node resource input ratio sequence. For task TASK-DR-001, extract its corresponding resource input data from the project initiation model. For example, if the planned input is 8 hours of development personnel A, 4 hours of testing personnel B, and 2 hours of server resources C, calculate the ratio of actual resource input to planned resource input to form a resource input ratio sequence. For example, if the actual input is 10 hours of development personnel A, 6 hours of testing personnel B, and 2 hours of server resources C, then the ratio of development personnel is 10 / 8 = 1.25, the ratio of testing personnel is 6 / 4 = 1.5, and the ratio of server resources is 2 / 2 = 1, with a sequence of {1.25, 1.5, 1}. Extract the anomaly distribution interval and compare it with the critical coefficient. The setting of the "critical coefficient" is determined according to the project resource management strategy and risk control requirements. For example, for human resources, the critical coefficient can be set to 1.2, indicating that the actual input... An input exceeding the planned amount by 20% is considered abnormal. For critical equipment resources, a critical coefficient can be set to 1.1, meaning that an actual input exceeding the planned amount by 10% is considered abnormal. Analyzing the resource input ratio sequence identifies ratios exceeding the normal range, forming abnormal distribution intervals. For example, comparing ratios of 1.25 and 1.5 with the critical coefficient of 1.2 reveals that both 1.25 and 1.5 exceed the critical coefficient of 1.2, forming an abnormal distribution interval. The direction of ratio deviation and node number are recorded; for example, recording a development personnel resource ratio of 1.2. A value of 5 indicates a "positive offset" (actual investment exceeds the planned investment). A tester resource ratio of 1.5 indicates a "positive offset," and the corresponding node number TASK-DR-001 is recorded to form a project resource matching offset indicator group. For example, the indicator group could be {Node Number: TASK-DR-001, Resource Type: Developers, Ratio: 1.25, Offset Direction: Positive Offset} or {Node Number: TASK-DR-001, Resource Type: Testers, Ratio: 1.5, Offset Direction: Positive Offset}.
[0110] S513: Based on the project resource matching offset index group, extract the time period of project model resource deployment and task execution, identify the difference between resource deployment cycle and operation cycle, and sort and label according to the progress benchmark to generate project progress monitoring structure indicators.
[0111] Based on the resource matching offset indicator group for project initiation, for example, {Node ID: TASK-DR-001, Resource Type: Developers, Ratio: 1.25, Offset Direction: Positive Offset}, {Node ID: TASK-DR-001, Resource Type: Testers, Ratio: 1.5, Offset Direction: Positive Offset}, extract the resource deployment and task execution time periods from the project initiation model. Extract the planned resource deployment time period for task TASK-DR-001 from the project initiation model; for example, the planned deployment time for developers is June 10th-June 15th, and for testers, it is June 14th-June 16th. Simultaneously, extract the actual execution time period for the task; for example, the actual execution time for developers is June 10th-June 17th, and for testers, it is June 15th-June 18th. Identify the difference between the resource deployment cycle and the operation cycle, comparing the planned resource deployment cycle with the actual operation cycle; for example, if the planned deployment for developers is 5 days, and the actual operation is 8 days, the difference is significant. The timeframe is 3 days; the testers planned to deploy resources for 3 days, but actually operated for 4 days, resulting in a 1-day difference. The results are sorted and labeled according to the progress baseline. The progress baseline is determined based on the project plan and critical path analysis. For example, for critical tasks, the progress baseline requires a 100% completion rate with no delays; for non-critical tasks, a small delay is allowed. The difference results are compared with the progress baseline and sorted. For example, if the developer resource deployment difference is 3 days, and the progress baseline allows a 1-day delay, the developer resource deployment difference is labeled as "Severe Delay." If the tester resource deployment difference is 1 day, and the progress baseline also allows a 1-day delay, the tester resource deployment difference is labeled as "Normal." Project progress monitoring structure indicators are generated. For example, the generated indicators are: {Task Number: TASK-DR-001, Developer Resource Deployment Difference: 3 days, Label: Severe Delay}, {Task Number: TASK-DR-001, Tester Resource Deployment Difference: 1 day, Label: Normal}.
[0112] The material data cleaning and grading screening system is used to perform the above-mentioned material data cleaning and grading screening methods. The system includes:
[0113] The integrity deviation extraction module obtains document structure integrity information, key field missing distribution patterns and logical vulnerability marker lists from digital project initiation materials, extracts document paragraph missing rate and logical connection strength, compares document integrity requirements with real-time distribution, and generates document integrity deviation tags.
[0114] The hierarchical screening mode classification module is based on document integrity deviation labels to locate the hierarchical screening mode that is suitable for the current document characteristics, extracts document node identifiers, process nodes and benchmark similarity, calculates the difference between on-site similarity and benchmark similarity, classifies and labels the difference type, and generates node mode adaptability identifiers.
[0115] The infeasibility identification module uses node pattern adaptive identification to filter task components with lagging similarity in function description, locate corresponding segments, extract interference periods for technical infeasibility evaluation, determine overlap with operation time periods, filter frequently interfering segments, and generate an infeasibility interference mapping set for project progress.
[0116] The compliance diagnosis module is based on the infeasibility interference mapping set of the project initiation progress, eliminates interference segment tasks, extracts compliance management specifications and operation data, matches task assignment and operation time periods, calculates the ratio of operation volume to task, identifies task-intensive and inefficient component tasks, and obtains the node operation execution deviation set.
[0117] The resource allocation analysis module locates the resource input records of tasks in the project initiation model based on the node operation execution deviation set, extracts the ratio of task quantity to resource quantity allocation, compares the difference between the operation cycle and the input cycle, maps the task resource usage and progress status, and forms a project initiation progress monitoring structure indicator.
[0118] The above are merely preferred embodiments of the present invention and are not intended to limit the present invention in any other way. Any person skilled in the art may make changes or modifications to the above-disclosed technical content to create equivalent embodiments that can be applied to other fields. However, any simple modifications, equivalent changes, and modifications made to the above embodiments based on the technical essence of the present invention without departing from the scope of the present invention shall still fall within the protection scope of the present invention.
Claims
1. A method for cleaning and classifying material data, characterized in that, Includes the following steps: S1: Obtain information on the document structure integrity, key field missing distribution patterns, and logical vulnerability marker list from the digital project initiation materials; extract the document paragraph missing rate and logical connection strength; and generate document integrity deviation tags. S2: Based on the document integrity deviation label, select a hierarchical screening mode that is suitable for the current document characteristics, extract the functional description similarity distribution range from the original project database, and perform deviation correction and classification processing with the benchmark similarity to obtain the functional duplication screening label. S3: Call the function repetition screening label, extract the function description similarity segment number, identify the technical feasibility assessment data within the segment, compare the technically infeasible items with the function description segment, record the number of overlapping time periods, and generate a list of technically infeasible items affecting segments; S4: Based on the list of technically infeasible items affecting the affected areas, filter out unmarked document compliance verification nodes, extract compliance management specifications and operational data, map the distribution of construction content and compliance requirements, determine whether there is a compliance deviation, and obtain the compliance deviation fluctuation group.
2. The material data cleaning and grading screening method according to claim 1, characterized in that, The document integrity deviation label includes deviation number, missing status label, logical correlation deviation amount, and pattern classification. The functional duplication screening label includes similarity level, fluctuation type, benchmark comparison result, and task association number. The list of technically infeasible item impact segments includes device type, impact time segment, number of overlapping time periods, and affected task number. The compliance deviation fluctuation group includes task distribution unevenness number, operation duration record, task completion deviation amount, and operation matching degree.
3. The material data cleaning and grading screening method according to claim 1, characterized in that, The specific steps for obtaining the document integrity deviation tag are as follows: S111: Obtain document structure integrity information, key field missing distribution patterns, and logical vulnerability marker list from the digital project initiation materials; extract document paragraph missing rate and logical correlation strength; match document integrity requirements with real-time distribution; compare time range with integrity requirements; and generate partition node document record time periods. S112: Based on the document record time period of the partition node, extract the overlapping time period between the document integrity status and the required time interval, calculate the ratio of overlapping time to the total required duration, filter nodes with an overlap ratio lower than the benchmark value, and obtain the partition node integrity coverage deviation rate based on the number of integrity status annotations. S113: Based on the partition node integrity coverage deviation rate, determine the deviation status of the node number, identify the node number whose deviation rate exceeds the node synchronization threshold, integrate the node number, integrity coverage information and deviation rate value, and generate a document integrity deviation label.
4. The material data cleaning and grading screening method according to claim 3, characterized in that, The specific steps for obtaining the functional repeatability screening label are as follows: S211: Based on the document integrity deviation label, identify the hierarchical screening mode and functional description similarity distribution characteristics that are adapted to the current document characteristics, extract the start and end fluctuation range of real-time functional description similarity, calculate the start and end fluctuation difference of functional description similarity segment, compare it with the benchmark similarity fluctuation range, and obtain the functional description similarity fluctuation deviation value. S212: Call the function description similarity fluctuation deviation value, combine the segment distribution, fluctuation trend and adjustment frequency, uniformly collect the graded screening mode deviation data, identify and calculate the adaptive deviation degree according to the segment number, determine the fluctuation direction according to the adjustment frequency, and obtain the functional repeatability screening label.
5. The material data cleaning and grading screening method according to claim 4, characterized in that, The specific steps for obtaining the list of affected sections due to the aforementioned technical infeasibility are as follows: S311: Call the function repetition screening label, filter the function description similarity deviation task segment number, extract the function description similarity time period according to the node association table, process the segment function description time period according to the time dimension, identify the function description similarity time period index table, and obtain the function description similarity task time period set. S312: Based on the set of time periods for the function description similarity task, collect the evaluation data of technical infeasibility items in the same period, identify the table of technical infeasibility item impact, determine the daily technical infeasibility item interference based on the interference threshold, match it with the time period for the function description similarity task, determine whether there is an abnormal association of function description similarity, and generate a list of technical infeasibility item impacted sections.
6. The material data cleaning and grading screening method according to claim 5, characterized in that, The specific steps for obtaining the compliance deviation volatility group are as follows: S411: Based on the list of affected sections by the aforementioned technical infeasibility items, filter out unmarked document compliance verification nodes, extract the task list, and obtain the task set of nodes not affected by technical infeasibility items; S412: Call the task set of nodes not affected by technical infeasibility, match compliance management specifications and operation data, extract the planned task volume by task number, count the number of operators and real-time operation time period, and generate a compliance execution status matching dataset. S413: Based on the compliance execution status matching dataset, evaluate the degree of matching between task execution efficiency and operation distribution, identify efficiency fluctuation nodes, calculate fluctuation identification difference value, mark tasks with fluctuations exceeding the benchmark value as abnormal nodes, identify compliance task execution matching deviations, and obtain compliance deviation fluctuation groups.
7. The material data cleaning and grading screening method according to claim 1, characterized in that, The method also includes step S5: S5: Call the compliance deviation fluctuation group, extract the abnormal fluctuation task group, calculate the ratio of resource input to task in the project establishment model, record the difference distribution between resource allocation cycle and task execution cycle, and generate project progress monitoring structure indicators. The project initiation progress monitoring structure indicators include resource allocation ratio, task intensity level, execution cycle difference, and project initiation efficiency indicators.
8. The material data cleaning and grading screening method according to claim 7, characterized in that, The specific steps for obtaining the project initiation progress monitoring structural indicators are as follows: S511: Call the compliance deviation fluctuation group, filter nodes that exceed the threshold, record the time interval and the magnitude of task volume change, and obtain the progress fluctuation abnormal identifier set; S512: Based on the project establishment model resources corresponding to the nodes in the progress fluctuation anomaly identifier set, identify the node resource input ratio sequence, extract the abnormal distribution interval and compare the critical coefficient, record the ratio offset direction and node number, and form a project establishment resource matching offset index group. S513: Based on the project initiation resource matching offset index group, extract the time period for project initiation model resource deployment and task execution, identify the difference between the resource deployment cycle and the operation cycle, and sort and label them according to the progress benchmark to generate project progress monitoring structure indicators.
9. A material data cleaning and grading screening system, characterized in that, The system is used to implement the material data cleaning and classification screening method according to any one of claims 1-8, the system comprising: The integrity deviation extraction module obtains document structure integrity information, key field missing distribution patterns and logical vulnerability marker lists from digital project initiation materials, extracts document paragraph missing rate and logical connection strength, compares document integrity requirements with real-time distribution, and generates document integrity deviation tags. The hierarchical screening mode classification module locates the hierarchical screening mode that is suitable for the current document characteristics based on the document integrity deviation label, extracts the document node identifier, process node and benchmark similarity, calculates the difference between the on-site similarity and the benchmark similarity, classifies and labels the difference type, and generates node mode adaptability identifier. The infeasibility identification module, based on the node pattern adaptive identifier, filters task components with lagging functional description similarity, locates the corresponding segment, extracts the interference period of technical infeasibility evaluation, determines the overlap with the operation time period, filters frequently interfered segments, and generates a set of interference mappings for infeasibility in project progress. The compliance diagnosis module, based on the infeasibility interference mapping set of the project initiation progress, eliminates interference segment tasks, extracts compliance management specifications and operation data, matches task assignments and operation time periods, calculates the ratio of operation volume to task, identifies task-intensive and inefficient component tasks, and obtains the node operation execution deviation set. Based on the node operation execution deviation set, the resource allocation analysis module locates the resource input records of the task in the project initiation model, extracts the ratio of task quantity to resource quantity allocation, compares the difference between the operation cycle and the input cycle, maps the task resource usage and progress status, and forms a project initiation progress monitoring structure index.
Citation Information
Patent Citations
Duplicate-checking comparison method for science and technology projects
CN105718506A
Medical equipment remote control system based on Internet of Things data
CN119916667A