Data collection method and system based on improved feature extraction and capture
Through multi-source data preprocessing and dynamic feature extraction, combined with parallel extraction windows and dynamic monitoring cycles, the problems of incomplete feature extraction and poor model adaptability in the prior art are solved, and efficient and accurate data collection is achieved.
Patent Information
- Application Number
- CN202510857235.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-25
- Publication Date
- 2025-07-25
AI Technical Summary
When the prior art faces semi-structured or unstructured data, feature extraction is incomplete or incorrectly extracted, and the model feature input dimension is limited, making it difficult to deal with data changes in complex scenarios, affecting the collection efficiency and accuracy.
Multi-source data preprocessing, dynamic feature extraction, bidirectional frequency analysis and matching capture methods are adopted to achieve real-time response to data feature changes and high-frequency priority matching by building parallel extraction windows and dynamic monitoring cycles.
It significantly improves the integrity and accuracy of data collection, supports automated collection of heterogeneous data, has modular design, clear data flow, and strong logical closed loop, and is suitable for data processing in complex scenarios.
Smart Images

Figure CN120372519A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of digital signal processing, and in particular relates to a data collection method and system based on improved feature extraction and capture. Background Art
[0002] With the development of information technology and the rapid expansion of data scale, data collection has become a basic link in processes such as data analysis, modeling training and intelligent decision-making. In many industries such as government affairs, finance, medical care, and transportation, the demand for automated collection of massive heterogeneous data continues to grow. Traditional data collection methods mostly rely on feature extraction technology based on keywords, structural templates, or rule matching to extract and aggregate target fields in data sources, and then complete the integration of multi-source data. However, due to factors such as the complex structure of the original data, inconsistent semantics, and large noise interference, the existing technology still has significant deficiencies in collection accuracy, robustness, and versatility.
[0003] Especially when facing semi-structured or unstructured data (such as logs, texts, images, etc.), traditional feature extraction methods do not have a strong understanding of contextual semantics, which often leads to incomplete feature capture or mis-extraction. In addition, although some systems have introduced model training methods to improve the collection effect, the model's feature input dimensions are limited and lack a dynamic adaptation mechanism, making it difficult to handle data changes in complex scenarios, affecting the overall collection efficiency and accuracy.
[0004] Therefore, how to build a data collection solution that has multi-dimensional feature perception, semantic understanding and adaptation capabilities to improve the extraction accuracy and collection efficiency in complex data environments has become one of the technical problems that need to be urgently solved in the current data processing field. Summary of the invention
[0005] In view of the above problems, the present invention aims to propose: a data collection method based on improved feature extraction and capture, comprising: S1. Data preprocessing: Acquire multi-source data and perform preprocessing to obtain multiple preprocessed sub-data and corresponding pre-feature information; S2, feature extraction and classification: extracting key features and secondary features from the pre-feature information, and determining a bidirectional feature set according to the acquisition frequency, wherein the bidirectional feature set includes a high-frequency feature set and a low-frequency feature set; S3, monitoring window construction: constructing a monitoring period, and setting a parallel extraction window within the monitoring period; S4, dynamic feature monitoring: at the start node of each parallel extraction window, real-time new data monitoring is performed, and at the end of the parallel extraction window, periodic feature evaluation is performed to obtain high-frequency feature items and low-frequency feature items; S5. Feature matching and collection output: All the high-frequency feature items are input into the capture feature queue, and the high-frequency feature items in the capture feature queue are associated and matched with the pre-processed sub-data one by one according to the matching degree, to obtain a data collection output set, and upload it to the data center.
[0006] As a preferred solution, step S1 includes: Obtain multi-source data, perform cleaning, denoising, and completion preprocessing operations, and obtain multiple preprocessed sub-data; Analyze and extract features from the preprocessed sub-data, and integrate them into pre-feature information; A standardized template is obtained, and the pre-characteristic information is input into the standardized template to obtain the pre-characteristic information in a unified format.
[0007] As a preferred solution, step S2 includes: Obtain all the pre-feature information, and perform feature vectorization and normalization processing; Obtaining a standard evaluation period, and dividing the standard evaluation period into a plurality of standard sampling periods; Several random samples are constructed in each standard sampling period, and feature screening is performed on each random sample to obtain the matching priority of key features. The matching priorities are arranged in order from large to small, and key features with the same rank are output together, and then secondary features are marked as non-priority items.
[0008] As a preferred solution, in step S2, the operation of determining the bidirectional feature set according to the acquisition frequency includes: The cutting amplitude is calculated based on the number of executions of the parallel extraction period within the standard evaluation period; Summarizing the frequencies of collecting the pre-feature information in the plurality of standard evaluation periods, and taking the average output frequency of the key features as the value condition; Get the judgment function, and combine the value condition with the judgment function to obtain the head and tail interval values; According to the cutting amplitude of the parallel extraction period, the key feature set corresponding to the head and tail interval values is marked as a low-frequency feature set, and other feature sets are marked as high-frequency feature sets.
[0009] As a preferred solution, the step of marking the key features corresponding to the head and tail interval values as a low-frequency feature set includes: Obtaining the cutting amplitudes of all parallel extraction periods from the standard evaluation period and setting the time length for them; According to the length of multiple standard evaluation periods, the gradient of the head and tail interval values is set; According to the head interval value of the gradient, the position of the corresponding key feature in the head and tail intervals is obtained and marked as a low-frequency feature set.
[0010] As a preferred solution, after obtaining the cutting amplitudes of all parallel extraction periods within the standard evaluation period and setting their time lengths, the start nodes of each of the parallel extraction periods are input into the parallel extraction channels, the source of the pre-features input into the channels is recorded, and they are extracted as monitoring data.
[0011] As a preferred solution, in step S4, the step of performing periodic feature evaluation at the end of the parallel extraction period to obtain high-frequency feature items and low-frequency feature items includes: Determine the dynamic monitoring period according to the update period of the multi-source data; Obtain the collection time points of the pre-feature information, and sequentially label each of the collection time points as the aging characteristics of the pre-feature information in chronological order; Measure the discontinuous sub-intervals within the dynamic monitoring period, and count the number of the discontinuous sub-intervals; Mark the key feature items whose aging characteristics are delayed beyond the discontinuous sub-intervals within the dynamic segmentation period as low-frequency feature items, and then mark the secondary feature items as high-frequency feature items.
[0012] As a preferred solution, after marking the key feature items whose delay exceeds the discontinuous sub-intervals as low-frequency feature items, the low-frequency feature items that repeatedly appear in multiple dynamic monitoring periods are screened out.
[0013] As a preferred solution, after screening out the low-frequency feature items that repeatedly appear in multiple dynamic monitoring periods, obtain the number of key feature marks of the low-frequency feature items and record it as the number of executions; calculate and output the execution ratio according to the total number of key feature items in each dynamic segmentation period; when the number of executions exceeds three and the execution ratio is greater than 0.5, mark it as an abnormal collection warning sample.
[0014] The present invention also provides a data collection system based on improved feature extraction and capture for implementing the data collection method based on improved feature extraction and capture, including: A data preprocessing module, configured to obtain multi-source data and perform preprocessing to obtain a plurality of preprocessed sub-data and corresponding pre-feature information; An information extraction module, configured to extract key features and secondary features, and construct a high-frequency feature set and a low-frequency feature set according to the collection frequency; A monitoring construction module, configured to construct a monitoring period and set a parallel extraction window; A feature item capture module, configured to perform data monitoring and periodic evaluation in the parallel extraction window to obtain high-frequency and low-frequency feature items; The association matching module is used to input high-frequency feature items into the capture queue and match them with the pre-processed sub-data, generate a collection output set and upload it to the data center.
[0015] Beneficial Effects The present invention provides a data collection method and system based on improved feature extraction and capture, which has the following beneficial effects: The present invention adopts a multi-stage process including multi-source data preprocessing, dynamic feature extraction, bidirectional frequency analysis and matching capture, which effectively solves the problems of poor adaptability to feature changes and easy omission of low-frequency features in existing collection methods. By constructing parallel extraction windows and dynamic monitoring cycles, real-time response to data feature changes and high-frequency priority matching mechanisms are achieved, thereby significantly improving the integrity and accuracy of data collection.
[0016] The system of the present invention integrates functional modules such as data preprocessing, feature extraction, monitoring construction, feature capture and association matching, and has the advantages of modular design, clear data flow, strong logic closed loop, etc. The system supports heterogeneous data access and standardized output, can automatically record feature tagging trajectories, supports low-frequency feature cleaning and abnormal warning mechanism, and ensures the stability, scalability and intelligence of the entire collection process.
[0017] The present invention is applicable to scenarios of automated collection of complex structures and multi-source heterogeneous data in industries such as government affairs, finance, medical care, and transportation. It can effectively support core data engineering tasks such as data center construction, data label update, and indicator management. Especially in actual environments where data changes frequently and features are highly heterogeneous, it has strong adaptability and robustness, providing stable and high-quality data support for upper-level modeling and intelligent decision-making. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] Figure 1 It is a schematic diagram of the method flow of the present invention; Figure 2 It is a schematic diagram of the system structure of the present invention. DETAILED DESCRIPTION
[0019] In order to deepen the understanding of the present invention, the present invention will be further described in detail below in conjunction with examples. The examples are only used to explain the present invention and do not constitute a limitation on the protection scope of the present invention.
[0020] Example 1 according to Figure 1 As shown, this embodiment provides a data collection method based on improved feature extraction and capture, including: S1. Data preprocessing: Acquire multi-source data and preprocess them to obtain multiple preprocessed sub-data and corresponding pre-feature information; specifically including: Obtain multi-source data, and perform preprocessing operations such as cleaning, denoising, and complementing to obtain multiple preprocessed sub-data; analyze and extract features from the preprocessed sub-data, and integrate them into pre-characteristic information; obtain a standardized template, and input the pre-characteristic information into the standardized template to obtain pre-characteristic information in a unified format.
[0021] S2. Feature extraction and classification: Extract key features and secondary features from the pre-characteristic information, and determine a two-way feature set according to the acquisition frequency, where the two-way feature set includes a high-frequency feature set and a low-frequency feature set.
[0022] The step of extracting key features and secondary features from the pre-characteristic information includes: Obtain all the pre-characteristic information, and perform feature vectorization and normalization processing; Obtain a standard evaluation period, and divide the standard evaluation period into multiple standard sampling periods; Construct several random samples within each standard sampling period, and perform feature screening on each random sample to obtain the matching priority of the key features, and arrange the matching priorities in descending order, and output the key features corresponding to the matching priorities with the same rank together, and then mark the secondary features as non-priority items.
[0023] The step of determining the two-way feature set according to the acquisition frequency includes: Calculate the cutting amplitude within the standard evaluation period for the number of executions of the parallel extraction period within the standard evaluation period; Summarize the frequencies of collecting pre-characteristic information within multiple standard evaluation periods, and take the average output frequency of all the key features as the value-taking condition; Obtain a judgment function, and combine the value-taking condition with the judgment function to obtain a head and tail interval value; According to the cutting amplitude of the parallel extraction period, mark the key features corresponding to the head and tail interval values as low-frequency features, and then mark the secondary features as high-frequency features.
[0024] The step of marking the key features corresponding to the head and tail interval values as low-frequency features according to the cutting amplitude of the parallel extraction period includes: Obtain all the cutting amplitudes within the standard evaluation period that belong to the parallel extraction period, and set a time length for them; Set a head and tail interval value gradient according to the lengths of multiple standard evaluation periods; According to the head interval value of the interval value gradient, obtain the position of the key feature corresponding to the head interval value in the head and tail interval values, and then mark it as a low-frequency feature.
[0025] S3, monitoring window construction: constructing a monitoring period, and setting a parallel extraction window within the monitoring period; S4, dynamic feature monitoring: at the start node of each parallel extraction window, real-time new data monitoring is performed, and at the end of the parallel extraction window, periodic feature evaluation is performed to obtain high-frequency feature items and low-frequency feature items; The step of performing periodic feature evaluation at the end of the parallel extraction period to obtain high-frequency feature items and low-frequency feature items includes: determining a dynamic monitoring period according to an update period of multi-source data; obtaining the collection time points of the preceding feature information, and marking the collection time points one by one as time-sensitive features of the preceding feature information in chronological order; calculating the discontinuous sub-intervals within the dynamic monitoring period, and counting the number of the discontinuous sub-intervals; marking key feature items whose time-sensitive feature delays exceed the discontinuous sub-intervals within the dynamic segmentation period as low-frequency feature items, screening out low-frequency features that exceed multiple dynamic monitoring periods, and then marking secondary feature items as high-frequency feature items.
[0026] S5. Feature matching and collection output: All the high-frequency feature items are input into the capture feature queue, and the high-frequency feature items in the capture feature queue are associated and matched with the pre-processed sub-data one by one according to the matching degree, to obtain a data collection output set, and upload it to the data center.
[0027] The number of key feature tags in the low-frequency features is obtained and recorded as the number of executions; then the execution ratio is calculated and output according to the total number of key feature tags in each dynamic segmentation cycle; when the number of executions exceeds three times and the execution ratio is greater than 0.5, it is marked as an abnormal collection warning sample.
[0028] Example 2 This embodiment provides a data collection method based on improved feature extraction and capture, which includes: Step S1: Acquire multi-source data and perform preprocessing to obtain a plurality of preprocessed sub-data and corresponding pre-feature information.
[0029] Data collection refers to the process of obtaining data from multiple different sources (such as database tables, files, real-time streams, etc.) and organizing these data together for further analysis, processing or storage. In this embodiment, multi-source data can be classified and summarized into multiple different data types, each of which contains multiple sub-data sources. For example, login user information includes user account, user phone number, user address, user's commonly used login platform, etc.
[0030] Clean and denoise multi-source data, remove errors, duplicates or invalid values in the data, and also supplement or correct missing data. In addition, the data format needs to be standardized and normalized so that the subsequent processing flow can effectively utilize this data.
[0031] It should be noted that the data preprocessing program usually acts on a small batch of data just accessed to the system, or is imported into the preprocessing module through an interface provided by R & D, and then taken over by the big data system. In the business system, data cleaning is generally automatically executed through a predefined rule set.
[0032] After the cleaning is completed, the relevant information is extracted as feature information and marked as pre-feature information. In the same data collection subsystem, different feature extraction modes and biases coexist. For example, although both user accounts and bank account numbers are represented as strings, the former tends to be Chinese characters, while the latter is a combination of Arabic numerals. Such features are used as basic references in the initial access stage.
[0033] After the feature extraction is completed, the pre-feature information also needs to be integrated, and the features with the same or similar meanings are merged. For example, when processing the names of logged-in users, it is necessary to determine whether users with the same name are the same user, and pre-features can be used to assist in matching, thereby improving the accuracy of data collection.
[0034] Step S2: Extract key features and secondary features from the pre-feature information, and determine a two-way feature set according to the collection frequency, where the two-way feature set includes a high-frequency feature set and a low-frequency feature set.
[0035] By extracting and analyzing the pre-feature information, key features reflecting data characteristics and secondary features composed of collaborative features can be identified. The construction of the two-way feature set depends on the collection frequency, that is, the usage frequencies of key features and secondary features are counted to divide the two types of feature sets.
[0036] The high-frequency feature set usually contains important fields that appear frequently and can reflect the characteristics of the main data, and is suitable for fast processing; while the low-frequency feature set retains some fields with low usage frequencies but with important significance in specific situations, and is suitable for refined analysis or specific tasks.
[0037] Therefore, the feature set can be flexibly selected according to the task scenario. For example: use the high-frequency feature set for quick queries and the low-frequency feature set for anomaly analysis.
[0038] Step S3: Construct a monitoring period and add a parallel extraction window within the monitoring period.
[0039] The monitoring period is dynamically constructed according to the data collection target and access form. As a sub-time period, the parallel extraction window can support multiple extraction tasks to be executed in parallel, improving the feature extraction efficiency.
[0040] At the same time, parameters such as the window running time and threshold need to be set to reasonably control the sample size, ensure load balancing of each channel, and avoid waste of operating resources.
[0041] Step S4 performs real-time monitoring of newly added data at the start node of each parallel extraction window, and conducts periodic feature evaluation at the end of the window, outputting high-frequency feature items and low-frequency feature items.
[0042] Due to the different access times of multi-source data, in order to avoid insufficient load of some extraction window tasks, a real-time data monitoring platform can be accessed to ensure that there is data input at the start node of each window. The newly added data is preferentially extracted as pre-features and added to the preprocessing queue.
[0043] After all window tasks are completed, the frequencies of all feature information are re-statistically output to update the high- and low-frequency feature items, ensuring the dynamics and timeliness of feature evaluation.
[0044] Step S5 inputs all high-frequency feature items into the captured feature queue, sorts them according to the matching degree, performs associated matching with the preprocessing sub-data, obtains the aggregated output set, and uploads it to the data middle platform.
[0045] The construction of the captured feature queue is comprehensively sorted according to factors such as user requirements and data attributes. The feature items in the queue are compared with the preprocessing data according to the matching degree to achieve an accurate mapping between key features and business fields.
[0046] During the matching process, operations such as data classification, dimensionality reduction, and duplicate removal can also be performed to further improve the data aggregation efficiency and quality.
[0047] Embodiment 3 This embodiment makes further limitations on the basis of Embodiment 2: As described in step S1 above, the step of obtaining multi-source data and performing preprocessing to obtain a plurality of preprocessing sub-data and corresponding pre-feature information includes: Step S101, obtaining multi-source data, and performing preprocessing operations such as cleaning, denoising, and complementing to obtain a plurality of preprocessing sub-data; Step S102, analyzing and extracting features from the preprocessing sub-data, and integrating them into pre-feature information; Step S103, obtaining a standardization template, and inputting the pre-feature information into the standardization template to obtain pre-feature information in a unified format.
[0048] In this embodiment, in step S1, preprocessing of multi-source data involves operations such as cleaning, denoising, and complementing. Specifically, the cleaning operation includes filtering out noise or invalid data, the denoising operation includes removing irrelevant interference factors, and the complementing operation includes processing missing information by comparison or filling.
[0049] In addition, data merging and updating can be achieved, and output standardization can be realized through text conversion and other means, which is convenient for users or data aggregation centers to extract key information.
[0050] As described in step S2 above, the step of extracting key features and secondary features from the precondition feature information includes: Step S201: Obtain all precondition feature information and perform feature vectorization and normalization processing; Step S202: Obtain the standard evaluation period and divide it into multiple standard sampling periods; Step S203: Construct a number of random samples within each standard sampling period, perform feature screening on them, obtain the matching priority of the key features, and arrange them in descending order. The key features with the same matching priority are output together, and the secondary features are marked as non-priority items.
[0051] In this embodiment, after vectorizing the precondition feature information in step S2, normalization processing is adopted to eliminate the feature dimension, and then the features are translated into a fixed interval through data mapping.
[0052] Generally, the data input situation is preferably described according to the standard evaluation period, and then it is divided into multiple sampling periods. After each evaluation, the random samples are repeatedly analyzed, the key features and secondary features are output, and their importance in the system is judged based on the occurrence frequency. If a key feature appears only once, it is the last option and is listed as a non-priority item.
[0053] As described in step S3 above, the step of determining the two-way feature set according to the collection frequency includes: Step S301: Count the execution times of the parallel extraction period within the standard evaluation period and calculate the cutting amplitude; Step S302: Summarize the precondition feature frequencies collected within multiple standard evaluation periods and calculate the average output frequency of the key features; Step S303: Obtain the judgment function and combine the value condition with the judgment function to obtain the value range of the head and tail intervals; Step S304: According to the cutting amplitude of the parallel extraction period, mark the key features corresponding to the value range of the head and tail intervals as low-frequency features, and mark the other features as high-frequency features.
[0054] In this embodiment, after each standard evaluation period ends, the execution times of the parallel extraction are statistically analyzed and the cutting section is output.
[0055] Determine by merging the output of key features within multiple time periods, counting the average frequency, and combining with a judgment function model. In practical applications, inefficient items with fewer repetitions or short active periods can be excluded; while excellent or important features can be highlighted to improve data query efficiency.
[0056] As described in step S4 above, the step of marking the key features corresponding to the value ranges of the head and tail intervals as low-frequency features according to the cutting amplitude of the parallel extraction time period includes: Step S401, obtain the parallel extraction cutting amplitude within all standard evaluation time periods and set its time length; Step S402, set the value gradient of the head and tail interval according to the lengths of multiple evaluation time periods; Step S403, obtain the position of the corresponding key feature within the interval according to the head interval value and mark it as a low-frequency feature.
[0057] As shown in this embodiment, cut the parallel extraction time period according to different standards to construct a shared channel with rich feature information, low memory usage, and reasonable distribution.
[0058] After that, generate a backend feedback according to the accumulation of the sample quantity, transfer the head of the special sample into the value range, and throughout the subsequent work plan, input it into the parallel extraction channel at the start node of each parallel extraction time period, record its pre-feature source, and extract it as monitoring data.
[0059] As in this embodiment, under the condition of importing multiple synchronization states at the node, record the actual processing dynamic content and extract the key time series, maximize the time utilization rate, and construct monitoring data for subsequent verification and analysis.
[0060] As described in step S5 above, the step of obtaining high-frequency feature items and low-frequency feature items by performing periodic feature evaluation at the end of the parallel extraction window includes: Step S501, determine the dynamic monitoring period according to the update period of multi-source data; Step S502, obtain the acquisition time points of pre-features and mark them as timeliness features in chronological order; Step S503, measure the discontinuous sub-intervals within the dynamic monitoring period and count the quantity; Step S504, mark the key features that exceed the sub-interval within this period as low-frequency features, and mark the remaining features as high-frequency features.
[0061] As in this embodiment, after each parallel extraction window ends, measure the frequency of key features. Use the discontinuous sub-intervals as the abscissa and the timeliness features as the ordinate to draw a histogram to analyze the feature distribution. Minor features can be introduced to supplement the aggregation coverage in the sparse interval of key features.
[0062] After marking the key features with aging characteristics delayed beyond the discontinuous sub-intervals as low-frequency features, the low-frequency features that repeatedly appear in multiple cycles are screened out.
[0063] As in this embodiment, after marking is completed, the low-frequency key features that repeatedly appear in multiple consecutive cycles should be screened out, and regular or irregular inspections can be carried out to clean up the low-active data. For the key features that do not meet the screening conditions, in the continuous missing state, they can be reset to high-frequency features after crossing the dynamic cycle to balance resource allocation.
[0064] After marking the low-frequency features, count the number of key feature markings and record it as the number of executions. Then, calculate the execution ratio based on the total number of key features in each dynamic cycle. When the number of executions exceeds three and the ratio is greater than 0.5, mark it as an abnormal aggregation warning sample.
[0065] In this embodiment, when a certain key feature is repeatedly marked as a low-frequency feature in multiple cycles, it can be combined with the standard index and the secondary feature response mechanism to mark it into the abnormal warning pool to prevent negative impacts on the system performance.
[0066] Embodiment 4 As Figure 2 shown, this embodiment provides a data aggregation system based on improved feature extraction and capture for implementing the data aggregation method based on improved feature extraction and capture, including: A data preprocessing module for obtaining multi-source data and performing preprocessing to obtain multiple preprocessed sub-data and corresponding prefrontal feature information; An information extraction module for extracting key features and secondary features and constructing a high-frequency feature set and a low-frequency feature set based on the collection frequency; A monitoring construction module for constructing a monitoring period and setting a parallel extraction window; A feature item capture module for performing data monitoring and cycle evaluation in the parallel extraction window to obtain high-frequency and low-frequency feature items; An association matching module for inputting high-frequency feature items into the capture queue and matching them with the preprocessed sub-data to generate an aggregation output set and upload it to the data center.
[0067] The above shows and describes the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited by the above embodiments. What is described in the above embodiments and the specification only illustrates the principles of the present invention. Without departing from the spirit and scope of the present invention, the present invention will have various changes and improvements, and these changes and improvements all fall within the scope of the present invention claimed. The scope of protection claimed by the present invention is defined by the appended claims and their equivalents.
Claims
1. An improved feature extraction and capture-based data collection method, characterized in that include: S1. Data preprocessing: Acquire multi-source data and perform preprocessing to obtain multiple preprocessed sub-data and corresponding pre-feature information; S2, feature extraction and classification: extracting key features and secondary features from the pre-feature information, and determining a bidirectional feature set according to the acquisition frequency, wherein the bidirectional feature set includes a high-frequency feature set and a low-frequency feature set; S3, monitoring window construction: constructing a monitoring period, and setting a parallel extraction window within the monitoring period; S4, dynamic feature monitoring: at the start node of each parallel extraction window, real-time new data monitoring is performed, and at the end of the parallel extraction window, periodic feature evaluation is performed to obtain high-frequency feature items and low-frequency feature items; S5. Feature matching and collection output: All the high-frequency feature items are input into the capture feature queue, and the high-frequency feature items in the capture feature queue are associated and matched with the pre-processed sub-data one by one according to the matching degree, to obtain a data collection output set, and upload it to the data center.
2. The data collection method based on improved feature extraction and capture according to claim 1, wherein Step S1 includes: Obtain multi-source data, perform cleaning, denoising, and completion preprocessing operations to obtain multiple preprocessed sub-data; Analyzing and extracting features from the preprocessed sub-data, and integrating them into pre-feature information; A standardized template is obtained, and the pre-characteristic information is input into the standardized template to obtain the pre-characteristic information in a unified format.
3. The data collection method based on improved feature extraction and capture according to claim 1, characterized in that Step S2 includes: Obtain all the pre-feature information, and perform feature vectorization and normalization processing; Obtaining a standard evaluation period, and dividing the standard evaluation period into a plurality of standard sampling periods; Several random samples are constructed in each standard sampling period, and feature screening is performed on each random sample to obtain the matching priority of key features. The matching priorities are arranged in order from large to small, and key features with the same rank are output together, and then secondary features are marked as non-priority items.
4. A data collection method based on improved feature extraction and capture according to claim 1, characterized in that In step S2, the operation of determining the bidirectional feature set according to the acquisition frequency includes: The cutting amplitude is calculated based on the number of executions of the parallel extraction period within the standard evaluation period; Summarizing the frequencies of collecting the pre-feature information in the plurality of standard evaluation periods, and taking the average output frequency of the key features as the value condition; Get the judgment function, and combine the value condition with the judgment function to obtain the head and tail interval values; According to the cutting amplitude of the parallel extraction period, the key feature set corresponding to the head and tail interval values is marked as a low-frequency feature set, and other feature sets are marked as high-frequency feature sets.
5. The data collection method based on improved feature extraction and capture according to claim 4, wherein The step of marking the key features corresponding to the head and tail interval values as a low-frequency feature set comprises: Obtaining the cutting amplitudes of all parallel extraction periods from the standard evaluation period and setting the time length for them; According to the length of multiple standard evaluation periods, the gradient of the head and tail interval values is set; According to the head interval value of the gradient, the position of the corresponding key feature in the head and tail intervals is obtained and marked as a low-frequency feature set.
6. The data collection method based on improved feature extraction and capture according to claim 5, characterized in that, After obtaining the cutting amplitudes of all parallel extraction periods within the standard evaluation period and setting the time length for them, the starting node of each parallel extraction period is input into the parallel extraction channel, the source of the preceding features input into the channel is recorded, and extracted as monitoring data.
7. A data collection method based on improved feature extraction and capture according to claim 1, characterized in that, In step S4, the step of performing periodic feature evaluation at the end of the parallel extraction period to obtain high-frequency feature items and low-frequency feature items includes: Determine the dynamic monitoring cycle based on the update cycle of multi-source data; Acquire the collection time points of the pre-feature information, and mark the collection time points one by one as time-effectiveness features of the pre-feature information in chronological order; Calculating the discontinuous subintervals within the dynamic monitoring period and counting the number of the discontinuous subintervals; The key feature items whose time-sensitivity feature delay exceeds the discontinuous subinterval within the dynamic segmentation period are marked as low-frequency feature items, and the secondary feature items are marked as high-frequency feature items.
8. A data collection method based on improved feature extraction and capture according to claim 7, characterized in that After marking the key feature items whose delay exceeds the discontinuous sub-interval as low-frequency feature items, the low-frequency feature items that appear repeatedly in multiple dynamic monitoring cycles are screened out.
9. A data collection method based on improved feature extraction and capture according to claim 8, characterized in that, After screening out the low-frequency feature items that appear repeatedly in multiple dynamic monitoring cycles, the number of key feature tags of the low-frequency feature items is obtained and recorded as the number of executions; based on the total number of key feature items in each dynamic segmentation cycle, the execution ratio is calculated and output; when the number of executions exceeds three times and the execution ratio is greater than 0.5, it is marked as an abnormal collection warning sample.
10. An improved feature extraction and capture-based data collection system for implementing the improved feature extraction and capture-based data collection method according to any one of claims 1 to 9, characterized in that, include: The data preprocessing module is used to obtain multi-source data and perform preprocessing to obtain multiple preprocessed sub-data and corresponding pre-feature information; Information extraction module, used to extract key features and secondary features, and construct high-frequency feature sets and low-frequency feature sets according to the acquisition frequency; A monitoring construction module, used to construct monitoring periods and set parallel extraction windows; The feature item capture module is used to perform data monitoring and period evaluation in the parallel extraction window to obtain high-frequency and low-frequency feature items; The association matching module is used to input high-frequency feature items into the capture queue and match them with the pre-processed sub-data, generate a collection output set and upload it to the data center.