Big data collecting and cleaning method and system based on AI

By adopting AI-based methods, grouping, vectorized encoding and dynamic compensation aggregation technology in big data acquisition and cleaning, the problem of insufficient efficiency and accuracy of large-scale data cleaning in the existing technology is solved, effectively identifying and clearing noise data, and improving the effect of data cleaning.

CN120030281AInactive Publication Date: 2025-05-23BEIJING BEIDOUWU TECHNOLOGY CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510110563.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-23
Publication Date
2025-05-23
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

When the prior art processes large-scale, high-dimensional and complex data sets, the efficiency and accuracy of data cleaning are insufficient, making it difficult to effectively identify and clear noise data.

Method used

Using AI-based big data acquisition and cleaning method, the data sample set is grouped by grouping the data sample sets by data type, a subset of data samples is extracted, and the data samples are vectorized and encoded using deep learning technology to capture the data sample features. Then, through hub search and dynamic compensation aggregation coding, the data sample prototype anchoring features are constructed to identify and clear noise samples.

Benefits of technology

Effectively identify and eliminate noise data in large-scale data sets, improve the accuracy and efficiency of data cleaning, and ensure the performance of data analysis and machine learning models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120030281A_ABST
    Figure CN120030281A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of data cleaning, and particularly discloses an AI-based big data collecting and cleaning method and system.The method comprises the steps that after to-be-processed data sample sets are grouped according to data types, a first to-be-processed data sample subset is extracted; a deep learning-based data processing technology is introduced to carry out vectorization coding on each to-be-processed data sample in the subset so as to extract data sample features, and hub search and dynamic compensation aggregation coding are carried out on each to-be-processed data sample feature so as to capture prototype anchoring features of the data sample set; and on the basis, the semantic offset degree of each to-be-processed data sample feature relative to a data set prototype anchoring feature is measured, so that noise samples are identified and cleaned. According to the invention, noise data in a large-scale data set can be effectively identified and eliminated, and the accuracy and efficiency of data cleaning are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of data cleaning technology, and more specifically, to an AI-based big data collection and cleaning method and system. Background Art

[0002] In the field of big data processing, data collection and cleaning are key steps to ensure data quality and the accuracy of analysis results. With the popularization of the Internet, the Internet of Things, and various smart devices, the amount of data has exploded, and the data types have become increasingly diverse, including but not limited to text, images, audio, video, etc. However, massive data is often mixed with a lot of noise, redundancy, and incomplete information. If this information is not effectively processed, it will seriously affect the performance of data analysis and machine learning models. Therefore, it is particularly important to develop efficient and automated data collection and cleaning methods.

[0003] Traditional big data collection and cleaning processes mostly rely on manually set rules and thresholds for screening and filtering. This method can achieve good results when dealing with small-scale, single-structured data sets, but its efficiency and accuracy are insufficient when faced with large-scale, high-dimensional and complex data sets.

[0004] In recent years, with the rapid development of artificial intelligence technology, using artificial intelligence technology to improve the data cleaning process has become a new research hotspot. Therefore, a big data collection and cleaning method and system based on AI is needed to improve the accuracy and efficiency of data cleaning. Summary of the invention

[0005] In order to solve the above technical problems, the present application is proposed. The embodiment of the present application provides an AI-based big data collection and cleaning method and system, which first groups the data sample set to be processed according to the data type, extracts the first subset of data samples to be processed, and introduces data processing technology based on deep learning to vectorize and encode each data sample to be processed in the subset to extract the data sample features, and then, by performing pivot search and dynamic compensation aggregation coding on each data sample feature to be processed, the prototype anchoring feature of the data sample set is captured, and then based on this, the semantic deviation degree of each data sample feature to be processed relative to the prototype anchoring feature of the data set is measured to achieve the identification and cleaning of noise samples. In this way, noise data in large-scale data sets can be effectively identified and eliminated, and the accuracy and efficiency of data cleaning can be improved.

[0006] According to one aspect of the present application, a big data collection and cleaning method based on AI is provided, which includes:

[0007] Obtain a sample set of data to be processed;

[0008] Based on the data type, the to-be-processed data sample set is grouped to obtain a set of to-be-processed data sample subsets;

[0009] Extracting a first subset of data samples to be processed from the set of subsets of data samples to be processed;

[0010] Performing vectorized encoding on each of the data samples to be processed in the first subset of data samples to be processed to obtain a set of embedded encoding vectors of the data samples to be processed;

[0011] Performing dynamic compensation aggregation based on feature hubs on the set of embedding coding vectors of the data sample to be processed to obtain a data sample prototype anchoring feature vector;

[0012] Extracting a first embedded coding vector of the data sample to be processed from the set of embedded coding vectors of the data sample to be processed as a query feature vector;

[0013] Based on the semantic offset between the query feature vector and the data sample prototype anchor feature vector, it is determined whether to delete the first data sample to be processed corresponding to the first data sample to be processed embedding encoding vector from the first data sample subset to be processed.

[0014] Preferably, the set of embedded coding vectors of the data sample to be processed is subjected to dynamic compensation aggregation based on feature hubs to obtain a data sample prototype anchor feature vector, including:

[0015] Inputting the set of embedding encoding vectors of the data samples to be processed into the hub feature extraction network to obtain the hub feature encoding vectors of the data samples to be processed;

[0016] Extracting complementary information of each of the embedded coding vectors of the data sample to be processed relative to the hub feature coding vector of the data sample to be processed in the set of embedded coding vectors of the data sample to be processed to obtain a set of embedded coding vectors of the data sample to be processed-hub feature complementary information;

[0017] Based on the hub feature coding vector of the data sample to be processed, the set of the data sample to be processed-hub feature complementary information embedding coding vector is significantly dynamically modulated to obtain a set of significantly modulated data sample to be processed-hub feature complementary information embedding coding vectors;

[0018] The data sample prototype anchoring feature vector is obtained by fusing the hub feature encoding vector of the data sample to be processed and the set of the significantly modulated data sample to be processed-hub feature complementary information embedding encoding vectors.

[0019] Preferably, extracting complementary information of each of the embedded coding vectors of the data sample to be processed in the set of embedded coding vectors of the data sample to be processed relative to the hub feature coding vector of the data sample to be processed to obtain a set of embedded coding vectors of the data sample to be processed-hub feature complementary information, comprises:

[0020] Performing point convolution coding based on Sigmoid activation function on the embedded coding vector of the data sample to be processed and the pivot feature coding vector of the data sample to be processed respectively to obtain a standardized embedded coding vector of the data sample to be processed and a standardized pivot feature coding vector of the data sample to be processed;

[0021] Taking the positional difference vector between the standardized embedded coding vector of the data sample to be processed and the standardized hub feature coding vector of the data sample to be processed as the weight vector, the embedded coding vector of the data sample to be processed and the hub feature coding vector of the data sample to be processed are subjected to complementary feature enhancement modulation and aggregation coding to obtain the embedded coding vector of the data sample to be processed-hub feature complementary information.

[0022] Preferably, based on the hub feature coding vector of the data sample to be processed, the set of the data sample to be processed-hub feature complementary information embedding coding vector is significantly dynamically modulated to obtain a set of significantly modulated data sample to be processed-hub feature complementary information embedding coding vectors, including:

[0023] Inputting each of the to-be-processed data sample-hub feature complementary information embedding coding vectors in the set of the to-be-processed data sample-hub feature complementary information embedding coding vectors into a complementary information significant identification module based on an attention mechanism to obtain a set of to-be-processed data sample complementary information attention weights;

[0024] Based on the set of attention weights of the complementary information of the data samples to be processed, the set of the complementary information embedded coding vectors of the data samples to be processed and the hub features is subjected to attention modulation to obtain the set of the significantly modulated complementary information embedded coding vectors of the data samples to be processed and the hub features.

[0025] Preferably, fusing the hub feature encoding vector of the data sample to be processed and the set of the significantly modulated data sample to be processed-hub feature complementary information embedding encoding vectors to obtain the data sample prototype anchoring feature vector comprises:

[0026] The hub feature encoding vector of the data sample to be processed and the set of the significantly modulated data sample to be processed-hub feature complementary information embedding encoding vectors are cascaded and fused to obtain the data sample prototype anchoring feature vector.

[0027] Preferably, determining whether to delete the first data sample to be processed corresponding to the first data sample to be processed embedding encoding vector from the first data sample to be processed subset based on the semantic offset between the query feature vector and the data sample prototype anchor feature vector comprises:

[0028] Extracting the semantic offset feature between the query feature vector and the data sample prototype anchor feature vector to obtain a prototype offset semantic encoding vector of the sample point to be processed;

[0029] Inputting the prototype offset semantic coding vector of the sample point to be processed into a sample identifier based on a classifier to obtain a recognition result, wherein the recognition result is used to indicate whether the first data sample to be processed is a noise sample;

[0030] In response to the identification result that the first data sample to be processed is a noise sample, the first data sample to be processed is deleted from the first subset of data samples to be processed.

[0031] Preferably, extracting the semantic offset feature between the query feature vector and the data sample prototype anchor feature vector to obtain the prototype offset semantic encoding vector of the sample point to be processed includes:

[0032] The position difference between the query feature vector and the data sample prototype anchor feature vector is calculated to obtain the prototype offset semantic encoding vector of the sample point to be processed.

[0033] Preferably, the prototype offset semantic coding vector of the sample point to be processed is input into a sample identifier based on a classifier to obtain a recognition result, and the recognition result is used to indicate whether the first data sample to be processed is a noise sample, including:

[0034] Using the fully connected layer of the sample identifier to perform fully connected encoding on the prototype offset semantic encoding vector of the sample point to be processed to obtain the prototype offset semantic fully connected encoding vector of the sample point to be processed;

[0035] Inputting the prototype offset semantic fully connected encoding vector of the sample point to be processed into the Softmax classification function of the sample identifier to obtain the probability value of the prototype offset semantic encoding vector of the sample point to be processed belonging to each classification label, wherein the classification label includes whether the first data sample to be processed is a noise sample and whether the first data sample to be processed is not a noise sample;

[0036] The classification label corresponding to the largest probability value among the probability values ​​is determined as the recognition result.

[0037] According to another aspect of the present application, there is provided an AI-based big data collection and cleaning system, which includes:

[0038] A data sample set acquisition module is used to acquire a data sample set to be processed;

[0039] A data sample set grouping module, used for grouping the data sample set to be processed based on data type to obtain a set of data sample subsets to be processed;

[0040] A sample subset extraction module, used to extract a first to-be-processed data sample subset from the set of to-be-processed data sample subsets;

[0041] A vectorized encoding module, used for performing vectorized encoding on each of the data samples to be processed in the first subset of data samples to be processed to obtain a set of embedded encoding vectors of the data samples to be processed;

[0042] A dynamic compensation aggregation module, used for performing dynamic compensation aggregation based on feature hubs on the set of embedded coding vectors of the data sample to be processed to obtain a data sample prototype anchor feature vector;

[0043] A query feature extraction module, used to extract a first embedded coding vector of the data sample to be processed from the set of embedded coding vectors of the data sample to be processed as a query feature vector;

[0044] A sample screening module is used to determine whether to delete the first data sample to be processed corresponding to the first data sample to be processed embedding coding vector from the first data sample to be processed subset based on the semantic offset between the query feature vector and the data sample prototype anchor feature vector.

[0045] This application has at least the following technical effects:

[0046] Compared with the prior art, the AI-based big data collection and cleaning method and system provided by the present application first groups the data sample set to be processed according to the data type, extracts the first subset of data samples to be processed, and introduces data processing technology based on deep learning to vectorize and encode each data sample to be processed in the subset to extract the data sample features. Then, by performing pivot search and dynamic compensation aggregation encoding on each data sample feature to be processed, the prototype anchoring feature of the data sample set is captured, and then based on this, the semantic deviation degree of each data sample feature to be processed relative to the prototype anchoring feature of the data set is measured to achieve the identification and cleaning of noise samples. The present application can effectively identify and eliminate noise data in large-scale data sets, and improve the accuracy and efficiency of data cleaning. BRIEF DESCRIPTION OF THE DRAWINGS

[0047] By describing the embodiments of the present application in more detail in conjunction with the accompanying drawings, the above and other purposes, features and advantages of the present application will become more apparent. The accompanying drawings are used to provide a further understanding of the embodiments of the present application and constitute a part of the specification. Together with the embodiments of the present application, they are used to explain the present application and do not constitute a limitation of the present application. In the accompanying drawings, the same reference numerals generally represent the same components or steps.

[0048] Figure 1 This is a flow chart of an AI-based big data collection and cleaning method according to an embodiment of the present application.

[0049] Figure 2 This is a data flow diagram of the AI-based big data collection and cleaning method according to an embodiment of the present application.

[0050] Figure 3 This is a flowchart of sub-step S5 of the AI-based big data collection and cleaning method according to an embodiment of the present application.

[0051] Figure 4 This is a flowchart of sub-step S53 of the AI-based big data collection and cleaning method according to an embodiment of the present application.

[0052] Figure 5 This is a flowchart of sub-step S7 of the AI-based big data collection and cleaning method according to an embodiment of the present application.

[0053] Figure 6 It is a block diagram of an AI-based big data acquisition and cleaning system according to an embodiment of the present application. DETAILED DESCRIPTION

[0054] As shown in this application and claims, unless the context clearly indicates an exception, the words "a", "an", "an" and / or "the" do not refer to the singular and may also include the plural. Generally speaking, the terms "include" and "comprise" only indicate the inclusion of the steps and elements that have been clearly identified, and these steps and elements do not constitute an exclusive list. The method or device may also include other steps or elements.

[0055] Although the present application makes various references to certain modules in the system according to the embodiments of the present application, any number of different modules can be used and run on the user terminal and / or server. The modules are only illustrative, and different aspects of the system and method can use different modules.

[0056] Flowcharts are used in the present application to illustrate the operations performed by the system according to the embodiments of the present application. It should be understood that the preceding or following operations are not necessarily performed accurately in order. On the contrary, various steps may be processed in reverse order or simultaneously as required. Meanwhile, other operations may also be added to these processes, or a certain step or several steps of operations may be removed from these processes.

[0057] Below, the exemplary embodiments according to the present application will be described in detail with reference to the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments of the present application, and it should be understood that the present application is not limited to the exemplary embodiments described here.

[0058] It should be noted that all data acquisition actions in this application are carried out in compliance with the relevant data protection laws and policies of the country where the application is located, and with the authorization given by the owner of the corresponding device.

[0059] Figure 1 This is a flow chart of an AI-based big data collection and cleaning method according to an embodiment of the present application. Figure 2 Schematic diagram of data flow of the AI-based big data collection and cleaning method according to an embodiment of the present application. Figure 1 and Figure 2 As shown, the AI-based big data acquisition and cleaning method includes the following steps: S1, obtaining a set of data samples to be processed; S2, grouping the data sample set to be processed based on the data type to obtain a set of data sample subsets to be processed; S3, extracting a first data sample subset to be processed from the set of data sample subsets to be processed; S4, vectorizing and encoding each data sample to be processed in the first data sample subset to be processed to obtain a set of embedded coding vectors of the data samples to be processed; S5, performing dynamic compensation aggregation based on feature hubs on the set of embedded coding vectors of the data samples to be processed to obtain a data sample prototype anchoring feature vector; S6, extracting a first data sample embedded coding vector to be processed from the set of embedded coding vectors of the data samples to be processed as a query feature vector; S7, determining whether to delete the first data sample to be processed corresponding to the first data sample embedded coding vector from the first data sample subset to be processed based on the semantic offset between the query feature vector and the data sample prototype anchoring feature vector.

[0060] In the above AI-based big data collection and cleaning method, the step S1 is to obtain a sample set of data to be processed. It should be understood that by collecting representative data samples, the characteristics and laws of the research object can be reflected, providing a high-quality data foundation for subsequent work, improving the accuracy and reliability of analysis, helping to fully understand the research object, and assisting decision-making.

[0061] In order to efficiently obtain large-scale data samples from the Internet environment, web crawler technology has been widely adopted. A web crawler is an automated program or script that systematically browses the World Wide Web, extracts required information by accessing web pages and parsing the content, and converts this information into a structured format for storage. When faced with publicly accessible data resources such as news websites, social media platforms, and e-commerce websites, web crawlers provide an efficient and economical way to collect data. Web crawlers automatically perform tasks at predetermined time intervals without human intervention; they support multi-threaded or multi-process concurrent operation to speed up data collection; and they allow different parsers to be configured to adapt to various types of web page structures. As Internet regulations continue to improve, web crawler activities need to comply with relevant laws and regulations, respect the robots.txt protocol of the website, and avoid infringing on personal privacy or violating copyright regulations. In other words, legal compliance issues must be considered when designing web crawlers, and appropriate measures must be taken to ensure that the data collection process is legal and reasonable.

[0062] Another important way to collect data is to obtain data directly from the data source through an application programming interface (API). Many online service providers provide official APIs that allow developers to query, retrieve, and download data programmatically. This method usually has higher efficiency and more stable performance because it is designed based on service-side support and can avoid legal risks and technical challenges caused by web crawler activities. With APIs, organizations can directly access specific data streams to ensure that they obtain the latest and most accurate data, while also facilitating the implementation of data management policies such as access control and logging. API calls provide a standardized way to interact with data, making data exchange between different systems easier and faster. However, not all data sources open API interfaces to the outside world, which limits the access to certain types of data. Even if there are API interfaces, they may be affected by factors such as rate limits and request frequency caps, resulting in the inability to meet the needs of large-scale data collection. Finally, the quality of API documentation varies. If the documentation is not clearly described or updated in a timely manner, it may bring additional work burden to developers.

[0063] For non-digital historical data or data that cannot be obtained through the public network, it can be collected through other channels. This involves mining of internal databases, digital conversion of physical documents, cooperative purchase of third-party data suppliers, etc. In view of these situations, a strict data governance framework must be established to ensure that the data collection process complies with relevant laws and regulations, respects the principle of personal privacy protection, and takes appropriate security measures to prevent data leakage. Internal database mining refers to the data resources stored in various information systems that already exist within the enterprise. Due to historical reasons, the data formats of many old systems are not uniform, and there are even a lot of redundant and erroneous information. Therefore, it is necessary to invest a certain amount of manpower and material resources for cleaning and standardization. For paper or other forms of non-electronic documents, digital conversion is an indispensable process. This process involves the application of technologies such as optical character recognition (OCR) and image scanning to convert physical files into electronic documents for subsequent data processing.

[0064] In addition to relying on their own strength to collect data, companies can also seek help from external partners to obtain the specific types of data they need through purchase or sharing. This method not only saves time and cost, but also quickly supplements the deficiencies in their own data system. For example, in market research, consumer behavior analysis, etc., professional data service providers often have accumulated a lot of industry knowledge and user portraits, and can provide companies with accurate insights and support. When choosing third-party data cooperation, it is necessary to pay attention to evaluating the credibility and service level of the partners to ensure the authenticity and reliability of the data provided; clarify the rights and obligations of both parties, especially in terms of data ownership, usage rights and confidentiality terms to reach a consensus; pay attention to the legality of the data source, and avoid illegal collection or transactions.

[0065] With the development of social networks and mobile Internet, more and more users are participating in content creation, forming a huge pool of user-generated content (UGC). This type of data is characterized by a wide variety, rapid updates, and wide coverage, covering a variety of forms such as text, pictures, and videos. UGC data can be collected by using the official API interfaces provided by social media platforms, such as Weibo and WeChat public accounts, to obtain publicly released posts, comments, and other content in accordance with the rules; or by developing special monitoring tools to continuously track specific keywords or topics and capture the latest hot topics of discussion. Regardless of the method, it is necessary to comply with the platform's regulations and must not infringe on the user's privacy and other legal rights.

[0066] In the above-mentioned AI-based big data collection and cleaning method, the step S2, based on the data type, groups the data sample set to be processed to obtain a set of data sample subsets to be processed. It should be understood that different types of data have different characteristics and processing methods. This application can perform adaptive processing on different types of data subsets by grouping the data sample set to be processed according to the data type, which helps to better adapt to the processing requirements of different types of data and improve the pertinence and efficiency of data cleaning.

[0067] Specifically, grouping data sample sets based on data types means identifying and separating different types of data elements. When faced with a mixed data set containing multiple data types, it is first necessary to define a set of standards or rules to guide the grouping operation. This usually involves building a classification framework that can automatically identify the specific type to which the data point belongs and assign it to the corresponding subset accordingly. To achieve this, metadata can be used as a basis for classification. Metadata can provide information about file format, encoding method, creation date, source platform, etc., which are key factors in determining data types. In addition, the classification logic can be further refined by combining content features such as text length, image resolution, audio sampling rate, etc. In this way, even in the absence of clear labels, preliminary data grouping can be effectively completed.

[0068] In some cases, a single data sample may contain multiple types of information at the same time, such as a picture with a descriptive title or a video clip with comments. For this type of composite data, a more sophisticated approach is needed to disassemble and classify it. A common practice is to decompose multimodal data into independent components and then group them according to their respective data types. This not only helps to simplify the processing process, but also ensures that each part receives appropriate attention. However, in some application scenarios, it may be necessary to maintain the integrity and relevance of the data. At this time, you should consider designing a grouping scheme that can reflect the characteristics of each component and the overall structure. For example, you can introduce a hierarchy or network diagram to establish the relationship between different levels or nodes, so that even after grouping, you can still trace the intrinsic connection between the original data.

[0069] In addition to classification by type, factors such as the time dimension and geographical distribution of the data also need to be considered. Time series data are usually sorted by timestamp to capture the order and trend changes of events; data related to geographic location can be clustered according to coordinate information to reveal spatial aggregation patterns. Such a grouping method can help reveal the dynamic characteristics and regional differences hidden in the data, and facilitate spatiotemporal analysis. It is worth noting that with the popularity of Internet of Things (IoT) devices, more and more data sources have begun to generate information flows with time and location tags, which poses new challenges to traditional static grouping methods. To this end, it is also possible to explore the possibility of dynamically adjusting the grouping strategy, that is, automatically optimizing the grouping results based on real-time updated data features to adapt to the ever-changing data environment.

[0070] In order to ensure the accuracy and stability of the grouping process, a quality control mechanism must be introduced. On the one hand, it is necessary to ensure the consistency of classification rules to avoid misclassification due to rule conflicts; on the other hand, it is necessary to regularly evaluate the grouping effect and correct possible problems in a timely manner. For example, some key indicators such as purity and coverage can be set to measure the consistency within each subset and the completeness of the coverage. In addition, visualization tools can be used to display the distribution of grouped data, intuitively identify outliers or outliers, and then take appropriate measures to deal with them.

[0071] In the above-mentioned AI-based big data collection and cleaning method, the step S3 extracts a first subset of data samples to be processed from the set of the data sample subsets to be processed. Specifically, after the data sample set to be processed is divided into multiple subsets of data samples to be processed, in order to perform more refined data analysis and noise cleaning for different sample subsets, the present application first extracts a first subset of data samples to be processed from the set of data sample subsets to be processed. It should be understood that in order to simplify the description, the present application only takes the first subset of data samples to be processed as an example for illustration, but the data cleaning method described in the present application is also applicable to other subsets of data samples to be processed.

[0072] In the above-mentioned AI-based big data acquisition and cleaning method, the step S4 vectorizes and encodes each of the data samples to be processed in the first subset of data samples to be processed to obtain a set of embedded coding vectors of the data samples to be processed. Specifically, the present application takes into account that the AI-based data processing algorithm needs to convert the data samples into a computer-recognizable format, that is, a numerical vector form. Therefore, the present application further vectorizes and encodes each of the data samples to be processed in the first subset of data samples to be processed to obtain a set of embedded coding vectors of the data samples to be processed. It should be understood that vectorized coding is a process of converting the original data samples into a numerical vector form, which can map the data samples from the original feature space to a high-dimensional vector space, so that the similarities and differences between the data samples can be maintained and reflected in the high-dimensional vector space, thereby facilitating subsequent data processing and cleaning. In an embodiment of the present application, different vectorized coding methods are required for different types of data sample subsets. For example, for a subset of data samples of text type, word embedding technology, such as Word2Vec, BERT, etc., can be used to convert text data into numerical vectors; for a subset of data samples of image type, deep learning models such as convolutional neural network (CNN) can be used to convert image data into numerical vectors; for a subset of data samples of audio type, audio feature extraction methods such as Mel frequency cepstral coefficient (MFCC) can be used to convert audio data into numerical vectors. In this way, by using adaptive vector encoding methods for different types of data sample subsets, it can be ensured that the data samples can retain their original feature information as much as possible after being converted into numerical vectors, providing strong support for subsequent data processing and cleaning.

[0073] In the above-mentioned AI-based big data acquisition and cleaning method, the step S5 performs dynamic compensation aggregation based on feature hubs on the set of embedded coding vectors of the data samples to be processed to obtain a data sample prototype anchor feature vector. It should be understood that in order to accurately determine whether each data sample in the first subset of data samples to be processed is a noise sample, the present application further constructs a data sample prototype anchor feature vector that can reflect the essential characteristics of the first subset of data samples to be processed by performing hub search and dynamic compensation aggregation coding on the set of embedded coding vectors of the data samples to be processed. In this way, in the subsequent data cleaning process, the data sample prototype anchor feature vector is used as a benchmark, and by comparing and analyzing the characteristics of each data sample with it, it can be determined whether the sample deviates from the essential characteristics of the data set, thereby identifying the noise sample. Among them, Figure 3 FIG. 5 is a flowchart of sub-step S5 of the AI-based big data collection and cleaning method according to an embodiment of the present application. Figure 3As shown, the step S5 includes the steps of: S51, inputting the set of embedded coding vectors of the data samples to be processed into the hub feature extraction network to obtain the hub feature coding vectors of the data samples to be processed; S52, extracting the complementary information of each embedded coding vector of the data samples to be processed in the set of embedded coding vectors of the data samples to be processed relative to the hub feature coding vector of the data samples to be processed to obtain a set of embedded coding vectors of the data samples to be processed-hub feature complementary information; S53, based on the hub feature coding vector of the data samples to be processed, performing significant dynamic modulation on the set of embedded coding vectors of the data samples to be processed-hub feature complementary information to obtain a set of significantly modulated embedded coding vectors of the data samples to be processed-hub feature complementary information; S54, fusing the hub feature coding vector of the data samples to be processed and the set of significantly modulated embedded coding vectors of the data samples to be processed-hub feature complementary information to obtain the prototype anchoring feature vector of the data sample.

[0074] Specifically, the step S51 is expressed by the formula:

[0075]

[0076] Among them, f hub (·) represents the hub feature extraction network, X represents the set of embedding encoding vectors of the data samples to be processed, and x 1 、x 2 、x i and x n represent the first, second, i-th and n-th embedded coding vectors of the data samples to be processed in the set of embedded coding vectors of the data samples to be processed, n is the number of embedded coding vectors of the data samples to be processed, W 1i and b 1i Respectively represent the weight parameter matrix and bias term of the hub feature extraction network, represents the hub feature relevance score transformation vector of the hub feature extraction network, e i Indicates that x i The corresponding hub feature relevance scoring factor, represents matrix multiplication, softmax represents the normalized exponential function, a i Indicates that x i The corresponding normalized hub feature relevance score factor, v h Represents the hub feature encoding vector of the data sample to be processed.

[0077] Specifically, this application first constructs a hub feature extraction network based on a neural network architecture, and with the help of the powerful representation ability of the neural network model, performs hub feature correlation analysis on the embedded coding vectors of each data sample to be processed. In this way, the network can deeply explore the potential correlations between the features of each data sample to be processed, and then identify and extract representative data sample feature representations in the set of embedded coding vectors of the data samples to be processed, that is, the hub feature coding vectors of the data samples to be processed.

[0078] Specifically, in a specific example of the present application, the step S52 includes: first, performing point convolution encoding based on the Sigmoid activation function on the embedded coding vector of the data sample to be processed and the hub feature coding vector of the data sample to be processed to obtain a standardized embedded coding vector of the data sample to be processed and a standardized hub feature coding vector of the data sample to be processed, which is expressed as follows:

[0079] x i ′=Sigmoid(W 11 (conv 1×1 (x i )))

[0080] v h ′=Sigmoid(W 21 (conv 1×1 (v h )))

[0081] Among them, Sigmoid represents the Sigmoid activation function, x i ' represents the standardized embedding encoding vector of the data sample to be processed, v h ' represents the hub feature encoding vector of the standardized data sample to be processed, W 11 and W 21 Represents different weight matrices, conv 1×1 (·) represents point convolutional coding;

[0082] Then, the position difference vector between the standardized data sample embedding coding vector and the standardized data sample hub feature coding vector is used as the weight vector, and the data sample embedding coding vector and the data sample hub feature coding vector are subjected to complementary feature enhancement modulation and aggregation coding to obtain the data sample-hub feature complementary information embedding coding vector, which is expressed as follows:

[0083]

[0084] xb i =Sigmoid(conv 1×1 (xi ·p i +v h ·p i ))

[0085] Among them, p i represents the differential encoding vector of the data sample to be processed-the hub feature, represents the difference operation, |·| is the absolute value, xb i Indicates that x i The corresponding data sample to be processed-hub feature complementary information is embedded into the encoding vector.

[0086] That is, in order to retain as much data sample detail information as possible during the hub feature extraction process, the present application further integrates and associates the unique features of each data sample to be processed around the hub feature coding vector of the data sample to be processed, and introduces a dynamic compensation mechanism to compensate and repair the lost data sample detail information, thereby achieving comprehensive coverage and accurate description of the data sample features. In this process, by calculating the complementary information of the embedded coding vector of each data sample to be processed relative to the hub feature coding vector of the data sample to be processed, the feature differences and unique contributions of each data sample feature to be processed relative to the global core information of the first subset of data samples to be processed are captured, and a set of embedded coding vectors of complementary information of the data sample to be processed-hub features is generated, so as to restore the feature components of each data sample embedded coding vector to be processed that are not considered and absorbed during the hub feature extraction process. In this way, the overall characteristics of the data sample can be more comprehensively reflected, avoiding information loss caused by over-simplification. The application of the dynamic compensation mechanism enables the model to take into account the individual characteristics of the data sample while maintaining efficient feature extraction, thereby improving the accuracy and reliability of the overall analysis results. The final generated set of complementary information embedding coding vectors serves as a supplement, which helps to build a richer and more complete data representation and support subsequent more sophisticated data processing and analysis tasks.

[0087] Figure 4 FIG. 5 is a flowchart of sub-step S53 of the AI-based big data collection and cleaning method according to an embodiment of the present application. Figure 4As shown, the step S53 includes the steps of: S531, respectively inputting each of the data sample to be processed-hub feature complementary information embedded coding vectors in the set of the data sample to be processed-hub feature complementary information embedded coding vectors into the complementary information significant identification module based on the attention mechanism to obtain a set of attention weights of the complementary information of the data sample to be processed; S532, based on the set of attention weights of the complementary information of the data sample to be processed, performing attention modulation on the set of the data sample to be processed-hub feature complementary information embedded coding vectors to obtain the set of the significantly modulated data sample to be processed-hub feature complementary information embedded coding vectors.

[0088] More specifically, the step S531 is expressed by the formula:

[0089]

[0090] Among them, g i Indicates xb i The corresponding significant identification factor, W 2i and b 2i They represent the weight parameter matrix and bias term of the complementary information saliency identification module, respectively. represents the complementary information saliency score conversion vector of the complementary information saliency identification module, exp(·) represents the exponential function operation with e as the base, and w i Indicates xb i The corresponding complementary information attention weight of the data sample to be processed.

[0091] That is, since the complementary information of each to-be-processed data sample feature relative to the global core information of the data set may constitute a beneficial supplement, it may also cause redundancy or interference. Therefore, the present application further utilizes the attention mechanism to evaluate the importance of each to-be-processed data sample-hub feature complementary information embedded coding vector, so as to dynamically adjust the weight distribution of each to-be-processed data sample-hub feature complementary information embedded coding vector in the information aggregation process, so that the model can focus on the most valuable information while filtering out irrelevant or potential interference factors. Through this method, not only the accuracy and robustness of the data sample representation are enhanced, but also the effects of subsequent data analysis and modeling tasks are improved.

[0092] More specifically, the step S532 is expressed by the formula:

[0093] Y={xb 1 ·w 1 ,xb 2 ·w 1 ,...,xb i ·w i ,...,xb n ·wn}

[0094] Among them, xb 1 、xb 2 、xb i and xb n Respectively represent x 1 、x 2 、x i and x n The corresponding data sample to be processed-hub feature complementary information embedding encoding vector, w 1 、w 2 、w i and w n Respectively represent xb 1 、xb 2 、xb i and xb n The corresponding complementary information attention weight of the data sample to be processed, Y represents the set of significantly modulated complementary information embedding encoding vectors of the data sample to be processed-hub feature.

[0095] That is, based on the calculated set of attention weights of the complementary information of the data samples to be processed, the set of the complementary information embedded coding vectors of the data samples to be processed and the hub features is modulated for attention, and the set of the complementary information embedded coding vectors of the data samples to be processed and the hub features is enhanced or suppressed, so that the modulated set of the complementary information embedded coding vectors of the data samples to be processed and the hub features can better reflect the key role and contribution of each data sample to be processed in the description of the anchoring features of the data sample prototype. Specifically, the attention weights guide which features should be strengthened to highlight their importance, and indicate which features should be suppressed to reduce the impact of redundant or interfering information. After such modulation, the final set of significantly modulated complementary information embedded coding vectors of the data samples to be processed and the hub features not only retains the core characteristics of the original data, but also optimizes the feature expression, so that the key characteristics and unique contributions of each sample can be better presented.

[0096] Specifically, in a specific example of the present application, the step S54 includes: cascading the hub feature encoding vector of the data sample to be processed and the set of the significantly modulated data sample to be processed-hub feature complementary information embedding encoding vector to obtain the data sample prototype anchor feature vector, which is expressed as:

[0097] v f =Concat{v h ; Y}

[0098] Among them, Concat{·;·} represents the cascade function, v f Represents the prototype anchor feature vector of the data sample.

[0099] That is, through the cascade operation, the set of significantly modulated data sample to be processed-hub feature complementary information embedded coding vectors is fused with the hub feature coding vector of the data sample to be processed, while comprehensively considering the global core information of the data sample set and taking into account the local detail information and unique contribution of each data sample to be processed, thereby improving the accuracy and comprehensiveness of feature expression. In this way, the precise construction of the prototype anchoring feature vector of the data sample can be achieved, ensuring that it can be used as a reliable benchmark in the subsequent data cleaning process, effectively distinguishing noise samples from valid samples, and thus improving the accuracy and efficiency of data cleaning.

[0100] In the above-mentioned AI-based big data collection and cleaning method, the step S6 extracts the first data sample embedded coding vector from the set of the data sample embedded coding vector as the query feature vector. Specifically, in order to achieve accurate cleaning of each data sample to be processed in the first data sample subset to be processed, the present application further extracts the first data sample embedded coding vector from the set of the data sample embedded coding vector as the query feature vector, and compares and analyzes it with the data sample prototype anchor feature vector to determine whether the corresponding first data sample to be processed is a noise sample.

[0101] In the above AI-based big data collection and cleaning method, the step S7 determines whether to delete the first data sample to be processed corresponding to the first data sample to be processed embedded coding vector from the first data sample subset based on the semantic offset between the query feature vector and the data sample prototype anchor feature vector. Figure 5 FIG. 1 is a flowchart of sub-step S7 of the AI-based big data collection and cleaning method according to an embodiment of the present application. Figure 5 As shown, the step S7 includes the steps of: S71, extracting the semantic offset feature between the query feature vector and the data sample prototype anchor feature vector to obtain the prototype offset semantic coding vector of the sample point to be processed; S72, inputting the prototype offset semantic coding vector of the sample point to be processed into a classifier-based sample identifier to obtain a recognition result, and the recognition result is used to indicate whether the first data sample to be processed is a noise sample; S73, in response to the recognition result that the first data sample to be processed is a noise sample, deleting the first data sample to be processed from the first data sample subset to be processed.

[0102] Specifically, in a specific example of the present application, the step S71 includes: calculating the position difference between the query feature vector and the data sample prototype anchor feature vector to obtain the prototype offset semantic coding vector of the sample point to be processed. It should be understood that the position difference refers to subtracting the eigenvalues ​​of the query feature vector and the data sample prototype anchor feature vector in each dimension one by one, so as to obtain a vector representation reflecting the degree of difference between the two in each feature dimension. Through this position difference calculation method, the offset of the query feature vector relative to the data sample prototype anchor feature vector can be accurately captured, thereby providing a strong basis for determining whether the data sample corresponding to the query feature vector is a noise sample.

[0103] Specifically, the step S72 includes: using the fully connected layer of the sample identifier to fully connect the semantic coding vector of the prototype offset of the sample point to be processed to obtain the fully connected semantic coding vector of the prototype offset of the sample point to be processed; inputting the fully connected semantic coding vector of the prototype offset of the sample point to be processed into the Softmax classification function of the sample identifier to obtain the probability value of the prototype offset semantic coding vector of the sample point to be processed belonging to each classification label, wherein the classification label includes the first data sample to be processed is a noise sample and the first data sample to be processed is not a noise sample; and determining the classification label corresponding to the largest of the probability values ​​as the recognition result. Specifically, the classifier-based sample identifier adopts a neural network architecture, and performs deep learning and feature analysis on the semantic coding vector of the prototype offset of the sample point to be processed, so as to automatically determine whether the first data sample to be processed corresponding to the query feature vector deviates from the essential characteristics of the data set, that is, whether it is a noise sample, according to the feature difference and offset degree information contained in the semantic coding vector of the prototype offset of the sample point to be processed. If the judgment result is yes, the first data sample to be processed is marked as a noise sample, and corresponding cleaning or correction operations are performed in the subsequent data processing process; if the judgment result is no, the first data sample to be processed is retained for subsequent data analysis and utilization. In this way, accurate identification and effective cleaning of noise samples in the data set can be achieved, thereby improving data quality and the accuracy of analysis results.

[0104] In particular, considering that the query feature vector and the data sample prototype anchor feature vector respectively represent the embedded coding features of the data sample to be processed corresponding to the first data sample to be processed and the full sample domain aggregated coding features based on the sequence hub of the first data sample subset to be processed, when calculating the positional difference between the query feature vector and the data sample prototype anchor feature vector to obtain the prototype offset semantic coding vector of the sample point to be processed, the difference in the number of samples will lead to insufficient representation of the long-distance interactive response of the fine-grained offset semantic field formed by the positional difference, thereby reducing the expression effect of the prototype offset semantic coding vector of the sample point to be processed and affecting the accuracy of the recognition result obtained by inputting it into the classifier-based sample identifier.

[0105] Therefore, in a preferred example of the present application, when the prototype offset semantic coding vector of the sample point to be processed is input into a sample identifier based on a classifier, the prototype offset semantic coding vector of the sample point to be processed is optimized, and the optimization includes the following steps:

[0106] First, the eigenvalues ​​of the prototype offset semantic coding vector of the sample point to be processed are arranged in ascending order to form a prototype offset semantic sequential coding vector of the sample point to be processed;

[0107] Secondly, in response to the absolute value of the difference between the i-th eigenvalue and the i+1-th eigenvalue of the prototype offset semantic order encoding vector of the sample point to be processed being less than or equal to the distance difference hyperparameter ε, the weighted sum between the i-th eigenvalue and the i+1-th eigenvalue is calculated as the optimized i+1-th eigenvalue, which is expressed as:

[0108] v′ i+1 =ω 1 ×v i +ω 2 ×v i+1

[0109] Among them, v i 、v i+1 They represent the i-th eigenvalue and the i+1-th eigenvalue of the semantic sequential encoding vector of the prototype offset of the sample point to be processed, respectively, 1 ,ω 2 Respectively represent the first weight parameter and the second weight parameter, v' i+1 Represents the i+1th eigenvalue of the optimized prototype offset semantic order encoding vector of the sample point to be processed;

[0110] Next, in response to the absolute value of the difference between the i-th eigenvalue and the i+1-th eigenvalue of the prototype offset semantic sequential coding vector of the sample point to be processed being greater than the distance difference hyperparameter ε, the square root of the sum of the squares of all eigenvalues ​​of the prototype offset semantic coding vector of the sample point to be processed is calculated, which is expressed as follows:

[0111]

[0112] Wherein, root represents the square root of the sum of squares of all eigenvalues ​​of the prototype offset semantic encoding vector of the sample point to be processed, v 1 、v 2 Representing each eigenvalue of the semantic order encoding vector of the prototype offset of the sample point to be processed;

[0113] The square root is multiplied by 2 and then divided by the square of the length of the semantic coding vector of the prototype offset of the sample point to be processed to obtain the primitive value of the semantic space of the prototype offset of the sample point to be processed, which is expressed by the formula:

[0114] base=2×root / L 2

[0115] Wherein, L represents the length of the semantic coding vector of the prototype offset of the sample point to be processed, and base represents the primitive value of the semantic space of the prototype offset of the sample point to be processed;

[0116] After multiplying the primitive value of the prototype offset semantic space of the sample point to be processed by the i-th eigenvalue, the weighted subtraction between the product and the i+1-th eigenvalue is calculated to obtain the optimized i+1-th eigenvalue, which is expressed as:

[0117] v i+1 '=ω 3 ×v i ×base-ω 4 ×v i+1

[0118] Among them, ω 3 ,ω 4 Respectively represent the third weight parameter and the fourth weight parameter, v i+1 ' represents the i+1th eigenvalue of the optimized prototype offset semantic order encoding vector of the sample point to be processed;

[0119] Based on v' 1 =v 1 , the i+1th eigenvalue v of the combinatorial optimization i+1 'To obtain the optimized prototype offset semantic encoding vector of the sample point to be processed.

[0120] In order to solve the problem that the feature set of the prototype offset semantic coding vector of the sample point to be processed has insufficient global interactive response representation capability due to the long distance exceeding the predetermined local distribution interval threshold under the predetermined feature value sequential distribution, the high-dimensional feature space primitive representation of the prototype offset semantic coding vector of the sample point to be processed based on self-inner product fusion is used to capture the complex structure of the global network interaction of its feature values, so as to reconstruct the interactive response relationship between the feature values ​​of the prototype offset semantic coding vector of the sample point to be processed by simulating the scale-based high-dimensional feature space potential primitives, so as to realize the coding reconstruction of the real sequence distribution behavior of the prototype offset semantic coding vector of the sample point to be processed under long distance, improve the coding expression effect of the prototype offset semantic coding vector of the sample point to be processed, and improve the accuracy of the recognition result obtained by inputting it into the classifier-based sample identifier.

[0121] Specifically, in step S73, in response to the recognition result that the first data sample to be processed is a noise sample, the first data sample to be processed is deleted from the first data sample subset to be processed. It should be understood that since the noise sample will seriously interfere with subsequent data processing and analysis, reducing the accuracy and reliability of the results, therefore, in response to the recognition result that the first data sample to be processed is a noise sample, the first data sample to be processed is deleted from the first data sample subset to be processed.

[0122] Specifically, after confirming that a sample is a noise sample, the next task is to remove the sample from the current data subset. To achieve this, technical details at the database management and file operation levels are usually involved. If the data is stored in a relational database, the DELETE command can be directly executed through SQL statements, specifying a unique identifier (such as a primary key) to locate and delete specific records. For non-relational databases or distributed file systems, you may need to use the corresponding API interface or command line tool to perform similar operations.

[0123] In addition to simple physical deletion, another common practice is to mark noise samples in the original data set without immediately removing them from the disk. The advantage of this approach is that it preserves the integrity of the original data, making it easier to review or recheck in the future. Specifically, a column of flags can be added to the data table to indicate the status of each sample, that is, whether it has been identified as noise. When performing data analysis or modeling, you only need to filter out samples that are marked as noise. This method is particularly suitable for large-scale data sets because the cost of rebuilding the entire data set is high, and logical deletion can significantly reduce resource consumption.

[0124] Considering the complex situations in practical applications, sometimes it is necessary to control the deletion process more finely. For example, in some scenarios, you may want to temporarily retain suspected noise samples for a period of time and wait for further verification before making a final decision. This requires the system to have flexible time management functions and allow the setting of delayed deletion policies. In addition, different processing rules can be defined for different types of noise samples. For example, obviously erroneous data can be directly deleted permanently, while edge cases can be archived in a separate data set for review by auditors. Such flexibility helps improve decision-making accuracy while reducing the risk of accidentally deleting useful information.

[0125] Before performing a deletion operation, the impact on the overall data set must be fully evaluated. Especially when it comes to key business indicators or sensitive information, any changes may lead to unforeseen problems. For this reason, it is recommended to simulate the entire process in a small-scale test environment to check whether there are potential risk points. Only after full verification can it be safely applied to the production environment. In addition, a complete historical record mechanism should be established to record in detail the specific time, reason and scope of impact of each deletion operation, so as to track the source of the problem in the future or meet audit needs.

[0126] To summarize, an AI-based big data acquisition and cleaning method based on an embodiment of the present application is explained, which first groups the data sample set to be processed according to the data type, extracts the first subset of data samples to be processed, and introduces deep learning-based data processing technology to vectorize and encode each data sample to be processed in the subset to extract data sample features. Then, by performing hub search and dynamic compensation aggregation coding on each data sample feature to be processed, the prototype anchoring feature of the data sample set is captured. Based on this, the degree of semantic deviation of each data sample feature to be processed relative to the prototype anchoring feature of the data set is measured to realize the identification and cleaning of noise samples, which can effectively identify and eliminate noise data in large-scale data sets and improve the accuracy and efficiency of data cleaning.

[0127] Furthermore, an AI-based big data collection and cleaning system is also provided.

[0128] Figure 6 FIG. 1 is a block diagram of an AI-based big data collection and cleaning system according to an embodiment of the present application. Figure 6As shown, according to the AI-based big data acquisition and cleaning system 100 of the embodiment of the present application, it includes: a data sample set acquisition module 110, which is used to acquire a data sample set to be processed; a data sample set grouping module 120, which is used to group the data sample set to be processed based on the data type to obtain a set of data sample subsets to be processed; a sample subset extraction module 130, which is used to extract a first data sample subset to be processed from the set of data sample subsets to be processed; a vectorized encoding module 140, which is used to vectorize and encode each data sample to be processed in the first data sample subset to be processed to obtain a set of embedded encoding vectors of the data sample to be processed ; A dynamic compensation aggregation module 150 is used to perform dynamic compensation aggregation based on feature hubs on the set of embedded coding vectors of the data samples to be processed to obtain a data sample prototype anchor feature vector; a query feature extraction module 160 is used to extract a first embedded coding vector of the data sample to be processed from the set of embedded coding vectors of the data samples to be processed as a query feature vector; a sample screening module 170 is used to determine whether to delete the first data sample to be processed corresponding to the first embedded coding vector of the data sample to be processed from the first subset of data samples to be processed based on the semantic offset between the query feature vector and the data sample prototype anchor feature vector.

[0129] The specific operations of each module in the above AI-based big data collection and cleaning system have been referenced above. Figures 1 to 5 The description of the AI-based big data acquisition cleaning method has been introduced in detail, and therefore, its repeated description will be omitted.

[0130] The basic principle of the present invention is described above in conjunction with specific embodiments. However, it should be pointed out that the advantages, strengths, effects, etc. mentioned in the present invention are only examples and not limitations, and it cannot be considered that these advantages, strengths, effects, etc. must be possessed by each embodiment of the present invention. In addition, the specific details of the above embodiments are only for the purpose of illustration and facilitation of understanding, rather than limitation, and the above details do not limit the present invention to being implemented by adopting the above specific details.

[0131] In the above embodiments, the description of each embodiment has its own emphasis. For the parts that are not described or recorded in detail in a certain embodiment, please refer to the relevant description of other embodiments. In the several embodiments provided by the present invention, it should be understood that the disclosed system and method can be implemented in other ways. For example, the system embodiment described above is only schematic. For example, the unit division is only a logical function division, and there may be other division methods in actual implementation. The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the scheme of this embodiment.

[0132] It will be apparent to those skilled in the art that the invention is not limited to the details of the exemplary embodiments described above and that the invention can be implemented in other specific forms without departing from the spirit or essential features of the invention. Therefore, the embodiments should be considered exemplary and non-limiting in all respects, and the scope of the invention is defined by the appended claims rather than the foregoing description, and it is intended that all variations falling within the meaning and scope of the equivalent elements of the claims be included in the invention. Any reference to a figure in a claim should not be considered as limiting the claim to which it relates.

[0133] In addition, it is obvious that the word "comprising" does not exclude other units or steps, and the singular does not exclude the plural. The multiple units stated in the system claims can also be implemented by one unit through software or hardware.

[0134] Finally, it should be noted that the above description has been given for the purpose of illustration and description. In addition, the above embodiments are only used to illustrate the technical solution of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to the preferred embodiments, it should be understood by those skilled in the art that the technical solution of the present invention can be modified or replaced by equivalents without departing from the spirit and scope of the technical solution of the present invention.

Claims

1. A big data collection and cleaning method based on AI, characterized in that: include: Obtain a sample set of data to be processed; Based on the data type, the to-be-processed data sample set is grouped to obtain a set of to-be-processed data sample subsets; Extracting a first subset of data samples to be processed from the set of subsets of data samples to be processed; Performing vectorized encoding on each of the data samples to be processed in the first subset of data samples to be processed to obtain a set of embedded encoding vectors of the data samples to be processed; Performing dynamic compensation aggregation based on feature hubs on the set of embedding coding vectors of the data sample to be processed to obtain a data sample prototype anchoring feature vector; Extracting a first embedded coding vector of the data sample to be processed from the set of embedded coding vectors of the data sample to be processed as a query feature vector; Based on the semantic offset between the query feature vector and the data sample prototype anchor feature vector, it is determined whether to delete the first data sample to be processed corresponding to the first data sample to be processed embedding encoding vector from the first data sample subset to be processed.

2. The AI-based big data collection and cleaning method according to claim 1, characterized in that: The set of embedded coding vectors of the data sample to be processed is subjected to dynamic compensation aggregation based on feature hubs to obtain a data sample prototype anchor feature vector, including: Inputting the set of embedding encoding vectors of the data samples to be processed into the hub feature extraction network to obtain the hub feature encoding vectors of the data samples to be processed; Extracting complementary information of each of the embedded coding vectors of the data sample to be processed relative to the hub feature coding vector of the data sample to be processed in the set of embedded coding vectors of the data sample to be processed to obtain a set of embedded coding vectors of the data sample to be processed-hub feature complementary information; Based on the hub feature coding vector of the data sample to be processed, the set of the data sample to be processed-hub feature complementary information embedding coding vector is significantly dynamically modulated to obtain a set of significantly modulated data sample to be processed-hub feature complementary information embedding coding vectors; The data sample prototype anchoring feature vector is obtained by fusing the hub feature encoding vector of the data sample to be processed and the set of the significantly modulated data sample to be processed-hub feature complementary information embedding encoding vectors.

3. The AI-based big data collection and cleaning method according to claim 2 is characterized in that: Extracting complementary information of each of the embedded coding vectors of the data sample to be processed in the set of the embedded coding vectors of the data sample to be processed relative to the hub feature coding vector of the data sample to be processed to obtain a set of embedded coding vectors of the data sample to be processed-hub feature complementary information, including: Performing point convolution coding based on Sigmoid activation function on the embedded coding vector of the data sample to be processed and the pivot feature coding vector of the data sample to be processed respectively to obtain a standardized embedded coding vector of the data sample to be processed and a standardized pivot feature coding vector of the data sample to be processed; Taking the positional difference vector between the standardized embedded coding vector of the data sample to be processed and the standardized hub feature coding vector of the data sample to be processed as the weight vector, the embedded coding vector of the data sample to be processed and the hub feature coding vector of the data sample to be processed are subjected to complementary feature enhancement modulation and aggregation coding to obtain the embedded coding vector of the data sample to be processed-hub feature complementary information.

4. The AI-based big data collection and cleaning method according to claim 3 is characterized in that: Based on the hub feature coding vector of the data sample to be processed, a set of the data sample to be processed-hub feature complementary information embedding coding vectors is modulated dynamically to obtain a set of significantly modulated data sample to be processed-hub feature complementary information embedding coding vectors, including: Inputting each of the to-be-processed data sample-hub feature complementary information embedding coding vectors in the set of the to-be-processed data sample-hub feature complementary information embedding coding vectors into a complementary information significant identification module based on an attention mechanism to obtain a set of to-be-processed data sample complementary information attention weights; Based on the set of attention weights of the complementary information of the data samples to be processed, the set of the complementary information embedded coding vectors of the data samples to be processed and the hub features is subjected to attention modulation to obtain the set of the significantly modulated complementary information embedded coding vectors of the data samples to be processed and the hub features.

5. The AI-based big data collection and cleaning method according to claim 4 is characterized in that: The method of fusing the hub feature encoding vector of the data sample to be processed and the set of the significantly modulated data sample to be processed-hub feature complementary information embedding encoding vectors to obtain the data sample prototype anchoring feature vector comprises: The hub feature encoding vector of the data sample to be processed and the set of the significantly modulated data sample to be processed-hub feature complementary information embedding encoding vectors are cascaded and fused to obtain the data sample prototype anchoring feature vector.

6. The AI-based big data collection and cleaning method according to claim 5, characterized in that: Determining whether to delete a first data sample to be processed corresponding to the first data sample to be processed embedding encoding vector from the first data sample to be processed subset based on a semantic offset between the query feature vector and the data sample prototype anchor feature vector includes: Extracting the semantic offset feature between the query feature vector and the data sample prototype anchor feature vector to obtain a prototype offset semantic encoding vector of the sample point to be processed; Inputting the prototype offset semantic coding vector of the sample point to be processed into a sample identifier based on a classifier to obtain a recognition result, wherein the recognition result is used to indicate whether the first data sample to be processed is a noise sample; In response to the identification result that the first data sample to be processed is a noise sample, the first data sample to be processed is deleted from the first subset of data samples to be processed.

7. The AI-based big data collection and cleaning method according to claim 6, characterized in that: Extracting the semantic offset feature between the query feature vector and the data sample prototype anchor feature vector to obtain the prototype offset semantic encoding vector of the sample point to be processed, including: The position difference between the query feature vector and the data sample prototype anchor feature vector is calculated to obtain the prototype offset semantic encoding vector of the sample point to be processed.

8. The AI-based big data collection and cleaning method according to claim 7, characterized in that: Inputting the prototype offset semantic coding vector of the sample point to be processed into a sample identifier based on a classifier to obtain a recognition result, wherein the recognition result is used to indicate whether the first data sample to be processed is a noise sample, including: Using the fully connected layer of the sample identifier to perform fully connected encoding on the prototype offset semantic encoding vector of the sample point to be processed to obtain the prototype offset semantic fully connected encoding vector of the sample point to be processed; Inputting the prototype offset semantic fully connected encoding vector of the sample point to be processed into the Softmax classification function of the sample identifier to obtain the probability value of the prototype offset semantic encoding vector of the sample point to be processed belonging to each classification label, wherein the classification label includes whether the first data sample to be processed is a noise sample and whether the first data sample to be processed is not a noise sample; The classification label corresponding to the largest probability value among the probability values ​​is determined as the recognition result.

9. The AI-based big data collection and cleaning system is characterized by: include: A data sample set acquisition module is used to acquire a data sample set to be processed; A data sample set grouping module, used for grouping the data sample set to be processed based on data type to obtain a set of data sample subsets to be processed; A sample subset extraction module, used to extract a first to-be-processed data sample subset from the set of to-be-processed data sample subsets; A vectorized encoding module, used for performing vectorized encoding on each of the data samples to be processed in the first subset of data samples to be processed to obtain a set of embedded encoding vectors of the data samples to be processed; A dynamic compensation aggregation module, used for performing dynamic compensation aggregation based on feature hubs on the set of embedded coding vectors of the data sample to be processed to obtain a data sample prototype anchor feature vector; A query feature extraction module, used to extract a first embedded coding vector of the data sample to be processed from the set of embedded coding vectors of the data sample to be processed as a query feature vector; A sample screening module is used to determine whether to delete the first data sample to be processed corresponding to the first data sample to be processed embedding coding vector from the first data sample to be processed subset based on the semantic offset between the query feature vector and the data sample prototype anchor feature vector.