Processing method and system for collecting, cleaning and labeling culture big data and medium

By evaluating the collection, cleaning and labeling technology of type division and adaptation of cultural big data, the problems of inefficiency and low accuracy in the existing technology are solved, and efficient and accurate data processing is achieved.

CN120296078APending Publication Date: 2025-07-11BEIJING BLANSTAR TECH CO LTD

Patent Information

Application Number
CN202510800031.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-16
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

When facing massive and complex cultural big data, the existing technology has problems such as inefficiency, low accuracy and high labor costs, and it is difficult to efficiently collect, clean and label.

Method used

By dividing cultural big data in type, using adaptive acquisition technology for data collection, evaluating the adaptability of the acquisition technology; cleaning the target data subset to evaluate the suitability of the cleaning technology; labeling the cleaned data subset to evaluate the fit of the labeling technology, and realizing intelligent processing.

Benefits of technology

It realizes efficient and accurate collection, cleaning and labeling of cultural big data, provides a reliable data foundation, and lays a solid foundation for subsequent analysis and application.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120296078A_ABST
    Figure CN120296078A_ABST
Patent Text Reader

Abstract

The invention provides a processing method and system for collecting, cleaning and labeling culture big data and a medium. The method comprises the steps that to-be-processed culture big data is subjected to type division, corresponding collection technologies are adopted for collection, corresponding target culture big data subsets are obtained, data collection evaluation indexes corresponding to the target culture big data subsets are extracted, the adaptation degree of the corresponding collection technologies is evaluated, and a data collection result is obtained; and processing the target culture big data subset by adopting a corresponding cleaning technology to obtain a corresponding cleaned data subset, extracting a data cleaning evaluation index corresponding to each cleaned data subset, evaluating the adaptability of the corresponding cleaning technology, processing the cleaned data subset by adopting a corresponding labeling technology, and obtaining a data cleaning result. The method comprises the following steps of: acquiring a corresponding labeled data subset, extracting a data labeling evaluation index corresponding to each labeled data subset, and evaluating the integrating degree of a corresponding labeling technology, so that the technology for collecting, cleaning and intelligently labeling the culture big data is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of cultural big data processing. Specifically, it relates to a processing method, system, and medium for cultural big data collection, cleaning, and annotation. Background Art

[0002] With the rapid development of information technology, the scale of cultural big data has increased explosively. Cultural big data covers various content types such as text, images, audio, and video. How to efficiently collect, clean, and annotate these data to obtain valuable information has become an urgent problem to be solved. Traditional data processing methods have defects such as low efficiency, low accuracy, and high labor costs when dealing with massive and complex cultural big data.

[0003] In view of the above problems, there is an urgent need for effective technical solutions. Summary of the Invention

[0004] The purpose of this application is to provide a processing method, system, and medium for cultural big data collection, cleaning, and annotation. It can classify the cultural big data to be processed, adopt corresponding collection technologies for collection to obtain corresponding target cultural big data subsets, extract data collection evaluation indicators, and evaluate the adaptability of the corresponding collection technologies. Then, adopt corresponding cleaning technologies to process the target cultural big data subsets to obtain corresponding cleaned data subsets, extract data cleaning evaluation indicators, and evaluate the applicability of the corresponding cleaning technologies. Next, adopt corresponding annotation technologies to process the cleaned data subsets to obtain corresponding annotated data subsets, extract data annotation evaluation indicators, and evaluate the fitness of the corresponding annotation technologies, so as to realize the technology of intelligent processing of cultural big data collection, cleaning, and annotation.

[0005] This application also provides a processing method for cultural big data collection, cleaning, and annotation, including the following steps: Classify the cultural big data to be processed, and adopt corresponding collection technologies for the data subsets of the type to obtain corresponding target cultural big data subsets; Extract the data collection evaluation indicators corresponding to each of the target cultural big data subsets, and evaluate the adaptability of the corresponding collection technologies; Adopt corresponding cleaning technologies to process the target cultural big data subsets to obtain corresponding cleaned data subsets; Extract the data cleaning evaluation indicators corresponding to each of the cleaned data subsets, and evaluate the applicability of the corresponding cleaning technologies; Adopt corresponding annotation technologies to process the cleaned data subsets to obtain corresponding annotated data subsets; Extract the data annotation evaluation metrics corresponding to each of the annotated data subsets, and evaluate the degree of fit with the corresponding annotation technology.

[0006] Optionally, in the method for collecting, cleaning, and annotating cultural big data described in this application, the method for classifying the cultural big data to be processed and collecting the corresponding data subsets of the classified data by using the corresponding collection technology to obtain the corresponding target cultural big data subsets includes: Classify the cultural big data to be processed according to the data content to obtain data subsets of different types; The types include text, image, audio, and video; Collect the corresponding data subsets of the types by using the corresponding collection technology to obtain the corresponding target cultural big data subsets; If the type is text, use web crawler technology for data collection; If the type is image, use image recognition technology for data collection; If the type is audio, use speech recognition technology for data collection; If the type is video, use video recognition technology for data collection.

[0007] Optionally, in the method for collecting, cleaning, and annotating cultural big data described in this application, the method for extracting the data collection evaluation metrics corresponding to each of the target cultural big data subsets and evaluating the adaptability of the corresponding collection technology includes: Extract the data collection evaluation metrics corresponding to each of the target cultural big data subsets; The data collection evaluation metrics include error rate, data coverage rate, proportion of data missing values, and collection speed; Process the error rate, data coverage rate, proportion of data missing values, and collection speed through a preset data collection ability evaluation model to obtain collection reliability data; Compare the collection reliability data with a preset collection reliability threshold; Evaluate the adaptability of the corresponding collection technology according to the comparison result.

[0008] Optionally, in the method for collecting, cleaning, and annotating cultural big data described in this application, the method for processing the target cultural big data subsets by using the corresponding cleaning technology to obtain the corresponding cleaned data subsets includes: Process the target cultural big data subsets by using the corresponding cleaning technology to obtain the corresponding cleaned data subsets; The cleaning technology includes one or more of data deduplication, missing value processing, outlier processing, and noise data processing.

[0009] Optionally, in the method for processing cultural big data collection, cleaning, and annotation described in this application, the step of extracting the data cleaning evaluation indicators corresponding to each of the cleaned data subsets and evaluating the applicability of the corresponding cleaning technology includes: Extracting the data cleaning evaluation indicators corresponding to each of the cleaned data subsets; The data cleaning evaluation indicators include the error rate reduction value, the data loss rate, and the cleaning running time; Processing according to the error rate reduction value, the data loss rate, and the cleaning running time through a preset data cleaning ability evaluation model to obtain cleaning reliability data; Comparing the cleaning reliability data with a preset cleaning reliability threshold; Evaluating the applicability of the corresponding cleaning technology according to the comparison result.

[0010] Optionally, in the method for processing cultural big data collection, cleaning, and annotation described in this application, the step of processing the cleaned data subsets with corresponding annotation technologies to obtain corresponding annotated data subsets includes: Processing the cleaned data subsets with corresponding annotation technologies to obtain corresponding annotated data subsets; The annotation technologies include one or more of crowdsourcing annotation, regular expression annotation, deep learning annotation, supervised learning annotation, semi-supervised learning annotation, and unsupervised learning annotation.

[0011] Optionally, in the method for processing cultural big data collection, cleaning, and annotation described in this application, the step of extracting the data annotation evaluation indicators corresponding to each of the annotated data subsets and evaluating the fitness of the corresponding annotation technology includes: Extracting the data annotation evaluation indicators corresponding to each of the annotated data subsets; The data annotation evaluation indicators include annotation precision, recall rate, annotation accuracy rate, and annotation speed; Processing according to the annotation precision, recall rate, annotation accuracy rate, and annotation speed through a preset data annotation ability evaluation model to obtain annotation reliability data; Comparing the annotation reliability data with a preset annotation reliability threshold; Evaluating the fitness of the corresponding annotation technology according to the comparison result.

[0012] In a second aspect, this application provides a processing system for cultural big data collection, cleaning, and annotation. The system includes: a memory and a processor. The memory includes a program for the method of processing cultural big data collection, cleaning, and annotation. When the program for the method of processing cultural big data collection, cleaning, and annotation is executed by the processor, the following steps are implemented: Classify the cultural big data to be processed, and collect the corresponding data subsets of the type using corresponding collection technologies to obtain the corresponding target cultural big data subsets; Extract the data collection evaluation indicators corresponding to each of the target cultural big data subsets, and evaluate the adaptability of the corresponding collection technologies; Process the target cultural big data subsets using corresponding cleaning technologies to obtain the corresponding cleaned data subsets; Extract the data cleaning evaluation indicators corresponding to each of the cleaned data subsets, and evaluate the applicability of the corresponding cleaning technologies; Process the cleaned data subsets using corresponding annotation technologies to obtain the corresponding annotated data subsets; Extract the data annotation evaluation indicators corresponding to each of the annotated data subsets, and evaluate the fitness of the corresponding annotation technologies.

[0013] Optionally, in the processing system for cultural big data collection, cleaning and annotation of the present application, the classifying the cultural big data to be processed, and collecting the corresponding data subsets of the type using corresponding collection technologies to obtain the corresponding target cultural big data subsets includes: Classify the cultural big data to be processed according to the data content to obtain data subsets of different types; The types include text, images, audio, and video; Collect the corresponding data subsets of the type using corresponding collection technologies to obtain the corresponding target cultural big data subsets; If the type is text, use web crawler technology for data collection; If the type is images, use image recognition technology for data collection; If the type is audio, use speech recognition technology for data collection; If the type is video, use video recognition technology for data collection.

[0014] In a third aspect, the present application also provides a computer-readable storage medium, in which a program for the processing method of cultural big data collection, cleaning and annotation is stored. When the program for the processing method of cultural big data collection, cleaning and annotation is executed by a processor, the steps of the processing method of cultural big data collection, cleaning and annotation as described in any one of the above are implemented.

[0015] As can be seen from the above, the processing method, system, and medium for collecting, cleaning, and annotating cultural big data disclosed by the present invention classify the cultural big data to be processed, adopt corresponding collection technologies for collection to obtain corresponding target cultural big data subsets, extract the data collection evaluation indicators corresponding to each target cultural big data subset, and evaluate the adaptability of the corresponding collection technologies. Then, corresponding cleaning technologies are adopted to process the target cultural big data subsets to obtain corresponding cleaned data subsets, the data cleaning evaluation indicators corresponding to each cleaned data subset are extracted, and the applicability of the corresponding cleaning technologies is evaluated. Next, corresponding annotation technologies are adopted to process the cleaned data subsets to obtain corresponding annotated data subsets, the data annotation evaluation indicators corresponding to each annotated data subset are extracted, and the fit of the corresponding annotation technologies is evaluated, thereby realizing the technology of intelligent processing of collecting, cleaning, and annotating cultural big data.

[0016] Other features and advantages of the present application will be described in the subsequent specification, and, in part, will be obvious from the specification, or can be understood by implementing the embodiments of the present application. The objectives and other advantages of the present application can be realized and obtained by the structures specifically pointed out in the written specification and the drawings. Brief Description of the Drawings

[0017] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the drawings required to be used in the embodiments of the present application. It should be understood that the following drawings only show some embodiments of the present application, and therefore should not be regarded as limiting the scope. For those of ordinary skill in the art, without creative efforts, other relevant drawings can also be obtained based on these drawings.

[0018] Figure 1 It is a flowchart of the processing method for collecting, cleaning, and annotating cultural big data provided by the embodiments of the present application; Figure 2 It is a flowchart of obtaining the target cultural big data subset of the processing method for collecting, cleaning, and annotating cultural big data provided by the embodiments of the present application; Figure 3 It is a flowchart of evaluating the adaptability of the collection technology of the processing method for collecting, cleaning, and annotating cultural big data provided by the embodiments of the present application; Figure 4 It is a flowchart of obtaining the cleaned data subset of the processing method for collecting, cleaning, and annotating cultural big data provided by the embodiments of the present application. Detailed Description of the Embodiments

[0019] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. The components of the embodiments of the present application described and illustrated in the accompanying drawings here can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present application provided in the accompanying drawings is not intended to limit the scope of the present application to be protected, but only represents the selected embodiments of the present application. All other embodiments obtained by those skilled in the art based on the embodiments of the present application without creative efforts belong to the scope of protection of the present application.

[0020] It should be noted that similar reference numerals and letters indicate similar items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings. At the same time, in the description of the present application, terms such as "first" and "second" are only used for differential description and cannot be understood as indicating or implying relative importance.

[0021] Please refer to Figure 1 , Figure 1 which is a flowchart of a processing method for cultural big data collection, cleaning, and annotation in some embodiments of the present application. This processing method for cultural big data collection, cleaning, and annotation is used in terminal devices, such as computers, mobile phone terminals, etc. This processing method for cultural big data collection, cleaning, and annotation includes the following steps: S11. Classify the cultural big data to be processed, and collect the corresponding target cultural big data subsets by using the corresponding collection technologies for the data subsets of the types; S12. Extract the data collection evaluation indicators corresponding to each of the target cultural big data subsets, and evaluate the adaptability of the corresponding collection technologies; S13. Process the target cultural big data subsets by using the corresponding cleaning technologies to obtain the corresponding cleaned data subsets; S14. Extract the data cleaning evaluation indicators corresponding to each of the cleaned data subsets, and evaluate the applicability of the corresponding cleaning technologies; S15. Process the cleaned data subsets by using the corresponding annotation technologies to obtain the corresponding annotated data subsets; S16. Extract the data annotation evaluation indicators corresponding to each of the annotated data subsets, and evaluate the fitness of the corresponding annotation technologies.

[0022] It should be noted that in the intelligent processing of cultural big data, a scientific and systematic processing process is the key to ensuring data quality and processing efficiency. First of all, it is necessary to conduct a detailed type division according to the content form of cultural big data, covering text type, image type, audio type, and video type. For different types of data, corresponding acquisition technologies need to be used for accurate acquisition. For text data, intelligent web crawler technology can be adopted. By analyzing the web page structure and setting keyword rules, during the acquisition process, technologies such as regular expressions and XPath can be used to accurately extract the required text content and process dynamic web page data to ensure the integrity of the acquisition. For image data, image grabbing tools and API interfaces can be used. Audio data acquisition can utilize audio recording software and streaming media data acquisition technology. During acquisition, attention should be paid to the conversion of audio formats and the integrity of data. Video data acquisition can be achieved through video download tools and video stream parsing technology, and at the same time, parameters such as the resolution and frame rate of the video are adaptively adjusted. Through the above targeted acquisition technologies, the corresponding target cultural big data subsets can be obtained, and these subsets form the basis for subsequent data processing. After the acquisition is completed, it is crucial to extract the data acquisition evaluation indicators corresponding to each target cultural big data subset. Through the quantitative analysis of these indicators, the adaptability of the corresponding acquisition technology is evaluated, and it is judged whether the acquisition technology can meet the acquisition requirements of this type of cultural big data. If the adaptability does not meet the standard, the acquisition technology needs to be optimized or replaced to ensure the quality of the acquired data; after the acquisition is completed, for each target cultural big data subset, corresponding cleaning technologies need to be used for processing to remove noise, errors, and redundant information in the data, and the corresponding cleaned data subsets are obtained. The cleaned data also needs to be strictly evaluated, and the data cleaning evaluation indicators corresponding to each cleaned data subset are extracted. Through the evaluation of these indicators, the applicability of the corresponding cleaning technology is judged. If the applicability is insufficient, the cleaning algorithm or process needs to be further optimized. After the cleaning is completed, it enters the data annotation link. The corresponding annotation technology is used to process the cleaned data subsets to endow the data with semantic information, and the corresponding annotated data subsets are obtained. Finally, the data annotation evaluation indicators corresponding to each annotated data subset are extracted to evaluate the compatibility of the corresponding annotation technology. Through the comprehensive evaluation of these indicators, it is comprehensively judged whether the annotation technology meets the annotation requirements of this type of cultural big data, providing a reliable data basis for subsequent data application and analysis, so as to achieve the technical goal of intelligent processing of cultural big data for acquisition, cleaning, and annotation.

[0023] Please refer to Figure 2 , Figure 2It is a flowchart for obtaining a target cultural big data subset in the processing method of cultural big data collection, cleaning, and annotation in some embodiments of this application. According to the embodiments of the present invention, the cultural big data to be processed is classified by type, and corresponding collection techniques are adopted for the data subsets of the types to obtain corresponding target cultural big data subsets, including: S21. Classify the cultural big data to be processed according to the data content to obtain data subsets of different types; S22. The types include text, image, audio, and video; S23. Adopt corresponding collection techniques for the data subsets of the types to obtain corresponding target cultural big data subsets; S24. If the type is text, use web crawler technology for data collection; S25. If the type is image, use image recognition technology for data collection; S26. If the type is audio, use speech recognition technology for data collection; S27. If the type is video, use video recognition technology for data collection.

[0024] It should be noted that in the processing flow of cultural big data, data collection, as the starting and crucial link, its accuracy and efficiency directly affect the quality of subsequent data cleaning, annotation, and in-depth analysis. Facing the massive and complex cultural big data to be processed, the primary task is to conduct detailed data type classification based on the data content to obtain different data subsets. These types mainly cover four categories: text, image, audio, and video. Each data type has a unique manifestation form and data characteristics, which requires the adoption of highly adaptable collection techniques to ensure that the collected data can completely and accurately reflect cultural information. Specifically, if it is text data, use web crawler technology for data collection; if it is image data, use image recognition technology for data collection; if it is audio data, use speech recognition technology for data collection; if it is video data, use video recognition technology for data collection. Through the above classification collection methods for different data types and the use of their respective adaptable collection techniques, corresponding target cultural big data subsets can be obtained. These high-quality data sets lay a solid foundation for the subsequent cleaning, annotation, and in-depth mining analysis of cultural big data, helping to achieve the efficient and intelligent processing of cultural big data.

[0025] Please refer to Figure 3 , Figure 3It is a flowchart for evaluating the adaptability of the acquisition technology in the processing method of cultural big data acquisition, cleaning and annotation in some embodiments of the present application. According to the embodiments of the present invention, extracting the data acquisition evaluation indicators corresponding to each of the target cultural big data subsets and evaluating the adaptability of the corresponding acquisition technology includes: S31. Extract the data acquisition evaluation indicators corresponding to each of the target cultural big data subsets; S32. The data acquisition evaluation indicators include error rate, data coverage rate, data missing value ratio, and acquisition speed; S33. Process through a preset data acquisition capability evaluation model according to the error rate, data coverage rate, data missing value ratio, and acquisition speed to obtain acquisition reliability data; S34. Compare the acquisition reliability data with a preset acquisition reliability threshold; S35. Evaluate the adaptability of the corresponding acquisition technology according to the comparison result.

[0026] It should be noted that after the classification and collection of cultural big data are completed and each target cultural big data subset is obtained, in order to ensure the effectiveness and applicability of the collection technology, a comprehensive and detailed evaluation of the collection process is required. The key initial step in this evaluation process is to extract the data collection evaluation indicators corresponding to each target cultural big data subset. These indicators can accurately reflect the quality and efficiency of data collection from multiple dimensions and provide a basis for subsequent judgment on whether the collection technology is suitable. The data collection evaluation indicators mainly cover several important aspects such as error rate, data coverage rate, proportion of data missing values, and collection speed. The error rate is a key indicator for measuring the deviation degree between the collected data and the real data. The calculation of the error rate is usually through comparing the difference between the collected data and the known real data and presenting it in the form of a proportion. The lower the error rate, the higher the accuracy of the collected data. The data coverage rate is used to measure the coverage degree of the collected data within the entire target data range. The level of the data coverage rate directly reflects whether the collection technology can comprehensively obtain the target data. A higher data coverage rate means that the collection technology can cover all aspects of the target data as much as possible and provide richer and more comprehensive information for subsequent data analysis and mining. The proportion of data missing values refers to the proportion of the missing part in the collected data. An excessively high proportion of data missing values will seriously affect the quality and usability of the data, so it needs to be strictly monitored and evaluated. The collection speed is also an important evaluation indicator, which reflects the amount of data that the collection technology can collect per unit time. The speed of data collection is affected by various factors. After these data collection evaluation indicators are extracted, next, it is necessary to comprehensively process these indicators by means of a preset data collection ability evaluation model. This evaluation model is constructed based on a large amount of historical data and practical experience. It can calculate the collection reliability data according to the mutual relationship between indicators such as error rate, data coverage rate, proportion of data missing values, and collection speed through complex algorithms and mathematical models. The collection reliability data is a comprehensive evaluation result, which can comprehensively reflect the reliability and effectiveness of the collection technology in the current collection task.After obtaining the acquisition reliability data, it is necessary to compare it with the preset acquisition reliability threshold. By comparing the acquisition reliability data with the threshold, it is possible to intuitively determine whether the acquisition technology meets the requirements. Finally, based on the comparison results, the adaptability of the corresponding acquisition technology is evaluated. If the acquisition reliability data is higher than the preset acquisition reliability threshold, it indicates that the acquisition technology performs well in the current acquisition task, has a high adaptability, and can effectively complete the data acquisition work. This technology can continue to be used for subsequent data acquisition. Conversely, if the acquisition reliability data is lower than the preset acquisition reliability threshold, it indicates that there may be some problems with the acquisition technology, such as inaccurate acquisition algorithms or insufficient performance of the acquisition tools. In this case, it is necessary to adjust and optimize the acquisition technology, or consider replacing it with other more suitable acquisition technologies to ensure that the cultural big data acquisition task can be completed with high quality and efficiency.;

[0027] Please refer to Figure 4 , Figure 4 FIG. is a flowchart for obtaining a cleaned data subset in a method for collecting, cleaning, and annotating cultural big data according to some embodiments of the present application. According to an embodiment of the present invention, processing the target cultural big data subset with a corresponding cleaning technique to obtain a corresponding cleaned data subset includes: S41. Processing the target cultural big data subset with a corresponding cleaning technique to obtain a corresponding cleaned data subset; S42. The cleaning technique includes one or more of data deduplication, missing value processing, outlier processing, and noise data processing.

[0028] It should be noted that after obtaining the target cultural big data subset, due to the wide and complex data sources, problems such as duplication, missing, outliers, and noise inevitably exist, seriously affecting the data quality and the accuracy of subsequent analysis and applications. Therefore, it is necessary to deeply process these target cultural big data subsets with corresponding cleaning techniques to obtain high-quality cleaned data subsets. The cleaning techniques cover various means such as data deduplication, missing value processing, outlier processing, and noise data processing. They cooperate with each other to jointly improve the usability and reliability of the data. Moreover, for different data types, the cleaning techniques adopted will be different, and one or more of data deduplication, missing value processing, outlier processing, and noise data processing can be selected according to actual needs.

[0029] According to an embodiment of the present invention, extracting the data cleaning evaluation indexes corresponding to each of the cleaned data subsets and evaluating the applicability of the corresponding cleaning technique includes: Extracting the data cleaning evaluation indexes corresponding to each of the cleaned data subsets; The data cleaning evaluation metrics include the error rate reduction value, the data loss rate, and the cleaning running time; Process the error rate reduction value, the data loss rate, and the cleaning running time through a preset data cleaning ability evaluation model to obtain cleaning reliability data; Compare the cleaning reliability data with a preset cleaning reliability threshold; Evaluate the applicability of the corresponding cleaning technology according to the comparison result.

[0030] It should be noted that after using the corresponding cleaning technology to process the target cultural big data subset and obtaining the cleaned data subset, in order to measure the quality and efficiency of the cleaning work and determine whether the adopted cleaning technology is suitable for the data characteristics, it is necessary to conduct a comprehensive and scientific evaluation of the cleaning process. The first step in this evaluation process is to extract the evaluation metrics corresponding to each cleaned data subset. These metrics can reflect the effect and performance of data cleaning from different dimensions. The data cleaning evaluation metrics mainly include the error rate reduction value, the data loss rate, and the cleaning running time. Among them, the calculation method of the error rate reduction value is the error rate of the data before cleaning minus the error rate of the data after cleaning. The data loss rate is expressed as the ratio of the amount of data deleted during the cleaning process to the total amount of data before cleaning. The cleaning running time refers to the time spent from the start of the cleaning operation to the completion of the entire cleaning process. After extracting these data cleaning evaluation metrics, the next step is to use a preset data cleaning ability evaluation model to comprehensively process these metrics. This evaluation model is constructed based on a large amount of historical data and practical experience. It can take into account the mutual relationship between the error rate reduction value, the data loss rate, and the cleaning running time, and through complex algorithms and mathematical models, convert these metrics into a single value, that is, the cleaning reliability data. The cleaning reliability data is a comprehensive evaluation result that can comprehensively reflect the reliability and effectiveness of the cleaning technology in the current cleaning task. After obtaining the cleaning reliability data, it is necessary to compare it with a preset cleaning reliability threshold. This threshold represents the minimum reliable degree that the cleaning technology needs to achieve in this cleaning task. By comparing the cleaning reliability data with the threshold, it can be intuitively judged whether the cleaning technology meets the requirements.

[0031] According to an embodiment of the present invention, the processing of the cleaned data subset by adopting a corresponding annotation technology to obtain a corresponding annotated data subset includes: Process the cleaned data subset by adopting a corresponding annotation technology to obtain a corresponding annotated data subset; The annotation technology includes one or more of crowdsourcing annotation, regular expression annotation, deep learning annotation, supervised learning annotation, semi-supervised learning annotation, and unsupervised learning annotation.

[0032] It should be noted that after the cleaning work of the cultural big data subsets is completed, in order to enable these data to better serve subsequent analysis, mining, and the training of artificial intelligence models, it is necessary to process the cleaned data subsets with corresponding annotation techniques to obtain annotated data subsets with clear semantic information. The selection and application of annotation techniques are crucial, as they directly affect the quality and efficiency of the annotated data. In actual operation, one or more of crowdsourcing annotation, regular expression annotation, deep learning annotation, supervised learning annotation, semi-supervised learning annotation, and unsupervised learning annotation can be flexibly selected according to the type of data.

[0033] According to the embodiments of the present invention, extracting the data annotation evaluation indicators corresponding to each of the annotated data subsets and evaluating the fitness of the corresponding annotation techniques includes: Extracting the data annotation evaluation indicators corresponding to each of the annotated data subsets; The data annotation evaluation indicators include annotation precision, recall rate, annotation accuracy rate, and annotation speed; Processing according to the annotation precision, recall rate, annotation accuracy rate, and annotation speed through a preset data annotation ability evaluation model to obtain annotation reliability data; Performing a threshold comparison between the annotation reliability data and a preset annotation reliability threshold; Evaluating the fitness of the corresponding annotation technique according to the comparison result.

[0034] It should be noted that after processing the cleaned data subset with appropriate annotation techniques and obtaining the annotated data subset, in order to ensure the quality of the annotation results and the effectiveness of the selected annotation techniques, a comprehensive and detailed evaluation of the annotation process is required; the starting point of this evaluation process is to extract the evaluation indicators corresponding to each annotated data subset, and these indicators can measure the effectiveness of the annotation work from multiple key dimensions; the data annotation evaluation indicators mainly cover annotation precision, recall rate, annotation accuracy rate, and annotation speed; among them, annotation precision refers to the proportion of data that is actually positive among all data annotated as positive; recall rate refers to the proportion of data that is actually positive and is correctly annotated as positive; annotation accuracy rate refers to the proportion of all correctly annotated data in the total annotated data; annotation speed is an important indicator to measure annotation efficiency, which refers to the amount of annotation work completed per unit time; after extracting these data annotation evaluation indicators, a preset data annotation ability evaluation model needs to be used to comprehensively process these indicators. This evaluation model is constructed based on a large amount of historical data and practical experience. It will consider the mutual relationship among annotation precision, recall rate, annotation accuracy rate, and annotation speed, and through a series of complex algorithms and mathematical models, convert these indicators into a single value, that is, annotation reliability data; after obtaining the annotation reliability data, the next step is to compare it with a preset annotation reliability threshold; the preset annotation reliability threshold is a standard value preset according to specific annotation tasks and business requirements; this threshold represents the minimum reliable degree that the annotation technique needs to reach in this annotation task. Its setting usually combines factors such as the project goals, the importance of the data, and the requirements of subsequent applications. By comparing the annotation reliability data with the threshold, it is possible to intuitively judge whether the annotation technique meets the requirements. If the annotation reliability data is higher than the preset annotation reliability threshold, it means that the annotation technique performs well in the current annotation task and has a high degree of fit, which means that the annotation technique can accurately and efficiently complete the annotation work and matches the characteristics of the data and the annotation requirements, and can continue to be used in subsequent annotation work; on the contrary, if the annotation reliability data is lower than the preset annotation reliability threshold, it indicates that there may be some problems with the annotation technique. It may be that the annotation algorithm is not accurate enough and cannot well adapt to the characteristics of the data; or it may be that the annotation process is not optimized enough, resulting in a slow annotation speed; it is necessary to adjust and optimize the annotation technique, or consider replacing it with other more suitable annotation techniques.

[0035] In a second aspect, the present invention also discloses a processing system for cultural big data collection, cleaning, and annotation, including a memory and a processor. The memory includes a processing method program for cultural big data collection, cleaning, and annotation. When the processing method program for cultural big data collection, cleaning, and annotation is executed by the processor, the following steps are implemented: Classify the cultural big data to be processed, and collect the corresponding data subsets of the classified data using corresponding collection techniques to obtain the corresponding target cultural big data subsets; Extract the data collection evaluation indicators corresponding to each of the target cultural big data subsets, and evaluate the adaptability of the corresponding collection techniques; Process the target cultural big data subsets using corresponding cleaning techniques to obtain the corresponding cleaned data subsets; Extract the data cleaning evaluation indicators corresponding to each of the cleaned data subsets, and evaluate the applicability of the corresponding cleaning techniques; Process the cleaned data subsets using corresponding annotation techniques to obtain the corresponding annotated data subsets; Extract the data annotation evaluation indicators corresponding to each of the annotated data subsets, and evaluate the fit of the corresponding annotation techniques.

[0036] It should be noted that in the intelligent processing of cultural big data, a scientific and systematic processing process is the key to ensuring data quality and processing efficiency. First, it is necessary to conduct a detailed type division of cultural big data according to its content form, covering text type, image type, audio type, and video type. For different types of data, corresponding acquisition technologies need to be used for accurate acquisition. For text data, intelligent web crawler technology can be adopted. By analyzing the web page structure and setting keyword rules, during the acquisition process, technologies such as regular expressions and XPath are used to accurately extract the required text content and process dynamic web page data to ensure the integrity of the acquisition. For image data, image grabbing tools and API interfaces can be used. Audio data acquisition can utilize audio recording software and streaming media data acquisition technology. When acquiring, attention should be paid to the conversion of audio formats and the integrity of the data. Video data acquisition can be achieved through video download tools and video stream parsing technology, and at the same time, parameters such as the resolution and frame rate of the video are adaptively adjusted. Through the above targeted acquisition technologies, the corresponding target cultural big data subsets can be obtained, and these subsets form the basis for subsequent data processing. After the acquisition is completed, it is crucial to extract the data acquisition evaluation indicators corresponding to each target cultural big data subset. Through the quantitative analysis of these indicators, the suitability of the corresponding acquisition technology is evaluated to determine whether the acquisition technology can meet the acquisition requirements of this type of cultural big data. If the suitability does not meet the standard, the acquisition technology needs to be optimized or replaced to ensure the quality of the acquired data; after the acquisition is completed, for each target cultural big data subset, corresponding cleaning technologies need to be adopted for processing to remove noise, errors, and redundant information in the data, and the corresponding cleaned data subsets are obtained. The cleaned data also needs to be strictly evaluated, and the data cleaning evaluation indicators corresponding to each cleaned data subset are extracted. Through the evaluation of these indicators, the applicability of the corresponding cleaning technology is judged. If the applicability is insufficient, the cleaning algorithm or process needs to be further optimized. After the cleaning is completed, it enters the data annotation link, and the corresponding annotation technology is adopted for processing the cleaned data subsets to endow the data with semantic information, and the corresponding annotated data subsets are obtained. Finally, the data annotation evaluation indicators corresponding to each annotated data subset are extracted to evaluate the fitness of the corresponding annotation technology. Through the comprehensive evaluation of these indicators, it is comprehensively judged whether the annotation technology meets the annotation requirements of this type of cultural big data, providing a reliable data basis for subsequent data application and analysis, so as to achieve the technical goal of intelligent processing of cultural big data for acquisition, cleaning, and annotation.

[0037] According to an embodiment of the present invention, the type division of the cultural big data to be processed is performed, and the corresponding acquisition technology is adopted for the data subsets of the type to obtain the corresponding target cultural big data subsets, including: The data type division of the cultural big data to be processed is performed according to the data content to obtain data subsets of different types; The types include text, image, audio, and video; For the data subsets of the above types, corresponding acquisition technologies are adopted for acquisition to obtain corresponding target cultural big data subsets; If the type is text, web crawler technology is used for data acquisition; If the type is image, image recognition technology is used for data acquisition; If the type is audio, speech recognition technology is used for data acquisition; If the type is video, video recognition technology is used for data acquisition.

[0038] It should be noted that in the processing flow of cultural big data, data acquisition is the starting and crucial link, and its accuracy and efficiency directly affect the quality of subsequent data cleaning, annotation, and in-depth analysis. Facing the massive and complex cultural big data to be processed, the primary task is to conduct a detailed data type classification based on the data content to obtain different data subsets. These types mainly cover four categories: text, image, audio, and video. Each data type has unique forms of expression and data characteristics, which requires the adoption of highly adaptable acquisition technologies to ensure that the acquired data can completely and accurately reflect cultural information. Specifically, if it is text data, web crawler technology is used for data acquisition; if it is image data, image recognition technology is used for data acquisition; if it is audio data, speech recognition technology is used for data acquisition; if it is video data, video recognition technology is used for data acquisition. Through the above classification acquisition methods for different data types and the use of their respective adaptable acquisition technologies, corresponding target cultural big data subsets can be obtained. These high-quality data sets lay a solid foundation for the subsequent cleaning, annotation, and in-depth mining and analysis of cultural big data, helping to achieve the efficient and intelligent processing of cultural big data.

[0039] According to the embodiments of the present invention, extracting the data acquisition evaluation indicators corresponding to each of the target cultural big data subsets and evaluating the adaptability of the corresponding acquisition technology includes: Extracting the data acquisition evaluation indicators corresponding to each of the target cultural big data subsets; The data acquisition evaluation indicators include error rate, data coverage rate, data missing value ratio, and acquisition speed; According to the error rate, data coverage rate, data missing value ratio, and acquisition speed, they are processed through a preset data acquisition ability evaluation model to obtain acquisition reliability data; Comparing the acquisition reliability data with a preset acquisition reliability threshold; Evaluating the adaptability of the corresponding acquisition technology according to the comparison result.

[0040] It should be noted that after completing the classification and collection of cultural big data and obtaining each target cultural big data subset, in order to ensure the effectiveness and applicability of the collection technology, a comprehensive and detailed evaluation of the collection process is required. The key starting step of this evaluation process is to extract the data collection evaluation indicators corresponding to each target cultural big data subset. These indicators can accurately reflect the quality and efficiency of data collection from multiple dimensions and provide a basis for subsequent judgment of whether the collection technology is suitable. The data collection evaluation indicators mainly cover several important aspects such as error rate, data coverage rate, proportion of data missing values, and collection speed. The error rate is a key indicator for measuring the deviation degree between the collected data and the real data. The calculation of the error rate is usually presented in the form of a ratio by comparing the difference between the collected data and the known real data. The lower the error rate, the higher the accuracy of the collected data; the data coverage rate is used to measure the coverage degree of the collected data within the entire target data range; the level of the data coverage rate directly reflects whether the collection technology can comprehensively obtain the target data. A higher data coverage rate means that the collection technology can cover all aspects of the target data as much as possible, providing richer and more comprehensive information for subsequent data analysis and mining; the proportion of data missing values refers to the proportion of the missing part in the collected data. A too high proportion of data missing values will seriously affect the quality and usability of the data, so it needs to be strictly monitored and evaluated; the collection speed is also an important evaluation indicator, which reflects the amount of data that the collection technology can collect per unit time. The speed of data collection is affected by various factors. After extracting these data collection evaluation indicators, the next step is to comprehensively process these indicators with the help of a preset data collection ability evaluation model. This evaluation model is constructed based on a large amount of historical data and practical experience. It can calculate the collection reliability data according to the mutual relationship between indicators such as error rate, data coverage rate, proportion of data missing values, and collection speed through complex algorithms and mathematical models. The collection reliability data is a comprehensive evaluation result, which can comprehensively reflect the reliability and effectiveness of the collection technology in the current collection task;After obtaining the acquisition reliability data, it is necessary to compare it with a preset acquisition reliability threshold. By comparing the acquisition reliability data with the threshold, it is possible to intuitively determine whether the acquisition technology meets the requirements. Finally, based on the comparison results, the adaptability of the corresponding acquisition technology is evaluated. If the acquisition reliability data is higher than the preset acquisition reliability threshold, it indicates that the acquisition technology performs well in the current acquisition task, has a high adaptability, and can effectively complete the data acquisition work. This technology can continue to be used for subsequent data acquisition. On the contrary, if the acquisition reliability data is lower than the preset acquisition reliability threshold, it indicates that there may be some problems with the acquisition technology, such as inaccurate acquisition algorithms or insufficient performance of the acquisition tools. In this case, it is necessary to adjust and optimize the acquisition technology, or consider replacing it with other more suitable acquisition technologies to ensure that the acquisition task of cultural big data can be completed with high quality and efficiency.;

[0041] According to an embodiment of the present invention, processing the target cultural big data subset with a corresponding cleaning technology to obtain a corresponding cleaned data subset includes: Processing the target cultural big data subset with a corresponding cleaning technology to obtain a corresponding cleaned data subset; The cleaning technology includes one or more of data deduplication, missing value processing, outlier processing, and noise data processing.

[0042] It should be noted that after obtaining the target cultural big data subset, due to the wide and complex data sources, problems such as duplication, missing values, outliers, and noise inevitably exist, seriously affecting the data quality and the accuracy of subsequent analysis and applications. Therefore, it is necessary to deeply process these target cultural big data subsets with corresponding cleaning technologies to obtain high-quality cleaned data subsets. The cleaning technology covers various means such as data deduplication, missing value processing, outlier processing, and noise data processing. They cooperate with each other to jointly improve the usability and reliability of the data. Moreover, for different data types, the cleaning technologies adopted will be different, and one or more of data deduplication, missing value processing, outlier processing, and noise data processing can be selected according to actual needs.

[0043] According to an embodiment of the present invention, extracting the data cleaning evaluation indicators corresponding to each of the cleaned data subsets and evaluating the applicability of the corresponding cleaning technology includes: Extracting the data cleaning evaluation indicators corresponding to each of the cleaned data subsets; The data cleaning evaluation indicators include the error rate reduction value, the data loss rate, and the cleaning running time; Processing the error rate reduction value, the data loss rate, and the cleaning running time through a preset data cleaning ability evaluation model to obtain cleaning reliability data; Compare the cleaning reliability data with a preset cleaning reliability threshold; Evaluate the applicability of the corresponding cleaning technology according to the comparison result.

[0044] It should be noted that after using the corresponding cleaning technology to process the target cultural big data subset and obtaining the cleaned data subset, in order to measure the quality and efficiency of the cleaning work and determine whether the adopted cleaning technology is suitable for the data characteristics, a comprehensive and scientific evaluation of the cleaning process is required. The first step in this evaluation process is to extract the evaluation indicators corresponding to each cleaned data subset. These indicators can reflect the effect and performance of data cleaning from different dimensions; the data cleaning evaluation indicators mainly include the error rate reduction value, the data loss rate, and the cleaning running time; among them, the calculation method of the error rate reduction value is the error rate of the data before cleaning minus the error rate of the data after cleaning, and the data loss rate is represented by the ratio of the amount of data deleted during the cleaning process to the total amount of data before cleaning. The cleaning running time refers to the time spent from the start of the cleaning operation to the completion of the entire cleaning process; after extracting these data cleaning evaluation indicators, the next step is to use a preset data cleaning ability evaluation model to comprehensively process these indicators. This evaluation model is constructed based on a large amount of historical data and practical experience. It can take into account the mutual relationship between the error rate reduction value, the data loss rate, and the cleaning running time. Through complex algorithms and mathematical models, these indicators are converted into a single value, that is, the cleaning reliability data. The cleaning reliability data is a comprehensive evaluation result that can comprehensively reflect the reliability and effectiveness of the cleaning technology in the current cleaning task; after obtaining the cleaning reliability data, it is necessary to compare it with the preset cleaning reliability threshold. This threshold represents the minimum reliable degree that the cleaning technology needs to achieve in this cleaning task; by comparing the cleaning reliability data with the threshold, it is possible to intuitively judge whether the cleaning technology meets the requirements.

[0045] According to an embodiment of the present invention, the processing of the cleaned data subset by adopting a corresponding annotation technology to obtain a corresponding annotated data subset includes: Process the cleaned data subset by adopting a corresponding annotation technology to obtain a corresponding annotated data subset; The annotation technology includes one or more of crowdsourcing annotation, regular expression annotation, deep learning annotation, supervised learning annotation, semi-supervised learning annotation, and unsupervised learning annotation.

[0046] It should be noted that after the cleaning work of the cultural big data subsets is completed, in order to enable these data to better serve subsequent analysis, mining, and the training of artificial intelligence models, it is necessary to process the cleaned data subsets with corresponding annotation techniques to obtain annotated data subsets with clear semantic information. The selection and application of annotation techniques are crucial, as they directly affect the quality and efficiency of the annotated data. In actual operation, one or more of crowdsourcing annotation, regular expression annotation, deep learning annotation, supervised learning annotation, semi-supervised learning annotation, and unsupervised learning annotation can be flexibly selected according to the type of data.

[0047] According to an embodiment of the present invention, extracting the data annotation evaluation indicators corresponding to each of the annotated data subsets and evaluating the fitness of the corresponding annotation technique includes: Extracting the data annotation evaluation indicators corresponding to each of the annotated data subsets; The data annotation evaluation indicators include annotation precision, recall rate, annotation accuracy rate, and annotation speed; Processing according to the annotation precision, recall rate, annotation accuracy rate, and annotation speed through a preset data annotation ability evaluation model to obtain annotation reliability data; Performing a threshold comparison between the annotation reliability data and a preset annotation reliability threshold; Evaluating the fitness of the corresponding annotation technique according to the comparison result.

[0048] It should be noted that after processing the cleaned data subset using appropriate annotation techniques and obtaining the annotated data subset, in order to ensure the quality of the annotation results and the effectiveness of the selected annotation techniques, a comprehensive and detailed evaluation of the annotation process is required; the starting point of this evaluation process is to extract the evaluation indicators corresponding to each annotated data subset, and these indicators can measure the effectiveness of the annotation work from multiple key dimensions; the data annotation evaluation indicators mainly cover annotation precision, recall rate, annotation accuracy rate, and annotation speed; among them, annotation precision refers to the proportion of data actually belonging to the positive class among all the data annotated as the positive class; the recall rate refers to the proportion of data actually belonging to the positive class that is correctly annotated as the positive class; the annotation accuracy rate refers to the proportion of all correctly annotated data in the total annotated data; the annotation speed is an important indicator to measure the annotation efficiency, which refers to the amount of annotation work completed per unit time; after extracting these data annotation evaluation indicators, a preset data annotation ability evaluation model needs to be used to comprehensively process these indicators. This evaluation model is constructed based on a large amount of historical data and practical experience. It will consider the mutual relationship among annotation precision, recall rate, annotation accuracy rate, and annotation speed, and through a series of complex algorithms and mathematical models, transform these indicators into a single value, that is, the annotation reliability data; after obtaining the annotation reliability data, the next step is to compare it with the preset annotation reliability threshold; the preset annotation reliability threshold is a standard value preset according to specific annotation tasks and business requirements; this threshold represents the minimum reliable degree that the annotation technique needs to reach in this annotation task, and its setting usually combines factors such as the project's goals, the importance of the data, and the requirements of subsequent applications. By comparing the annotation reliability data with the threshold, it is possible to intuitively judge whether the annotation technique meets the requirements. If the annotation reliability data is higher than the preset annotation reliability threshold, it indicates that the annotation technique performs well in the current annotation task and has a high degree of fit, which means that the annotation technique can accurately and efficiently complete the annotation work and matches the characteristics of the data and the annotation requirements, and can continue to be used in subsequent annotation work; on the contrary, if the annotation reliability data is lower than the preset annotation reliability threshold, it indicates that there may be some problems with the annotation technique. It may be that the annotation algorithm is not accurate enough and cannot well adapt to the characteristics of the data; or it may be that the annotation process is not optimized enough, resulting in a slow annotation speed; it is necessary to adjust and optimize the annotation technique, or consider replacing it with other more suitable annotation techniques.

[0049] The third aspect of the present invention provides a readable storage medium, in which a program for the processing method of cultural big data collection, cleaning, and annotation is stored. When the program for the processing method of cultural big data collection, cleaning, and annotation is executed by a processor, the steps of the processing method of cultural big data collection, cleaning, and annotation as described in any one of the above are implemented.

[0050] The processing method, system, and medium for collecting, cleaning, and annotating cultural big data disclosed by the present invention classify the cultural big data to be processed, adopt corresponding collection technologies for collection to obtain corresponding target cultural big data subsets, extract the data collection evaluation indicators corresponding to each target cultural big data subset, and evaluate the adaptability of the corresponding collection technologies. Then, corresponding cleaning technologies are adopted to process the target cultural big data subsets to obtain corresponding cleaned data subsets, the data cleaning evaluation indicators corresponding to each cleaned data subset are extracted, and the applicability of the corresponding cleaning technologies is evaluated. Next, corresponding annotation technologies are adopted to process the cleaned data subsets to obtain corresponding annotated data subsets, the data annotation evaluation indicators corresponding to each annotated data subset are extracted, and the fit degree of the corresponding annotation technologies is evaluated, thereby realizing the technology of intelligent processing of collecting, cleaning, and annotating cultural big data.

[0051] In several embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined, or can be integrated into another system, or some features can be ignored, or not executed. In addition, the coupling, direct coupling, or communication connection between the various components shown or discussed with each other can be through some interfaces. The indirect coupling or communication connection of devices or units can be electrical, mechanical, or other forms.

[0052] The units described as separate components above may or may not be physically separated. The components shown as units may or may not be physical units; they can be located in one place or distributed to multiple network units; some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0053] In addition, in each embodiment of the present invention, the various functional units can all be integrated in one processing unit, or each unit can be separately regarded as one unit, or two or more units can be integrated in one unit; the above-mentioned integrated units can be implemented in the form of hardware, or in the form of a combination of hardware and software functional units.

[0054] Those of ordinary skill in the art can understand that all or part of the steps of implementing the above method embodiments can be completed by hardware related to program instructions. The foregoing program can be stored in a readable storage medium. When the program is executed, it performs the steps including those of the above method embodiments. The foregoing storage medium includes: various media such as a removable storage device, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disc that can store program code.

[0055] Alternatively, if the above integrated unit of the present invention is implemented in the form of a software functional module and sold or used as an independent product, it can also be stored in a readable storage medium. Based on such an understanding, the technical solution of the embodiments of the present invention, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. The software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the methods described in the various embodiments of the present invention. The foregoing storage medium includes: various media such as a removable storage device, a ROM, a RAM, a magnetic disk, or an optical disc that can store program code.

Claims

1. A processing method for cultural big data collection, cleaning, and annotation, characterized in that, It includes the following steps: Classify the cultural big data to be processed, and collect the corresponding data subsets of the said type using the corresponding collection techniques to obtain the corresponding target cultural big data subsets; Extract the data collection evaluation indicators corresponding to each of the said target cultural big data subsets, and evaluate the adaptability of the corresponding collection techniques; Process the said target cultural big data subsets using the corresponding cleaning techniques to obtain the corresponding cleaned data subsets; Extract the data cleaning evaluation indicators corresponding to each of the said cleaned data subsets, and evaluate the applicability of the corresponding cleaning techniques; Process the said cleaned data subsets using the corresponding annotation techniques to obtain the corresponding annotated data subsets; Extract the data annotation evaluation indicators corresponding to each of the said annotated data subsets, and evaluate the fitness of the corresponding annotation techniques.

2. The processing method for cultural big data collection, cleaning and annotation according to claim 1, wherein, The classifying the cultural big data to be processed, and collecting the corresponding data subsets of the said type using the corresponding collection techniques to obtain the corresponding target cultural big data subsets includes: Classify the cultural big data to be processed according to the data content to obtain data subsets of different types; The said types include text, image, audio, and video; Collect the corresponding data subsets of the said type using the corresponding collection techniques to obtain the corresponding target cultural big data subsets; If the said type is text, use web crawler technology for data collection; If the said type is image, use image recognition technology for data collection; If the said type is audio, use speech recognition technology for data collection; If the said type is video, use video recognition technology for data collection.

3. The processing method for cultural big data collection, cleaning and annotation according to claim 2, wherein, The extracting the data collection evaluation indicators corresponding to each of the said target cultural big data subsets, and evaluating the adaptability of the corresponding collection techniques includes: Extract the data collection evaluation indicators corresponding to each of the said target cultural big data subsets; The said data collection evaluation indicators include error rate, data coverage rate, proportion of data missing values, and collection speed; Process according to the error rate, data coverage rate, proportion of data missing values, and collection speed through a preset data collection ability evaluation model to obtain collection reliability data; Compare the collection reliability data with a preset collection reliability threshold; Evaluate the adaptability of the corresponding collection techniques according to the comparison result.

4. The processing method for cultural big data collection, cleaning and annotation according to claim 3, characterized in that, The processing the said target cultural big data subsets using the corresponding cleaning techniques to obtain the corresponding cleaned data subsets includes: Process the said target cultural big data subsets using the corresponding cleaning techniques to obtain the corresponding cleaned data subsets; The said cleaning techniques include one or more of data deduplication, missing value processing, outlier processing, and noise data processing.

5. The processing method for cultural big data collection, cleaning and annotation according to claim 4, wherein The extracting the data cleaning evaluation indicators corresponding to each of the said cleaned data subsets, and evaluating the applicability of the corresponding cleaning techniques includes: Extract the data cleaning evaluation indicators corresponding to each of the said cleaned data subsets; The said data cleaning evaluation indicators include error rate reduction value, data loss rate, and cleaning operation time; Process through a preset data cleaning ability evaluation model according to the error rate reduction value, data loss rate, and cleaning running time to obtain cleaning reliability data; Compare the cleaning reliability data with a preset cleaning reliability threshold; Evaluate the applicability of the corresponding cleaning technology according to the comparison result.

6. The processing method for cultural big data collection, cleaning and annotation according to claim 5, characterized in that, Processing the cleaned data subset with a corresponding annotation technology to obtain a corresponding annotated data subset, including: Processing the cleaned data subset with a corresponding annotation technology to obtain a corresponding annotated data subset; The annotation technology includes one or more of crowdsourcing annotation, regular expression annotation, deep learning annotation, supervised learning annotation, semi-supervised learning annotation, and unsupervised learning annotation.

7. The processing method for cultural big data collection, cleaning and annotation according to claim 6, characterized in that, Extracting the data annotation evaluation indicators corresponding to each of the annotated data subsets and evaluating the suitability of the corresponding annotation technology, including: Extracting the data annotation evaluation indicators corresponding to each of the annotated data subsets; The data annotation evaluation indicators include annotation precision, recall rate, annotation accuracy rate, and annotation speed; Process through a preset data annotation ability evaluation model according to the annotation precision, recall rate, annotation accuracy rate, and annotation speed to obtain annotation reliability data; Compare the annotation reliability data with a preset annotation reliability threshold; Evaluate the suitability of the corresponding annotation technology according to the comparison result.

8. A processing system for the collection, cleaning, and annotation of cultural big data, characterized in that, The system includes: a memory and a processor. The memory includes a program for the processing method of cultural big data collection, cleaning, and annotation. When the program for the processing method of cultural big data collection, cleaning, and annotation is executed by the processor, the following steps are implemented: Classify the cultural big data to be processed, and collect the corresponding data subset of the type with a corresponding collection technology to obtain a corresponding target cultural big data subset; Extract the data collection evaluation indicators corresponding to each of the target cultural big data subsets and evaluate the suitability of the corresponding collection technology; Process the target cultural big data subset with a corresponding cleaning technology to obtain a corresponding cleaned data subset; Extract the data cleaning evaluation indicators corresponding to each of the cleaned data subsets and evaluate the applicability of the corresponding cleaning technology; Process the cleaned data subset with a corresponding annotation technology to obtain a corresponding annotated data subset; Extract the data annotation evaluation indicators corresponding to each of the annotated data subsets and evaluate the suitability of the corresponding annotation technology.

9. The processing system for cultural big data collection, cleaning and annotation according to claim 8, characterized in that, The classifying the cultural big data to be processed and collecting the corresponding data subset of the type with a corresponding collection technology to obtain a corresponding target cultural big data subset includes: Classify the cultural big data to be processed according to the data content to obtain data subsets of different types; The types include text, image, audio, and video; Collect the corresponding data subset of the type with a corresponding collection technology to obtain a corresponding target cultural big data subset; If the type is text, use web crawler technology for data collection; If the type is image, use image recognition technology for data collection; If the type is audio, speech recognition technology is used for data collection; If the type is video, video recognition technology is used for data collection.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes a processing method, system, and medium program for cultural big data collection, cleaning, and annotation. When the processing method, system, and medium program for cultural big data collection, cleaning, and annotation are executed by a processor, the steps of the processing method for cultural big data collection, cleaning, and annotation according to any one of claims 1 to 7 are implemented.

Citation Information

Patent Citations

  • Auxiliary teaching method and device and storage medium

    CN114926044A

  • Data annotation method and device, computer equipment and storage medium

    CN118916733A

  • Big data processing method and system and storage medium

    CN120068003A

  • Method and apparatus for determining code generation quality and efficiency evaluation values based on multiple indicators

    US20230141348A1

  • Data set construction method, mobile terminal and readable storage medium

    WO2019233297A1

Cited By

  • Data acquisition method and device, server and medium

    CN120564430A