Data auditing and labeling system and method based on marketing
By capturing and classifying marketing data in real time in the data audit and annotation system, optimizing the data processing queue, removing data redundancy, and performing data annotation, the problem of insufficient data deduplication and merging capabilities in the existing technology is solved, and more efficient data processing and more accurate analysis results are achieved.
Patent Information
- Application Number
- CN202510036323.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-09
- Publication Date
- 2025-05-23
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The limited capabilities of the prior art in data deduplication and data merging have affected the accuracy of data analysis and the effectiveness of machine learning models, and the problems of data redundancy and mislabeling are common.
The marketing-based data review and labeling system is adopted, and marketing data is captured in real time through the data capture module and classified and dynamically scored, giving priority to high-scoring events; the priority sorting module is used to adjust the data queue, monitor processing efficiency and optimize the processing logic; the data is identified and removed through hashing; finally, the data is annotated and optimized through the labeling optimization module.
It improves the pertinence and efficiency of data processing, reduces data redundancy, improves data quality and analysis accuracy, enhances the effectiveness of machine learning models, and provides stronger support for data-driven decision-making.
Smart Images

Figure CN120030440A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data review and annotation, and in particular to a data review and annotation system and method based on marketing. Background Art
[0002] The field of data audit and annotation technology involves the use of various methods and tools to process, verify and annotate data for use in machine learning, data analysis and other forms of automated processing. The core activities in this field include data collection, cleaning, classification and annotation, where annotation is the process of giving data specific labels or classifications, which is used to train machine learning models to identify and process similar data. The marketing-based data audit and annotation system is a system specially designed to process marketing-related data, aiming to optimize the accuracy and availability of data through the audit and annotation process. It is usually used to identify and classify data such as market dynamics, consumer behavior, and competition to support more accurate market analysis and decision making. Existing technologies have limited capabilities in deduplication and data merging, which affects the accuracy of data analysis and the effectiveness of machine learning models. For example, the problems of data redundancy and incorrect annotation are still prevalent, which not only consumes a lot of processing resources, but also may lead to wrong business decisions. Therefore, improvements are needed. Summary of the invention
[0003] The purpose of the present invention is to solve the shortcomings of the prior art and to propose a data review and annotation system and method based on marketing.
[0004] In order to achieve the above-mentioned purpose, the present invention adopts the following technical scheme: A data review and annotation system based on marketing includes: The data capture module captures new events from the marketing data stream in real time, classifies the new events, and generates event classification results; based on the event classification results, dynamically scores the events according to their types and urgency to obtain event priority scores; The priority sorting module adjusts the data queue based on the event priority score, pushes high-scoring events to the front end of the processing queue, and generates an adjusted data queue; based on the adjusted data queue, monitors the data processing efficiency, adjusts the queue processing logic according to real-time feedback, and obtains an optimized processing strategy; A data deduplication module hashes the data items in the adjusted data queue, identifies duplicate data, and generates a deduplication data set; based on the deduplication data set, merges or removes similar data to obtain an optimized labeled data set; The labeling optimization module determines the labels of the data items according to the optimized labeling data set and generates the labeling results.
[0005] Preferably, the steps of obtaining the event classification result are: Capture new events from marketing data streams in real time, identify market dynamics and new product release information, apply text parsing and keyword extraction to obtain preliminary classification data; Based on the preliminary classification data, use a decision tree or support vector machine to perform data segmentation and feature weight assignment, identify the attributes and types of various events, and generate detailed classification results; Based on the detailed classification results, the importance and urgency of the event are analyzed, and the event classification results are formed in combination with the market influence and customer urgency of the event.
[0006] Preferably, the steps for obtaining the event priority score are: Extracting the type and urgency information of each event from the event classification results, analyzing the preliminary score of each event type and urgency through event data attributes, and obtaining preliminary score data; Based on the preliminary scoring data, combined with the market impact and feedback of the event, the priority of the event is evaluated to obtain weighted scoring data; Based on the weighted scoring data, the event priority score is calculated using the following formula: in, For events The priority score of For events The weight of type and urgency, For events exist Type of score, For events The relative importance index of is the number of event types.
[0007] Preferably, the steps of acquiring the adjusted data queue are: Based on the event priority score, all events are sorted by score, and the priority of each event in the data queue is determined by comparing the event priority score, and a list of events sorted by priority is generated; Based on the priority-ordered event list, the data processing queue is reconfigured to dynamically push high-scoring events to the front end of the data queue to form an adjusted data queue.
[0008] Preferably, the steps of obtaining the optimized processing strategy are: From the adjusted data queue, tracking the processing time and queue waiting time of each event, calculating the processing efficiency and load, and obtaining preliminary efficiency monitoring data; Based on the preliminary efficiency monitoring data, the data processing efficiency is calculated using the following formula: in, Represents data processing efficiency, For the The processing time of an event, For the The waiting time of an event in the queue, is the total number of events; Based on the data processing efficiency, the processing logic of the data queue is re-evaluated and optimized to obtain an optimized processing strategy.
[0009] Preferably, the steps of obtaining the deduplicated data set are: Based on the adjusted data queue, extract each data item, apply the SHA-256 hash function to generate a unique hash identifier for each data item, and obtain a hash identifier list; Based on the hash identification list, perform a repeatability check, identify and mark repeated data items by comparing hash values, integrate all repeat marks, and obtain a marked repeat data list; According to the duplicated data list, all duplicated data items are removed from the original data queue to obtain a deduplicated data set.
[0010] Preferably, the steps of obtaining the optimized annotated data set are: Based on the deduplicated data set, the feature vector of each data item is extracted, and the similarity between the data items is calculated. The calculation formula is: in, Represents a data item and The similarity between is a parameter that adjusts the similarity sensitivity. Represents the Euclidean distance between two eigenvectors; Based on the similarity, data items with similarity higher than a threshold are merged, and repeated or similar data items are removed to obtain an optimized labeled data set.
[0011] Preferably, the steps of obtaining the annotation results are: Selecting data items from the optimized labeled data set, applying a predefined labeling rule to each data item, applying a label to each item, and obtaining a preliminary labeled data set; Based on the preliminary labeled data set, quality checks are performed, the labeling results are reviewed, and incorrect labels are adjusted.
[0012] The present invention provides a data review and annotation method based on marketing, comprising the following steps: Monitor marketing data streams in real time, capture new events, and classify them according to their content features to obtain event classification results; Based on the event classification results, various events are dynamically scored according to urgency and impact through the set scoring rules to generate event priority scores; high-scoring events in the event priority score are pushed to the front of the processing queue, and the queue is monitored in real time. The queue processing logic is adjusted according to the data processing rate and feedback information to generate an optimized processing strategy; Perform deduplication processing on the data queue in the optimization processing strategy, use hashing to identify and mark duplicate data items, and generate a deduplication data set; Process the data in the deduplication dataset, merge duplicate or similar data, remove redundant information, and obtain an optimized labeled dataset; Based on the optimized labeled data set, a label is assigned to each piece of data, and the data is labeled according to the data content and the previously set labeling rules to obtain the labeling results.
[0013] Compared with the prior art, the advantages and positive effects of the present invention are: In the present invention, by capturing marketing data in real time and classifying and prioritizing events, the pertinence and efficiency of data processing are improved. By dynamically scoring event types and urgency, real-time adjustment of data queues is allowed to ensure that key events are given priority, which not only speeds up response time but also enhances the adaptability of the processing process. By identifying and removing duplicate data through hashing technology, the quality of the data set is improved and redundancy is reduced, which is particularly critical for subsequent data analysis and machine learning model training. Overall, through this comprehensive data processing and optimization process, data processing efficiency is maximized and data quality is improved, providing stronger support for data-based decision-making. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] Figure 1 It is a system flow chart of the present invention. DETAILED DESCRIPTION
[0015] In order to make the purpose, technical solution and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.
[0016] See also Figure 1 The present invention provides a technical solution: a data review and annotation system based on marketing includes: The data capture module captures new events from the marketing data stream in real time, classifies the new events, and generates event classification results. Based on the event classification results, dynamic scoring is performed according to event type and urgency to obtain event priority scores. The priority sorting module adjusts the data queue based on the event priority score, pushes high-scoring events to the front end of the processing queue, and generates an adjusted data queue; based on the adjusted data queue, it monitors the data processing efficiency, adjusts the queue processing logic according to real-time feedback, and obtains an optimized processing strategy; The data deduplication module hashes the data items in the adjusted data queue, identifies duplicate data, and generates a deduplication data set; based on the deduplication data set, similar data is merged or removed to obtain an optimized labeled data set; The annotation optimization module determines the labels of data items based on the optimized annotation data set and generates annotation results.
[0017] The steps to obtain event classification results are: Capture new events from marketing data streams in real time, identify market dynamics and new product release information, apply text parsing and keyword extraction to obtain preliminary classification data; Based on the preliminary classified data, use decision trees or support vector machines to perform data segmentation and feature weight assignment, identify the attributes and types of various events, and generate detailed classification results; Based on the refined classification results, the importance and urgency of the event are analyzed, and the event classification results are formed by combining the market impact and customer urgency of the event.
[0018] Specifically, referring to the collected industry terminology list and market activity information, the text content read in real time from the marketing data stream is analyzed one by one. First, each data is compared with a pre-established text length range, for example, the standard range of 50 to 2000 characters. If it is less than 50 characters, it is marked as a record with insufficient information. If it is more than 2000 characters, it is split into multiple paragraphs. Then, based on the agreed keyword database, related terms such as "new products", "promotional activities", and "user feedback" are searched one by one, and an index is built based on character position and number of occurrences. Keywords with a number of occurrences significantly greater than the average frequency are counted. The threshold for comparing the average frequency is the mean value obtained through industry-related text training. For example, 10,000 sample texts containing marketing vocabulary are selected, the average number of occurrences of each keyword is calculated and this mean is set as the threshold. If the number of keyword occurrences exceeds the threshold, it is marked as a high-frequency keyword. Through this marking, the context fragment of the corresponding term can be called in subsequent steps and associated with market dynamics and new product release information, thereby laying the foundation for data segmentation processing, and finally obtaining preliminary classified data based on multiple sets of statistical information records.
[0019] According to the preliminary classification data, text samples marked as high-frequency keywords and text samples marked as low-frequency keywords are selected as training sets. When using support vector machine for model training, each record in the training set is converted into a vector expression, including the count value of the keyword and the contextual semantic encoding of several words before and after. At the same time, the kernel function is set to a linear kernel and the penalty coefficient is set to 1 to ensure the discrimination of different features. For the feature weight allocation stage, the mean and variance of all features in the training set are calculated first, and then the features with a mean greater than 50% of the average level and a significantly high variance are marked as key features. The weight of these components will be increased later. The weight of this part will be doubled in actual calculation. The value comes from the evaluation results of historical training effects and is stored in the existing experimental records. Through multiple training and validation set tests, the hyperplane that can be used to distinguish different event types is gradually converged. At this time, data segmentation is performed again. The comprehensive judgment of accuracy and recall rate is maintained above the pre-agreed standard (such as at least 85% accuracy in the training set). If it is lower than 85%, the weight distribution needs to be readjusted. Finally, the attributes and types of various events are identified and corresponding labels are established to generate detailed classification results.
[0020] Combined with the refined classification results, the importance and urgency of records under different event types are analyzed. First, the market influence index and customer urgency value of each event are read in the record. For example, the market influence range between 0 and 100 points and the customer urgency range between 0 and 10 are compared respectively. If the market influence exceeds 60, it is determined that it has a large driving factor for brand attention. The value of 60 is set based on the average value and variance calculation summarized by previous promotional activities. If the customer urgency is greater than 5, it indicates that there is a high density of customer response. The standard of 5 is also set with reference to the average value obtained from the historical after-sales response time statistics. By reading these data and calculating the weighted sum, the importance and urgency are recorded in separate numerical fields respectively, and then the two fields are merged to classify the event. For example, first check whether the importance exceeds more than half of the average benchmark value, and then check whether the urgency is in the range of more than 5. The above information is combined to obtain the final event classification result.
[0021] The steps to obtain the event priority score are: Extract the type and urgency information of each event from the event classification results, analyze the preliminary score of each event type and urgency through event data attributes, and obtain preliminary score data; Based on the preliminary scoring data, combined with the market impact and feedback of the event, the priority of the event is evaluated to obtain weighted scoring data; Based on the weighted scoring data, the event priority score is calculated using the following formula: in, For events The priority score of For events The weight of type and urgency, For events exist Type of score, For events The relative importance index of is the number of event types.
[0022] Specifically, the type and urgency value of each event is read from the event classification result, and the event data attributes corresponding to the type and urgency value are screened and recorded. For example, the key attributes of each event are listed in the record table in numerical form and a corresponding index is established. Then, each event is compared and classified according to the type identifier and the urgency value range. The type identifier is compared with the valid number range between 0 and 10, and the urgency value is compared with the urgency range between 0 and 5. If the urgency value is higher than 5, it is marked as an extreme emergency event. The value comes from the statistical analysis of the processing time of various events in the market activity history database. By recording The time difference from the occurrence of an event to the completion of feedback is calculated and the average value is calculated. Then, based on the average value, the urgency data exceeding the 60% percentile is defined as a range above 5. If the type identifier exceeds 10, it is marked as an invalid type and requires subsequent additional inspection. In this quantitative way, several records are screened out and compared. Then, the key attributes corresponding to the type identifier and the urgency value are called in the record. Events with a high frequency of occurrence and an urgency significantly higher than the average range are separately sorted and numerically tracked. The sorting results are reviewed once, the reviewed data are summarized, and the preliminary score value of each event is recorded to finally obtain the preliminary score data.
[0023] Based on the preliminary scoring data, the market impact information and feedback information associated with the event are recorded. The market impact information can be judged by referring to the sales volume increase ratio of the product in the range of 0% to 50%. If the sales volume increase ratio exceeds 20%, it is recorded as a high-impact event. The 20% threshold is obtained by counting the daily average sales volume in the past six months in the product sales database and sorting out the actual sales growth rate, and is determined after comparing the ratio of the basic growth rate to the average growth rate of six months. The feedback information is compared with the standard range of the number of feedbacks between 0 and 200. If the number of feedbacks exceeds 120, it is marked as a frequent feedback event. Then, the score of each event is numerically accumulated with the market impact information and feedback information. For example, the event score value is added to the sales volume increase ratio and the number of feedbacks, and the three items are added according to the established ratio in the record. The adjustment coefficient is determined by comparing the statistical data of previous historical events for many times. Different adjustment coefficients correspond to different impact levels. Here, the reference value can be approximately between 1.1 and 1.5. The final coefficient is obtained by actually comparing the sales volume increase and the feedback density. The result is recorded in the weighted scoring field to finally obtain the weighted scoring data.
[0024] The formula is useful in that it takes into account the type and urgency of events when measuring event priorities through the weighted accumulation of the numerator and the square root of the sum of squares of the denominator, thereby numerically comparing different types of events on a unified scale. The parameter acquisition step is to calculate the type frequency between 0 and 1 based on the type identification and urgency value obtained above, combined with the statistical weight of each type in the actual record, and calculate the cumulative proportion of the number of occurrences of more than 100 event types to obtain the type weight sequence and record it in among; The parameter acquisition step is to extract the comprehensive score of each event under different types through the weighted score data obtained above, and limit the score value to between 0 and 100. The higher the value, the more prominent the event can reflect the characteristics of this type. The steps to obtain the parameters are to form a relative importance index for the market impact and feedback data. The specific method is to calculate the mean increase and feedback mean of this type of event from the previously recorded sales increase ratio and feedback quantity, and then divide the mean by the total mean of all events in the same statistical interval to obtain a relative importance index with a numerical range of 0.5 to 2.0, which is recorded in among; Represents the number of event types. According to previous statistics, its value is 5, which is obtained from the statistics of all type mapping lists participating in the calculation.
[0025] Calculation process: Let ,Pick These weight values are derived from the cumulative proportion of the number of occurrences of event types, and are obtained by combining the increasing proportion of the urgency attribute in the data of different industries; for a certain event , its ratings under five types They are 40, 55, 70, 30, and 50 respectively. These scores come from the weighted score data obtained above. The relative importance index of each type , these values come from the relative ratio statistics of factors such as the mean sales growth rate and the mean number of feedbacks; substitute the above values into the numerator: , the calculation process is as follows: , , , , ; Adding the above five terms gives the sum of the numerator: ; Denominator: ; The items are: Square them separately and add them together: = Take the square root to get the denominator: ; From this we get The calculation result is: ; The results show that the event The priority score value is 2.50. If the score exceeds 2.00, it means that the event is at a higher priority. Combined with the type and urgency characteristics obtained above, further follow-up operations can be performed.
[0026] The steps for obtaining the adjusted data queue are: Based on the event priority score, all events are ranked according to their scores. By comparing the event priority score of each event, the priority in the data queue is determined, and a list of events sorted by priority is generated. Based on the priority-ordered event list, the data processing queue is reconfigured to dynamically push high-scoring events to the front end of the data queue to form an adjusted data queue.
[0027] Specifically, based on the event priority scores obtained above, the priority values of all events are compared, and these values are read one by one and arranged in a table. The priority score of each event is compared with the reference threshold by scanning one by one. The threshold can be selected between 2.0 and 3.0. The basis is to count the cumulative 300 events in the past year and calculate the median and corresponding quartile of their priority scores. The threshold is set by comprehensively considering the distribution characteristics. If the priority score is higher than this threshold, it is marked as a high priority event, and it is isolated from the event partitions with medium or low priorities and recorded in the sorted lists of different areas. In the sorted list, the score differences are further compared for the high priority events. If the priority score of an event is significantly higher than that of other high priority events, it can be regarded as a special key event. With reference to the highest segmentation threshold in the previous statistical process, for example, 3.5 or above, a secondary screening is performed, and the differentiated scores of the special event and the ordinary high priority event are recorded, and its position in the sorted list is reconfirmed. In this way, the scores are arranged in order. After completing the above records, the sorted lists are integrated and numbered in descending order, and finally a list of events sorted by priority is generated.
[0028] Based on the event list sorted by priority, when reconfiguring the data processing queue, the event identifier and corresponding priority value in the list are first read, and the events with scores higher than the previously set threshold of 2.0 are automatically positioned to the front of the queue. In the specific implementation, the insertion order is controlled according to the absolute difference in their values. For example, if an event is scored 3.2 and another event is scored 2.1, 3.2 is preferentially inserted to the front position. The insertion order is determined by comparing the score range between high-priority events and ordinary events. For events with scores lower than 2.0, the original order is temporarily maintained and they are recorded as events that can be processed later and can be delayed. Then, a test will be performed during the queue rearrangement process. If a recent change in the score is found or the previously accumulated data is revised, the scores of these events are re-compared to see if they are still above 2.0. If there is a significant change, the position of the queue is updated again. Finally, the high-scoring events are concentrated at the front of the queue and all the rankings are recorded to form an adjusted data queue.
[0029] The steps to obtain the optimized processing strategy are: From the adjusted data queue, track the processing time and queue waiting time of each event, calculate the processing efficiency and load, and obtain preliminary efficiency monitoring data; Based on the preliminary efficiency monitoring data, the data processing efficiency is calculated using the following formula: in, Represents data processing efficiency, For the The processing time of an event, For the The waiting time of an event in the queue, is the total number of events; Based on data processing efficiency, re-evaluate and optimize the processing logic of the data queue to obtain an optimized processing strategy.
[0030] Specifically, from the adjusted data queue obtained earlier, record and read the processing time and queue waiting time of each event one by one, then compare the processing time in the range of 0 seconds to 180 seconds, record the processing time of each event in the interval, and record the queue waiting time in the range of 0 seconds to 300 seconds. If the waiting time of an event exceeds 300 seconds, it is marked as an abnormal delay event, and compared with the waiting peak distribution in the historical data in the subsequent steps to determine whether additional investigation and comparison are needed. At the same time, in this process, refer to the event category to perform a detailed inspection on the processing time and queue waiting time. For example, if the event type is a serious alarm, compare it with the urgent interval of 0 seconds to 60 seconds. If the processing is completed within this interval, it is marked as an event that can be responded to quickly. In the processing process, all records are arranged in order according to the timestamp, and then each time item is compared with the average generated in the previous period. The average benchmark value is compared with the average benchmark value, which is accumulated from the operation and maintenance data of several months. For example, the average processing time for routine events is about 45 seconds, and the average waiting time is about 100 seconds. The event is compared with the two average values. If the processing time or waiting time of an event deviates from the average value by more than 30 seconds, it is re-verified. During the verification process, the actual operation steps are compared with the previously established operation specifications to confirm whether there are process blockages or repeated links. If it cannot pass the automatic rule screening, the details are tracked by re-reviewing the relevant records. After completing this comparison, the key values of all events are merged and summarized to form a separate integrated table for subsequent efficiency and load calculation processes. The total number of events, the processing duration of each event, and the total queue time are retrieved from the integrated table to finally obtain preliminary efficiency monitoring data.
[0031] The formula is useful in that it combines the logarithmic function relationship between processing time and queue waiting time, and incorporates processing time and queue delay into the calculation, so that the impact of waiting time on efficiency can be highlighted in a large-scale queue environment; The steps to obtain the parameters are to measure the efficiency performance in all event records as a whole, sum up the individual ratios calculated for each event and divide by the total number of events. The required information comes from the previously summarized processing duration and total queue time. The parameter acquisition step is to record the actual duration from the start of execution to the completion of execution item by item in the processing of each event, and obtain the value by subtracting the timestamp. The value unit is seconds and fluctuates in the range of 0 seconds to 300 seconds. The parameter acquisition step is to record the queuing time of each event in the queue one by one, and compare the timestamp of the event starting to queue with the timestamp of the formal entry into the processing to obtain the value, which is also in seconds. Under some high-load conditions, the queuing time may reach more than 180 seconds. This value can be obtained by comparing with the actual data; The step of obtaining the parameter is to count the total number of all events recorded in the adjusted data queue, which is usually between 100 and 1000, and determine its value according to the number of events collected in the specific scenario.
[0032] Calculation process: In a certain operation, get the total number of events , each event and Recorded as: Event 1: Second, Second; Event 2: Second, Second; Event 3: Second, Second; Event 4: Second, Second; Event 5: Second, Second; Substitute the above parameters into the formula: Calculate for each event separately: Event 1: ; Event 2: ; Event 3: ; Event 4: ; Event 5: ; Adding these five values gives: ; Then divide by 5: ; The result shows that the data processing efficiency in this scenario is approximately 9.99. If the value is large, it means that each event has a certain efficiency advantage in the ratio of queuing time to processing time. If the value is too low, it means that there is a high consumption of waiting and processing time. Further queue optimization strategies can be carried out based on this value in the future.
[0033] Based on the obtained data processing efficiency value, when re-evaluating and optimizing the processing logic of the data queue, you can first read the waiting time distribution of different event types, locate the events with a waiting time of more than 120 seconds to the abnormal interval, and compare the actual processing time of the event to see if it also exceeds the recorded benchmark. In this process, if it is found that the waiting time deviates significantly from the average level, it needs to be recorded separately and compared with the historical average data. The historical average data is usually in the range of 100 seconds to 150 seconds. This range is obtained by counting the data of the past quarter or longer period and calculating the quantiles. After confirming the deviation, these events can be prioritized or moved forward in order in the queue, and the recently processed events can be moved forward in order. The actual measured duration of the event is put into the cumulative database for further comparison. The benchmark value of the average queuing and processing time is incrementally updated through the comparison results of each operation. During configuration optimization, the processing order of multiple types of events is adjusted item by item, and the events included in the high-load or abnormal interval are marked and their detailed processing process is recorded. The final processing completion status of the marked event is reviewed in a specific time period. If similar high waiting and high processing time overlap in subsequent queue monitoring, the queue scheduling method is reviewed again, and the real situation is obtained through continuous comparison of timestamps and continuous monitoring records. Finally, the entire scheduling plan is summarized and updated to the queue optimization strategy for the next stage.
[0034] The steps to obtain the deduplicated data set are: Based on the adjusted data queue, extract each data item, apply the SHA-256 hash function to generate a unique hash identifier for each data item, and obtain a hash identifier list; Based on the hash identifier list, perform a duplication check, identify and mark duplicate data items by comparing hash values, integrate all duplicate marks, and obtain a list of marked duplicate data; According to the list of marked duplicate data, all duplicate data items are removed from the original data queue to obtain a deduplicated data set.
[0035] Specifically, based on the adjusted data queue obtained above, each data item is read one by one, and its identification information is compared with the corresponding data content after reading, and then the data content is processed using the SHA-256 algorithm. The algorithm performs multiple rounds of iterative compression on the string or byte sequence and obtains a hash value of fixed length. Each iteration refers to the initial vector value and constant table specified in the SHA-256 specification for calculation, divides the input data into 512-bit blocks and calculates them segment by segment, merges the intermediate hash values generated in the middle, performs logical operations and bit shift operations on each segmented data according to the algorithm flow, and finally obtains a 256-bit hash value, which can be expressed as a hexadecimal string after generation. If a data item is too long, it is split into multiple segments and processed in sequence. The hash value of each segment output is then spliced to form the final hash identifier. Especially when the data item contains multiple groups of fields, the fields need to be combined in a fixed order first to avoid hash value deviations caused by different fields in different orders. During the combination process, the text field may need to be converted into bytes before being processed by the algorithm. At the same time, the numeric field is first converted into a string and then processed in blocks according to the requirements of the SHA-256 algorithm. If the data item contains a specific binary file, the block conversion method is continued to ensure that the hash value can be calculated for all the content. After the hash calculation is completed for all items, each hash value is recorded as a hash identifier list, and finally the hash identifier list is obtained.
[0036] Based on the hash identifier list, when performing a repeatability check, each hash value is compared one by one. If the hash value of a data item is exactly the same as the hash value of a registered record, it is marked as suspected duplicate data. In the process, each hash value is sorted in dictionary order to facilitate rapid search for adjacent similar items. The presence of duplication is further determined by binary search. For duplicate items, they are marked in the duplicate data marker column. If a hash value appears more than three times, it is classified as multiple duplications, and its corresponding data items are assigned to the multiple item partition of the duplicate list when recording. At this time, some data segmented into binary files are additionally compared to confirm whether the difference is caused by different field orders or different contents. If it is determined that the hash values are consistent, it is considered a complete duplication and the duplication mark is recorded. For the situation where the hash value has appeared multiple times in the database, its duplication level needs to be confirmed according to three preset ranges of 0 to 3 times, 4 to 6 times, and 7 times and above. If it reaches 7 times or above, an identifier that requires further manual verification is added to the tag, and then all duplicate entries are integrated and a list of marked duplicate data is generated, and finally a list of marked duplicate data is obtained.
[0037] According to the obtained list of marked duplicate data, the duplicate items are removed one by one from the original data queue according to the listed duplicate item numbers. In this process, the multiple data items whose hash values are determined to be completely duplicates are finally checked. If it is confirmed that the hash values are the same, the duplicate items are directly deleted and only one piece of data is retained. The remaining items with the same hash value are excluded from the data queue. If the hash values are partially consistent but some field information is found to have slight differences during the subsequent segment verification, the record is temporarily marked as a possible duplicate and the user is allowed to do further review later. For ordinary duplicate items, only regular deletion is required and the deletion is recorded. In addition to the action in the log, the remaining entries after deletion are reordered according to the original order of the data queue, and the final result of this process is confirmed by the reduction of the total number of entries. If the number of reduced entries is greater than the preset high duplication threshold, such as 25% of the data, an additional re-examination is performed to read the source of the high duplication threshold. The source here refers to the range of repetition rates between 15% and 25% under normal circumstances obtained based on historical data statistics. If the deletion value exceeds 25%, a deep check can be performed later. After completing the above investigation, the deduplication data set is obtained.
[0038] The steps to obtain the optimized labeled dataset are: Based on the deduplicated data set, the feature vector of each data item is extracted and the similarity between the data items is calculated using the following formula: in, Represents a data item and The similarity between is a parameter that adjusts the similarity sensitivity. Represents the Euclidean distance between two eigenvectors; Based on the similarity, data items with similarity higher than the threshold are merged, and duplicate or similar data items are removed to obtain an optimized labeled data set.
[0039] Specifically, the formula is useful in that it maps the Euclidean distance value to interval, and use Adjust the sensitivity so that a higher similarity value is obtained when the difference in feature vectors is small, thereby improving the distinction between highly similar data; The steps for obtaining the parameters are as follows: first, select several pairs of representative data items in the deduplication data set, calculate the Euclidean distance distribution between these data items, and form a set of intervals in to , and then perform quantile statistics on the distribution and observe that approximately The distance value at the quantile is recorded multiple times and compared with the historical data. Finally, the value near this quantile is used as a reference, and combined with the experience of professionals in similarity judgment, the distance value at the quantile is recorded multiple times and compared with the historical data. Set to ; The steps to obtain the parameters are as follows: Article and Each data item is converted into a numerical feature vector. For example, the key attributes of each data item are taken in a fixed order. values, and the Euclidean distance formula Calculations are performed with values ranging from to The specific value is generated by the numerical processing results of each attribute; The steps to obtain the parameters are as follows: After the Euclidean distance is entered into the above formula and calculated one by one, a similarity value will be generated between each piece of data and Floating within a range; Calculation process: Select , Order Vector of data , No. Vector of data , then the Euclidean distance ; The distance Substituting into the formula: in , so: The results show that the similarity between the two data is approximately , when the similarity is greater than It indicates that the data are highly similar in feature vectors and can be subsequently considered as duplicate or similar data for merging or removal.
[0040] Based on the similarity value record, when comparing the similarities between all data pairs, read the previously calculated similarities one by one and record them in the similarity distribution list, and compare each similarity with the preset threshold. The threshold is usually in the range of 0.85 to 0.90. These values are selected from the statistical results of similar distribution of previous data. For example, after collecting 1000 deduplicated data, the distance between them is calculated and mapped to similarity, and then the average and median of high similarity are counted and the threshold is determined based on subjective experience. For example, when the frequency of a certain type of data is most dense around 0.85, 0.85 is selected as the starting threshold. When multiple actual measurements show that some data are more concentrated, the threshold can be fine-tuned to 0.90. Whenever the similarity exceeds 0.85, the threshold is adjusted to 0.90. Entries that exceed the threshold are merged in the list, and their related attributes are compared to confirm that they are indeed highly similar or duplicate records. Then, entries with particularly significant duplication are managed in a hierarchical manner. For example, if the similarity reaches 0.98 or above, it will be marked as a complete duplicate. During the recording process, pay attention to maintaining the original index and give priority to retaining the merged record while deleting redundant items. Entries whose similarity is close to the threshold but has not yet reached it are retained in subsequent procedures for further investigation. If the actual test results after repeated verification show that the similarity of these entries is still in a higher distribution range, the next step of the secondary merging process is entered. After all comparisons and screening, the duplicate or highly similar data will be removed one by one to obtain an optimized labeled data set.
[0041] The steps to obtain the annotation results are: Select data items from the optimized labeled data set, apply a predefined labeling rule to each data item, apply a label to each item, and obtain a preliminary labeled data set; Based on the preliminary labeled data set, perform quality checks, review the annotation results, and adjust incorrect labels.
[0042] Specifically, based on the optimized annotated data set, each data item is read and compared with the predefined labeling rules one by one. These labeling rules can select corresponding identification content from the market category list and customer preference list compiled in the early stage, and then match these identification contents according to the number and name. In order to ensure that the labeling rules can be correctly called on different data items, it is necessary to first extract the corresponding mapping relationship from the labeling rule library. For example, ten categories numbered from 1 to 10 can be defined for the category to which the product belongs, and ten intention levels numbered from 11 to 20 can be set for user behavior intentions. Each time a piece of data is read, its key fields are compared with these numbers one by one. If the data item contains "newly listed products", " and meets the matching rule number 1, label 1 is applied to the corresponding entry. If the data item also contains "marketing promotion" and matches the promotion intention number 15, label 15 can be accumulated. Multiple labels can be matched for each data item. During the matching process, if some fields differ too much from the value range set in the label library, check again whether the value range of the field between 0 and 100 or 0 and 1000 meets the control standard defined by the label. If the value does deviate from the label definition, it is marked as an unlabeled adaptation item. At this time, continue to search for other matching numbers. After all the comparisons are completed, several information records with multiple labels will be generated, and these information records will be summarized to form a preliminary annotation data set.
[0043] Based on the preliminary annotated data set, when performing quality checks, we first review the correspondence between each entry and the label one by one, and determine whether there is a label out-of-bounds or duplicate naming by comparing the label number in the set range of 1 to 10 or 11 to 20. Then, we compare the actual field value with the valid range defined by the label, for example, comparing the numeric field with the regular range between 0 and 100, and checking whether the text keyword is included in the label keyword list. If it is found that the label number of a record is obviously inconsistent with the data content, it will be marked as a suspected wrong label, and then the suspected wrong labels will be centrally checked and the text keywords will be checked. The segment value is re-matched with the label rule library. If it is confirmed that a field is incorrectly referenced, the field content is retrieved for analysis and checked in sequence in combination with the category or intent number obtained previously. In some scenarios, it is necessary to check whether there are spelling or coding anomalies in the data. For example, if the spelling of a word is too different from the predefined keyword, it will not match or match incorrectly. After completing the above verification and troubleshooting, adjustments are made to the labels that are confirmed to be incorrect, replacing them with correct numbers or deleting redundant numbers. The modified entries are then recorded again and all the audited data are summarized to finally obtain the labeling results after adjusting the incorrect labels.
[0044] The present invention provides a data review and annotation method based on marketing, comprising the following steps: Monitor marketing data streams in real time, capture new events, and classify them according to their content features to obtain event classification results; Based on the event classification results, various events are dynamically scored according to urgency and impact through the set scoring rules to generate event priority scores; high-scoring events in the event priority score are pushed to the front of the processing queue, and the queue is monitored in real time. The queue processing logic is adjusted according to the data processing rate and feedback information to generate an optimized processing strategy; Perform deduplication processing on the data queue in the optimization processing strategy, use hashing to identify and mark duplicate data items, and generate a deduplication data set; Process the data in the deduplication dataset, merge duplicate or similar data, remove redundant information, and obtain an optimized labeled dataset; Based on the optimized labeled data set, a label is assigned to each piece of data, and the data is labeled according to the data content and the previously set labeling rules to obtain the labeling results.
Claims
1. A data review and annotation system based on marketing, characterized in that: The system comprises: The data capture module captures new events from the marketing data stream in real time, classifies the new events, and generates event classification results; based on the event classification results, dynamically scores the events according to their types and urgency to obtain event priority scores; The priority sorting module adjusts the data queue based on the event priority score, pushes high-scoring events to the front end of the processing queue, and generates an adjusted data queue; based on the adjusted data queue, monitors the data processing efficiency, adjusts the queue processing logic according to real-time feedback, and obtains an optimized processing strategy; A data deduplication module hashes the data items in the adjusted data queue, identifies duplicate data, and generates a deduplication data set; based on the deduplication data set, merges or removes similar data to obtain an optimized labeled data set; The labeling optimization module determines the labels of the data items according to the optimized labeling data set and generates the labeling results.
2. The data review and annotation system based on marketing according to claim 1, characterized in that: The steps for obtaining the event classification result are: Capture new events from marketing data streams in real time, identify market dynamics and new product release information, apply text parsing and keyword extraction to obtain preliminary classification data; Based on the preliminary classification data, use a decision tree or support vector machine to perform data segmentation and feature weight assignment, identify the attributes and types of various events, and generate detailed classification results; Based on the detailed classification results, the importance and urgency of the event are analyzed, and the event classification results are formed in combination with the market influence and customer urgency of the event.
3. The data review and annotation system based on marketing according to claim 1, characterized in that: The steps for obtaining the event priority score are as follows: Extracting the type and urgency information of each event from the event classification results, analyzing the preliminary score of each event type and urgency through event data attributes, and obtaining preliminary score data; Based on the preliminary scoring data, combined with the market impact and feedback of the event, the priority of the event is evaluated to obtain weighted scoring data; Based on the weighted scoring data, the event priority score is calculated using the following formula: in, For events The priority score of For events The weight of type and urgency, For events exist Type of score, For events The relative importance index of is the number of event types.
4. The data review and annotation system based on marketing according to claim 1, characterized in that: The steps of obtaining the adjusted data queue are: Based on the event priority score, all events are sorted by score, and the priority of each event in the data queue is determined by comparing the event priority score, and a list of events sorted by priority is generated; Based on the priority-ordered event list, the data processing queue is reconfigured to dynamically push high-scoring events to the front end of the data queue to form an adjusted data queue.
5. The data review and annotation system based on marketing according to claim 1, characterized in that: The steps for obtaining the optimized processing strategy are: From the adjusted data queue, tracking the processing time and queue waiting time of each event, calculating the processing efficiency and load, and obtaining preliminary efficiency monitoring data; Based on the preliminary efficiency monitoring data, the data processing efficiency is calculated using the following formula: in, Represents data processing efficiency, For the The processing time of an event, For the The waiting time of an event in the queue, is the total number of events; Based on the data processing efficiency, the processing logic of the data queue is re-evaluated and optimized to obtain an optimized processing strategy.
6. The data review and annotation system based on marketing according to claim 1, characterized in that: The steps for obtaining the deduplication data set are: Based on the adjusted data queue, extract each data item, apply the SHA-256 hash function to generate a unique hash identifier for each data item, and obtain a hash identifier list; Based on the hash identification list, perform a repeatability check, identify and mark repeated data items by comparing hash values, integrate all repeat marks, and obtain a marked repeat data list; According to the duplicated data list, all duplicated data items are removed from the original data queue to obtain a deduplicated data set.
7. The data review and annotation system based on marketing according to claim 1, characterized in that: The steps for obtaining the optimized labeled data set are: Based on the deduplicated data set, the feature vector of each data item is extracted, and the similarity between the data items is calculated. The calculation formula is: in, Represents a data item and The similarity between is a parameter that adjusts the similarity sensitivity. Represents the Euclidean distance between two eigenvectors; Based on the similarity, data items with similarity higher than a threshold are merged, and repeated or similar data items are removed to obtain an optimized labeled data set.
8. The data review and annotation system based on marketing according to claim 1, characterized in that: The steps for obtaining the annotation results are: Selecting data items from the optimized labeled data set, applying a predefined labeling rule to each data item, applying a label to each item, and obtaining a preliminary labeled data set; Based on the preliminary labeled data set, quality checks are performed, the labeling results are reviewed, and incorrect labels are adjusted.
9. A data review and annotation method based on marketing, characterized in that: The data review and annotation system based on marketing according to any one of claims 1 to 8 is implemented, comprising the following steps: Monitor marketing data streams in real time, capture new events, and classify them according to their content features to obtain event classification results; Based on the event classification results, various events are dynamically scored according to urgency and impact through the set scoring rules to generate event priority scores; high-scoring events in the event priority score are pushed to the front of the processing queue, and the queue is monitored in real time. The queue processing logic is adjusted according to the data processing rate and feedback information to generate an optimized processing strategy; Perform deduplication processing on the data queue in the optimization processing strategy, use hashing to identify and mark duplicate data items, and generate a deduplication data set; Process the data in the deduplication dataset, merge duplicate or similar data, remove redundant information, and obtain an optimized labeled dataset; Based on the optimized labeled data set, a label is assigned to each piece of data, and the data is labeled according to the data content and the previously set labeling rules to obtain the labeling results.
Citation Information
Cited By
Marketing promotion online platform-oriented data entry method
CN120580005A
Online sports event management system based on reliable transmission of event information
CN120768955A
Online sports event activity management system based on reliable transmission of event information
CN120768955B