Massive data automatic label generation method based on large model and rule engine
By combining large models with rule engines, the problems of low efficiency, inconsistent quality, and slow update iteration in traditional tagging systems are solved. This enables efficient, accurate, and scalable tag generation and updates, supports cross-domain consistency, and improves the accuracy and timeliness of data analysis.
Patent Information
- Application Number
- CN202511906487.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-17
- Publication Date
- 2026-03-17
- Estimated Expiration
- 2045-12-17
AI Technical Summary
Traditional manual annotation methods are inefficient, inconsistent in quality, slow in updating and iteration when building a label system, and lack cross-domain uniformity. They cannot meet the processing needs of massive multi-source heterogeneous data, and relying solely on large model label generation has the problem of poor domain adaptability.
By combining large-scale models with rule engines, we construct an efficient, accurate, and scalable tag system through multi-source data collection and preprocessing, tag system framework construction, large-scale model tag generation, intelligent clustering and tag refinement, rule engine constraints and optimization, tag quality assessment, and automated update mechanisms.
It automates label generation, improves efficiency and quality, ensures that the label system keeps pace with product iterations, supports cross-domain uniformity and standardization, reduces costs, and improves the accuracy and timeliness of data analysis.
Smart Images

Figure CN121350628B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of data processing technology, specifically relating to a method for automatically generating labels for massive amounts of data based on a large model and a rule engine. Background Technology
[0002] In the field of customer experience management, the tagging system is the core foundation for data analysis, user insights, and business optimization. Its quality and efficiency directly affect a company's responsiveness to market demands and the accuracy of its decisions. Traditional tagging systems primarily rely on manual annotation, which has revealed several prominent problems in practical applications:
[0003] First, manual annotation is extremely inefficient and struggles to handle the demands of processing massive amounts of data. With the rapid development of internet technology, the amount of customer data accumulated by enterprises is exploding, including structured transaction data, semi-structured form data, and unstructured text comments and customer service conversations. Manual annotation methods are insufficient to complete the tagging of large-scale data within a reasonable timeframe, failing to meet the needs of enterprises to build large-scale tagging systems.
[0004] Secondly, poor label quality consistency severely impacts the reliability of data analysis results. Differences in the professional skills, knowledge background, and subjective judgment of labelers can lead to different labels being assigned to the same type of data, resulting in a chaotic labeling system. Data analysis results based on this system deviate from reality and fail to provide effective support for business decisions.
[0005] Furthermore, the tagging system is updated and iterated slowly, becoming out of sync with the product iteration pace. In the context of increasingly fierce market competition, companies are constantly shortening their product iteration cycles and continuously launching new features and services. Traditional manually defined tagging systems struggle to adapt quickly to these changes, causing the tagging system to lag behind the actual product status. This results in outdated analysis results and an inability to promptly capture user feedback on new products and services.
[0006] At the same time, manually predefined tagging systems have inherent limitations. It is difficult for humans to fully predict the potential valuable information in the data, and predefined tagging systems often have omissions or inaccuracies, resulting in some important user needs and data characteristics not being identified, thus missing business optimization opportunities.
[0007] Furthermore, the labeling systems across different fields and industries lack uniformity and standardization. Manually defined labels vary significantly across different industries and business scenarios, making cross-field comparative analysis of data difficult and impacting companies' ability to grasp overall market trends and optimize cross-business collaboration.
[0008] With the increasing development of large-scale model technology, new ideas have been provided to address the pain points of traditional tagging system construction. However, relying solely on large-scale models for tag generation still has significant shortcomings: large-scale models have poor domain adaptability, are prone to generating semantically biased tags in specific industry scenarios, lack effective rule constraints, make it difficult to guarantee tag quality and consistency, and cannot achieve dynamic updates to the tagging system, making it difficult to adapt to the needs of rapid product iteration. Therefore, there is an urgent need for a solution that combines large-scale model technology with rule engines to achieve an efficient, accurate, and scalable automated tag generation method. Summary of the Invention
[0009] To address the shortcomings of existing technologies, this invention provides a method for automated tag generation of massive amounts of heterogeneous data from multiple sources without human intervention. This method utilizes large-scale models and rule engines to automatically generate tags for these massive amounts of data, constructing a large-scale tag system containing thousands of tags. Simultaneously, it ensures that the tag system can be rapidly updated with product iterations, maintaining timeliness.
[0010] To solve the above-mentioned technical problems, the present invention adopts the following technical solution:
[0011] A method for automated tag generation of massive amounts of data based on large models and rule engines includes the following steps:
[0012] S1. Multi-source data acquisition and preprocessing: Data is acquired through a unified data interface, and then processed through data cleaning, data standardization, and data format conversion to obtain standard input data;
[0013] S2. Tag System Framework Construction: Automatically construct a multi-level tag framework based on domain knowledge graph, design a model of hierarchical, relational and mapping relationships between tags, and establish a metadata management mechanism that includes tag definitions, scope of application and version control;
[0014] S3. Large Model Tag Generation: Adopting a targeted prompting engineering strategy, the system extracts the user's original and authentic opinions through specific prompt words, generates initial tags based on a multi-turn conversational tag extraction mechanism and contextual understanding reasoning ability, and combines knowledge distillation technology to realize the domain knowledge transfer of the general large model.
[0015] S4. Intelligent Clustering and Tag Refinement: Under the constraints of preset prompt words, the original user voice is intelligently clustered to generate an initial tag set. Similar tags are merged through semantic similarity analysis, and representative tags are selected through tag importance evaluation.
[0016] S5. Rule Engine Constraints and Optimization: Based on semantic rules, business rules, and logical rules, a domain-specific rule library is built. Conflicts are resolved through rule priority management, and a rule filtering mechanism is applied to select high-quality and practical tags, realizing an automated cycle of rule verification, updating, and feedback optimization.
[0017] S6. Label quality assessment and optimization: Design a multi-dimensional label quality assessment index system and establish a label quality verification mechanism based on small samples. Then, construct a label consistency assessment model through the multi-dimensional label quality assessment index system and the small sample verification mechanism.
[0018] S7. Automatic expansion and update of the tag system: The tag system is continuously optimized based on the incremental learning mechanism. Through tag version management, usage effect feedback loop mechanism and automatic upgrade mechanism, the tag system is updated synchronously with product iteration.
[0019] Preferably, the data collected through the unified data interface in S1 includes one or more of structured data, semi-structured data, and unstructured data.
[0020] Preferably, in step S1, data cleaning and data standardization processing includes eliminating data noise, removing redundant information, and unifying the data into a standard input format that is compatible with large models and rule engines after data format conversion.
[0021] Preferably, in S2, the multi-level tag framework includes a first-level tag, a second-level tag, and a final-level tag.
[0022] Preferably, in S3, the multi-turn conversational tag extraction mechanism can optimize the semantic accuracy of tags based on contextual relationships.
[0023] Preferably, in step S5, the rule priority management mechanism resolves multiple rule conflicts by presetting rule weights, and the feedback optimization loop dynamically adjusts the rule base based on the tag usage effect.
[0024] Preferably, in step S6, the multi-dimensional label quality evaluation index system includes label accuracy, consistency, and coverage.
[0025] Preferably, in S7, the automated upgrade mechanism periodically scans and identifies high-frequency unmapped content when user voice data that cannot be mapped to existing tags is recorded, and then automatically adds tags to the tag system.
[0026] By adopting the above technical solution, the present invention has the following beneficial effects:
[0027] (1) High degree of automation, greatly improving the efficiency of tag generation. This invention realizes full-process automation of tag system construction, without the need for business personnel to manually annotate or define tags. Every step from data collection and preprocessing to tag generation, optimization and updating is completed automatically through technical means. The tag generation speed is significantly improved compared with traditional manual annotation, and the data processing capability is also improved accordingly. It supports the tagging processing of massive data and can easily cope with the processing needs of massive multi-source heterogeneous data, meeting the efficiency requirements of large-scale tag system construction for enterprises;
[0028] (2) Significantly improved tag quality ensures the accuracy of data analysis. By combining the semantic understanding capability of the large model with the rigid constraints of the rule engine, the quality of the tags generated by this invention is greatly improved. It can more comprehensively capture the core features and potential user needs in the data, avoid the omission problem of manually predefined tags, provide a high-quality tag foundation for data analysis, and ensure the accuracy and reliability of the analysis results.
[0029] (3) The data-driven tag system fits the real business needs. This invention adopts the "data-driven" tag generation method, which generates tags based on real user voice data. It completely abandons the limitations of the traditional "manual predefined" mode. The tag system directly reflects the real needs of users and the actual status of products. It can automatically discover valuable information hidden in the data and generate tag categories that are difficult for humans to predict. This makes the tag system more in line with the actual business and provides more accurate directional guidance for enterprise business optimization.
[0030] (4) Adaptive tag upgrade mechanism to maintain the timeliness of the tag system. The tag system of the present invention has the ability to automatically expand and update. Through incremental learning mechanism, version management and automatic upgrade mechanism, it can quickly adapt to product iteration and business changes. When the product launches new functions, new services or user needs change, the system can automatically identify new user feedback patterns and add or update tags in a timely manner, which effectively solves the problem of lagging updates in the traditional tag system, ensures that the tag system is always synchronized with the actual business status and the analysis results have continuous timeliness;
[0031] (5) Supports the construction of large-scale, cross-domain tag system and realizes standardized management. This invention can successfully construct a multi-level tag system with more than 5,000 tags. Through the design of a unified tag system framework, relationship model and rule base, it realizes the standardization and uniformity of cross-domain and cross-industry tag system, solves the problem of lack of unified standards in traditional tag system, improves the comparability and consistency of data in different domains, and provides the possibility for cross-business and cross-industry data analysis and collaborative optimization for enterprises.
[0032] (6) Significant cost-effectiveness and promotion of business optimization. This invention significantly reduces the cost of building and maintaining the labeling system, reduces manual labeling costs, and improves the efficiency of label application, enabling enterprises to invest more resources in core business development. Through a precise labeling system, enterprises can quickly identify user pain points, product defects and service shortcomings, discover key nodes for business improvement in a timely manner, promote product iteration and service optimization, and ultimately improve customer satisfaction.
[0033] In summary, this invention achieves automated tag generation of massive, multi-source, heterogeneous data without human intervention, through large-scale model and rule engine technology, and constructs a large-scale tag system containing thousands of tags. At the same time, it ensures that the tag system can be quickly updated with product iterations to maintain timeliness, and the generated tags are accurate and scalable. Attached Figure Description
[0034] Figure 1 This is a flowchart illustrating the present invention. Detailed Implementation
[0035] The technical solution of the present invention will now be clearly and completely described with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments.
[0036] The components of the embodiments of the invention described and shown in the accompanying drawings can typically be arranged and designed in a variety of different configurations. Therefore, the following detailed description of the embodiments of the invention provided in the drawings is not intended to limit the scope of the claimed invention, but merely to illustrate selected embodiments of the invention.
[0037] Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0038] In the description of this invention, it should be noted that the terms "center," "upper," "lower," "left," "right," "vertical," "horizontal," "inner," and "outer," etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings. They are used only for the convenience of describing the invention and for simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, they should not be construed as limitations on the invention. Furthermore, the terms "first," "second," and "third" are used for descriptive purposes only and should not be construed as indicating or implying relative importance.
[0039] In the description of this invention, it should be noted that, unless otherwise explicitly specified and limited, the terms "installation," "connection," and "linking" should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral connection; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium; and they can refer to the internal connection of two components. Those skilled in the art can understand the specific meaning of the above terms in this invention based on the specific circumstances.
[0040] Example 1
[0041] In this embodiment, the present invention proposes a method for automated tag generation of massive data based on large models and rule engines. Without the intervention of business personnel, it achieves intelligent analysis and automated tag generation of multi-source heterogeneous data through a "data-driven" approach that combines large model technology with rule engines. It also constructs a closed-loop automated upgrade mechanism for the tag system, which can be widely applied to scenarios such as customer service conversation analysis, social media opinion monitoring, e-commerce evaluation analysis, return and exchange reason identification, and product quality control.
[0042] This invention primarily addresses the following core problems in the prior art:
[0043] The construction of the tag system is inefficient: Traditional manual annotation methods cannot cope with massive amounts of multi-source heterogeneous data, and it is difficult to meet the needs of building a large-scale tag system containing thousands of tags. The construction cycle is long and the efficiency is low.
[0044] Inconsistent label quality: Differences in the professional level of labelers lead to inconsistent label quality, affecting the accuracy and reliability of data analysis results;
[0045] Lagging Tag System Updates: With frequent product iterations, the traditional tag system cannot keep up with the update speed, resulting in the tag system becoming out of touch with the actual product status and the analysis results losing their timeliness.
[0046] Limitations of predefined tags: Manually predefined tag systems are incomplete and inaccurate, making it difficult to discover potentially valuable information in the data and limiting the scope for business optimization;
[0047] Poor cross-domain label consistency: The lack of standardized label system construction methods makes it difficult to unify labels across industries and scenarios, affecting data comparison analysis and collaborative applications.
[0048] The core objective of this invention is to achieve automated tag generation of massive, multi-source, heterogeneous data through deep integration of large models and rule engines without the intervention of business personnel, thereby building a large-scale, high-quality, and dynamically updatable tag system, while ensuring the cross-domain uniformity and standardization of the tag system.
[0049] Specifically, such as Figure 1 As shown, in one embodiment of the present invention, the automated tag generation method for massive data based on large models and rule engines includes the following steps:
[0050] S1. Multi-source data acquisition and preprocessing: Data is acquired through a unified data interface, and then processed through data cleaning, data standardization, and data format conversion to obtain standard input data;
[0051] Data is the foundation of tag generation. This step aims to achieve efficient collection and standardized processing of multiple types of data, providing high-quality data input for subsequent tag generation. The data collected through the unified data interface in S1 includes one or more of structured data, semi-structured data, and unstructured data. In S1, data cleaning and data standardization processing includes eliminating data noise, removing redundant information, and unifying the data into a standard input format that is compatible with large models and rule engines after data format conversion.
[0052] The more specific branching steps are as follows:
[0053] S1.1 Unified Data Interface Design: Design a universal data collection interface that supports the comprehensive collection of structured data (such as transaction records and user information tables), semi-structured data (such as form data and log files), and unstructured data (such as customer service dialogues, social media comments, and product review texts), breaking down format barriers between different data sources and achieving one-stop data acquisition.
[0054] S1.2 Data Cleaning and Standardization: Data cleaning algorithms are used to eliminate noise in the data, including removing invalid data, correcting erroneous data, and handling missing values. At the same time, the data is standardized to unify the naming conventions, encoding formats, and units of measurement, and to remove redundant and duplicate information to ensure data consistency and availability.
[0055] S1.3 Data Format Conversion: Convert the cleaned multi-source heterogeneous data into a unified standard input format. This format must be compatible with the requirements of subsequent large model label generation and rule engine processing to ensure that the data can be efficiently parsed and analyzed, and to avoid low processing efficiency or result deviation due to format differences.
[0056] S2. Tag System Framework Construction: Automatically construct a multi-level tag framework based on domain knowledge graph, design a model of hierarchical, relational and mapping relationships between tags, and establish a metadata management mechanism that includes tag definitions, scope of application and version control;
[0057] The tag system framework is the skeleton of tag generation. This step uses a domain knowledge graph-driven approach to construct a well-defined, hierarchical, and clearly structured tag system framework, providing support for the orderly generation and management of tags. In step S2, the multi-level tag framework includes first-level tags, second-level tags, and last-level tags. The tag metadata management mechanism supports dynamic control throughout the entire tag lifecycle. More specific steps are as follows:
[0058] S2.1 Automatic Construction of Multi-Level Tag Framework: Based on domain knowledge graphs, it automatically extracts core industry concepts, business scenario elements, and user demand dimensions to construct a multi-level tag framework that includes first-level tags, second-level tags, and last-level tags. For example, in the e-commerce field, first-level tags can be divided into "product-related," "service-related," and "user-related," while second-level tags can be further subdivided into "product quality," "delivery service," and "price-sensitive," and last-level tags correspond to specific user feedback points, such as "short battery life," "delivery delay," and "price fluctuation."
[0059] S2.2 Tag Relationship Model Design: Define three core relationships between tags: hierarchical relationship (e.g., the lowest-level tag belongs to the second-level tag, and the second-level tag belongs to the first-level tag), association relationship (e.g., "short battery life" is associated with "product quality"), and mapping relationship (e.g., the mapping between the user's original statement "one charge only lasts one day" and the lowest-level tag "short battery life"). Through the relationship model (i.e., the hierarchical, association, and mapping relationship model between tags), the structured organization and efficient retrieval of tags are achieved.
[0060] S2.3 Establishment of Tag Metadata Management Mechanism: Establish a comprehensive tag metadata management system to record the core information of each tag, including tag definition (clarifying the connotation and extension of the tag), scope of application (limiting the application scenarios and data types of the tag), and version control (recording the tag's creation time, update history, obsolescence status, etc.), to achieve dynamic management and control of the entire life cycle of the tag, and provide a basis for the maintenance and updating of the tag system.
[0061] S3. Large Model Tag Generation: Adopting a targeted prompting engineering strategy, the system extracts the user's original and authentic opinions through specific prompt words, generates initial tags based on a multi-turn conversational tag extraction mechanism and contextual understanding reasoning ability, and combines knowledge distillation technology to realize the domain knowledge transfer of the general large model.
[0062] This step is the core of tag generation. Leveraging the natural language understanding and intelligent reasoning capabilities of a large model, it extracts genuine user feedback from standardized data to generate an initial tag set. In step S3, the multi-turn conversational tag extraction mechanism optimizes the semantic accuracy of tags based on contextual relationships. The knowledge distillation technique adapts the general large model to specific business scenarios through domain data training. More detailed steps are as follows:
[0063] S3.1 Targeted Prompt Engineering Strategy Design: Based on the characteristics of different industries and data types, design differentiated prompt engineering strategies, clarify the structure, expression, and core instructions of prompt words, and ensure that the large model can accurately understand the needs and standards for tag generation. For example, for customer service dialogue data, prompt words should emphasize extracting core information such as user consultation content, complaint points, and emotional tendencies. For product evaluation data, prompt words should focus on key dimensions such as product functions, user experience, and quality issues.
[0064] S3.2 Extraction of Original User Opinions: Guided by specific prompt words 1, the large model accurately extracts the core opinions of users from massive amounts of real data, avoiding information omissions or biases. For example, from the user comment "Fingerprint recognition is extremely insensitive. It often alarms and locks after two or three attempts without unlocking. Someone at home has to open the door from the inside. In less than six months, various problems have occurred," core feedback points such as "insensitive fingerprint recognition," "unable to open the door," and "six months" are extracted.
[0065] S3.3 Multi-turn Conversational Tag Extraction Mechanism Construction: For conversational data (such as customer service conversations), a multi-turn conversational tag extraction mechanism is constructed. The large model can understand the user's complete needs and feedback by combining the context, avoiding inaccurate tags caused by isolated analysis of single-sentence dialogues. For example, if a user first asks "Can I return or exchange it after installation?" and then adds "I'm worried that it won't be suitable after installation", the large model can generate tags such as "Inquiry about return and exchange rules" and "Concerns about installation compatibility" by comprehensively considering the context.
[0066] S3.4 Enhanced Contextual Understanding and Reasoning Ability: Through the advantages of pre-training and fine-tuning of the large model, its contextual understanding and logical reasoning ability are enhanced, enabling it to identify implicit information and deep needs in the user's original voice. For example, if a user feedback is "This lock is too complicated to operate, and the elderly will not know how to use it at all", the large model can infer labels such as "poor ease of operation" and "insufficient adaptability to elderly users".
[0067] S3.5 Knowledge Distillation Technology Application: To address the issue of poor domain adaptability of general-purpose large-scale models, knowledge distillation technology is employed to transfer the core capabilities of general-purpose large-scale models to specific domains. Specifically, the general-purpose large-scale model is used as the teacher model, and labeled data from specific industries (a small amount of high-quality manually labeled data) is used as training samples to train a domain-specific student model. This allows the student model to retain the language understanding capabilities of the general-purpose large-scale model while also possessing the precise analytical capabilities for industry scenarios, thereby improving the domain adaptability and accuracy of label generation.
[0068] S4. Intelligent Clustering and Tag Refinement: Under the constraints of preset prompt words, the original user voice is intelligently clustered to generate an initial tag set. Similar tags are merged through semantic similarity analysis, and representative tags are selected through tag importance evaluation.
[0069] The initial labels generated by the large model may have problems such as repetition, redundancy, and semantic similarity. This step optimizes the initial label set through intelligent clustering and refinement, improving the standardization and representativeness of the labels. In S4, the initial label set is several thousand in size, and after refinement, about one thousand high-quality and practical labels are selected. The more specific steps in this step are as follows:
[0070] S4.1 Intelligent Clustering under Preset Prompt Constraints: Under the constraint of the agreed prompt word 2, the large model intelligently clusters the initial tags corresponding to tens of thousands of original user voices, grouping tags with similar semantics and consistent core meanings into one category. For example, tags such as "fingerprint recognition is not sensitive", "low fingerprint unlock success rate", and "fingerprint recognition often fails" are clustered into the "fingerprint recognition abnormal" category.
[0071] S4.2 Initial Tag Set Generation: Through clustering, an initial tag set of several thousand categories is automatically generated, covering various core dimensions of user feedback to ensure the comprehensiveness of the tags.
[0072] S4.3 Semantic Similarity Analysis and Tag Merging: Semantic similarity algorithms (such as cosine similarity, Euclidean distance, etc.) are used to quantitatively analyze the initial tags, calculate the semantic similarity between tags, and merge tags with similarity higher than a preset threshold to eliminate redundant tags. For example, "slow delivery speed" and "long delivery time" have high semantic similarity and are merged into the tag "low delivery efficiency".
[0073] S4.4 Tag Importance Assessment and Screening: Establish a tag importance assessment index system, including tag frequency, user attention, business impact, etc., score the importance of the merged tags, and screen out the most representative tags to lay the foundation for subsequent rule engine processing.
[0074] S5. Rule Engine Constraints and Optimization: Based on semantic rules, business rules, and logical rules, a domain-specific rule library is built. Conflicts are resolved through rule priority management, and a rule filtering mechanism is applied to select high-quality and practical tags, realizing an automated cycle of rule verification, updating, and feedback optimization.
[0075] To ensure the stability and consistency of tag quality, this step uses the rigid constraints of the rule engine to further filter and optimize the refined tags, eliminating tags that do not meet business specifications and generating high-quality, practical tags. In S5, the rule priority management mechanism resolves multi-rule conflicts through preset rule weights, and the feedback optimization loop dynamically adjusts the rule base based on tag usage effects. More specifically:
[0076] S5.1 Domain-Specific Rule Base Construction: Based on industry knowledge, business norms, and common sense, a domain-specific rule base is constructed, containing three core types of rules: semantic rules (ensuring accurate and unambiguous label semantics), business rules (conforming to industry business logic and enterprise business needs), and logical rules (avoiding logical contradictions between labels). For example, semantic rules stipulate that "short battery life" must clearly correspond to issues related to the product's power system; business rules stipulate that labels must be related to business indicators that the enterprise focuses on; and logical rules stipulate that "product quality is qualified" and "product has quality defects" cannot occur simultaneously.
[0077] S5.2 Rule Priority Management Mechanism Design: To address potential rule conflicts, a rule priority management mechanism is designed, assigning different weights to different rules. When multiple rules apply simultaneously, the rule with the higher weight is executed first. For example, business rules have higher priority than semantic rules, ensuring that tags meet the core business needs of the enterprise.
[0078] S5.3 Application of rule filtering mechanism: The refined tags are input into the rule engine and filtered through the multi-dimensional constraints of the rule base to remove tags that do not conform to the rules and select about a thousand high-quality and practical tags to ensure the accuracy, consistency and practicality of the tags.
[0079] S5.4 Implementation of Automated Rule Validation and Update Mechanism: Establish an automated rule validation mechanism to verify the effectiveness of rules through actual application data. When a rule is found to have vulnerabilities or is inapplicable, the rule update process is automatically triggered. At the same time, a rule feedback optimization loop is constructed to dynamically adjust the rule base based on the effect of tag usage and business changes, thereby continuously improving the constraint effect of the rule engine.
[0080] S6. Label quality assessment and optimization: Design a multi-dimensional label quality assessment index system and establish a label quality verification mechanism based on small samples. Then, construct a label consistency assessment model through the multi-dimensional label quality assessment index system and the small sample verification mechanism.
[0081] To ensure the reliability of the labeling system, this step establishes a multi-dimensional label quality assessment system to comprehensively evaluate and optimize the generated high-quality and practical labels, and promptly identify and correct label problems. In step S6, the multi-dimensional label quality assessment index system includes label accuracy, consistency, and coverage, with a label accuracy of no less than 95%. The more detailed steps in this step are as follows:
[0082] S6.1 Multi-dimensional Tag Quality Evaluation Index System Design: This system designs an evaluation index system covering core dimensions such as tag accuracy, consistency, coverage, timeliness, and relevance. Specifically, accuracy measures the degree of matching between tags and user feedback; consistency measures the uniformity of tags across different data samples; coverage measures the extent to which tags cover user feedback dimensions; timeliness measures the speed at which tags respond to business changes; and relevance measures the degree of association between tags and business needs.
[0083] S6.2 Establishment of a label quality verification mechanism based on small samples: Considering the high cost of full data verification, a label quality verification mechanism based on small samples is established. A certain proportion of labeled data is randomly selected for manual review. The overall label quality is estimated through the small sample verification results to ensure that the label accuracy rate is not less than 95%.
[0084] S6.3 Tag Conflict Detection and Automatic Correction: A logical reasoning algorithm is used to detect conflict relationships between tags, such as the conflict between "product function is normal" and "product function is faulty". When a conflicting tag is detected, the correction process is automatically triggered. The accuracy of the tag is re-evaluated by combining the user's original data and the rule base, and the conflicting tag is removed or corrected.
[0085] S6.4 Construction of Label Consistency Assessment Model: Construct a label consistency assessment model, calculate the label consistency score of different data samples corresponding to the same type of feedback, and when the score is lower than the preset threshold, analyze the reasons for inconsistency (such as ambiguous label definition, incomplete rules, etc.) and optimize the label system or rule base accordingly to ensure improved label consistency.
[0086] S7. Automatic expansion and update of the tag system: The tag system is continuously optimized based on the incremental learning mechanism. Through tag version management, usage effect feedback loop mechanism and automatic upgrade mechanism, the tag system is updated synchronously with product iteration.
[0087] To address the issue of lagging tag system updates, this step establishes an automatic expansion and update mechanism for the tag system. This ensures that the tag system can dynamically adjust with product iterations and business changes, maintaining timeliness. In step S7, the automatic upgrade mechanism periodically scans and identifies high-frequency unmapped content when user voice data cannot be mapped to existing tags, and then automatically adds tags to the tag system. More detailed steps in this step are as follows:
[0088] S7.1 Incremental Learning Mechanism Design: Incremental learning technology is adopted to continuously input newly generated data into the model. The model is incrementally trained based on the original training results, continuously learning new user feedback patterns and business characteristics, and supporting the continuous optimization and expansion of the tag system.
[0089] S7.2 Rapid Adaptation to New Domains and Scenarios: Through dynamic updates of the domain knowledge graph and flexible configuration of the rule base, the tag system can quickly adapt to new domains and scenarios. When enterprises expand into new businesses or enter new industries, there is no need to rebuild the tag system. They only need to update the domain knowledge graph and rule base based on the existing framework to quickly generate tags that are suitable for the new business.
[0090] S7.3 Establishment of Tag Version Management and Iteration Mechanism: Establish a tag version management system to record the version information of each tag, including the initial version, updated version, obsolete version, etc., and support the retrospective and comparison of tag versions; at the same time, establish a tag iteration mechanism, set the tag update cycle, and regularly conduct a comprehensive evaluation and optimization of the tag system to ensure that the tag system keeps pace with business development.
[0091] S7.4 Tag Usage Effect Feedback Loop Construction: Collect data on the usage effect of tags in actual applications, including tag call frequency, accuracy of data analysis results, and support effect on business decisions. Based on the feedback data, identify the shortcomings of the tag system, such as low usage rate of some tags or lack of coverage of some user feedback, and optimize and supplement the tags accordingly.
[0092] S7.5 Implementation of Automatic Tag Upgrade Mechanism: In practical applications, the newly generated user voice data is mapped to defined tags through the prompt word guidance model. Voice data that cannot be mapped is recorded and categorized for storage. The system periodically scans these unmapped data. When the frequency of a certain type of unmapped content reaches a preset threshold, its core meaning is automatically identified, and it is added as a new tag to the tag system. The tag relationship model and rule base are also updated to ensure that the tag system can be automatically upgraded with product iteration and changes in user needs, keeping it synchronized with the actual business status.
[0093] Through the above specific implementation process, this invention has successfully constructed an automated tagging system, realizing efficient tagging processing of massive user feedback data. The tagging system can be automatically updated with user needs and product iterations, providing accurate and timely support for enterprise product optimization, service improvement and market decision-making.
[0094] This invention has a wide range of applications, with a core focus on the tagging of multi-source heterogeneous data in the field of customer experience management. Specifically, it can be applied to the following scenarios:
[0095] (a) Customer Service Conversation VOC Analysis
[0096] In customer service conversation analysis scenarios, this invention can automatically classify and tag customer service conversation content, quickly identifying customer inquiries, complaints, emotional attitudes, and needs. It can also be used for customer service quality assessment, analyzing customer service personnel's efficiency in resolving customer issues and their service attitude through tag analysis; automatically extracting and classifying customer complaints, identifying key issues in the service process, and providing a basis for optimizing customer service processes; furthermore, it can analyze factors influencing customer satisfaction, helping companies to improve customer service quality in a targeted manner.
[0097] (II) Social Media VOC Analysis
[0098] In social media opinion monitoring scenarios, this invention can automatically classify and tag social media user comments, quickly capture user feedback on brands and products; identify trending topics and opinion trends, and promptly detect the spread of positive or negative opinions; analyze factors influencing brand reputation, and understand users' core concerns and dissatisfactions with the brand; conduct competitive product comparison analysis, identify competitors' strengths and weaknesses by tagging competitor user feedback; monitor user concerns in real time, provide early warnings of potential public opinion risks, and provide data support for optimizing social media marketing strategies.
[0099] (III) E-commerce VOC Analysis
[0100] In e-commerce platform operation scenarios, this invention can automatically classify and tag user reviews, subdivide product review dimensions such as product quality, functionality, price, and logistics; identify user purchase decision factors and understand the core points that users care about most during the purchase process; automatically extract user experience problems of e-commerce platforms, such as cumbersome page operations and complex payment processes; analyze the differences between product descriptions and actual experiences to help companies optimize product detail page descriptions; and identify factors affecting sales conversion rates, providing direction for e-commerce platform operation strategy optimization and product improvement.
[0101] (iv) Analysis of reasons for returns and exchanges
[0102] In return and exchange management scenarios, this invention can automatically classify and tag the reasons for returns and exchanges, quickly identify the core factors leading to returns and exchanges, such as product quality defects, unsuitable sizes, and delivery damage; identify return and exchange trends, provide early warnings for high-frequency return and exchange reasons, and help enterprises promptly identify serious problems with products or services; analyze the correlation between return and exchange reasons and product characteristics, such as specific functions of a certain type of product that easily lead to returns and exchanges; generate suggestions for optimizing the return and exchange process, such as simplifying the return and exchange review process and improving product quality inspection standards; and build a return and exchange decision support system to provide data support for enterprises to formulate return and exchange rules.
[0103] (v) Product quality analysis
[0104] In product quality control scenarios, this invention can automatically classify and label product defects, identify various quality problems such as hardware failures, functional defects, and unreasonable designs; automatically identify the root causes of quality problems, such as manufacturing process defects and poor raw material quality; identify product improvement directions, providing a basis for product design and manufacturing process optimization; monitor product quality trends, tracking and analyzing changes in quality problems; and analyze the correlation between quality problems and production processes, helping enterprises accurately locate weak links in the production process and improve overall product quality.
[0105] In addition to the core application scenarios mentioned above, this invention can also be extended to customer feedback analysis, data mining, and business optimization in multiple industries such as finance, healthcare, and education, and has broad industry applicability and promotional value.
[0106] Using the above solution, the present invention improves the label generation speed by more than 100 times, from the monthly level of manual annotation to the hourly level of the system; improves the data processing capacity by 50 times, supporting the processing of tens of millions of data points per day; achieves a label accuracy rate of over 95%, exceeding the average level of manual annotation; improves label consistency by 40%, solving the problem of inconsistency in annotation by multiple people; improves label coverage by 35%, achieving more comprehensive data feature capture; reduces label maintenance costs by 70%; shortens the label system iteration cycle by 60%; successfully constructs a multi-level label system containing 5,000+ labels; supports automatic label generation for 30+ industries and 100+ scenarios; reduces manual annotation costs by 90%; improves label application efficiency by 65%; and improves customer satisfaction by 18%.
[0107] This embodiment does not impose any limitation on the shape, material, structure, etc. of the present invention. Any simple modifications, equivalent changes and alterations made to the above embodiments based on the technical essence of the present invention shall fall within the protection scope of the technical solution of the present invention.
Claims
1. A method for generating labels for massive data automatically based on a large model and a rule engine, characterized in that, Comprise the following steps: S1, multi-source data acquisition and pretreatment: collecting data through a unified data interface, and then obtaining standard input data through data cleaning, data standardization processing and data format conversion; S2, label system framework construction: automatically constructing a multi-level label framework based on a domain knowledge graph, designing a hierarchical, correlation and mapping relationship model between labels, and establishing a metadata management mechanism including label definition, application scope and version control; S3, large model label generation: adopting a targeted prompt engineering strategy, extracting user original sound real opinions through specific prompt words, generating initial labels based on a multi-round dialogue label extraction mechanism and context understanding reasoning ability, and realizing domain knowledge transfer of general large models through knowledge distillation technology; S4, intelligent clustering and label refining: intelligently clustering original user sound under the constraint of preset prompt words to generate an initial label set, merging similar labels through semantic similarity analysis, and selecting representative labels through label importance evaluation; S5, rule engine constraint and optimization: constructing a domain-specific rule library based on semantic rules, business rules and logical rules, solving conflicts through rule priority management, filtering high-quality practical labels through rule filtering mechanism, and realizing rule automation verification, update and feedback optimization cycle; S6, label quality evaluation and optimization: designing a multi-dimensional label quality evaluation index system and establishing a small sample-based label quality verification mechanism, and then constructing a label consistency evaluation model through the multi-dimensional label quality evaluation index system and the small sample verification mechanism; S7, label system automatic expansion and update: realizing continuous optimization of the label system based on an incremental learning mechanism, and synchronously updating the label system with product iteration through label version management, usage effect feedback loop mechanism and automatic upgrade mechanism.
2. The large model and rule engine-based mass data automatic label generation method according to claim 1, characterized in that: The data collected through the unified data interface in S1 includes one or more of structured data, semi-structured data and unstructured data.
3. The large model and rule engine-based mass data automatic label generation method according to claim 2, characterized in that: In S1, data cleaning and data standardization processing include eliminating data noise and removing redundant information, and after data format conversion, the data is unified into a standard input format suitable for large models and rule engines.
4. The large model and rule engine-based mass data automatic label generation method according to claim 1, characterized in that: In S2, the multi-level label framework includes primary labels, secondary labels and final labels.
5. The large model and rule engine-based mass data automatic label generation method according to claim 1, characterized in that: In S3, the multi-round dialogue label extraction mechanism can optimize label semantic accuracy based on context association.
6. The large model and rule engine-based mass data automatic label generation method according to claim 1, characterized in that: In S5, the rule priority management mechanism solves multi-rule conflicts through preset rule weights, and the feedback optimization cycle dynamically adjusts the rule library based on label usage effect.
7. The large model and rule engine-based mass data automatic label generation method according to claim 1, characterized in that: In S6, the multi-dimensional label quality evaluation index system includes label accuracy, consistency and coverage. 8.The large model and rule engine based mass data automatic label generation method according to claim 1, characterized in that: In S7, the automatic upgrade mechanism periodically scans and identifies high-frequency unmapped content when recording user original sound data that cannot be mapped to existing labels, and then automatically adds labels to the label system.
Citation Information
Patent Citations
Film and television content label processing method and terminal based on large model
CN120873232A
Data label generation and quality control method and system based on pre-training large model
CN121117627A