A data classification and grading method and system based on a large language model
By uploading classification-guided terms to an information platform and locking multiple databases, performing ternary reconstruction and parallel processing of a large language model, and generating hierarchical identification codes, the problem of insufficient accuracy in classification and grading of multi-source heterogeneous data in traditional methods is solved, achieving accurate and efficient data classification and grading.
Patent Information
- Application Number
- CN202511158011.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-19
- Publication Date
- 2026-02-10
- Estimated Expiration
- 2045-08-19
AI Technical Summary
Traditional methods struggle to fully capture the semantic, structural, and usage pattern characteristics of multi-source heterogeneous data when classifying and grading it, resulting in insufficient classification and grading accuracy and failing to meet the precise processing needs in complex scenarios.
By uploading classification-guided terms to an information platform, performing ternary reconstruction to determine the ternary classification pattern, locking multiple databases and performing data encoding and virtual layer mapping, using a large language model to perform three-thread parallel patterned classification processing, generating hierarchical identification codes, and combining the ternary classification results with the hierarchical identification codes.
It achieves accurate and efficient classification and grading of multi-source heterogeneous data, meeting the needs for accurate classification and grading in complex scenarios.
Smart Images

Figure CN120950688B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of data management and processing, in particular to a data classification and grading method and system based on a large language model. BACKGROUND
[0002] In the process of digital data processing, when classifying and grading multi-party databases, traditional methods often rely on single-dimensional rules or fixed architectures. However, in the face of multi-source heterogeneous data, these methods are difficult to fully capture the semantic, structural and usage pattern characteristics of the data, leading to difficulties in data coding integration, loose association between virtual layer mapping and original database. In the classification process, there is a lack of multi-thread parallel mechanism, the classification mode is fixed, and it is difficult to adapt to dynamic data changes, resulting in insufficient accuracy of classification and grading results, and unable to meet the precise processing needs in complex scenarios. SUMMARY
[0003] The present application provides a data classification and grading method and system based on a large language model, which is used to solve the technical problems of insufficient classification and grading accuracy and low efficiency of multi-source heterogeneous data in digital data processing.
[0004] In a first aspect, the present application provides a data classification and grading method based on a large language model, the method comprising: uploading classification-oriented terms on an information platform, performing ternary reconstruction of the classification-oriented terms according to a derivative engine, and determining a ternary classification mode, wherein the ternary reconstruction includes a data end, a system end and a response end; locking a plurality of pre-divided databases in the information platform, performing data coding and virtual layer mapping on each data unit as a data pool; initializing a large language model according to the ternary classification mode, performing three-thread parallel pattern classification processing on the data pool, determining a ternary classification result, performing phase grading on the ternary classification result, and generating a grading identification code, wherein the ternary data class is classified as a first level, and the ternary data class is classified as a same frequency data in the ternary classification result; fusing the ternary classification result and the grading identification code as a classification and grading result, and establishing a mapping between the classification and grading result and the multi-party database.
[0005] The second aspect of this application provides a data classification and grading system based on a large language model. The system includes: a ternary classification pattern determination module, used to upload classification-guided terms to an information platform, perform ternary reconstruction of the classification-guided terms according to a derivative engine, and determine the ternary classification pattern, wherein the ternary reconstruction includes a data end, a system end, and a response end; a data pool construction module, used to lock a pre-divided multi-party database in the information platform, perform data encoding and virtual layer mapping on each data unit, and serve as a data pool; a grading identifier code construction module, used to initialize a large language model according to the ternary classification pattern, perform three-thread parallel patterned classification processing on the data pool, determine the ternary classification result, perform phase grading on the ternary classification result, and generate a grading identifier code, wherein a three-phase data class is the first level, and the three-phase data class is the data classification class with the same frequency in the ternary classification result; and a classification and grading result acquisition module, used to fuse the ternary classification result and the grading identifier code as the classification and grading result, and establish a mapping between the classification and grading result and the multi-party database.
[0006] One or more technical solutions provided in this application have at least the following technical effects or advantages:
[0007] This application achieves accurate classification and grading of data in a multi-source database by uploading classification-guided terms to an information platform and locking a pre-divided multi-source database. It then obtains the ternary classification results through operations such as ternary reconstruction to determine the ternary classification mode, data encoding and virtual layer mapping to form a data pool, large language model initialization, and three-thread parallel patterned classification processing. Phase grading information is calculated and a grading identifier code is generated. The ternary classification results are then fused with the grading identifier code, and fine-tuned as needed according to directional requirements. This results in more accurate and efficient classification and grading of digital data, achieving precise and efficient classification and grading of multi-source heterogeneous data, and meeting the technical requirements for accurate classification and grading in complex scenarios. Attached Figure Description
[0008] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0009] Figure 1 This is a flowchart illustrating a data classification and grading method based on a large language model provided in an embodiment of this application.
[0010] Figure 2 This is a schematic diagram of the structure of a data classification and grading system based on a large language model provided in an embodiment of this application.
[0011] Figure labeling: Module 1 for determining the ternary classification pattern, Module 2 for constructing the data pool, Module 3 for constructing the hierarchical identification code, and Module 4 for obtaining the classification and grading results. Detailed Implementation
[0012] This application provides a data classification and grading method and system based on a large language model to solve the technical problems of insufficient accuracy and low efficiency in the classification and grading of multi-source heterogeneous data during digital data processing.
[0013] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them. All other embodiments obtained by those skilled in the art based on the embodiments of this application without creative effort are within the scope of protection of this application.
[0014] It should be noted that the terms "first," "second," etc., in the specification and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or server that includes a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or modules not explicitly listed or inherent to such processes, methods, products, or devices.
[0015] Example 1, as Figure 1 As shown, a data classification and grading method based on a large language model is described, wherein the method includes:
[0016] Step A100: Upload the classification-oriented terms to the information platform, and perform a three-element reconstruction of the classification-oriented terms based on the derivative engine to determine the three-element classification mode. The three-element reconstruction includes the data end, the system end, and the response end.
[0017] In this embodiment, the information platform is used to upload classification-guided terms, store and manage a pre-divided multi-party database, and support the derivative engine in performing ternary reconstruction of classification-guided terms to determine a ternary classification pattern. Classification-guided terms are terms used for subsequent ternary reconstruction to determine the ternary classification pattern; their semantic information, after interpretation, can serve as the basis for the derivative engine to generate the classification pattern.
[0018] Specifically, the process of uploading classification-guided terms to the information platform requires a clear operational procedure and system processing mechanism. Classification-guided terms are the core input items used for subsequent data classification and grading, and are typically professional terms with business semantics, such as user transaction data, equipment operation logs, and enterprise supply chain information. These terms must contain key semantic elements that can be recognized by the system. For example, user transaction data must cover information such as transaction amount, time, and account type, so that the system can extract core features through semantic interpretation technology.
[0019] Uploading is typically done through the user interface of the information platform. Users first need to log in to the platform and access the term management module through the navigation bar or function entry. In the upload interface, users can submit terms in two ways: one is to directly enter the term name and a brief description in the text box, such as user transaction data, including transaction amount, time, and account information; the other is to upload a pre-organized structured file, such as an Excel spreadsheet, and the system will automatically parse the file content and extract the terms.
[0020] Subsequently, a ternary reconstruction of the classification-oriented term is performed to determine the ternary classification mode. First, the semantic information of the term is identified and interpreted. Then, the derivative engine is triggered to generate a first classification mode based on the automatic hierarchical rules of the data end, a second classification mode based on the graph system architecture, and a third classification mode based on the mutual information response triggered by event elements. Finally, these three modes are integrated. The specific steps are explained in detail in A110-A130.
[0021] By semantically interpreting classification-oriented terms, generating and integrating multi-dimensional classification patterns, a comprehensive ternary classification model was obtained, providing a precise standard for subsequent data classification and grading.
[0022] Step A200: Locate the pre-divided multi-party database in the information platform, and perform data encoding and virtual layer mapping on each data unit to form a data pool.
[0023] In this embodiment, the multi-party database refers to multiple pre-divided databases within an information platform, each database being identified by an encoding based on its storage location and access path. The data pool is formed by performing data encoding and virtual layer mapping on each data unit in the multi-party database, and is a data set acted upon by the virtual layer.
[0024] Optionally, firstly, the pre-divided multi-party databases within the information platform are identified. These databases are each identified by a code based on their storage location and access path. By recognizing these codes, these multi-party databases can be accurately identified and located, preparing for subsequent data encoding and virtual layer mapping of each data unit to form a data pool.
[0025] Next, data encoding and virtual layer mapping are performed on each data unit to serve as a data pool. First, multi-party databases with encoding based on storage location and access path are identified. Then, data encoding methods are introduced to traverse these databases to encode the data units and integrate them into an encoded database. Finally, a virtual layer is established as a data pool based on the encoded database, and this virtual layer has a mapping relationship with the multi-party database. The specific steps are explained in detail in A210-A230.
[0026] Step A300: Initialize the large language model according to the ternary classification pattern, perform three-thread parallel pattern classification processing on the data pool, determine the ternary classification result, perform phase grading on the ternary classification result, and generate a grading identifier code, wherein the three-phase data class is the first level, and the three-phase data class is the data classification class with the same frequency in the ternary classification result.
[0027] In one embodiment of this application, a large language model is first initialized according to the ternary classification pattern, including performing transfer calls and partitioned lightweight training on the large language model to determine the initialized model, and then using the model to perform parallel classification processing on the data pool to determine the ternary classification result. The specific steps are described in detail in A310-A320.
[0028] Next, the data pool performs three-thread parallel patterned classification processing to determine the ternary classification results, including three-thread parallel classification processing of each partition processing domain and class association of the results based on mutual information relationships. The specific steps are described in detail in A330-A340.
[0029] Finally, the three-phase and two-phase data classes in the ternary classification results are determined by class association identification to classify the data levels, and then a hierarchical identification code is generated accordingly. The specific steps are explained in detail in A350-A380.
[0030] Step A400: Merge the ternary classification result with the hierarchical identification code as the classification and grading result, and establish a mapping between the classification and grading result and the multi-party database.
[0031] Specifically, when integrating ternary classification results and hierarchical identification codes, the correspondence between the two must first be clarified. The ternary classification result includes classification labels for each data unit across three dimensions: data end, system end, and response end. For example, in an e-commerce platform, the ternary classification result for user real-name authentication information is core data end + strongly related system end + key support response end; the ternary classification result for product browsing records is general data end + weakly related system end + auxiliary support response end. The hierarchical identification code is a level identifier determined based on the data class phase, such as L1 for the first level, L2 for the second level, and L3 for the third level. These identifiers are then bound to their corresponding data classes and integrated into hierarchical identification codes.
[0032] Next, the ternary classification result of each data unit is bound and merged with the corresponding hierarchical identification code. Taking user real-name authentication information as an example, the merged classification and hierarchical result is: core data + strong correlation + key support | L1-001; the merged result for product browsing records is: general data + weak correlation + auxiliary support | L2-002. This process is performed on all data units in the data pool to ensure that each unit has a complete classification and hierarchical result containing three-dimensional classification labels and hierarchical identifiers.
[0033] Subsequently, a mapping is established between the classification and grading results and the multi-party databases. Each original data unit in the multi-party database has a unique code based on its storage location and access path; for example, user real-name authentication information is stored at address 1, and product browsing records are stored at address 2. Through the existing mapping relationship between the data pool and the multi-party databases, virtual layer units in the data pool correspond to original database units. The fused classification and grading results are associated with the storage codes of the original data units to form a mapping table. For example, the mapping table records address 1 → core data + strong correlation + key support | L1-001, address 2 → general data + weak correlation + auxiliary support | L2-002, achieving precise pointing of the classification and grading results to the original data.
[0034] By integrating the ternary classification results with the hierarchical identification code to generate complete classification and grading information, and by using mapping relationships to link to the original data in multiple databases, the classification and grading results are accurately matched with the source data, providing a clear and traceable basis for data querying, management and application.
[0035] Furthermore, step A100 in the method provided in this application embodiment includes:
[0036] A110: Identify the classification-guided terms and determine the semantic information of the terms through interpretation.
[0037] A120: Trigger the derivative engine to determine the first classification mode by performing a transformation based on automatic data-side hierarchical rules for the semantic information of the term, determine the second classification mode by performing an architecture transformation based on the graph system, and determine the third classification mode by using mutual information response triggered by event elements.
[0038] A130: Integrate the first classification mode, the second classification mode, and the third classification mode to determine the ternary classification mode.
[0039] Specifically, when processing category-oriented terms, the first step is to accurately identify the terms and extract their core semantic information using semantic interpretation technology. The process of identifying and interpreting the semantic information of category-oriented terms relies on the collaboration of natural language processing technology and a domain knowledge base. First, the uploaded category-oriented terms are preliminarily analyzed using a term structure recognition module. This involves word segmentation of the term text to extract core vocabulary units; for example, user transaction data is segmented into three basic units: user, transaction, and data. Next, entity recognition technology is used to locate the key entities, identifying users as the subject, transactions as the behavior, and data as the object, thus clarifying the basic constituent elements of the term.
[0040] Subsequently, a pre-trained semantic parsing model is invoked to perform deep semantic mining on the entries. This model, combined with a domain knowledge base, such as a transaction data specification library in the financial field, performs semantic role labeling on the decomposed lexical units, identifying the implicit attribute dimensions in user transaction data. For example, by analyzing the semantic relationships of transactions, it determines that they must contain core attributes such as transaction amount, transaction time, and transaction account; through the user's subject characteristics, it infers related information such as user identity and user level. At the same time, the system compares the parsed semantic elements with preset domain semantic templates, such as matching the template structure of [subject] + [behavior] + [data type], to check for semantic missingness. If it is found that the transaction data does not clearly include whether it contains transaction status, such as success or failure, the potential attribute dimension is automatically supplemented.
[0041] Furthermore, the semantic parsing model is built on a pre-trained large language model as its basic framework, integrating domain knowledge bases, such as a financial transaction data specification library, as the core training data source. The training process adopts a transfer call and partitioned lightweight training approach. First, the basic model is transferred and the training domain is divided according to the domain semantic processing requirements. Then, supervised training is conducted using term samples labeled with attribute dimensions and semantic roles, such as user transaction data examples containing attributes such as transaction amount. Through iterative optimization using verification feedback data matching domain semantic templates, the model is equipped with the ability to accurately identify the implicit attributes of terms and fill in semantic gaps.
[0042] After multiple rounds of semantic verification and completion, structured semantic information of terms is formed. For example, the interpretation result of user transaction data is clearly defined as user behavior record data containing attributes such as user identity, transaction amount, transaction time, transaction account, and transaction status. This semantic information will serve as the basis for the derivative engine to generate classification patterns.
[0043] Next, the derivative engine is triggered to generate classification patterns from three dimensions based on the interpreted semantic information. From the data perspective, based on automatic hierarchical rule conversion, user transaction data is divided into core data (e.g., account information), important data (e.g., transaction details), and general data (e.g., transaction timestamps) according to data importance, thus forming the first classification pattern. From the system perspective, based on the architecture conversion of the graph system, a relationship graph of user transaction data, users, transaction types, and payment channels is constructed, and data is divided into strongly correlated data (e.g., transaction details and user accounts) and weakly correlated data (e.g., transaction time and marketing activities) according to the closeness of the relationship between data and business, thus forming the second classification pattern. From the response perspective, based on mutual information response triggered by event elements, for transaction anomaly verification events, the role of data in the event is determined, such as key evidence data (e.g., account transaction records) and auxiliary explanatory data (e.g., transaction location records), thus forming the third classification pattern.
[0044] Next, the generated first, second, and third classification patterns are integrated to determine a ternary classification pattern, as shown in Table 1. Since the source basis for all three classification patterns is the initial classification guide terminology, in subsequent classification processing, data portions with consistent classification results—i.e., the three-phase data classes—are identified as data classes with high mutual information and are designated as Level 1 data. Based on this, further subdivisions can be made to achieve hierarchical classification. Simultaneously, when other hierarchical requirements exist, subordinate processing is performed based on the above three-level division to provide a clear basis for data hierarchical classification.
[0045] By semantically interpreting classification-guided terms, generating and integrating multi-dimensional classification patterns, a ternary classification model covering data attributes, relationships, and scenario effects was formed, laying the foundation for accurate classification and grading of subsequent data.
[0046] Table 1: Integrated Table of Ternary Classification Models
[0047]
[0048] Furthermore, step A200 in the method provided in this application embodiment includes:
[0049] A210: Identify the multi-party databases, wherein each database is identified by an encoding based on its storage location and access path.
[0050] A220: Introducing a data encoding method, traversing the multiple databases to perform data unit encoding, and integrating them into an encoded database.
[0051] A230: A virtual layer is established based on the encoded database, and the virtual layer is used as a data pool, wherein the virtual layer has a mapping relationship with the multi-party database.
[0052] Optionally, when processing data units to build a data pool, the system first identifies the pre-divided multi-party databases within the information platform. These databases exhibit heterogeneity due to their storage location, data architecture, etc. For example, some databases are deployed on local servers using a relational architecture; some are deployed in the cloud using a document-oriented architecture; and others are distributed storage using a time-series architecture. This heterogeneity in storage location and data architecture increases the difficulty of data processing. Each database has a unique code based on its storage location and access path. By identifying these codes, the system accurately locates and associates all multi-party databases to be processed.
[0053] Next, a data encoding method is introduced to encode each data unit. Specifically, this is achieved by generating a multi-dimensional feature vector for each data unit. The semantic feature vector extracts the core meaning of the data unit; for example, the semantic feature vector of a record in a user database contains information such as user ID, registration channel, and membership level. The structural feature vector reflects the organization of the data unit in the original database; for example, the structural feature vector of a record in a relational database indicates its table and field position, while the structural feature vector of an entry in a document database indicates its nesting level. The usage pattern vector counts the usage of the data unit; for example, the usage pattern vector of a transaction record records its frequency of use in daily settlement and monthly report scenarios. After traversing all multi-databases and completing the above encoding for each data unit, these encoding results are integrated into a unified format encoded database to standardize the data dimensions and resolve differences caused by heterogeneity.
[0054] Next, a virtual layer is established based on the encoded database and used as a data pool. This virtual layer acts as an intermediary between the source (multi-party databases) and the processing end (subsequent classification and grading operations). It stores the unified encoded data and is linked to the original multi-party databases through mapping relationships. For example, one encoded unit in the virtual layer corresponds to the 128th record in the user table, while another encoded unit corresponds to the 5th entry in the document collection. This mapping ensures that the data in the virtual layer is traceable to its original source, and the processing end only needs to operate on the virtual layer, without directly interacting with the heterogeneous source databases.
[0055] By identifying heterogeneous multi-party databases, generating multi-dimensional feature vectors for unified encoding, and establishing a virtual layer with mapping relationships, the heterogeneity problem of multi-party databases is solved, forming a data pool that can be directly used for subsequent processing.
[0056] Furthermore, step A220 in the method provided in this application embodiment includes:
[0057] A221: A data encoding method is to generate a multidimensional feature vector for each data unit, wherein the multidimensional feature vector includes at least a semantic feature vector, a structural feature vector, and a usage pattern vector.
[0058] Specifically, when encoding data units, the encoding object is first clearly defined as each data unit in the multi-party database. These data units may cover various types, such as basic user information, transaction records, and device operating parameters. For example, the multi-party database of an e-commerce platform contains data units such as user ID, name, and mobile phone number from the user registration information table, as well as data units such as order number, transaction amount, and payment method from the transaction log table, and data units such as waybill number, shipping address, and delivery time from the logistics information table.
[0059] For each data unit, the first step is to generate a semantic feature vector. Natural language processing techniques are used to parse the content of the data unit, extracting and quantifying its core semantic information. Taking a data unit with transaction amount of 500 yuan, transaction time of 2025-07-20, and payment method of bank card as an example, the semantic feature vector will include dimensions such as the transaction amount, timestamp, and payment method category. 500 yuan is mapped to a numerical feature, 2025-07-20 is converted to a time feature, and the bank card corresponds to a pre-defined category code, forming a vector representation that reflects the core meaning of the data.
[0060] Next, a structural feature vector is generated, focusing on the storage structure and location information of the data unit in the original database. For example, if the transaction data unit mentioned above comes from a transaction table in a relational database, the structural feature vector will include information such as table name identifier, field index, and storage path encoding, thus reflecting the organization of the data in the original architecture. For data units in a document-based database, such as positive reviews and fast logistics in a user review document, the structural feature vector will include structural information such as document ID, paragraph position, and nesting level.
[0061] Then, a usage pattern vector is generated. By analyzing the historical call records of the statistical data units, usage characteristics such as usage scenarios, call frequency, and associated operations are transformed into vectors. For example, if the aforementioned transaction data unit is called once during daily settlement and once during monthly report generation, and is often used in conjunction with user ID and order number data units, the usage pattern vector will include quantitative characteristics such as daily call frequency, monthly call frequency, and the number of associated data units.
[0062] Finally, the semantic feature vector, structural feature vector, and usage pattern vector are combined to form a multidimensional feature vector for each data unit, which serves as a unified data encoding method. For example, the multidimensional feature vector of the aforementioned transaction data unit is a combination of the three vectors: semantic, structural, and usage pattern. This preserves the core meaning of the data while reflecting its structural attributes and usage patterns.
[0063] By generating multi-dimensional feature vectors containing semantics, structure, and usage patterns for each data unit, standardized encoding of data units of different types and structures is achieved, providing a unified and rich feature foundation for the subsequent construction and classification and hierarchical processing of the data pool.
[0064] Furthermore, step A300 in the method provided in this application embodiment includes:
[0065] A310: Based on the ternary classification pattern, perform transfer calling and partitioned lightweight training on the large language model to determine the initialized large language model.
[0066] A320: For the data pool, perform model-parallel classification processing using the initialized large language model to determine the ternary classification result.
[0067] Specifically, after determining the initial large language model, parallel classification processing is performed on the data pool. At this point, the data pool has integrated all encoded data units from multiple databases. Each data unit exists as a multi-dimensional feature vector containing semantics, structure, and usage patterns. For example, in the data pool of an e-commerce platform, there is an encoded vector for user registration information, with semantic features including user ID and registration time, structural features including the third field stored in the user table, and usage pattern features including an average of 2 calls per day. There is also an encoded vector for transaction records, with semantic features including order number and transaction amount, structural features including the fifth field stored in the transaction table, and usage pattern features including an average of 5 calls per day.
[0068] The initialized large language model has been partitioned according to a ternary classification model. The three partitions correspond to the first, second, and third classification models, respectively, and are capable of parallel processing. During processing, the model simultaneously inputs all data units from the data pool into the three partitions. The first partition analyzes the semantic feature vector of each data unit based on the automatic data-side classification rules. For example, user data containing account passwords is marked as core data, and data containing transaction time is marked as general data. The second partition analyzes the structural feature vector of the data unit and its association with other data based on the graph architecture. For example, data closely associated with order numbers and user IDs is marked as strongly associated data. The third partition uses event element mutual information response and pattern vectors to determine the role of data in the scenario. For example, transaction amount data frequently accessed in abnormal transaction verification is marked as key data.
[0069] After the three partitions complete processing synchronously, each outputs the classification results for its corresponding dimension. These three types of results are collected, and the classification labels for the same data unit across the three dimensions are integrated to form a ternary classification record for that data unit. For example, the ternary classification result for a transaction data item might be: core data class + strongly correlated data class + key data class. After traversing all data units in the data pool, a complete ternary classification result is finally formed, covering the classification information of each data unit across the three dimensions of data source, system, and response.
[0070] By leveraging the partitioned parallel processing capability of the initialized large language model, multi-dimensional classification was performed synchronously on the encoded data in the data pool, efficiently obtaining ternary classification results covering three dimensions, laying the foundation for subsequent phase classification.
[0071] Furthermore, step A310 in the method provided in this application embodiment includes:
[0072] A311: Migrate call the large language model, execute model processing domain partitioning at the classification mode scale, and determine the partition processing domain.
[0073] A312: Using the one-to-one correspondence between the ternary classification pattern and the partition processing domain, perform classification pattern writing and data-driven lightweight supervised training to determine the initialized large language model.
[0074] Specifically, when performing transfer calls and lightweight training on partitioned large language models, transfer calls are performed first. A pre-trained general-purpose large language model is selected, leveraging its learned semantic understanding, logical reasoning, and other basic capabilities to avoid the resource consumption of training from scratch. Next, model processing domain partitioning is performed on a classification mode scale. If a ternary classification mode includes three modes—data-side, system-side, and response-side—that is, the classification mode scale is 3, then the processing domain of the large language model is divided into 3 independent partitioned processing domains. Each partition corresponds to the processing requirements of one classification mode. For example, the first partition focuses on handling tasks related to automatic classification rules on the data side, the second focuses on graph system architecture transformation, and the third focuses on mutual information response of event elements.
[0075] Next, a one-to-one correspondence is established between the ternary classification model and the partitioned processing domains. The first classification model, which is based on the core features of the automatic data-side hierarchical rules, such as the classification criteria and data attribute label system, is written into the first partitioned processing domain; the second classification model, which is based on the core features of the graph system architecture, such as entity association rules and hierarchical structure templates, is written into the second partition; and the third classification model, which is based on the core features of the mutual information response of event elements, such as event type and data association weight, and response trigger threshold, is written into the third partition, so that each partition clearly defines its own processing objective.
[0076] Subsequently, data-driven, lightweight supervised training is performed. Annotated data adapted to each classification mode is collected: For the first partition, user transaction data samples labeled with core, important, and general data categories are prepared (e.g., data containing account passwords is marked as core data); for the second partition, enterprise data and departmental association samples with strong and weak correlations are prepared (e.g., production data and production departments are marked as strongly correlated); for the third partition, key and auxiliary event-related data samples are prepared (e.g., financial statements in audit events are marked as key). This data is input into the corresponding partition, and only some parameters within the partition are adjusted, such as attention mechanism weights and output layer mapping relationships. Through multiple iterations, the processing accuracy of each partition for the target classification mode is improved. Finally, the parameters of each partition are integrated to determine the initialized large language model.
[0077] By transferring and calling a general large language model, partitioning by classification mode, writing pattern features, and performing lightweight training, an initial large language model that can accurately adapt to the ternary classification mode was obtained, providing a model foundation for subsequent efficient processing of data pool classification tasks.
[0078] Furthermore, step A300 in the method provided in this application embodiment includes:
[0079] A330: For the data pool, each partition processing domain performs pattern classification processing under three-thread parallelism to determine the ternary classification result.
[0080] A340: Use mutual information relationships to perform class association on the ternary classification results.
[0081] In this embodiment of the application, mutual information relationship is a relationship used to measure the degree of association between data classes. It can be used to trigger mutual information response based on event elements to determine the third classification mode, and it can also be used to class associate the ternary classification results.
[0082] In one embodiment, the data units in the data pool exist in the form of multi-dimensional feature vectors, covering features such as semantics, structure, and usage patterns. For example, in the data pool of a financial platform, there are encoded data units such as user A's account balance records, user B's transfer details, and system transaction logs. The feature vector of each unit clearly reflects its core information and attributes.
[0083] Next, each partition processing domain, corresponding to the three dimensions of the ternary classification pattern, starts an independent thread to perform pattern classification processing. The first thread focuses on the automatic classification rules of the data end, analyzing the semantic feature vectors of data units, such as marking the account balance record of user A containing account password and ID number as the core data class; the second thread, based on the graph system architecture, parses the structural feature vectors of data units and analyzes their relationship with other data, such as marking the transfer details of user A, which are closely related to user A's account balance record, as the strongly related data class; the third thread, based on the mutual information response of event elements, combines the use of pattern vectors to determine the role of data in the scenario, such as marking the transaction logs of the system frequently called in abnormal transaction verification as the key supporting data class. The three threads run synchronously, completing the classification of all data units in the data pool. Each data unit obtains a classification label in three dimensions, which together constitute the ternary classification result. For example, the ternary classification result of user A's account balance record is: core data class + strongly related data class + key supporting data class.
[0084] After classification, the ternary classification results are associated using mutual information. The degree of association is measured by calculating the mutual information value between the classification labels of different data units; a higher mutual information value indicates a stronger association. First, three-dimensional labels are extracted from the ternary classification results for each data unit: one label each for the data side, the system side, and the response side. These labels are used as the basis for calculation. When calculating the mutual information value, the co-occurrence of labels for any two data units is first statistically analyzed. For example, if data unit X is labeled as "core data," "strong association," and data unit Y is labeled as "core data," "strong association," and "key support," the frequency of identical labels for the two data units at the data side, system side, and response side is calculated, along with the probability of their overall label combinations appearing together.
[0085] Then, based on the information theory formula for calculating mutual information, I(X,Y)=H(X)+H(Y)-H(X,Y), where H(X) and H(Y) are the label entropies of X and Y respectively, and H(X,Y) is the joint entropy of the two, the mutual information value is calculated by substituting the statistically obtained probability values. The value ranges from 0 to 1, with the closer to 1 indicating a stronger association. After calculating the mutual information value, a threshold such as 0.7 is set for class association: if the mutual information value of two data units is higher than the threshold, they are determined to be strongly associated and classified into the same association class; if it is lower than the threshold, they are weakly associated and belong to different categories. For example, the mutual information value of user account information and transaction order details is 0.85, classifying them into the same class; the mutual information value of user account information and product browsing records is 0.4, classifying them into different classes. This achieves class association of the ternary classification results based on mutual information relationships. Among them, data units with completely identical or highly homogeneous classification labels are divided into three data classes, serving as the first-level basis for subsequent hierarchical classification.
[0086] The data pool is efficiently classified by parallel processing of three threads. The class association of the results is completed by combining mutual information relationship, and finally a structured ternary classification result is obtained, which provides an accurate classification basis for the phase classification of data.
[0087] Furthermore, step A300 in the method provided in this application embodiment includes:
[0088] A350: By performing class association recognition, the three-phase data classes in the ternary classification results are determined, and the first data level is determined.
[0089] A360: Identify the biphasic data class in the ternary classification results and determine the second data level.
[0090] A370: Determine the three-phase data class in the ternary classification results and determine the third data level.
[0091] A380: Generate a hierarchical identification code based on the first data level, the second data level, and the third data level.
[0092] Optionally, when performing class association identification, the analysis is based on the frequency and association characteristics of each data class in the ternary classification results. The strength of the association between classes is determined by comparing the label matching degree and frequency of different data units across the three classification dimensions: data, system, and response. For example, in the ternary classification results of an e-commerce platform, user real-name authentication information is marked as a core data class on the data side, strongly associated on the system side, and key support on the response side, and the frequency of these three labels is completely consistent. Similarly, transaction order details are marked as core data class, strongly associated, and key support in the three dimensions respectively, with frequencies synchronized with user account information. These two data classes belong to the same frequency ternary data classes.
[0093] After identifying the three data categories in the ternary classification results, they are designated as the first data level. For example, the user account information and transaction order details mentioned above are classified as first level because they are synchronous and closely related in the three dimensions, representing the highest data importance and relevance.
[0094] Next, biphasic data types are identified, which are datasets where the labels in two of the three classification dimensions have the same frequency, while the third dimension shows differences. For example, product browsing records are marked as important data types on the data side and as weakly correlated on the system side, with consistent frequencies in these two dimensions. However, on the response side, they are marked as auxiliary support, which is out of sync with the first two dimensions. This type of data is identified as a biphasic data type and classified as the second data level, with its importance and correlation being lower than the first level.
[0095] Then, another type of three-phase data class is identified, namely, the same-frequency data that differs from the first level in terms of correlation characteristics. For example, system log information is marked as general data class, no correlation, and no response in all three dimensions. Although it is the same frequency, its characteristics are different from the first level, so it is classified as the third data level, which represents the lowest importance.
[0096] Based on the first, second, and third data levels mentioned above, a unique code identifier is assigned to each level, such as L1 for the first level, L2 for the second level, and L3 for the third level. These identifiers are then bound to the corresponding data classes and integrated into hierarchical identification codes. For example, the hierarchical identification code for user real-name authentication information is L1-001, for product browsing records it is L2-002, and for system log information it is L3-003.
[0097] By class association identification, three-phase and two-phase data classes are distinguished and classified into levels, and finally a hierarchical identification code with clear hierarchy is generated. This achieves accurate quantitative classification of the ternary classification results and provides a clear identification basis for the subsequent integration and application of classification and grading results.
[0098] Furthermore, step A400 in the method provided in this application embodiment includes:
[0099] A510: Determine whether there is a need for targeted classification and grading.
[0100] A520: If it exists, fine-tune the classification and grading results according to the specified targeted classification and grading requirements.
[0101] In one embodiment, after generating the basic classification and grading results, a targeted demand determination mechanism is initiated to identify any special demands that exceed the main architecture classification. The determination process is implemented in two ways: first, by receiving targeted demand parameters actively input by users through an information platform interface, such as an instruction from an e-commerce platform operator that order data during the 618 promotion period should be marked as high priority; second, by connecting to the rule base of the actual project's business system to automatically search for pre-defined special scenario rules, such as the requirement for additional risk level labeling for credit approval data in financial institutions' compliance rules. The system performs semantic analysis on these inputs or retrieved information, extracting key demand elements, such as the adjustment object (order data, credit approval data); the adjustment direction (priority increase, adding risk level labels); and the applicable scenario (618 promotion, credit approval process), thereby determining whether there are valid targeted classification and grading demands.
[0102] If a specific need is identified, the system will perform targeted fine-tuning based on the original classification and grading results. Taking a certain medical platform as an example, the basic classification of patient outpatient records results in a ternary classification of "Important Data + Strong Correlation + Auxiliary Support," with a grading identifier code of L2-005. The specific need, however, is that outpatient records require an increased protection level and privacy labeling according to medical privacy regulations. In this case, the system will first locate the patient outpatient record entry in the classification and grading results, adjust its grading identifier code to L1-005 to increase the level, and simultaneously add a privacy protection attribute to the response label of the ternary classification result, changing it to "Important Data + Strong Correlation + Auxiliary Support | Privacy Protection." During the fine-tuning process, the system will synchronously update the mapping relationship between this data and multiple databases to ensure that records in the original storage location can be associated with the adjusted classification and grading information; for example, updating the result corresponding to storage path address 3 to the adjusted value.
[0103] For multiple targeted requests, the system will process them sequentially according to priority. For example, if a company simultaneously needs to fine-tune the encryption level of its financial data and the access permission level of its personnel data, the system will, based on a preset priority (financial data takes precedence over personnel data), complete the fine-tuning of the financial data first, and then process the personnel data to avoid conflicts. After the fine-tuning is completed, a fine-tuning report will be generated, recording the comparison before and after the adjustment, the targeted requests on which the adjustment was based, and the effective date, for subsequent auditing and traceability.
[0104] By identifying and fine-tuning targeted classification and grading needs, the classification and grading results are made more closely aligned with specific business scenarios and special rules on the basis of the main architecture, thereby improving the flexibility and practical application value of the method.
[0105] In summary, the data classification and grading method based on a large language model provided in this application has the following technical effects:
[0106] This application determines the ternary classification model by uploading classification-guided terms to an information platform and performing ternary reconstruction through a derivative engine. It locks multiple databases to encode data units and maps them to a data pool. Based on the ternary classification model, it initializes a large language model to perform three-thread parallel classification processing on the data pool to obtain ternary classification results. It generates a graded identification code based on the phase classification, and establishes a mapping with multiple databases after fusing the results and identification codes. It can also be fine-tuned according to targeted needs, thereby accurately realizing the classification and grading of data, making the data classification and grading results more accurate and reliable. It achieves accurate and efficient classification and grading of multi-source heterogeneous data, and meets the technical effect of accurate classification and grading of digital data processing in complex scenarios.
[0107] Example 2, as Figure 2 As shown, based on the same inventive concept as in Embodiment 1 above, this application provides a data classification and grading system based on a large language model, the system comprising:
[0108] The ternary classification mode determination module 1 is used to upload classification-oriented terms to the information platform, perform ternary reconstruction of the classification-oriented terms according to the derivative engine, and determine the ternary classification mode. The ternary reconstruction includes the data end, the system end and the response end.
[0109] Data pool construction module 2 is used to lock the pre-divided multi-party database in the information platform, and perform data encoding and virtual layer mapping on each data unit to serve as a data pool.
[0110] The hierarchical identification code construction module 3 is used to initialize the large language model according to the ternary classification mode, perform three-thread parallel pattern classification processing on the data pool, determine the ternary classification result, perform phase classification on the ternary classification result, and generate hierarchical identification codes, wherein the three-phase data class is the first level, and the three-phase data class is the data classification class with the same frequency in the ternary classification result.
[0111] The classification and grading result acquisition module 4 is used to fuse the ternary classification result with the grading identifier code as the classification and grading result, and to establish a mapping between the classification and grading result and the multi-party database.
[0112] Furthermore, the ternary classification pattern determination module 1 is used to perform the following steps:
[0113] Identify the classification-guided terms and determine the semantic information of the terms through interpretation; trigger the derivative engine to determine the first classification mode by performing a transformation based on automatic data-side hierarchical rules for the semantic information of the terms, determine the second classification mode by performing an architecture transformation based on the graph system, and determine the third classification mode by mutual information response triggered by event elements; integrate the first classification mode, the second classification mode and the third classification mode to determine the ternary classification mode.
[0114] Furthermore, the data pool construction module 2 is used to perform the following steps:
[0115] The multi-party databases are identified, wherein each database is identified by an encoding based on its storage location and access path; a data encoding method is introduced, and the multi-party databases are traversed to perform data unit encoding and integrate them into an encoded database; a virtual layer is established based on the encoded database, and the virtual layer is used as a data pool, wherein the virtual layer and the multi-party databases have a mapping relationship.
[0116] Furthermore, the data pool construction module 2 is used to perform the following steps:
[0117] The data encoding method uses a multidimensional feature vector generated for each data unit, wherein the multidimensional feature vector includes at least a semantic feature vector, a structural feature vector, and a usage pattern vector.
[0118] Furthermore, the hierarchical identification code construction module 3 is used to perform the following steps:
[0119] Based on the ternary classification pattern, the large language model is subjected to transfer learning and partitioned lightweight training to determine the initialized large language model; for the data pool, the initialized large language model is used to perform model-parallel classification processing to determine the ternary classification result.
[0120] Furthermore, the hierarchical identification code construction module 3 is used to perform the following steps:
[0121] The large language model is invoked during migration, and the model processing domain is partitioned at the level of classification mode to determine the partition processing domain. Based on the one-to-one correspondence between the ternary classification mode and the partition processing domain, classification mode writing and data-driven lightweight supervised training are performed to determine the initialized large language model.
[0122] Furthermore, the hierarchical identification code construction module 3 is used to perform the following steps:
[0123] For the data pool, each partition processing domain performs pattern classification processing in parallel using three threads to determine the ternary classification result; and class association is performed on the ternary classification result using mutual information relationships.
[0124] Furthermore, the hierarchical identification code construction module 3 is used to perform the following steps:
[0125] By performing class association recognition, the three-phase data class in the ternary classification result is determined, and the first data level is determined; the two-phase data class in the ternary classification result is determined, and the second data level is determined; the three-phase data class in the ternary classification result is determined, and the third data level is determined; a hierarchical identification code is generated based on the first data level, the second data level, and the third data level.
[0126] Furthermore, the classification and grading result acquisition module 4 is used to perform the following steps:
[0127] Determine whether there is a need for targeted classification and grading; if so, fine-tune the classification and grading results according to the targeted classification and grading needs.
[0128] The data classification and grading system based on a large language model provided in this embodiment of the invention can execute the data classification and grading method based on a large language model provided in any embodiment of the invention, and has the corresponding functional modules and beneficial effects of the execution method.
[0129] Although this application makes various references to certain modules in the system according to the embodiments of this application, any number of different modules can be used and run on user terminals and / or servers. The various units and modules included are only divided according to functional logic, but are not limited to the above division, as long as the corresponding functions can be achieved; in addition, the specific names of each functional unit are only for easy distinction between each other and are not used to limit the scope of protection of this invention.
[0130] The specific embodiments described above do not constitute a limitation on the scope of protection of this application. Those skilled in the art should understand that various modifications, combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this application should be included within the scope of protection of this application. In some cases, the actions or steps described in this application can be performed in a different order than that shown in the embodiments and still achieve the desired results. Furthermore, the processes depicted in the accompanying drawings do not necessarily require a specific or sequential order to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.
Claims
1. A data classification and grading method based on a large language model, characterized in that, The method includes: Upload classification-oriented terms to the information platform, and perform a three-element reconstruction of the classification-oriented terms based on the derivative engine to determine the three-element classification mode. The three-element reconstruction includes the data end, the system end, and the response end. Lock the pre-divided multi-party database in the information platform, and perform data encoding and virtual layer mapping on each data unit to serve as a data pool; The large language model is initialized according to the ternary classification pattern. The data pool is processed by three parallel threads to determine the ternary classification result. The ternary classification result is then classified by phase and a classification identifier code is generated. The three-phase data class is the first level, and the three-phase data class is the data class with the same frequency in the ternary classification result. The ternary classification result and the hierarchical identification code are combined to form a classification and grading result, and a mapping between the classification and grading result and the multi-party database is established. The process of performing ternary reconstruction of classification-guided terms to determine the ternary classification pattern includes: Identify the classification-guided terms and determine their semantic information through interpretation; The derivative engine is triggered to determine the first classification mode by performing a transformation based on automatic data-side hierarchical rules for the semantic information of the term, the second classification mode by performing an architecture transformation based on the graph system, and the third classification mode by using mutual information response triggered by event elements. Integrate the first classification mode, the second classification mode, and the third classification mode to determine the ternary classification mode; The process of performing data encoding and virtual layer mapping on each data unit as a data pool includes: Identify the multiple databases, wherein each database is identified by an encoding based on its storage location and access path; A data encoding method is introduced, and the data unit encoding is performed by traversing the multiple databases and integrating them into an encoded database; A virtual layer is established based on the encoded database, and the virtual layer is used as a data pool, wherein the virtual layer has a mapping relationship with the multi-party database.
2. The data classification and grading method based on a large language model as described in claim 1, characterized in that, The data encoding method uses a multidimensional feature vector generated for each data unit, wherein the multidimensional feature vector includes at least a semantic feature vector, a structural feature vector, and a usage pattern vector.
3. The data classification and grading method based on a large language model as described in claim 1, characterized in that, Initialize the large language model according to the aforementioned ternary classification pattern, including: Based on the ternary classification model, the large language model is subjected to transfer invocation and partitioned lightweight training to determine the initialized large language model; For the data pool, perform model-parallel classification processing using the initialized large language model to determine the ternary classification result.
4. The data classification and grading method based on a large language model as described in claim 3, characterized in that, The large language model is subjected to transfer learning and lightweight training by partitioning, including: The migration invokes the large language model, performs model processing domain partitioning at the classification mode scale, and determines the partition processing domain; By using the one-to-one correspondence between the ternary classification pattern and the partition processing domain, classification pattern writing and data-driven lightweight supervised training are performed to determine the initialized large language model.
5. The data classification and grading method based on a large language model as described in claim 4, characterized in that, The data pool performs patterned classification processing in parallel using three threads to determine the ternary classification result, including: For the data pool, each partition processing domain performs pattern classification processing in parallel using three threads to determine the ternary classification result; The class associations are performed on the ternary classification results based on mutual information relationships.
6. The data classification and grading method based on a large language model as described in claim 5, characterized in that, By performing class association identification, the three-phase data classes in the ternary classification results are determined, and the first data level is determined; Identify the biphasic data class in the ternary classification results and determine the second data level; Determine the three-phase data classes in the ternary classification results and determine the third data level; A hierarchical identification code is generated based on the first data level, the second data level, and the third data level.
7. The data classification and grading method based on a large language model as described in claim 1, characterized in that, Determine whether there is a need for targeted classification and grading; If so, the classification and grading results are fine-tuned according to the specified targeted classification and grading requirements.
8. A data classification and grading system based on a large language model, characterized in that, The system for implementing the data classification and grading method based on a large language model according to any one of claims 1-7 includes: The ternary classification mode determination module is used to upload classification-guided terms to the information platform, and perform ternary reconstruction of the classification-guided terms according to the derivative engine to determine the ternary classification mode. The ternary reconstruction includes the data end, the system end and the response end. The data pool construction module is used to lock the pre-divided multi-party database in the information platform, and perform data encoding and virtual layer mapping on each data unit to serve as a data pool. The hierarchical identification code construction module is used to initialize the large language model according to the ternary classification mode, perform three-thread parallel pattern classification processing on the data pool, determine the ternary classification result, perform phase classification on the ternary classification result, and generate hierarchical identification codes, wherein the three-phase data class is the first level, and the three-phase data class is the data classification class with the same frequency in the ternary classification result; The classification and grading result acquisition module is used to fuse the ternary classification result with the grading identifier code as the classification and grading result, and to establish a mapping between the classification and grading result and the multi-party database.
Citation Information
Patent Citations
Public data automatic classification and grading method and system
CN119128607A
Text classification method, device and system based on large language model, storage medium and product
CN119647408A