Text governance method and device, equipment and medium
By employing a dual-mechanism approach to collect and construct knowledge graphs, the problem of low accuracy in text governance is addressed. This approach enables intelligent processing of external rule files and precise updates of internal rules, adapting to multiple business scenarios and improving the efficiency and accuracy of text governance.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- CHINA MERCHANTS FINANCE HLDG CO LTD
- Filing Date
- 2025-12-31
- Publication Date
- 2026-04-17
AI Technical Summary
Existing technologies have low accuracy in text governance, cannot achieve precise matching, real-time response and cross-scenario adaptation, rely on manual operation which is prone to errors, are difficult to build knowledge graphs, and have a lag in response to external rule changes.
External rule files are collected in real time through a dual mechanism of scheduled tasks and event-triggered tasks. Files that do not conform to the category are filtered out and labeled with the level. A knowledge graph is constructed, the correlation coefficient between business tags and data summaries is calculated, and changes in external rules are monitored and internal rule files are updated accurately.
It enables efficient processing of external rule files and accurate matching of internal rules, improving the accuracy and response speed of text governance, adapting to multiple business scenarios, and reducing compliance risks.
Smart Images

Figure CN121880283A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data processing technology, and in particular to a text processing method, apparatus, device, and medium. Background Technology
[0002] Currently, external compliance documents are characterized by their massive volume, multiple sources, and dynamic nature. Group enterprises, with their businesses spanning multiple industries and regions, urgently need to effectively integrate external documents with internal regulations, and have an urgent need for precise matching, real-time response, and cross-scenario adaptation in compliance management. Therefore, to meet the needs of group enterprises for precise adaptation, real-time response, and cross-scenario expansion in compliance management, intelligent innovation is required in the processing of external compliance documents and the mechanism for linking internal and external regulations to improve the accuracy of text governance.
[0003] Existing technologies collect external rule files through manual retrieval or passive notification, lacking both scheduled and event-triggered mechanisms, resulting in poor timeliness. They rely on manual filtering of files that do not fit the category and labeling them by file level, which is cumbersome, error-prone, and lacks consistency in labeling. External rule files are manually stored in an internal database, lacking a standardized process. Relying on manual experience to judge the correlation between business tags and files to match business lines lacks quantitative calculation, leading to significant matching bias. Without a knowledge graph, relying solely on manual recording of relationships makes traceability difficult. After changes to external rules, manual identification of corresponding internal rule files and related business lines is slow and prone to omissions; updates cannot accurately filter files requiring updating based on file level, resulting in low accuracy in text governance. Summary of the Invention
[0004] This invention provides a text processing method, apparatus, device, and medium to solve the problem of low accuracy in text processing.
[0005] Firstly, a text governance method is provided, including: The system uses a dual mechanism of pre-set timed tasks and event-triggered tasks to collect external rule files of the pre-set enterprise in real time. Files that do not conform to the preset file category in the external rule files are filtered out, and the file level of each filtered external rule file is marked according to the preset level label to obtain the file level of each external rule file; The external rule files are stored in the internal database of the preset enterprise according to the file level to obtain internal rule files; Extract the data digest of each internal rule file, calculate the correlation coefficient between the business tags of different business lines within the preset enterprise and the data digest, and determine the business line corresponding to each internal rule file based on the correlation coefficient; A knowledge graph is constructed using each internal rule file, its corresponding data summary, business tag, and file level as nodes, and the aforementioned correlation coefficient as edges. When a change is detected in the external rule file, the changed content of the external rule file is collected, and the target file corresponding to the changed external rule file in the internal rule file is identified. Business tags with a correlation coefficient greater than a preset threshold between the data digest of the target file are selected to obtain a set of associated tags. Collect the internal rule files of the business line corresponding to each business tag within the associated tag set; The file level of the external rule file that has changed is taken as the change level. Based on the change content, the internal rule files corresponding to all business tags in the associated tag set whose file level is lower than the change level are updated.
[0006] Secondly, a text governance device is provided, comprising: The external rule file acquisition module is used to acquire the external rule files of a preset enterprise in real time through a dual mechanism of preset timed scheduling tasks and event-triggered tasks; The file level labeling module is used to filter out files in the external rule files that do not conform to the preset file category, and to label the file level of each filtered external rule file according to the preset level label to obtain the file level of each external rule file. The external rule file storage module is used to store the external rule file into the internal database of the preset enterprise according to the file level, so as to obtain the internal rule file. The correlation coefficient calculation module is used to extract the data summary of each internal rule file, calculate the correlation coefficient between the business tags of different business lines within the preset enterprise and the data summary, and determine the business line corresponding to each internal rule file based on the correlation coefficient. The knowledge graph construction module is used to construct a knowledge graph with each internal rule file, the data summary corresponding to each internal rule file, the business tag and the file level as nodes, and the correlation coefficient as edges. The associated tag set determination module is used to collect the changed content of the external rule file when the external rule file is detected to have changed, identify the target file in the internal rule file that corresponds to the changed external rule file, select business tags with a correlation coefficient greater than a preset threshold between the data digest of the target file, and obtain an associated tag set. The internal rule file acquisition module is used to acquire the internal rule files of the business line corresponding to each business tag in the associated tag set; The internal rule file update module is used to take the file level of the changed external rule file as the change level, and update the internal rule files corresponding to all business tags in the associated tag set whose file level is lower than the change level according to the changed content.
[0007] Thirdly, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the above-described text governance method.
[0008] Fourthly, a computer-readable storage medium is provided, which stores a computer program that, when executed by a processor, implements the steps of the above-described text governance method.
[0009] The solutions implemented by the aforementioned text governance methods, devices, equipment, and media can collect external rule files of enterprises through the client, intelligently filter out non-compliant files and mark them at different levels for storage, calculate the correlation coefficient between business tags and data summaries to match business lines, construct a knowledge graph, accurately locate target files and related business lines when external rules change, update internal rules according to level, integrate automated processing, intelligent labeling, and dynamic cross-referencing innovation, significantly improve the efficiency and accuracy of external rule processing, accelerate the response speed of internal rule updates, adapt to multiple business scenarios, reduce compliance risks, and solve the problem of low accuracy in text governance. Attached Figure Description
[0010] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the description of the embodiments of the present invention will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0011] Figure 1 This is a schematic diagram of an application environment for a text governance method according to an embodiment of the present invention; Figure 2 This is a flowchart illustrating a text governance method in one embodiment of the present invention; Figure 3 This is a schematic diagram of the structure of a text processing device in one embodiment of the present invention; Figure 4 This is a schematic diagram of the structure of a computer device according to an embodiment of the present invention; Figure 5 This is another structural schematic diagram of a computer device according to one embodiment of the present invention. Detailed Implementation
[0012] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0013] The text processing method provided in this embodiment of the invention can be applied to, for example, Figure 1 In this application environment, the client communicates with the server via a network. The server can collect external rule files from the enterprise through the client, intelligently filter out non-compliant files, label them with levels and store them in the database, calculate the correlation coefficient between business tags and data summaries to match business lines, construct a knowledge graph, accurately locate target files and related business lines when external rules change, update internal rules according to levels, and integrate automated processing, intelligent labeling, and dynamic cross-referencing innovation to significantly improve the efficiency and accuracy of external rule processing, accelerate the response speed of internal rule updates, adapt to multiple business scenarios, reduce compliance risks, and solve the problem of low accuracy in text governance. The client can be, but is not limited to, various personal computers, laptops, smartphones, tablets, and portable wearable devices. The server can be implemented using a standalone server or a server cluster consisting of multiple servers. The following detailed description of the invention uses specific embodiments.
[0014] Please see Figure 2 As shown, Figure 2 A flowchart illustrating the text governance method provided in this embodiment of the invention includes the following steps: S1. Real-time collection of external rule files of the preset enterprise through a dual mechanism of preset timed scheduling tasks and event-triggered tasks.
[0015] In this embodiment of the invention, the preset timed scheduling task is a task that retrieves relevant data incrementally at fixed time intervals; the event-triggered task is a task that retrieves data in real time when the relevant data source pushes an update notification.
[0016] In detail, the external rule document is a document formed by external regulations related to the enterprise's industry or business that the enterprise must comply with, such as industry norms and industry standards.
[0017] When it is necessary to collect external rule files such as text and numbers of preset enterprises (such as enterprises, groups, etc.), the external rule files can be collected on a scheduled basis through preset timed tasks, and / or triggered by event-triggered tasks to collect the external rule files in a triggered manner, so as to improve the efficiency and accuracy of data collection.
[0018] In this embodiment of the invention, the real-time collection of external rule files of a preset enterprise through a dual mechanism of preset timed scheduling tasks and event-triggered tasks includes: Identify the execution cycle of a preset timed scheduling task, and periodically collect data from the preset enterprise according to the execution cycle to obtain periodic data fragments; Identify the triggering conditions of the preset event triggering task, use the triggering conditions to listen for the target event of the preset enterprise, and generate the event triggering data fragment of the preset enterprise when the target event is detected; The periodic data segments and the event-triggered data segments are combined into an external rule file for the preset enterprise.
[0019] In detail, the execution cycle is usually a fixed time interval that can be set. The background function of the scheduled task can be called by a preset SQL statement, and the execution cycle of the scheduled task can be determined according to the background function. The periodic data segment can be file data that is collected and pre-processed according to a set period.
[0020] In this embodiment of the invention, an automated data collection process can be initiated at fixed time intervals according to the scheduled task. The automated data collection process connects to the data provider's interface to automatically and incrementally retrieve relevant external rule files. During the retrieval process, metadata can be extracted using regular expressions, and structured cleaning operations such as format unification and redundancy removal can be performed simultaneously. At the same time, unnecessary files (such as files that are not in the target format) in the retrieved relevant external rule files are filtered out, and finally, file data that meets the requirements is obtained.
[0021] Specifically, the triggering condition can be an update push notification from an external rule file provider; the target event can be a pre-defined change in an external rule file, such as addition, revision, or repeal; and the event-triggered data fragment can be the change data and related information of the corresponding external rule file collected after the target event is detected. The system interfaces with compliance-related data provision channels, continuously monitoring the status of external rule files based on the set update push notification triggering conditions. Once a change (addition, revision, or repeal) is detected, the data collection process is immediately initiated, simultaneously performing a structured cleaning operation with standardized formatting and redundancy removal. Key metadata is extracted using regular expressions, ultimately generating an event-triggered data fragment containing the changed content and related attributes.
[0022] Furthermore, duplicate data can be deduplicated for both periodic and event-triggered data fragments, key metadata of each type of data can be extracted using regular expressions, and the data format can be standardized to improve the quality of file data.
[0023] S2. Filter out files in the external rule files that do not conform to the preset file category, and label the file level of each filtered external rule file according to the preset level label to obtain the file level of each external rule file.
[0024] In this embodiment of the invention, when the preset document category is a policy document, the corresponding documents to be filtered out are non-policy documents (e.g., if the collected external rule documents are multiple external rule documents of Unit A, and the preset document category is a policy document used to record the management regulations of Unit A, then the corresponding documents to be filtered out are non-policy documents in the external rule documents that record the position information of Unit A or record the annual budget of Unit A).
[0025] This invention can identify and determine the file category of each file in the external rule file through preset text processing models such as BERT classification model and transformer model. The file can be related to the unit's compliance management and has standardized guidance attributes, which can adapt to the compliance requirements of multiple business scenarios of the unit and exclude redundant files that are not standardized guidance.
[0026] In this embodiment of the invention, when the external rule file is a policy document within the compliance management scenario of a group unit, the preset hierarchical tags can include a four-dimensional tag system of "hierarchy-industry-business-clause type". The file level is determined by a dual-track tagging mechanism of "rule template + model correction". First, the hierarchical level and applicable priority of the file are initially determined based on established rules (such as file name format and numbering rules). Then, the ambiguity is corrected by an optimized intelligent model. At the same time, the specific clauses in the file are broken down, and key clauses such as "must comply" and "prohibited operation" are accurately identified by a specialized intelligent model. Key information such as the responsible party and behavioral requirements are extracted, and finally, a file with multi-dimensional tags is generated.
[0027] In detail, the process involves extracting the file ID and file signature from each filtered external rule file, concatenating the file ID and file signature, and then converting the concatenated content into a text vector corresponding to each external rule file. A preset set of hierarchical labels is obtained, and each external rule file is selected as a target file. The matching degree between the text vector of the target file and each hierarchical label in the hierarchical label set is calculated, and the hierarchical label with the highest matching degree is selected as the file hierarchical label of the target file. A preset text recognition model is used to extract the text features of the target file, and the effectiveness priority of the target file is queried from a preset effectiveness priority table based on the text features. The file hierarchical label and effectiveness priority are used as file levels to mark the target files until all files in the external rule files have been marked.
[0028] In this embodiment of the invention, the step of labeling the file level of each filtered external rule file according to a preset level label to obtain the file level of each external rule file includes: Extract the file number and file signature contained in each external rule file after filtering, and concatenate the file number and file signature to convert them into a text vector corresponding to each external rule file; Obtain a preset set of hierarchical tags, select any external rule file from the external rule files one by one as the target file, calculate the matching degree between the text vector of the target file and each hierarchical tag in the set of hierarchical tags, and select the hierarchical tag with the highest matching degree as the file hierarchical tag of the target file; The text features of the target file are extracted using a preset text recognition model, and the effectiveness priority of the target file is queried from a preset effectiveness priority table based on the text features. The target file is marked using the file level label and the effectiveness priority as the file level until all files in the external rule file have been marked.
[0029] In detail, the file number can be a unique identifier for the external rule file; the file signature can be an identifier for the publishing entity of the external rule file; regular expressions are used to parse the filtered external rule files to accurately locate the text segments corresponding to the file number and the file signature, so as to extract the two types of information. Furthermore, the concatenated text can be semantically encoded using a fine-tuned BERT model to convert the concatenated file number and the file signature into a text vector corresponding to each external rule file.
[0030] In this embodiment of the invention, the preset hierarchical tag set can be tags uploaded by the user in advance to represent the hierarchy, industry, business scenario, and clause type of the external rule file. Each tag in the hierarchical tag set is in vector form. The matching degree between the text vector of the target file and each hierarchical tag in the hierarchical tag set can be calculated using the Sentence-BERT model, as shown in the following formula: Where A represents the text vector of the target file, B represents a single hierarchical tag vector within the hierarchical tag set, A・B is the dot product of vector A and vector B, ||A|| is the L2 norm of vector A, and ||B|| is the L2 norm of vector B. The semantic matching degree between the two can be directly obtained through this formula, with a value range of [-1, 1]. The closer the value is to 1, the stronger the fit between the target file and the hierarchical tag, thus accurately selecting the hierarchical tag with the highest matching degree.
[0031] In this embodiment of the invention, the text recognition model can be a BERT-BiLSTM-CRF model. The text recognition model is used to perform clause segmentation, identification of key clauses such as mandatory / prohibited clauses, and extraction of core elements such as responsible parties and behavioral norms in the target document to obtain a data summary of the target document. Then, the text summary is converted into a summary vector, and the effectiveness priority of the summary vector in a preset effectiveness priority table is queried. The effectiveness priority table is pre-constructed based on the hierarchical features, issuing entity features, and applicable scope features of the external rule document, and contains a mapping table of different feature combinations and corresponding effectiveness priority levels (such as priority ordering from level 1 to level 5).
[0032] In this embodiment of the invention, in addition to marking the target file using the file hierarchy tags and the effectiveness priority, industry tags and business tags of the target file can also be determined based on the industry and related business of the target file, and the target file can be further marked using industry tags and business tags to achieve multi-type file marking. In detail, the determined file level labels and the effectiveness priority obtained after extracting text features through the BERT-BiLSTM-CRF model are integrated into complete file level information. This file level information is then bound to the unique identifier of the target file to complete the labeling operation of the target file. Then, unlabeled files in the external rule file are selected as new target files in turn, and the above integration, binding, and labeling operations are repeated until all files in the external rule file are labeled.
[0033] Specifically, by extracting the file number and file signature of the filtered external rule files and concatenating them into a text vector, the matching degree is calculated with a preset set of hierarchical labels to determine the file hierarchical label. The text recognition model is used to extract the text features of the target file and query the validity priority. Finally, the file hierarchical label and validity priority are used as the file level to complete the labeling of all external rule files. This realizes the automation and accuracy of external rule file level labeling, greatly reduces the time spent on manual operation, improves the labeling efficiency and accuracy, and effectively reduces the omission rate and error rate of external rule level and validity priority determination. This lays a solid foundation for subsequent dynamic reconciliation of internal and external rules and compliance risk warning, and adapts to the multi-scenario compliance management needs of group enterprises.
[0034] S3. Store the external rule file into the internal database of the preset enterprise according to the file level to obtain the internal rule file.
[0035] In this embodiment of the invention, the internal database can be a dedicated data storage system that stores external related file data after automated collection, cleaning, classification and multi-level tagging. It typically has version traceability, metadata management and related retrieval functions. The internal rule file can be a standardized data file with multi-dimensional tags and validity priority information formed by storing external related files in the data storage system after the above-mentioned full process and classification according to the corresponding level. It can typically establish dynamic association with the enterprise's internal management terms and support traceability and iteration.
[0036] In this embodiment of the invention, storing the external rule file into the internal database of the preset enterprise according to the file level to obtain the internal rule file includes: The external rule files are classified and archived according to the file level to obtain hierarchical external rule files; Obtain the internal database storage specifications of the preset enterprise, and allocate corresponding storage paths to the hierarchical external rule files according to the internal database storage specifications; The hierarchical external rule files are stored in the internal database of the preset enterprise according to the storage path, and a file-level association index of the preset enterprise is generated. The hierarchical external rule files are validated for entry into the database based on the file-level association index to obtain internal rule files.
[0037] In detail, the graded external rule files can be external related files that have undergone hierarchical determination, are marked with applicable priority, and are accompanied by four-dimensional tags of level, industry, business, and clause type. During execution, the system first uses a preset rule template and the relevant document identification format to initially screen the level and effectiveness of external rule files. Then, a fine-tuned BERT model is used to correct ambiguous scenarios and improve the labeling accuracy. At the clause level, an entity augmentation dictionary is introduced based on the BERT-BiLSTM-CRF model to accurately identify key clauses and extract core elements. At the same time, the system automatically performs effectiveness-related level sorting and conflict screening to complete the classification and archiving, resulting in graded files.
[0038] Specifically, internal database storage specifications can be enterprise / group-defined storage rules based on the hierarchical levels, four-dimensional tags, and metadata information of relevant external documents. These rules include file classification storage requirements and retrieval and traceability standards. Storage paths can be folder paths divided according to hierarchical level, industry tag, business, and clause type, directly linking to the file's four-dimensional tags and metadata for convenient and rapid retrieval and version traceability. The system retrieves the enterprise / group's pre-set internal database storage specifications, extracts the hierarchical levels, industry-business-clause type four-dimensional tags, and metadata of the hierarchical external rule files, and precisely matches these file attributes with the dimensions of the storage path. According to the classification requirements in the storage specifications, corresponding hierarchical folders and tag subfolders are assigned to each hierarchical file, completing the storage path allocation and ensuring that file storage complies with the specifications and supports subsequent traceability and association.
[0039] Furthermore, the file-level association index can be index data that records the hierarchical level, four-dimensional tags, storage path, metadata, and corresponding internal related clauses of hierarchically related external files. This supports rapid retrieval of files and their associated business scenarios and internal clauses. Following the allocated storage path, hierarchically related external files are uploaded to the corresponding folders in the enterprise / group internal database. The system automatically extracts the file's hierarchical level, industry-business-clause type four-dimensional tags, and metadata, matches the file with its corresponding business tags and internal related clauses, and records this association information in a structured manner to generate a file-level association index. This allows for rapid location of files and related content through the index.
[0040] Furthermore, the file-level association index is retrieved to extract the hierarchical level, four-dimensional tags, metadata, and association information of the hierarchical external related files. This information is then compared one by one with the content of the index records to verify whether the file hierarchical labeling is accurate, the four-dimensional tags are complete, and the metadata is standardized. At the same time, the association between the file and business tags and internal related clauses is verified to conform to the preset rules. After confirming that there are no information conflicts and that it fully complies with the storage specifications, the verification is deemed to have passed, and finally, an internal rule file is generated.
[0041] S4. Extract the data digest of each internal rule file, calculate the correlation coefficient between the business tags of different business lines within the preset enterprise and the data digest, and determine the business line corresponding to each internal rule file based on the correlation coefficient.
[0042] In this embodiment of the invention, the data summary typically includes the unique identifier of the internal rule, the relevant identifier of the external rule referenced, the corresponding business module information, and the revision time record. It may also cover core information such as associated business tags and adapted business scenarios, which can accurately reflect the key attributes and relationships of the internal rule and provide basic data support for subsequent dynamic reconciliation and linkage of internal and external rules.
[0043] In detail, enterprises / groups can first preset the core attribute items to be extracted from internal rule files, then capture key information such as exclusive identifiers, referenced external rule identifiers, corresponding business modules, and revision times from internal rule files, record the extraction process and results with a full-process operation log, and finally integrate these captured key information with the generated business tags to form a data summary for each internal rule file.
[0044] In this embodiment of the invention, the business lines can be different business areas or business segments divided within an enterprise, covering specific business categories in multiple industries and scenarios; the business tags are usually generated in combination with enterprise characteristics, including identifiers in dimensions such as industry, region, and specific business scenarios, and can be customized by the enterprise or its subsidiaries according to business needs; the correlation coefficient can be obtained through semantic vector calculation and other technical means, and is a numerical indicator used to quantify the degree of matching between business tags and data summaries in terms of applicable scenarios and semantic association.
[0045] In this embodiment of the invention, calculating the correlation coefficient between the business tags of different business lines within the preset enterprise and the data digest includes: Extract the business tag feature vector of the business tag and the data digest feature vector of the data digest, respectively; The similarity correlation value between the business tag feature vector and the data summary feature vector is calculated using a preset similarity algorithm; Identify the business weight coefficients corresponding to the business tags of different business lines within the preset enterprise according to preset business priority rules; The similarity correlation value and the business weight coefficient are weighted and calculated to obtain the correlation coefficient between the business tag and the data summary.
[0046] In detail, business tag feature vectors are typically numerical vectors transformed from the industry, region, and business scenario dimensions contained in the business tags after semantic processing; data summary feature vectors can be numerical vectors transformed from key information such as unique identifiers, external reference identifiers, and business modules in the data summary after semantic encoding. Enterprises use the Sentence-BERT model to semantically process the dimensional information of business tags and generate business tag feature vectors, and simultaneously use the same model to semantically encode the key information in the data summary and generate data summary feature vectors, thus extracting both types of vectors through semantic processing and encoding techniques.
[0047] Specifically, the preset similarity algorithm can be a sentence embedding model (Sentence-BERT) based on the Siamese BERT network; the similarity association value is usually a value that quantifies the degree of semantic matching between the business tag feature vector and the data summary feature vector. Enterprises can call the Sentence-BERT model, input the extracted business tag feature vector and data summary feature vector into the model, obtain the semantic vector similarity result of the two through model calculation, and directly output the corresponding similarity association value to complete the quantitative calculation of the similarity between vectors.
[0048] Furthermore, the pre-defined business priority rules can be formulated by the enterprise based on its own business characteristics and compliance requirements. These rules are used to determine the importance of different businesses and cover dimensions such as business scenario adaptability and compliance risk correlation. The business weight coefficient is usually calculated by an algorithm, which quantifies the numerical indicators of the importance of different business tags. The enterprise first clarifies the pre-defined business priority rules based on its own business characteristics and compliance requirements, extracts relevant information such as enterprise characteristics and business scenarios corresponding to the business tags of each business line, calls the meta-learning algorithm, inputs the extracted information into the algorithm, and generates differentiated business weight coefficients through algorithm calculation, thus completing the identification of the business weight coefficients corresponding to different business tags.
[0049] Furthermore, enterprises can retrieve the similarity correlation values calculated by Sentence-BERT and the business weight coefficients generated by the meta-learning algorithm, clarify the calculation rules for weighted operations, and perform weighted calculations on the two types of values through the system's built-in calculation function to obtain the correlation coefficient between the business tag and the data summary.
[0050] In this embodiment of the invention, the enterprise can retrieve the correlation coefficient corresponding to each internal rule file, compare the coefficient with the judgment criteria, filter out the business tags with the highest matching degree, and lock the business line corresponding to each internal rule file based on the correspondence between the business tags and the business lines.
[0051] S5. Construct a knowledge graph using each internal rule file, the corresponding data summary, business tag, and file level as nodes, and the correlation coefficient as edges.
[0052] In this embodiment of the invention, the core information of the enterprise's internal rule files is first extracted to generate corresponding data summaries. Business tags are extracted based on features such as the enterprise's registered industry and organizational structure. The internal rule files are labeled with file levels and priorities according to a preset rule base. The internal rule files, corresponding data summaries, business tags, and file levels are established as independent nodes of the knowledge graph and assigned their respective attributes (such as the ID and revision time of the internal rule files, key clause information of the data summaries, etc.). The correlation coefficient between the internal rule files and business tags is calculated through a tag matching mechanism. The semantic similarity between the internal rule files and corresponding data summaries is calculated using the Sentence-BERT model as the correlation coefficient. The correlation coefficient weights between the internal rule files and other nodes are determined according to the file level priority rules. The obtained correlation coefficients are used as edges connecting the corresponding nodes to build a complete knowledge graph.
[0053] In detail, building a knowledge graph enables an enterprise to form a clear network of relationships among its internal rule documents, corresponding data summaries, business tags, and document levels. By leveraging the correlation coefficient, it can quickly pinpoint compliance requirements corresponding to business scenarios, accurately identify potential compliance risks, reduce the workload and errors of manual retrieval and sorting, and significantly improve compliance management efficiency. It also supports dynamic tracking of changes in the relationships between rules, ensuring that compliance decisions have clear and traceable basis. Furthermore, it can focus on core compliance priorities based on document level priorities, avoiding redundant work and helping enterprises achieve intelligent, precise, and dynamic optimization of compliance management, thereby reducing compliance management costs.
[0054] When a change is detected in the external rule file, S6 is executed to collect the changed content of the external rule file, and the target file corresponding to the changed external rule file in the internal rule file is identified from the knowledge graph. Business tags with a correlation coefficient greater than a preset threshold between the data digest of the target file are selected to obtain a set of associated tags.
[0055] In this embodiment of the invention, the changes may be the addition of new entries, modification or deletion of existing entries in the rule file. They typically include adjustments to key entries that are restrictive or prohibitive, changes related to applicable priority, and updates to core elements that are adapted to the enterprise's industry and business scenarios. They also cover relevant modification information that affects the scope of application and effectiveness of the rule file.
[0056] In detail, enterprises / groups acquire external rule file change information through multi-source data interfaces via a dual mechanism of scheduled incremental fetching and event-triggered update notification reception. They utilize a BERT classification model with industry feature embedding layers to accurately filter irrelevant files, extract metadata using regular expressions and construct a version traceability chain, break down the smallest clause units of the changed files using regular expression templates and semantic segmentation models, identify key change clauses and core elements such as mandatory and prohibitive clauses using NER models, retrieve related internal rule reference points based on internal and external rule knowledge graphs with spatiotemporal attributes, filter valid change clauses based on priority logic, and push revision reminders through system messages, emails, and SMS, attaching screenshots of the changed clauses, the chapters to be revised, and explanations of applicable priorities. Simultaneously, they record a full-process operation log to ensure traceability, completing the collection and coordinated processing of change content.
[0057] In this embodiment of the invention, the target file may be an internal rule file that precisely matches the changed external rule file in terms of industry or business tags, or that reaches a similarity threshold through semantic vector calculation. It is usually an internal compliance management file that contains content related to the changed external rule, is adapted to the specific business scenario of the enterprise, and is associated with the key clauses of the change, such as the obligation and prohibition, and needs to be revised based on priority logic to respond to the change.
[0058] In this embodiment of the invention, identifying the target file corresponding to the changed external rule file in the internal rule file within the knowledge graph includes: Match the internal rule file nodes corresponding to the changed external rule files within the knowledge graph; Identify the correlation coefficient corresponding to each edge in the internal rule file node, and calculate the correlation strength of the internal rule file node based on the correlation coefficient; The internal rule file corresponding to the internal rule file node whose association strength meets the preset strength threshold is determined as the target file.
[0059] In detail, internal rule file nodes can be nodes containing attributes such as their own identifier, associated external rule identifiers, the business module to which they belong, and the revision time. They are typically associated with tags such as industry, region, and business scenario. Enterprises / groups first construct an internal and external rule knowledge graph with spatiotemporal attributes, defining the core relationships between internal rule file nodes and external rule files and business tags. Basic relationships are established through precise matching using industry or business tags. Sentence-BERT is used to calculate the semantic vector similarity of clauses to supplement the associated scenarios. Internal rule file nodes associated with changed external rule files are retrieved from the knowledge graph. Valid associated nodes are then filtered based on priority logic to complete the matching operation.
[0060] Specifically, the association strength can be a quantitative value reflecting the degree of close association between internal rule file nodes and changed external rule files, based on the degree of fit of precise matching of comprehensive industry and business tags, as well as the similarity of clause semantic vectors calculated by Sentence-BERT.
[0061] Furthermore, the preset strength threshold can be a quantitative standard set by the enterprise / group based on compliance management needs, semantic association accuracy requirements, and historical association effects. It usually refers to the matching degree of industry and business tags, the similarity of clause semantic vectors calculated by Sentence-BERT, and takes into account risk control needs to determine specific values. The enterprise / group first sets the preset strength threshold according to its own business scenario, obtains the association coefficient of each edge through precise matching of industry or business tags and Sentence-BERT calculation of clause semantic vector similarity, integrates these coefficients according to preset weights to obtain the association strength of each internal rule file node, compares the association strength of each node with the preset strength threshold, and filters out internal rule file nodes whose association strength reaches or exceeds the threshold. The internal rule files corresponding to these nodes are determined as target files, and the entire comparison and filtering process and results are recorded to support compliance audit traceability.
[0062] In this embodiment of the invention, the preset threshold can be a quantitative standard set by the enterprise in combination with business adaptation needs, semantic association accuracy requirements and historical matching results. It is usually determined by referring to the matching fit between industry or business tags and target file data summary, and the semantic vector similarity calculated by Sentence-BERT. The associated tag set can be a combination of tags that include dimensions such as industry, region, and business scenario, and whose correlation coefficient with target file data summary reaches the preset threshold. It is usually a summary of various tags that can accurately adapt to the business scenario to which the target file belongs and reflect its core adaptation needs.
[0063] In detail, enterprises extract data summaries from target files and obtain basic correlation coefficients through precise matching of industry or business tags. They then use Sentence-BERT to calculate the semantic vector similarity between the data summaries and each business tag to supplement the correlation coefficients. The coefficients obtained from the two methods are integrated, and the integrated correlation coefficients are compared with a preset threshold. Business tags with coefficients greater than the preset threshold are selected and then aggregated to obtain a set of related tags. The entire process relies on the business tag nodes and their relationships in a knowledge graph with spatiotemporal attributes.
[0064] S7. Collect the internal rule files of the business line corresponding to each business tag in the associated tag set.
[0065] In this embodiment of the invention, enterprises can rely on internal and external rule knowledge graphs containing spatiotemporal attributes to retrieve the corresponding business lines in the knowledge graph for each business tag in the associated tag set. Based on the "association" relationship between the business tags and internal rule files, the internal rule files corresponding to each business line are extracted. Core attributes such as the self-identifier, the business module to which the internal rule files belong, and the revision time of the internal rule files are collected simultaneously. The entire collection process and results are recorded to support compliance audit traceability.
[0066] In detail, we can first identify the specific business line corresponding to each business tag within the associated tag set (e.g., the "Finance-Credit Approval" tag corresponds to the credit business line, the "Transportation-Logistics" tag corresponds to the logistics business line, and the "Medical-Pharmaceutical Procurement" tag corresponds to the procurement business line). Relying on the internal and external rule knowledge graph with spatiotemporal attributes, based on the "association" between business tags and internal rule documents and the business module attributes to which the internal rule documents belong, we can accurately locate the internal rule documents corresponding to each business line (e.g., the credit business line corresponds to the credit compliance management system, the logistics business line corresponds to the logistics operation compliance specifications, and the procurement business line corresponds to the procurement compliance management methods, etc. These documents all contain compliance requirements and operational standards adapted to the corresponding business scenarios). We can collect metadata such as the self-identifier, the external rule identifiers referenced, the revision time, and the core clause content of these internal rule documents, and simultaneously record the collected business tags, corresponding business lines, internal rule documents, and metadata information to form a traceable collection log, ensuring that compliance audits are verifiable, and completing the accurate collection of internal rule documents for each business tag's corresponding business line.
[0067] Specifically, relying on a knowledge graph with spatiotemporal attributes, the system accurately matches the corresponding business line according to each business tag in the associated tag set, locates and collects internal rule files adapted to the business scenario, and simultaneously collects metadata such as the file's own identifier, core clauses, and revision time to form a traceable log. This collection method can significantly improve the accuracy and efficiency of internal rule file retrieval, reduce the redundant cost of manual screening, avoid missing key rule files that are strongly related to the business, and provide accurate data support for the subsequent verification, revision and optimization of internal rules and changed external rules. At the same time, it ensures that the collection process is auditable and traceable, strengthens the standardization and business adaptability of enterprise compliance management, and helps compliance requirements to be implemented quickly.
[0068] S8. Take the file level of the changed external rule file as the change level, and update the internal rule files corresponding to all business tags in the associated tag set whose file level is lower than the change level according to the changed content.
[0069] In this embodiment of the invention, the change level can be a level identifier determined based on the hierarchical attributes of the compliance-related external documents themselves. It usually combines the hierarchical characteristics of the documents with the logic of validity association to reflect the applicable priority of the documents in the compliance system. This provides a core judgment basis for the triggering priority of internal regulation revision warnings after changes in external regulations, and adapts to the differences in compliance impact brought about by changes in external documents at different levels.
[0070] In detail, enterprises can collect external compliance documents that have changed through a multi-source interface aggregation mechanism. They can use preset rule templates to initially screen the corresponding level and validity-related attributes of the documents based on features such as document name. Then, they can use a fine-tuned BERT model to correct fuzzy scenarios and improve accuracy. Subsequently, they can use a validity hierarchy rule library to automatically encode the screened level attributes, mark the corresponding applicable priority, and use the marked applicable priority as the change level.
[0071] In this embodiment of the invention, the internal rule file corresponding to the business tag whose file level is lower than the change level can be an internal compliance management related file of an enterprise that is associated with a specific industry, region or business scenario tag and whose external compliance file associated with it has a lower priority than the change level (the priority of the external change compliance file). Usually, due to changes in the content of the external compliance file, it is necessary to adjust it synchronously to conform to the compliance logic of "priority of higher-level rules".
[0072] In detail, "priority of higher-level rules" can be the core applicable logic in compliance management. It is usually reflected in the fact that compliance-related documents at different levels are sorted according to the preset priority of effectiveness. When there are conflicts or changes in the content of documents, the clauses of the document with higher priority of effectiveness shall be applied first. In the scenario of external regulation change linkage, the relevant internal compliance documents will also be revised first according to this logic, filtering out invalid reminders caused by conflicts of lower-level documents, and ensuring the accuracy of compliance risk prevention and control.
[0073] In this embodiment of the invention, updating the internal rule files corresponding to all business tags with file levels lower than the change level within the associated tag set according to the changed content includes: Filter out target business tags whose file level is lower than the change level from the set of associated tags; The internal rule file corresponding to the target business tag is identified as the file to be updated. Based on the changes, analyze the key points of the clause updates in the document to be updated, as well as the unadapted related clauses. The unadapted related clauses are updated according to the updated clause guidelines to obtain the updated internal rule file.
[0074] In detail, the target business tag can be an industry, region, or business scenario related tag that is associated with external change compliance documents and whose applicable priority is lower than the change level (the applicable priority of external change compliance documents). The enterprise retrieves the set of related tags through the internal and external regulations knowledge graph, determines the applicable priority of the documents corresponding to each business tag according to the preset validity hierarchy rule library, compares these priorities with the change level one by one, and filters out the tags with a priority lower than the change level, which are the target business tags.
[0075] Specifically, the files to be updated can be internal compliance management documents that are related to the selected target business tags and need to be adjusted synchronously due to external changes in compliance documents. Enterprises can retrieve the associated data of internal rule files corresponding to the target business tags through internal and external compliance knowledge graphs, and verify the binding validity of internal rule files and target business tags with the help of the "tag matching + semantic vector" dual association mechanism. Internal rule files that are effectively bound and have not adapted to external changes are identified as files to be updated.
[0076] Furthermore, key points for clause updates can be related to key clauses in external compliance documents, core elements (such as responsible parties and codes of conduct) and adaptation requirements that need to be adjusted in internal documents to be updated; unadapted related clauses can be clauses in the documents to be updated that are related to the changes in external regulations but do not cite the changed clauses or do not meet the requirements of the changes. Enterprises can use the BERT-BiLSTM-CRF model to break down the changes in external regulations into the smallest clause units, extract the core elements of mandatory / prohibitory key clauses, calculate the semantic vector similarity between the changed external regulations clauses and each clause in the documents to be updated using the Sentence-BERT model, combine the "external regulation-internal regulation" relationship in the internal and external regulation knowledge graph, screen out clauses in the documents to be updated that are related to the changes, compare the core elements and effectiveness requirements before and after the changes in external regulations, identify the key information to be adjusted as key points for clause updates, and mark clauses that do not match the core elements of the changes or do not meet the new effectiveness requirements as unadapted related clauses.
[0077] Furthermore, the updated internal rule document can be an internal compliance management document that incorporates key points of clause updates, corrects incompatible related clauses, and fully meets the requirements of external change compliance documents. Enterprises can access the revision suggestion template provided with the system, fill in the key points of clause updates (such as the adjustment requirements of core elements like responsible parties and behavioral norms) to the corresponding chapters where incompatible related clauses are located, use regular expression templates to accurately locate the clause modification positions, refer to the core elements of external regulations extracted by the BERT-BiLSTM-CRF model to calibrate the modified content, use Sentence-BERT to calculate the semantic vector similarity between the modified clause and the external regulation changes to ensure compliance, save the modified content and generate an updated internal rule document with a version traceability chain, and simultaneously record update operation logs to support compliance audit traceability.
[0078] As can be seen, the above solution involves collecting external rule documents from enterprises, intelligently filtering out non-compliant documents and labeling them with levels for storage, calculating the correlation coefficient between business tags and data summaries to match business lines, constructing a knowledge graph, accurately locating target documents and related business lines when external rules change, updating internal rules according to levels, and integrating automated processing, intelligent labeling, and dynamic cross-referencing innovations. This significantly improves the efficiency and accuracy of external rule processing, accelerates the response speed of internal rule updates, adapts to multiple business scenarios, reduces compliance risks, and can solve the problem of low accuracy in text governance.
[0079] It should be understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.
[0080] In one embodiment, a text processing device 100 is provided, which corresponds one-to-one with the text processing methods described in the above embodiments. For example... Figure 3 As shown, the text governance device 100 includes an external rule file acquisition module 101, a file-level annotation module 102, an external rule file storage module 103, a correlation coefficient calculation module 104, a knowledge graph construction module 105, a correlation tag set determination module 106, an internal rule file acquisition module 107, and an internal rule file update module 108. Detailed descriptions of each functional module are as follows: The external rule file acquisition module 101 is used to acquire the external rule files of a preset enterprise in real time through a dual mechanism of preset timed scheduling tasks and event-triggered tasks. The file level labeling module 102 is used to filter out files in the external rule files that do not conform to the preset file category, and to label the file level of each filtered external rule file according to the preset level label to obtain the file level of each external rule file. The external rule file storage module 103 is used to store the external rule file into the internal database of the preset enterprise according to the file level, so as to obtain the internal rule file. The correlation coefficient calculation module 104 is used to extract the data summary of each internal rule file, calculate the correlation coefficient between the business tags of different business lines within the preset enterprise and the data summary, and determine the business line corresponding to each internal rule file based on the correlation coefficient. The knowledge graph construction module 105 is used to construct a knowledge graph with each internal rule file, the data summary corresponding to each internal rule file, the business tag and the file level as nodes, and the correlation coefficient as edges. The associated tag set determination module 106 is used to collect the changed content of the external rule file when the external rule file is detected to have changed, identify the target file in the internal rule file that corresponds to the changed external rule file, select business tags with a correlation coefficient greater than a preset threshold between the data digest of the target file, and obtain an associated tag set. The internal rule file acquisition module 107 is used to acquire the internal rule file of the business line corresponding to each business tag in the associated tag set; The internal rule file update module 108 is used to take the file level of the changed external rule file as the change level, and update the internal rule files corresponding to all business tags in the associated tag set whose file level is lower than the change level according to the change content.
[0081] In one embodiment, the external rule file acquisition module 101, when executing a dual mechanism of real-time acquisition of external rule files of a preset enterprise through a preset timed scheduling task and an event-triggered task, is used to: Identify the execution cycle of a preset timed scheduling task, and periodically collect data from the preset enterprise according to the execution cycle to obtain periodic data fragments; Identify the triggering conditions of the preset event triggering task, use the triggering conditions to listen for the target event of the preset enterprise, and generate the event triggering data fragment of the preset enterprise when the target event is detected; The periodic data segments and the event-triggered data segments are combined into an external rule file for the preset enterprise.
[0082] In detail, the execution cycle is usually a fixed time interval that can be set. The background function of the scheduled task can be called by a preset SQL statement, and the execution cycle of the scheduled task can be determined according to the background function. The periodic data segment can be file data that is collected and pre-processed according to a set period.
[0083] In one embodiment, the file level labeling module 102, when performing the process of labeling the file level of each filtered external rule file according to a preset level label to obtain the file level of each external rule file, is used to: Extract the file number and file signature contained in each external rule file after filtering, and concatenate the file number and file signature to convert them into a text vector corresponding to each external rule file; Obtain a preset set of hierarchical tags, select any external rule file from the external rule files one by one as the target file, calculate the matching degree between the text vector of the target file and each hierarchical tag in the set of hierarchical tags, and select the hierarchical tag with the highest matching degree as the file hierarchical tag of the target file; The text features of the target file are extracted using a preset text recognition model, and the effectiveness priority of the target file is queried from a preset effectiveness priority table based on the text features. The target file is marked using the file level label and the effectiveness priority as the file level until all files in the external rule file have been marked.
[0084] In one embodiment, the external rule file storage module 103, when executing the process of storing the external rule file into the internal database of the preset enterprise according to the file level to obtain the internal rule file, is used to: The external rule files are classified and archived according to the file level to obtain hierarchical external rule files; Obtain the internal database storage specifications of the preset enterprise, and allocate corresponding storage paths to the hierarchical external rule files according to the internal database storage specifications; The hierarchical external rule files are stored in the internal database of the preset enterprise according to the storage path, and a file-level association index of the preset enterprise is generated. The hierarchical external rule files are validated for entry into the database based on the file-level association index to obtain internal rule files.
[0085] In one embodiment, the correlation coefficient calculation module 104, when calculating the correlation coefficient between the business tags of different business lines within the preset enterprise and the data digest, is used to: Extract the business tag feature vector of the business tag and the data digest feature vector of the data digest, respectively; The similarity correlation value between the business tag feature vector and the data summary feature vector is calculated using a preset similarity algorithm; Identify the business weight coefficients corresponding to the business tags of different business lines within the preset enterprise according to preset business priority rules; The similarity correlation value and the business weight coefficient are weighted and calculated to obtain the correlation coefficient between the business tag and the data summary.
[0086] In one embodiment, the association tag set determination module 106, when performing the task of identifying a target file corresponding to a modified external rule file from an internal rule file within the knowledge graph, is configured to: Match the internal rule file nodes corresponding to the changed external rule files within the knowledge graph; Identify the correlation coefficient corresponding to each edge in the internal rule file node, and calculate the correlation strength of the internal rule file node based on the correlation coefficient; The internal rule file corresponding to the internal rule file node whose association strength meets the preset strength threshold is determined as the target file.
[0087] In one embodiment, the internal rule file update module 108, when updating the internal rule files corresponding to all business tags in the associated tag set whose file level is lower than the change level according to the change content, is used to: Filter out target business tags whose file level is lower than the change level from the set of associated tags; The internal rule file corresponding to the target business tag is identified as the file to be updated. Based on the changes, analyze the key points of the clause updates in the document to be updated, as well as the unadapted related clauses. The unadapted related clauses are updated according to the updated clause guidelines to obtain the updated internal rule file.
[0088] This invention provides a text governance device that collects external rule files from enterprises, intelligently filters out non-compliant files and labels them with levels for storage, calculates the correlation coefficient between business tags and data summaries to match business lines, constructs a knowledge graph, accurately locates target files and related business lines when external rules change, and updates internal rules according to levels. It integrates automated processing, intelligent labeling, and dynamic cross-referencing innovation, significantly improving the efficiency and accuracy of external rule processing, accelerating the response speed of internal rule updates, adapting to multiple business scenarios, reducing compliance risks, and solving the problem of low accuracy in text governance.
[0089] For specific limitations regarding the text processing device, please refer to the limitations of the text processing method above, which will not be repeated here. Each module in the aforementioned text processing device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in the computer device in hardware form, or stored in the memory of the computer device in software form, so that the processor can call and execute the operations corresponding to each module.
[0090] In one embodiment, a computer device is provided, which may be a server, and its internal structure diagram may be as follows: Figure 4 As shown, the computer device includes a processor, memory, network interface, and database connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile and / or volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and database. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The network interface is used to communicate with external clients via a network connection. When the computer program is executed by the processor, it implements the functions or steps of a text governance method on the server side.
[0091] In one embodiment, a computer device is provided, which may be a client, and its internal structure diagram may be as follows: Figure 5 As shown, the computer device includes a processor, memory, network interface, display screen, and input devices connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The network interface is used to communicate with an external server via a network connection. When the computer program is executed by the processor, it implements the functions or steps of a text governance method on the client side.
[0092] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to perform the following steps: The system uses a dual mechanism of pre-set timed tasks and event-triggered tasks to collect external rule files of the pre-set enterprise in real time. Files that do not conform to the preset file category in the external rule files are filtered out, and the file level of each filtered external rule file is marked according to the preset level label to obtain the file level of each external rule file; The external rule files are stored in the internal database of the preset enterprise according to the file level to obtain internal rule files; Extract the data digest of each internal rule file, calculate the correlation coefficient between the business tags of different business lines within the preset enterprise and the data digest, and determine the business line corresponding to each internal rule file based on the correlation coefficient; A knowledge graph is constructed using each internal rule file, its corresponding data summary, business tag, and file level as nodes, and the aforementioned correlation coefficient as edges. When a change is detected in the external rule file, the changed content of the external rule file is collected, and the target file corresponding to the changed external rule file in the internal rule file is identified. Business tags with a correlation coefficient greater than a preset threshold between the data digest of the target file are selected to obtain a set of associated tags. Collect the internal rule files of the business line corresponding to each business tag within the associated tag set; The file level of the external rule file that has changed is taken as the change level. Based on the change content, the internal rule files corresponding to all business tags in the associated tag set whose file level is lower than the change level are updated.
[0093] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon, the computer program performing the following steps when executed by a processor: The system uses a dual mechanism of pre-set timed tasks and event-triggered tasks to collect external rule files of the pre-set enterprise in real time. Files that do not conform to the preset file category in the external rule files are filtered out, and the file level of each filtered external rule file is marked according to the preset level label to obtain the file level of each external rule file; The external rule files are stored in the internal database of the preset enterprise according to the file level to obtain internal rule files; Extract the data digest of each internal rule file, calculate the correlation coefficient between the business tags of different business lines within the preset enterprise and the data digest, and determine the business line corresponding to each internal rule file based on the correlation coefficient; A knowledge graph is constructed using each internal rule file, its corresponding data summary, business tag, and file level as nodes, and the aforementioned correlation coefficient as edges. When a change is detected in the external rule file, the changed content of the external rule file is collected, and the target file corresponding to the changed external rule file in the internal rule file is identified. Business tags with a correlation coefficient greater than a preset threshold between the data digest of the target file are selected to obtain a set of associated tags. Collect the internal rule files of the business line corresponding to each business tag within the associated tag set; The file level of the external rule file that has changed is taken as the change level. Based on the change content, the internal rule files corresponding to all business tags in the associated tag set whose file level is lower than the change level are updated.
[0094] It should be noted that the functions or steps that can be implemented by the computer-readable storage medium or computer device described above can be referred to the relevant descriptions on the server side and client side in the foregoing method embodiments. To avoid repetition, they will not be described one by one here.
[0095] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in a variety of forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), RAMbus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.
[0096] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is used as an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.
[0097] It should be noted that if any software tools or components not belonging to our company appear in the embodiments of this application, they are merely for illustrative purposes and do not represent actual use.
[0098] The above-described embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included within the protection scope of the present invention.
Claims
1. A text processing method, characterized in that, The method includes: The system uses a dual mechanism of pre-set timed tasks and event-triggered tasks to collect external rule files of the pre-set enterprise in real time. Files that do not conform to the preset file category in the external rule files are filtered out, and the file level of each filtered external rule file is marked according to the preset level label to obtain the file level of each external rule file; The external rule files are stored in the internal database of the preset enterprise according to the file level to obtain internal rule files; Extract the data digest of each internal rule file, calculate the correlation coefficient between the business tags of different business lines within the preset enterprise and the data digest, and determine the business line corresponding to each internal rule file based on the correlation coefficient; A knowledge graph is constructed using each internal rule file, its corresponding data summary, business tag, and file level as nodes, and the aforementioned correlation coefficient as edges. When a change is detected in the external rule file, the changed content of the external rule file is collected, and the target file corresponding to the changed external rule file in the internal rule file is identified. Business tags with a correlation coefficient greater than a preset threshold between the data digest of the target file are selected to obtain a set of associated tags. Collect the internal rule files of the business line corresponding to each business tag within the associated tag set; The file level of the external rule file that has changed is taken as the change level. Based on the change content, the internal rule files corresponding to all business tags in the associated tag set whose file level is lower than the change level are updated.
2. The text processing method as described in claim 1, characterized in that, The method of real-time collection of external rule files of a preset enterprise through a dual mechanism of preset timed scheduling tasks and event-triggered tasks includes: Identify the execution cycle of a preset timed scheduling task, and periodically collect data from the preset enterprise according to the execution cycle to obtain periodic data fragments; Identify the triggering conditions of the preset event triggering task, use the triggering conditions to listen for the target event of the preset enterprise, and generate the event triggering data fragment of the preset enterprise when the target event is detected; The periodic data segments and the event-triggered data segments are combined into an external rule file for the preset enterprise.
3. The text processing method as described in claim 1, characterized in that, The step of labeling the file level of each filtered external rule file according to a preset level label to obtain the file level of each external rule file includes: Extract the file number and file signature contained in each external rule file after filtering, and concatenate the file number and file signature to convert them into a text vector corresponding to each external rule file; Obtain a preset set of hierarchical tags, select any external rule file from the external rule files one by one as the target file, calculate the matching degree between the text vector of the target file and each hierarchical tag in the set of hierarchical tags, and select the hierarchical tag with the highest matching degree as the file hierarchical tag of the target file; The text features of the target file are extracted using a preset text recognition model, and the effectiveness priority of the target file is queried from a preset effectiveness priority table based on the text features. The target file is marked using the file level label and the effectiveness priority as the file level until all files in the external rule file have been marked.
4. The text processing method as described in claim 1, characterized in that, The step of storing the external rule file into the internal database of the preset enterprise according to the file level to obtain the internal rule file includes: The external rule files are classified and archived according to the file level to obtain hierarchical external rule files; Obtain the internal database storage specifications of the preset enterprise, and allocate corresponding storage paths to the hierarchical external rule files according to the internal database storage specifications; The hierarchical external rule files are stored in the internal database of the preset enterprise according to the storage path, and a file-level association index of the preset enterprise is generated. The hierarchical external rule files are validated for entry into the database based on the file-level association index to obtain internal rule files.
5. The text processing method as described in claim 1, characterized in that, The calculation of the correlation coefficient between the business tags of different business lines within the preset enterprise and the data summary includes: Extract the business tag feature vector of the business tag and the data digest feature vector of the data digest, respectively; The similarity correlation value between the business tag feature vector and the data summary feature vector is calculated using a preset similarity algorithm; Identify the business weight coefficients corresponding to the business tags of different business lines within the preset enterprise according to preset business priority rules; The similarity correlation value and the business weight coefficient are weighted and calculated to obtain the correlation coefficient between the business tag and the data summary.
6. The text processing method as described in claim 1, characterized in that, The step of identifying the target file corresponding to the modified external rule file in the internal rule file within the knowledge graph includes: Match the internal rule file nodes corresponding to the changed external rule files within the knowledge graph; Identify the correlation coefficient corresponding to each edge in the internal rule file node, and calculate the correlation strength of the internal rule file node based on the correlation coefficient; The internal rule file corresponding to the internal rule file node whose association strength meets the preset strength threshold is determined as the target file.
7. The text processing method as described in claim 1, characterized in that, The step of updating the internal rule files corresponding to all business tags in the associated tag set whose file level is lower than the change level according to the change content includes: Filter out target business tags whose file level is lower than the change level from the set of associated tags; The internal rule file corresponding to the target business tag is identified as the file to be updated. Based on the changes, analyze the key points of the clause updates in the document to be updated, as well as the unadapted related clauses. The unadapted related clauses are updated according to the updated clause guidelines to obtain the updated internal rule file.
8. A text processing device, characterized in that, include: The external rule file acquisition module is used to acquire the external rule files of a preset enterprise in real time through a dual mechanism of preset timed scheduling tasks and event-triggered tasks; The file level labeling module is used to filter out files in the external rule files that do not conform to the preset file category, and to label the file level of each filtered external rule file according to the preset level label to obtain the file level of each external rule file. The external rule file storage module is used to store the external rule file into the internal database of the preset enterprise according to the file level, so as to obtain the internal rule file. The correlation coefficient calculation module is used to extract the data summary of each internal rule file, calculate the correlation coefficient between the business tags of different business lines within the preset enterprise and the data summary, and determine the business line corresponding to each internal rule file based on the correlation coefficient. The knowledge graph construction module is used to construct a knowledge graph with each internal rule file, the data summary corresponding to each internal rule file, the business tag and the file level as nodes, and the correlation coefficient as edges. The associated tag set determination module is used to collect the changed content of the external rule file when the external rule file is detected to have changed, identify the target file in the internal rule file that corresponds to the changed external rule file, select business tags with a correlation coefficient greater than a preset threshold between the data digest of the target file, and obtain an associated tag set. The internal rule file acquisition module is used to acquire the internal rule files of the business line corresponding to each business tag in the associated tag set; The internal rule file update module is used to take the file level of the changed external rule file as the change level, and update the internal rule files corresponding to all business tags in the associated tag set whose file level is lower than the change level according to the changed content.
9. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the text governance method as described in any one of claims 1 to 7.
10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the text governance method as described in any one of claims 1 to 7.