Methods for synthesizing abnormal factual case data for the coal industry

By constructing anomaly case data for the coal industry, the problem of insufficient data in the industry was solved, high-quality training data was generated, and the recognition accuracy and reasoning ability of large models were improved.

CN120196625BActive Publication Date: 2026-05-26CHINA COAL RES INST +1
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
CHINA COAL RES INST
Filing Date
2025-05-26
Publication Date
2026-05-26

AI Technical Summary

Technical Problem

The lack of standardized data in the coal industry makes it impossible to train large models that can be used to trace and identify abnormal cases.

Method used

By acquiring a dictionary set of coal industry data, constructing a first corpus based on the job constraint file, classifying and extracting key points, simulating abnormal fact cases, and generating high-quality abnormal fact case data.

Benefits of technology

It provides high-quality training data, improves the accuracy and reasoning ability of large models in identifying anomalous fact cases, and solves the problems of insufficient anomalous fact cases and diverse formats.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120196625B_ABST
    Figure CN120196625B_ABST
Patent Text Reader

Abstract

This application proposes a method for synthesizing abnormal fact case data for the coal industry, relating to the field of data processing technology in smart mines. The method includes: acquiring a dictionary set for the coal industry; determining a first corpus set for the coal industry based on the operational constraint files within the dictionary set; classifying the first corpus set to obtain a second corpus set for different business domains within the coal industry; extracting key points from each corpus in the second corpus set for each business domain to obtain a key point set for the corpus; and simulating abnormal fact cases based on the second corpus set and the key point set of the corpus set to obtain a dataset of abnormal fact cases for the business domain. This application synthesizes simulated abnormal fact cases for the coal industry based on standardized item corpora from operational constraints, thereby providing data support for the training of large-scale models and improving the training effect of related large-scale models in the coal industry.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of data processing technology for smart mines, and in particular to a method for synthesizing abnormal fact cases for the coal industry. Background Technology

[0002] With the breakthrough development of large-scale model technology in recent years, more and more coal enterprises are accelerating the vertical application of technology, gradually building a large-scale model system that integrates the characteristics of the coal industry, and effectively combining it with the coal industry. In particular, the coal industry has an urgent need for large-scale models in tracing and identifying abnormal fact cases (such as illegal or irregular events), but the lack of standardized data related to the coal industry makes it impossible to train large-scale models that can be applied to the tracing and identification of abnormal fact cases. Summary of the Invention

[0003] The purpose of this application is to at least partially solve one of the technical problems in the related art.

[0004] Therefore, the first objective of this application is to propose a method for synthesizing abnormal factual case data for the coal industry, which can increase standardized data related to the coal industry, thereby providing data support for the training of large models.

[0005] To achieve the above objectives, the first aspect of this application proposes a method for synthesizing abnormal factual case data for the coal industry, comprising:

[0006] Obtain a dictionary set for the coal industry, and determine a first corpus set for the coal industry based on the operational constraint files of the coal industry in the dictionary set;

[0007] The first corpus set is classified to obtain a second corpus set for different business areas of the coal industry;

[0008] For each business domain, key points are extracted from each corpus in the second corpus set of the business domain to obtain the key point set of the corpus.

[0009] Anomaly case simulation is performed based on the second corpus set and the key point set of the corpus to obtain a data set of anomaly case cases in the business domain.

[0010] This application provides a method for synthesizing abnormal fact case data for the coal industry. The method involves acquiring operational constraint documents from the coal industry, performing text mining on these documents to determine a first corpus set, classifying this first corpus set to obtain a second corpus set for each business domain, and further extracting key points from each corpus in the second corpus set to obtain a key point set for each corpus. Based on the second corpus set and the key point sets of the second corpus set, abnormal fact case simulations are performed to obtain a dataset of abnormal fact cases for each business domain. This application utilizes text mining and data cleaning tools to organize and filter standardized entry corpora from relevant laws, regulations, and standards in the coal industry suitable for simulating abnormal fact cases. These standardized entry corpora provide a data foundation for the batch simulation of abnormal fact cases in the coal industry, addressing the issues of insufficient and diverse formats of abnormal fact cases in the coal industry. Furthermore, based on the key points of the standardized entry corpus, abnormal fact cases can be automatically simulated, providing high-quality training data for large-scale coal industry models to identify abnormal fact cases. This improves the training effect of large-scale models, thereby enhancing their inference accuracy and accelerating their adoption in the coal industry.

[0011] Additional aspects and advantages of this application will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of this application. Attached Figure Description

[0012] The above and / or additional aspects and advantages of this application will become apparent and readily understood from the following description of the embodiments taken in conjunction with the accompanying drawings, wherein:

[0013] Figure 1 A flowchart illustrating a method for synthesizing abnormal factual case data for the coal industry, provided in an embodiment of this application;

[0014] Figure 2 A flowchart illustrating another method for synthesizing abnormal fact case data for the coal industry, provided in an embodiment of this application;

[0015] Figure 3 A flowchart illustrating another method for synthesizing abnormal fact case data for the coal industry, provided in an embodiment of this application;

[0016] Figure 4 This is a flowchart illustrating another method for synthesizing abnormal factual case data for the coal industry, provided as an embodiment of this application. Detailed Implementation

[0017] The embodiments of this application are described in detail below. Examples of these embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and intended to explain this application, and should not be construed as limiting this application.

[0018] With the breakthrough development of large-scale modeling technology in recent years, more and more coal enterprises are accelerating the vertical integration of technology, gradually building large-scale model systems that integrate the characteristics of the coal industry, and effectively combining them with actual coal business operations. These coal industry-related large-scale models need to possess a large amount of laws, regulations, and industry expertise, as well as strong language understanding and generation capabilities, and be able to process and analyze massive amounts of industry knowledge and data. As large-scale models become more deeply integrated with the coal industry, they are no longer just used as a means of querying industry knowledge and data, but are also beginning to extend into practical scenarios such as the identification of illegal events (or instances of illegal activities) and the summarization of production reports.

[0019] To address the need for automated identification and recognition of illegal events in the coal industry using large-scale models, the lack of stable and reliable data acquisition methods for such tasks limits the development and application of such models. Against this backdrop, establishing methods for simulating or synthesizing anomalous factual cases (e.g., illegal factual cases) is crucial for training, fine-tuning, and testing large-scale models related to the coal industry. Taking the identification of anomalous factual cases (e.g., illegal factual cases) as an example, these cases can be used to assess the large-scale model's grasp and understanding of relevant coal industry laws and regulations. This helps coal-related enterprises or institutions select suitable large-scale models and improve their own model capabilities.

[0020] The starting point for establishing an abnormal fact case acquisition or synthesis system in the coal industry is that there is currently a lack of data related to abnormal fact cases (such as illegal acts), and the publicly available data formats are not uniform, and there is a lack of targeted analysis and description. The existing data related to abnormal fact cases (such as illegal acts) is insufficient to support the data volume requirements of large models.

[0021] The following section explains the method for synthesizing abnormal factual case data for the coal industry, with accompanying figures.

[0022] Figure 1 This is a flowchart illustrating a method for synthesizing abnormal factual case data for the coal industry, provided in an embodiment of this application. The entity executing this method can be an electronic device or a server; there are no further limitations here.

[0023] like Figure 1 As shown, the method for synthesizing abnormal factual case data for the coal industry may include, but is not limited to, the following steps:

[0024] S101, obtain the dictionary set of the coal industry, and determine the first corpus set of the coal industry based on the operation constraint files of the coal industry in the dictionary set.

[0025] In some embodiments, operational constraints documents for the coal industry may include, but are not limited to, the following documents: legal documents for the coal industry, regulatory documents for the coal industry, and standard documents for the coal industry.

[0026] In some embodiments, the dictionary set may include, but is not limited to, operational constraint documents such as legal documents, regulatory documents, and standard documents of the coal industry.

[0027] In some embodiments, a structured corpus for the coal industry can be constructed based on legal documents, regulatory documents, and / or standard documents in the coal industry from a dictionary set.

[0028] In some embodiments, a text mining engine can be used to perform semantic parsing and understanding of original documents such as legal documents, regulatory documents, and / or standard documents related to the coal industry in a dictionary set, extracting valid clause information from them. Further, the text mining engine extracts the valid clause information to obtain raw corpus, which is then subjected to structured processing to obtain a structured corpus. It is understood that the structured corpus includes structured data corresponding to the original corpus.

[0029] In some embodiments, after obtaining the structured corpus, the original corpus extracted from the job constraint file can be finely filtered to remove noisy data and retain high-quality corpus suitable for simulating abnormal fact cases. This retained high-quality corpus suitable for simulating abnormal fact cases is used to constitute the first corpus set.

[0030] S102, classify the first corpus set to obtain the second corpus set of different business areas of the coal industry.

[0031] In some embodiments, the business scope of the coal industry may include, but is not limited to: coal exploration, underground mining, open-pit mining, coal processing, coal transportation, and coal storage.

[0032] In some embodiments, a corpus may be mapped to multiple business domains. A multi-classification model can be pre-trained to classify and identify the corpus in the first corpus set, thereby obtaining a classification label for each corpus, where the classification label is used to indicate the business domain to which the corpus belongs.

[0033] In some embodiments, after obtaining the classification label of each corpus in the first corpus set, the corpus in the first corpus set can be mapped to different business domains based on the classification label, thereby obtaining a second corpus set corresponding to each business domain.

[0034] In some embodiments, the classification labels of the corpus can be standardized and stored.

[0035] S103, for each business domain, extract key points from each corpus in the second corpus set of the business domain to obtain the key point set of the corpus.

[0036] In some embodiments, key points can be extracted from each corpus in the second corpus set for each business domain to obtain at least one key point for each corpus, and the at least one key point for each corpus can be filtered to obtain a set of key points for the corpus.

[0037] In some embodiments, a large model is used to perform semantic analysis on the corpus in order to extract as many valuable key points as possible from the corpus. It is understood that the extracted valuable key points can be used to simulate abnormal fact cases, that is, the extracted valuable key points may be potential violations in the task.

[0038] In some embodiments, valuable key points extracted from the corpus may include key points that are weakly related to the industry or have a high degree of overlap. Key points that are weakly related to the industry or have a high degree of overlap can be removed from the valuable key points to obtain the key point combination of the corpus.

[0039] In some embodiments, the key points in the key point set of the corpus can be organized, for example, these key points can be formatted to obtain key points in a standardized format, and key points in a protected format can be stored.

[0040] S104. Based on the second corpus set and the key point set of the corpus, simulate abnormal fact cases to obtain a data set of abnormal fact cases in the business domain.

[0041] In some embodiments, exceptional fact cases may include, but are not limited to: events that violate legal documents of the coal industry, events that violate regulatory documents of the coal industry, and events that violate standard documents of the coal industry.

[0042] In some implementations, after obtaining the set of key points for each corpus, the descriptive information of the abnormal fact cases to be simulated can be configured. Optionally, the descriptive information may include, but is not limited to, the time, place, subject, and content of the abnormal fact case.

[0043] Furthermore, based on the corpus and its key points, combined with the pre-configured descriptive information of abnormal fact cases, simulations of abnormal fact cases are performed to obtain a dataset of abnormal fact cases in the business domain.

[0044] It is understandable that the dataset of abnormal fact cases in the business domain can include not only simulated abnormal fact cases, but also abnormal fact cases collected in actual operations.

[0045] In this embodiment, operational constraint documents for the coal industry can be obtained. Text mining is then performed on these documents to determine a first corpus set for the coal industry. This first corpus set is then categorized to obtain a second corpus set for each business domain. Furthermore, key points are extracted from each corpus in the second corpus set to obtain a key point set for the corpus. Based on the second corpus set and the key point sets, abnormal fact case simulations are performed to obtain a dataset of abnormal fact cases for each business domain. In this application, text mining and data cleaning tools are used to organize and filter standardized entry corpora suitable for simulating abnormal fact cases from relevant laws, regulations, and standards in the coal industry. These standardized entry corpora provide a data foundation for the batch simulation of abnormal fact cases in the coal industry, addressing the issues of insufficient and diverse formats of abnormal fact cases in the industry. Furthermore, based on the key points of the standardized entry corpus, abnormal fact cases can be automatically simulated, providing high-quality training data for large-scale coal industry models to identify abnormal fact cases. This improves the training effect of large-scale models, thereby enhancing their inference accuracy and accelerating their adoption in the coal industry.

[0046] Figure 2 This is a flowchart illustrating another method for synthesizing abnormal factual case data in the coal industry, provided as an embodiment of this application. Figure 2 As shown, the method for synthesizing abnormal factual case data for the coal industry may include, but is not limited to, the following steps:

[0047] S201, retrieve the dictionary set for the coal industry.

[0048] In some embodiments, legal documents, regulatory documents, and standard documents related to the coal industry are obtained and used as operational constraint documents for the coal industry. Further, a dictionary set is established using these operational constraint documents. Doc ={ D 1, D 2, …, D n},in, D iThis refers to an original job constraint document, which can be a regulatory document, a legal document, or a standard document. Optionally, the original job constraint document can be a PDF file or a Word file; this application does not limit the file type of the job constraint document.

[0049] S202, perform text mining on the job constraint files to obtain the original structured corpus.

[0050] In some embodiments, a text mining model R can be used to extract data from the original job constraint file. D i Extracting raw data one by one ,in, This represents the original structured corpus. This indicates the origin of the original job constraint file. D i The extracted first j One original corpus.

[0051] S203, for documents belonging to job constraint files D i corpus According to the corpus and job constraint files D i Determine the overall information density and job constraint files. D i The primary indicator.

[0052] Understandably, the corpus It refers to any raw data in a structured corpus.

[0053] In some embodiments, there is a mapping relationship between each corpus and the original job constraint file, and the job constraint file to which each corpus belongs can be determined based on the mapping relationship.

[0054] In some embodiments, for job constraint files D i corpus It is possible to obtain corpus of Parameters, where, Parameters are used to represent the corpus Chinese regulations and key terms in work constraint documents D i Frequency of occurrence in.

[0055] In some embodiments, obtaining the job constraint file D iFurthermore, after obtaining the overall information density, the Frobenius norm of the overall information density can be determined.

[0056] In some embodiments, job constraint files can be calculated through word frequency analysis. D i The frequency of each word in the document is used to determine the task constraints. D i The overall information density.

[0057] In some embodiments, job constraint files can be obtained. D i The term frequency-inverse document frequency (TF-IDF) of each word is used to determine the job constraint file. D i The overall information density.

[0058] In some embodiments, a job constraint file can be calculated. D i Entropy of content, based on job constraint file D i Entropy of content, determining job constraint files D i The overall information density.

[0059] Furthermore, based on the corpus of and job constraint files D i The Frobenius norm of the overall information density can be used to determine the job constraint file. D i The primary indicator.

[0060] In some embodiments, the job constraint file can be determined based on the following formula (1). D i The first indicator:

[0061]

[0062] in, IR ( D i ) represents the job constraint file D i The primary indicator; oh This refers to the weighting coefficient of expert experience. Represents job constraint file D i The Frobenius norm of the overall information density; nIndicates from the job constraint file D i The number of corpora extracted, n The value of is a natural number greater than or equal to 1.

[0063] In some embodiments, via job constraint files D i The first indicator represents the job constraint file. D i Information retention rate.

[0064] S204, Obtaining Corpus The inverse document frequency of the term, positional weight, and complexity evaluation value.

[0065] S205, determine the corpus based on word frequency inverse document frequency, positional weight value, and complexity evaluation value. The second indicator.

[0066] In some embodiments, a corpus can be obtained. Term Frequency (TF) and Inverse Document Frequency (IDF) of the corpus, further, for the corpus Multiplying TF and IDF results in the corpus. The TF_IDF can be represented as .

[0067] In some embodiments, the corpus can be located. In the job constraint file D i The position in the corpus, according to the corpus In the job constraint file D i The position in the corpus can be used to calculate the corpus. The corresponding positional weight values. Optionally, a positional weight function Pos can be pre-constructed for the corpus, and the corpus can be determined based on this positional weight function Pos. The corresponding positional weight value. In other words, It can represent corpus The positional weight value.

[0068] In some embodiments, the corpus can be determined based on information entropy. The complexity evaluation value of the corpus. Optionally, a complexity evaluation function Ent is pre-constructed for the corpus, and the complexity evaluation function Ent is used to determine the complexity evaluation value of the corpus. The corresponding complexity evaluation value. That is to say, It can represent corpus The complexity assessment value. It should be noted that the higher the information entropy, the greater the uncertainty of the information, and the key point... z The higher the complexity, the better.

[0069] Furthermore, after obtaining the corpus After obtaining the inverse document frequency, positional weight, and complexity evaluation value of the word frequency, it is possible to base the corpus on... The corpus is determined by the word frequency inverse document frequency, positional weight, and complexity evaluation value. The second indicator.

[0070] Alternatively, the following formula (2) can be used to determine the corpus. The second indicator:

[0071]

[0072] in, Representation corpus The second indicator. Optionally, the second indicator can represent the corpus. The information value.

[0073] a、b、c These are the weighting coefficients, which can be pre-calibrated parameters.

[0074] For corpus The inverse document frequency of the word frequency.

[0075] For corpus The positional weight value.

[0076] For corpus Complexity assessment based on information entropy.

[0077] S206, according to the job constraint file D i The first indicator and corpus The second metric is to filter the structured corpus to obtain the first corpus set.

[0078] In some embodiments, based on the corpus Second indicator and work constraint document D i The first indicator is to determine the corpus. The retention probability.

[0079] In some embodiments, the corpus can be determined using the following formula (3). retention probability:

[0080]

[0081] in, Representation corpus The probability of retention; s For the sigmoid function, l These are the weighting coefficients. t This is the information threshold.

[0082] After obtaining the corpus retention probability Then, the corpus retention probability The data is compared with a set retention threshold to determine whether to retain the corpus. .

[0083] In some embodiments, in response to the corpus retention probability If the original corpus is greater than or equal to the set retention threshold, the original corpus will be retained. .

[0084] In some embodiments, in response to the corpus retention probability If the value is less than the set retention threshold, the corpus will be deleted from the original structured corpus. .

[0085] Furthermore, based on the original corpus preserved in the structured corpus, the first corpus set is obtained.

[0086] Understandably, based on the noise filtering condition that the retention probability is greater than or equal to the set retention threshold, the original text in the original structured corpus can be standardized and filtered. That is, the original text belonging to the structured corpus will be filtered if the retention probability meets the filtering condition. The original corpus can be preserved, and further, based on the preserved original corpus, a first corpus set that conforms to simulated abnormal fact cases can be constructed. .

[0087] S207. The first corpus is classified to obtain the second corpus of different business areas in the coal industry.

[0088] For a detailed description of step S207, please refer to the relevant records in the various embodiments of this application, which will not be repeated here.

[0089] S208. For each business domain, extract key points from each corpus in the second corpus set of the business domain to obtain the key point set of the corpus.

[0090] For a detailed description of step S208, please refer to the relevant records in the various embodiments of this application, which will not be repeated here.

[0091] S209, based on the second corpus set and the key point set of the corpus, perform abnormal fact case simulation to obtain a data set of abnormal fact cases in the business domain.

[0092] For a detailed description of step S209, please refer to the relevant records in the various embodiments of this application, which will not be repeated here.

[0093] In this embodiment, for the corpus of entries in regulations, standards, and other documents compiled in batches, a screening and filtering mechanism can be used to remove entries with weak knowledge content and high repetition, which helps improve the accuracy and professionalism of the original corpus. Furthermore, by evaluating indicators such as word frequency inverse document frequency, positional weight, and information entropy, the overall information retention of the documents and the information value of individual entries are calculated respectively. Combining these two methods to comprehensively evaluate the corpus retention calculation, entries with weak knowledge content, high repetition, and high noise are removed. This process not only improves the quality and efficiency of the extracted original corpus but also ensures the application effectiveness of the obtained simulated abnormal fact cases, enabling them to better reflect the actual needs and capabilities of the industry, thereby optimizing the accuracy of the large-scale coal industry model in identifying illegal and irregular behaviors.

[0094] Figure 3 This is a flowchart illustrating another method for synthesizing abnormal factual case data in the coal industry, provided as an embodiment of this application. Figure 3 As shown, the method for synthesizing abnormal factual case data for the coal industry may include, but is not limited to, the following steps:

[0095] S301, Obtain the dictionary set for the coal industry, and determine the first corpus set for the coal industry based on the operational constraint files of the coal industry in the dictionary set.

[0096] For a detailed description of step S301, please refer to the relevant records in the various embodiments of this application, which will not be repeated here.

[0097] S302, classify the first corpus set to obtain the second corpus set of different business areas of the coal industry.

[0098] In some embodiments, the defined corpus classification set can be represented as Class ={ C 1, C 2, …, C n} are used to represent specific business areas in the coal industry, such as underground mining and open-pit mining, respectively. C i It refers to a specific business area.

[0099] In some embodiments, a multi-classification model can be invoked to classify the original corpus in the first corpus set, thereby obtaining at least one classification label for each original corpus.

[0100] For example, the multi-class classification model is C. The original corpus in the first corpus is classified using the multi-class classification model C. By performing classification operations, we can obtain the original corpus. At least one category label. That is, the original corpus Input into multi-class classification model C, and output the original corpus through multi-class classification model C. One or more classification labels. The original corpus is output through a multi-classification model C. The process of classification labeling can be characterized as and .

[0101] It is understandable that any corpus within the first corpus set could be mapped to multiple specific business domains. For example, any corpus within the first corpus set... It can be categorized into at least one specific business area. That is to say... It can be categorized into at least one specific business area.

[0102] S303, for any corpus in the second corpus set Extracting corpus through the first intelligent agent The key point.

[0103] S304, Obtaining Key Points and Corpus Similarity, complexity assessment of key points, and compliance deviation.

[0104] S305 calculates the score for each key point by weighting the similarity, complexity assessment, and compliance deviation.

[0105] S306, Analyze the corpus based on the score values ​​of key points. Filter the key points to obtain the corpus The set of key points.

[0106] In some embodiments, for any corpus in the second corpus set It can extract corpus The key point is that, optionally, a pre-built first agent can be invoked to extract the corpus. The key points in it.

[0107] For example, the first intelligent agent is Z, and the corpus is processed by the first intelligent agent Z. Key point extraction is performed to obtain the corpus. at least one key point z Optionally, the corpus Input into the first intelligent agent Z, through the first intelligent agent Z, can obtain from the corpus Extract one or more key points z The corpus was extracted using the first intelligent agent Z. The process of key points in the middle can be characterized as and .

[0108] In some embodiments, each key point can be obtained. z corpus Similarity, complexity assessment of key points, and compliance deviation. Furthermore, regarding key points... z The key points are obtained by weighting the corresponding similarity, complexity assessment values, and compliance deviation. z The rating value.

[0109] Alternatively, the key points can be determined using the following formula (4). z Rating:

[0110]

[0111] It should be noted that the key points z To obtain the corpus through the first intelligent agent Z The key points extracted from the corpus, and from the corpus The number of extracted key points is greater than or equal to 2. In other words, the following condition must be met: and .

[0112] Among them, in formula (4) l 1. l 2. l 3 represents the weighting coefficient, a parameter that can be pre-calibrated.

[0113] Sim is the cosine similarity calculation function used to calculate key points. z Its source corpus Cosine similarity can be used to characterize key points. z Its source corpus Semantic relevance between them.

[0114] Ent represents the complexity evaluation function based on information entropy, used to evaluate key points. z The complexity.

[0115] Compl represents the compliance deviation calculation function, used to assess key points. zThe compliance deviation degree can optionally be calculated using a pre-trained function based on compliance knowledge points, which can be used to analyze key points. z The degree of compliance deviation is scored. Understandably, key points... z The higher the score, the greater the compliance deviation, indicating key points. z The greater the likelihood of non-compliance; key points z The lower the score, the smaller the compliance deviation, indicating a key point. z The lower the likelihood of non-compliance.

[0116] It is understandable that the corpus can be calculated using formula (4). The score value of each key point extracted from the corpus is further used to analyze the score value. Each key point extracted is filtered to remove key points with weak relevance, unclear clarity, or semantic repetition from all key points, thus obtaining the corpus. Key points set: and n ≥1, of which the corpus Any key point in the key point set This is a high-value key point for simulating anomalous fact cases.

[0117] S307, based on the second corpus set and the key point set of the corpus, perform abnormal fact case simulation to obtain a data set of abnormal fact cases in the business domain.

[0118] For a detailed description of step S307, please refer to the relevant records in the various embodiments of this application, which will not be repeated here.

[0119] In this embodiment, a second corpus is obtained from documents such as coal industry regulations and standards. A first intelligent agent extracts key points from the corpus of the second corpus and scores the key points. Based on the score, the key points are filtered and screened, which can retain key points with strong relevance, high clarity, and no semantic repetition. This is beneficial for generating more accurate and effective simulated abnormal fact cases, thereby increasing the data on abnormal fact cases in the coal industry. Moreover, using documents such as regulations and standards as the basis for simulating abnormal fact cases can solve the problem of inconsistent understanding of abnormal fact cases.

[0120] Figure 4 This is a flowchart illustrating a method for synthesizing abnormal factual case data in the coal industry, provided as an embodiment of this application. Figure 4 As shown, the method for synthesizing abnormal factual case data for the coal industry may include, but is not limited to, the following steps:

[0121] S401, Obtain the dictionary set for the coal industry, and determine the first corpus set for the coal industry based on the operational constraint files for the coal industry in the dictionary set.

[0122] For a detailed description of step S401, please refer to the relevant records in the various embodiments of this application, which will not be repeated here.

[0123] S402, classify the first corpus set to obtain the second corpus set of different business areas in the coal industry.

[0124] For a detailed description of step S402, please refer to the relevant records in the various embodiments of this application, which will not be repeated here.

[0125] S403, extract key points from each corpus in the second corpus set of the business domain to obtain the key point set of the corpus.

[0126] For a detailed description of step S403, please refer to the relevant records in the various embodiments of this application, which will not be repeated here.

[0127] S404, Obtain the preset description information of the abnormal fact case to be simulated.

[0128] In some embodiments, the descriptive information includes the time, place, actor, and content of the unusual fact case.

[0129] Optionally, the descriptive information of the anomalous fact cases to be simulated can form a quadruple data set. For example, the quadruple data set of the anomalous fact cases to be simulated can be labeled as Ω=( time , pos , character , act ),in, time Indicates the time of occurrence of the abnormal fact case to be simulated; pos Indicates the location where the abnormal fact case to be simulated occurred; character The actor representing the abnormal factual case to be simulated; act This represents the behavioral content of the abnormal fact case to be simulated.

[0130] S405, invoke the second intelligent agent, which generates simulated abnormal fact cases of arbitrary key points based on the description information, corpus and arbitrary key points extracted from the corpus.

[0131] In some embodiments, descriptive information, corpus, and arbitrary key points extracted from the corpus can be input into a second intelligent agent, which can then output simulated abnormal fact cases of arbitrary key points based on the descriptive information, corpus, and arbitrary key points extracted from the corpus.

[0132] For example, the second agent is named Generate, and the arbitrary corpus is... From the corpus The extracted key points are ,Will , and the quadruple data Ω=( time , pos , character , act The data is input into the second agent, Generate, which processes the data and can output simulated abnormal fact cases. In this application, the process of simulating anomalous fact cases based on the output of the second agent's `Generate` function can be represented as follows: .

[0133] S406, based on simulated abnormal fact cases at arbitrary key points, obtains a dataset of abnormal fact cases in the business domain.

[0134] In some embodiments, existing abnormal fact cases in the business domain are determined. Further, the semantic similarity between simulated abnormal fact cases and existing abnormal fact cases at any key point is obtained. In response to a semantic similarity less than a set similarity threshold, the database of the business domain is updated based on the simulated abnormal fact cases to obtain a data set of abnormal fact cases. In response to a semantic similarity greater than or equal to the set similarity threshold, the simulated abnormal fact cases are deleted.

[0135] In some embodiments, when the semantic similarity of simulated anomalous fact cases is determined to be less than a set similarity threshold, the simulated anomalous fact cases can be cached in an array so that they can be read from the array for subsequent parsing processes.

[0136] For example, the simulated abnormal fact case is calculated using the following formula (5). Existing anomaly cases in the business domain q Semantic similarity:

[0137]

[0138] in, express and q semantic similarity; and They represent and q semantic vector; i This indicates the set similarity threshold.

[0139] Optionally, if the semantic similarity is greater than or equal to the similarity threshold i Then determine the simulated abnormal fact case. Cases of Existing Abnormal Facts q Repetition can eliminate the need to simulate abnormal fact cases. .

[0140] Optionally, if the semantic similarity is less than the similarity threshold i Then determine the simulated abnormal fact case. Cases of Existing Abnormal Facts q No repetition, can simulate abnormal fact cases. Store in array .

[0141] In some embodiments, to ensure the interpretability of the content of simulated abnormal fact cases, corresponding parsing results need to be generated for each simulated abnormal fact case. This involves determining the corpus (i.e., the source corpus) and key points associated with the simulated abnormal fact case, and then performing content analysis on the simulated abnormal fact case based on the corpus and key points associated with it, thereby obtaining the parsing results for the simulated abnormal fact case.

[0142] In some embodiments, a third-party intelligent agent can be invoked to perform content analysis on simulated anomalous fact cases, obtaining the parsing results of the simulated anomalous fact cases. Optionally, the corpus and key points associated with the simulated anomalous fact cases, along with the simulated anomalous fact cases themselves, are input into the third-party intelligent agent. The third-party intelligent agent then performs content analysis on the simulated anomalous fact cases based on the associated corpus and key points, obtaining the parsing results of the simulated anomalous fact cases. This application not only solves the problems of insufficient standardized data and lack of analytical descriptions but also improves the quality and efficiency of anomalous fact case simulation, providing high-quality data for the identification of anomalous fact cases in large-scale models related to the coal industry.

[0143] In some embodiments, simulated anomalous fact cases with a semantic similarity less than a set similarity threshold can be read from the array, and the simulated anomalous fact cases, along with their associated corpus and key points, can be input into a third agent for parsing to obtain the parsing results of the simulated anomalous fact cases.

[0144] For example, the third agent is Analysis, which will simulate abnormal fact cases. The associated corpus and key points The data is input into the third-party intelligent agent Analysis, which then outputs the analysis results. In this application, the third-party intelligent agent Analysis can be used to analyze... The process of conducting content analysis is represented as: .

[0145] Furthermore, based on the simulated abnormal fact cases, the parsing results, and the corpus associated with the simulated abnormal fact cases, triples are constructed, and the triples are format-converted to obtain standardized triples. The database is then updated based on the standardized triples to obtain a dataset of abnormal fact cases.

[0146] An example illustration will be provided, simulating anomaly cases. Analysis results The associated source corpus This forms a triple. Furthermore, the triples are standardized using a specific JSON format, and the standardized triples are stored. For example, the standardized triples can be stored in a group. .

[0147] In some embodiments, before updating the database based on simulated anomalous fact cases to obtain a dataset of anomalous fact cases, a large model can be invoked to automatically validate the simulated anomalous fact cases. Optionally, the large model can combine the parsing results of the simulated anomalous fact cases, the associated corpus, and key points to automatically validate the simulated anomalous fact cases.

[0148] In some embodiments, after a simulated anomalous fact case passes automatic verification, it is sent to a review device. The device can display the simulated anomalous fact case and monitor the review data input by reviewers, providing feedback on the data. The implementing entity of this application can receive the review data from the review device and further determine a score for the simulated anomalous fact case based on the review data. If the score exceeds a set score threshold, the classification labels of the corpus associated with the simulated anomalous fact case are obtained. Based on the classification labels and the standardized triples of the simulated anomalous fact case, a quadruple of the simulated anomalous fact case is obtained, and finally, the quadruple of the simulated anomalous fact case is stored in the database corresponding to the business domain.

[0149] In some embodiments, the quadruple of simulated anomalous fact cases includes the simulated anomalous fact case, the parsing result of the simulated anomalous fact case, the associated corpus, and the classification labels of the corpus.

[0150] In some embodiments, the review data for simulating abnormal fact cases may include, but is not limited to, review values ​​for dimensions such as content rationality, consistency with laws and regulations, and relevance to the coal industry.

[0151] In some embodiments, the review values ​​for each dimension in the review data are weighted to obtain the score for the simulated abnormal fact case.

[0152] For example, a large-scale model related to the coal industry (such as a mine model) is used to simulate abnormal fact cases. Automatic verification is performed, meaning that large-scale models related to the coal industry can be based on simulated cases of abnormal facts. The analysis results Source Corpus Corresponding key points , for simulated abnormal fact cases Conduct legal consistency assessments and verify their relevance to the coal industry to ensure the accuracy and reliability of simulated abnormal fact cases.

[0153] Optionally, a consistency threshold can be pre-configured. i g and correlation threshold i l Large-scale models related to the coal industry can output simulated cases of abnormal facts. The consistency and relevance parameters will be used to simulate anomalous fact cases. Consistency parameters and consistency thresholds i g Comparison, and simulation of abnormal fact cases. Correlation parameters and correlation thresholds i l Comparison in simulating abnormal fact cases The consistency parameter and correlation parameter are obtained through the consistency threshold. i g and correlation threshold i l After the determination, large-scale models related to the coal industry can simulate abnormal fact cases. Automatic verification.

[0154] Optionally, after automatic verification, the simulated abnormal case can be manually reviewed. The simulated abnormal case can be sent to the review equipment, where domain experts can then conduct a detailed examination. Understandably, experts will use their professional knowledge and experience to comprehensively assess the content's rationality, legal and regulatory consistency, and relevance to the coal industry of the simulated abnormal case.

[0155] Furthermore, the review equipment can feed back the review data from domain experts to the implementing entity of this application, which can then calculate the score for the simulated anomalous fact cases based on the scoring formula. Optionally, the review equipment can calculate the score for the simulated anomalous fact cases based on the scoring formula, obtain the score for the simulated anomalous fact cases, and feed back the score for the simulated anomalous fact cases to the implementing entity of this application.

[0156] For example, the scoring formula for simulating abnormal fact cases is: Source=αa+βr+γu ,in α , β , c These are the weighting coefficients. a The value is for the review of reasonableness. r For consistent review values, u This is the relevance assessment value.

[0157] After determining the scoring values ​​for simulated abnormal fact cases. Source Then, in the rating value Source Exceeding the predetermined threshold Source min Only then can it pass manual verification of simulated abnormal fact cases.

[0158] Optionally, simulated anomalous fact cases, verified both automatically and manually, are grouped according to a four-tuple. The data is processed in the specified format to obtain a quadruple of simulated abnormal fact cases, which is then stored in the database of the business domain.

[0159] Based on the above embodiments, after obtaining a set of abnormal fact cases in different business areas, an anomaly identification model for the business area can be trained based on the set of abnormal fact cases to obtain a target anomaly identification model for the business area. Then, based on the target anomaly identification model, abnormal work behavior can be identified, thereby achieving the purpose of standardizing the work behavior of the coal industry.

[0160] In this embodiment, verification can retain simulated abnormal fact cases that meet the requirements, which can significantly improve the simulation efficiency and quality of abnormal fact cases in the coal industry. Moreover, the various modules work together, from the extraction and screening of the original corpus, to the synthesis and parsing of abnormal fact cases, and finally to the formatted storage in the database after double verification, forming a standardized and highly reusable abnormal fact case acquisition process.

[0161] Furthermore, by organically combining automatic verification with manual review, the rationality, consistency with the corpus, and relevance to the coal industry of the synthesized simulated anomalous fact cases are ensured. This automated process not only improves the efficiency of event synthesis but also enhances the standardization of anomalous fact cases, making them more scientific and consistent. This leads to more accurate training and inference for large-scale models related to the coal industry, providing technical support for the intelligent transformation of the coal industry.

[0162] To implement the aforementioned method for synthesizing abnormal factual case data for the coal industry, this application also provides an apparatus for synthesizing abnormal factual case data for the coal industry. This apparatus includes: a first corpus acquisition module, a second corpus acquisition module, a key point acquisition module, and an event simulation module.

[0163] The first corpus acquisition module is used to acquire a dictionary set of the coal industry and, based on the operation constraint files of the coal industry in the dictionary set, determine the first corpus set of the coal industry.

[0164] The second corpus acquisition module is used to classify the first corpus set to obtain a second corpus set for different business areas of the coal industry.

[0165] The key point acquisition module is used to extract key points from each corpus in the second corpus set of the business domain for each business domain, so as to obtain the key point set of the corpus.

[0166] The event simulation module is used to simulate abnormal fact cases based on the second corpus set and the key point set of the corpus, and obtain a data set of abnormal fact cases in the business domain.

[0167] In some embodiments, the first corpus acquisition module is further configured to:

[0168] Text mining is performed on the aforementioned task constraint files to obtain the original structured corpus;

[0169] For those belonging to the job constraint file D i corpus According to the corpus and the job constraint file D i The overall information density is used to determine the job constraint file. D i The first indicator; wherein the corpus This refers to any raw data within the structured corpus.

[0170] Obtain the corpus The inverse document frequency, positional weight, and complexity evaluation value of the term frequency;

[0171] The corpus is determined based on the inverse document frequency of the term, the positional weight value, and the complexity evaluation value. The second indicator;

[0172] According to the job constraint file D i The first indicator and the corpus The second metric is used to filter the structured corpus to obtain the first corpus set.

[0173] In some embodiments, the first corpus acquisition module is further configured to:

[0174] According to the corpus The second indicator and the job constraint file D i The first indicator is used to determine the corpus. The probability of retention;

[0175] Response to the corpus If the retention probability is greater than or equal to the set retention threshold, the original corpus is retained. ;

[0176] Response to the corpus If the retention probability is less than a set retention threshold, the corpus is deleted from the structured corpus. ;

[0177] The first corpus set is obtained based on the original corpus preserved in the structured corpus.

[0178] In some embodiments, the key point acquisition module is further configured to:

[0179] For any corpus in the second corpus set The corpus is extracted by the first intelligent agent. Key points;

[0180] Obtain the key points and the corpus The similarity, the complexity assessment value of the key points, and the compliance deviation;

[0181] The similarity, complexity assessment, and compliance deviation are weighted and calculated to obtain the score of the key point;

[0182] Based on the scoring value, the corpus The corpus is obtained by filtering the key points. The set of key points.

[0183] In some embodiments, the event simulation module is further configured to:

[0184] Obtain the pre-defined descriptive information of the abnormal fact case to be simulated, wherein the descriptive information includes the time, place, subject of the behavior, and content of the behavior of the abnormal fact case;

[0185] The second intelligent agent is invoked, and the second intelligent agent generates simulated abnormal fact cases of the arbitrary key points based on the description information, the corpus, and the arbitrary key points extracted from the corpus.

[0186] Based on the simulated abnormal fact cases of any of the aforementioned key points, a dataset of abnormal fact cases in the business domain is obtained.

[0187] In some embodiments, the event simulation module is further configured to:

[0188] Identify existing anomaly cases in the aforementioned business area;

[0189] Obtain the semantic similarity between the simulated abnormal fact cases and the existing abnormal fact cases;

[0190] In response to the semantic similarity being less than a set similarity threshold, the database of the business domain is updated based on the simulated abnormal fact cases to obtain a data set of the abnormal fact cases.

[0191] In some embodiments, the event simulation module is further configured to:

[0192] Content analysis is performed on the simulated abnormal fact cases, the corpus associated with the simulated abnormal fact cases, and key points to obtain the analysis results of the simulated abnormal fact cases;

[0193] Based on the simulated abnormal fact cases, the parsing results, and the corpus associated with the simulated abnormal fact cases, triples are constructed;

[0194] The triples are format-converted to obtain standardized triples, and the database is updated based on the standardized triples to obtain the dataset of the abnormal fact cases.

[0195] In some embodiments, the event simulation module is further configured to:

[0196] Before updating the database based on the standardized triples to obtain the dataset of the abnormal fact cases, the large model is invoked;

[0197] The large model automatically verifies the simulated abnormal fact cases based on the analysis results of the simulated abnormal fact cases, the corpus associated with the simulated abnormal fact cases, and the key points.

[0198] In some embodiments, the event simulation module is further configured to:

[0199] After confirming that the simulated abnormal fact case has passed automatic verification, the simulated abnormal fact case is sent to the review device;

[0200] Receive review data of the simulated abnormal fact cases fed back by the review device;

[0201] Based on the review data, determine the score for the simulated abnormal fact case;

[0202] In response to the fact that the score of the simulated abnormal fact case is greater than a set score threshold, the classification label of the corpus associated with the simulated abnormal fact case is obtained;

[0203] Based on the classification labels and the standardized triples of the simulated anomalous fact cases, the quadruples of the simulated anomalous fact cases are obtained;

[0204] The quadruple of the simulated abnormal fact case is stored in the database corresponding to the business domain.

[0205] This application provides an apparatus for synthesizing abnormal factual case data for the coal industry. It can acquire operational constraint documents from the coal industry, perform text mining on these documents to determine a first corpus set for the coal industry, classify the first corpus set to obtain a second corpus set for each business domain, further extract key points from each corpus in the second corpus set to obtain a key point set for the corpus, and simulate abnormal factual cases based on the second corpus set and the key point sets to obtain a data set of abnormal factual cases for each business domain. In this application, through text mining and data cleaning tools, standardized entry corpora suitable for simulating abnormal factual cases from relevant laws, regulations, and standards in the coal industry can be organized and screened. These standardized entry corpora can lay a data foundation for the batch simulation of abnormal factual cases in the coal industry, compensating for the lack of abnormal factual cases and their diverse formats in the industry. Furthermore, based on the key points of the standardized entry corpora, abnormal factual cases can be automatically simulated, providing high-quality data for the identification of abnormal factual cases in large-scale coal industry models. This is beneficial for training large-scale models and significantly improves the effectiveness and accuracy of large-scale models in the coal industry.

[0206] To implement the above embodiments, this application also proposes an electronic device, including: a processor and a memory communicatively connected to the processor; the memory stores computer execution instructions; the processor executes the computer execution instructions stored in the memory to implement the method provided in the foregoing embodiments.

[0207] To implement the above embodiments, this application also proposes a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, are used to implement the methods provided in the foregoing embodiments.

[0208] To implement the above embodiments, this application also proposes a computer program product, including a computer program that, when executed by a processor, implements the methods provided in the foregoing embodiments.

[0209] The collection, storage, use, processing, transmission, provision, and application of user personal information involved in this application all comply with the provisions of relevant laws and regulations and do not violate public order and good morals.

[0210] It should be noted that personal information collected from users should be used for legitimate and reasonable purposes and should not be shared or sold outside of these legitimate uses. Furthermore, such collection / sharing should only be conducted after receiving the user's informed consent, including but not limited to notifying the user to read the user agreement / user notice and sign an agreement / authorization that includes authorization of relevant user information before the user uses the function. In addition, any necessary steps must be taken to protect and safeguard access to such personal information data and ensure that others with access to personal information data comply with their privacy policies and procedures.

[0211] This application is intended to provide an implementation scheme for users to selectively prevent the use or access to their personal information data. Specifically, this application is intended to provide hardware and / or software to prevent or block access to such personal information data. Once personal information data is no longer needed, risks can be minimized by restricting data collection and deleting data. Furthermore, where applicable, such personal information is de-identified to protect user privacy.

[0212] In the foregoing descriptions of the embodiments, the terms "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., refer to specific features, structures, materials, or characteristics described in connection with that embodiment or example, which are included in at least one embodiment or example of this application. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Moreover, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of different embodiments or examples.

[0213] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of that feature. In the description of this application, "multiple" means at least two, such as two, three, etc., unless otherwise explicitly specified.

[0214] Any process or method description in the flowchart or otherwise herein can be understood as representing a module, segment, or portion of code comprising one or more executable instructions for implementing custom logic functions or processes, and the scope of the preferred embodiments of this application includes additional implementations in which functions may be performed not in the order shown or discussed, including substantially simultaneously or in reverse order depending on the functions involved, as should be understood by those skilled in the art to which embodiments of this application pertain.

[0215] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (such as a computer-based system, a processor-included system, or other system that can fetch and execute instructions from, an instruction execution system, apparatus, or device). For the purposes of this specification, "computer-readable medium" can be any means that can contain, store, communicate, propagate, or transmit programs for use by, or in conjunction with, an instruction execution system, apparatus, or device. More specific examples (a non-exhaustive list) of computer-readable media include: an electrical connection having one or more wires (electronic device), a portable computer disk drive (magnetic device), random access memory (RAM), read-only memory (ROM), erasable and editable read-only memory (EPROM or flash memory), fiber optic devices, and portable optical disc read-only memory (CDROM). Alternatively, the computer-readable medium may be paper or other suitable media on which the program can be printed, since the program can be obtained electronically, for example, by optically scanning the paper or other medium, followed by editing, interpreting, or otherwise processing as necessary, and then stored in a computer memory.

[0216] It should be understood that various parts of this application can be implemented using hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented using software or firmware stored in memory and executed by a suitable instruction execution system. For example, if implemented in hardware as in another embodiment, it can be implemented using any of the following techniques known in the art, or a combination thereof: discrete logic circuits having logic gates for implementing logical functions on data signals, application-specific integrated circuits (ASICs) having suitable combinational logic gates, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.

[0217] Those skilled in the art will understand that all or part of the steps of the methods in the above embodiments can be implemented by a program instructing related hardware. The program can be stored in a computer-readable storage medium, and when executed, the program includes one or a combination of the steps of the method embodiments.

[0218] Furthermore, the functional units in the various embodiments of this application can be integrated into a processing module, or each unit can exist physically separately, or two or more units can be integrated into a module. The integrated module can be implemented in hardware or as a software functional module. If the integrated module is implemented as a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium.

[0219] The storage medium mentioned above can be a read-only memory, a disk, or an optical disk, etc. Although embodiments of this application have been shown and described above, it is understood that the above embodiments are exemplary and should not be construed as limiting this application. Those skilled in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of this application.

Claims

1. A method for synthesizing abnormal factual case data for the coal industry, characterized in that, The method includes: Obtain a dictionary set for the coal industry, and based on the job constraint files for the coal industry in the dictionary set, determine the first corpus set for the coal industry. The determination of the first corpus set includes: Text mining is performed on the aforementioned task constraint files to obtain the original structured corpus; For those belonging to the job constraint file D i corpus According to the corpus of Parameters and the job constraint file D i The overall information density is used to determine the job constraint file. D i The first indicator; wherein the corpus For any raw corpus in the structured corpus, the Parameters are used to represent the corpus Chinese regulations and key terms in work constraint documents D i Frequency of occurrence in; Obtain the corpus The inverse document frequency, positional weight, and complexity evaluation value of the term frequency; The corpus is determined based on the inverse document frequency of the term, the positional weight value, and the complexity evaluation value. The second indicator; According to the job constraint file D i The first indicator and the corpus The second indicator determines the corpus The retention probability is: in, σ For the sigmoid function, λ These are the weighting coefficients. τ Information threshold, Representation corpus The second indicator, IR ( D i ) represents the job constraint file D i The primary indicator; According to the corpus The retention probability is used to filter the structured corpus to obtain the first corpus set; The first corpus set is classified to obtain a second corpus set for different business areas of the coal industry; For each business domain, key points are extracted from each corpus in the second corpus set of the business domain to obtain the key point set of the corpus. Anomaly case simulation is performed based on the second corpus set and the key point set of the corpus to obtain a data set of anomaly case cases in the business domain.

2. The method according to claim 1, characterized in that, The corpus The retention probability is used to filter the structured corpus to obtain the first corpus set, including: Response to the corpus If the retention probability is greater than or equal to the set retention threshold, the original corpus is retained. ; Response to the corpus If the retention probability is less than a set retention threshold, the corpus is deleted from the structured corpus. ; The first corpus set is obtained based on the original corpus preserved in the structured corpus.

3. The method according to claim 1, characterized in that, Keypoint extraction is performed on the second corpus to obtain the keypoint set corresponding to the business domain, including: For any corpus in the second corpus set The corpus is extracted by the first intelligent agent. Key points; Obtain the key points and the corpus The similarity, the complexity assessment value of the key points, and the compliance deviation; The similarity, complexity assessment, and compliance deviation are weighted and calculated to obtain the score of the key point; Based on the scoring value, the corpus The corpus is obtained by filtering the key points. The set of key points.

4. The method according to claim 1, characterized in that, The step of simulating abnormal fact cases based on the second corpus set and the key point set of the corpus to obtain a data set of abnormal fact cases in the business domain includes: Obtain the pre-defined descriptive information of the abnormal fact case to be simulated, wherein the descriptive information includes the time, place, subject of the behavior, and content of the behavior of the abnormal fact case; The second intelligent agent is invoked, and the second intelligent agent generates simulated abnormal fact cases of the arbitrary key points based on the description information, the corpus, and the arbitrary key points extracted from the corpus. Based on the simulated abnormal fact cases of any of the aforementioned key points, a dataset of abnormal fact cases in the business domain is obtained.

5. The method according to claim 4, characterized in that, The simulated abnormal fact cases based on the arbitrary key points generate a dataset of abnormal fact cases in the business domain, including: Identify existing anomaly cases in the aforementioned business area; Obtain the semantic similarity between the simulated abnormal fact cases and the existing abnormal fact cases; In response to the semantic similarity being less than a set similarity threshold, the database of the business domain is updated based on the simulated abnormal fact cases to obtain a data set of the abnormal fact cases.

6. The method according to claim 5, characterized in that, The process of updating the database of the business domain based on the simulated abnormal fact cases yields a data set of abnormal fact cases, including: A third intelligent agent is invoked to perform content analysis on the simulated abnormal fact cases, the corpus associated with the simulated abnormal fact cases, and key points, so as to obtain the parsing results of the simulated abnormal fact cases; Based on the simulated abnormal fact cases, the parsing results, and the corpus associated with the simulated abnormal fact cases, triples are constructed; The triples are format-converted to obtain standardized triples, and the database is updated based on the standardized triples to obtain the dataset of the abnormal fact cases.

7. The method according to claim 6, characterized in that, Before updating the database based on the standardized triples to obtain the dataset of anomalous fact cases, the method further includes: Call the large model; The large model automatically verifies the simulated abnormal fact cases based on the analysis results of the simulated abnormal fact cases, the corpus associated with the simulated abnormal fact cases, and the key points.

8. The method according to claim 7, characterized in that, The process of updating the database based on the standardized triples to obtain the dataset of anomalous fact cases includes: After confirming that the simulated abnormal fact case has passed automatic verification, the simulated abnormal fact case is sent to the review device; Receive review data of the simulated abnormal fact cases fed back by the review device; Based on the review data, determine the score for the simulated abnormal fact case; In response to the fact that the score of the simulated abnormal fact case is greater than a set score threshold, the classification label of the corpus associated with the simulated abnormal fact case is obtained; Based on the classification labels and the standardized triples of the simulated anomalous fact cases, the quadruples of the simulated anomalous fact cases are obtained; The quadruple of the simulated abnormal fact case is stored in the database corresponding to the business domain.

9. The method according to claim 8, characterized in that, The method further includes: Cache the simulated abnormal fact cases whose semantic similarity is less than a set similarity threshold in the array; The simulated abnormal fact cases are read from the array and cached, and the simulated abnormal fact cases, the corpus associated with the simulated abnormal fact cases, and the key points are input into the third agent for content analysis.