Method for synthesizing abnormal fact case data for coal industry
Through the abnormal fact case data synthesis method for the coal industry, the problem of lack of standardized data in the coal industry was solved, high-quality training data was generated, the recognition accuracy and reasoning ability of large models were improved, and the intelligent transformation of the coal industry was promoted.
Patent Information
- Application Number
- CN202510678525.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-26
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2045-05-26
AI Technical Summary
The lack of standardized data in the coal industry has led to the inability to train large models that can be applied in the traceability and discrimination of abnormal fact cases.
A method of synthesis of abnormal fact case data for the coal industry is proposed. High-quality training data is generated by obtaining dictionary sets of coal industry, determining corpus sets, classifying corpus, extracting key points and simulating abnormal fact cases.
By generating high-quality abnormal fact case data and supporting training of large models, the accuracy and reasoning capabilities of large models in identifying abnormal fact case cases are improved, and the intelligent transformation of the coal industry is promoted.
Smart Images

Figure CN120196625A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of data processing in intelligent mines, and in particular to a method for synthesizing abnormal factual cases for the coal industry. Background Art
[0002] With the breakthrough development of big model technology in recent years, more and more coal enterprises are accelerating the verticalization of technology, gradually building a big model system that integrates the characteristics of the coal industry, and making a strong combination with the coal industry. In particular, the coal industry has an urgent need for big models in tracing and identifying abnormal factual cases (such as illegal or irregular events), but the lack of standardized data related to the coal industry makes it impossible to train a big model that can be used in tracing and identifying abnormal factual cases. Summary of the invention
[0003] The purpose of this application is to solve one of the technical problems in the related art at least to some extent.
[0004] To this end, the first purpose of this application is to propose a method for synthesizing abnormal fact case data for the coal industry, which can increase standardized data related to the coal industry, thereby providing data support for the training of large models.
[0005] To achieve the above-mentioned purpose, the first embodiment of the present application proposes a method for synthesizing abnormal fact case data for the coal industry, comprising: Acquire a dictionary set of the coal industry, and determine a first corpus set of the coal industry based on a job constraint file of the coal industry in the dictionary set; Classifying the first corpus set to obtain a second corpus set in different business fields of the coal industry; For each business field, extract key points from each corpus in the second corpus set of the business field to obtain a key point set of the corpus; Abnormal fact case simulation is performed based on the second corpus set and the key point set of the corpus to obtain a data set of abnormal fact cases in the business field.
[0006] The synthesis method of abnormal fact case data for the coal industry provided by this application obtains the operation constraint files of the coal industry, performs text mining on the operation constraint files to determine the first corpus set of the coal industry, classifies the first corpus set to obtain the second corpus set for each business area. Further, key points are extracted from each corpus in the second corpus set to obtain the key point set of the corpus. Based on the second corpus set and the key point set of the corpus, abnormal fact case simulation is carried out to obtain the data set of abnormal fact cases in the business area. In this application, through tools such as text mining and data cleaning, standardized item corpora suitable for simulating abnormal fact cases in relevant laws, regulations, and standards of the coal industry can be sorted out and screened. These standardized item corpora can lay a data foundation for the batch simulation of abnormal fact cases in the coal industry, and solve the problems of insufficient abnormal fact cases and diverse formats in the coal industry. Further, based on the key points of the standardized item corpora, abnormal fact cases can be automatically simulated, providing high-quality training data for the recognition of abnormal fact cases by large models in the coal industry, which is beneficial to improving the training effect of large models, and thus beneficial to the inference accuracy of large models, and can accelerate the popularization in the coal industry.
[0007] Additional aspects and advantages of this application will be given in part in the following description, become apparent in part from the following description, or be understood through the practice of this application. Brief Description of the Drawings
[0008] The above and / or additional aspects and advantages of this application will become apparent and easy to understand from the following description of the embodiments in conjunction with the drawings, where: Figure 1 is a schematic flow chart of a synthesis method of abnormal fact case data for the coal industry provided by an embodiment of this application; Figure 2 is a schematic flow chart of another synthesis method of abnormal fact case data for the coal industry provided by an embodiment of this application; Figure 3 is a schematic flow chart of another synthesis method of abnormal fact case data for the coal industry provided by an embodiment of this application; Figure 4 is a schematic flow chart of another synthesis method of abnormal fact case data for the coal industry provided by an embodiment of this application. Detailed Description of the Embodiments
[0009] The embodiments of this application will be described in detail below. Examples of the embodiments are shown in the drawings, where the same or similar reference numerals denote the same or similar elements or elements with the same or similar functions throughout. The embodiments described below with reference to the drawings are exemplary and are intended to explain this application and should not be construed as a limitation of this application.
[0010] With the breakthrough development of large model technology in recent years, more and more coal enterprises have accelerated the vertical implementation of technology, gradually built a large model system integrating the characteristics of the coal industry, and effectively combined it with the actual coal business. These large models related to the coal industry need to possess a large amount of laws, regulations, and industry expertise in the coal industry, as well as strong language understanding and generation capabilities, and be able to process and analyze massive amounts of industry knowledge and data. With the gradual deep integration of large models and the coal industry, they have begun to not only serve as a means of querying industry knowledge and retrieving data, but also extend to practical work scenarios such as the identification of illegal events (which can also be referred to as instances of illegal facts) and the summary of production reports.
[0011] Regarding the need to automatically identify and recognize illegal events in the coal industry through large models, there is a lack of a stable and reliable data acquisition method for such tasks in the related technologies, which has to a certain extent restricted the development and application of large models related to the coal industry. Against this background, it is particularly important to establish a method for simulating or synthesizing abnormal fact cases (such as illegal fact cases), which is of extremely important significance for the training, fine-tuning, and testing of large models related to the coal industry. Taking the identification task of abnormal fact cases (such as illegal fact cases) as an example, abnormal fact cases (such as illegal fact cases) can be used to detect the mastery and understanding of relevant laws and regulations in the coal industry by large models, thereby helping coal-related enterprises or institutions select large models that meet the requirements and assisting in enhancing the capabilities of their own models.
[0012] The starting point for establishing a system for obtaining or synthesizing abnormal fact cases in the coal industry is that there is currently a lack of data related to abnormal fact cases (such as illegal facts), the publicly available related data formats are not unified, and there is a lack of targeted analysis and parsing descriptions. The existing data related to abnormal fact cases (such as illegal facts) is insufficient to support the data volume requirements of large models.
[0013] The following explains the method for synthesizing abnormal fact case data for the coal industry in conjunction with the accompanying drawings.
[0014] Figure 1 FIG. is a schematic flowchart of a method for synthesizing abnormal fact case data for the coal industry provided by an embodiment of the present application. The execution subject of the method for synthesizing abnormal fact case data for the coal industry can be an electronic device or a server, and is not limited here.
[0015] As Figure 1 shown, the method for synthesizing abnormal fact case data for the coal industry may include but is not limited to the following steps: S101. Obtain the dictionary set of the coal industry, and determine the first corpus set of the coal industry based on the operation constraint files in the dictionary set for the coal industry.
[0016] In some embodiments, the operation constraint files of the coal industry may include, but are not limited to, the following files: legal documents of the coal industry, regulatory documents of the coal industry, and standard documents of the coal industry.
[0017] In some embodiments, the dictionary set may include, but is not limited to, operation constraint files such as legal documents of the coal industry, regulatory documents of the coal industry, and standard documents of the coal industry.
[0018] In some embodiments, a structured corpus for the coal industry may be constructed based on the legal documents, regulatory documents, and / or standard documents of the coal industry in the dictionary set.
[0019] In some embodiments, through a text mining engine, semantic parsing and understanding may be performed on the original documents such as legal documents, regulatory documents, and / or standard documents of the coal industry in the dictionary set, and effective clause information may be extracted therefrom. Further, through the text mining engine, the effective clause information is extracted to obtain the original corpus, and the original corpus is structurally processed to obtain a structured corpus. It can be understood that the structured corpus includes the structured data corresponding to the original corpus.
[0020] In some embodiments, after obtaining the structured corpus, the original corpus extracted from the operation constraint files may be refined and filtered to eliminate noise data and retain high-quality corpus suitable for simulating abnormal fact cases. These retained high-quality corpus suitable for simulating abnormal fact cases are used to constitute the first corpus set.
[0021] S102. Classify the first corpus set to obtain the second corpus set for different business areas of the coal industry.
[0022] In some embodiments, the business areas of the coal industry may include, but are not limited to: coal exploration, underground mining, open-pit mining, coal processing, coal transportation, coal storage, etc.
[0023] In some embodiments, a corpus may be mapped to multiple business areas. A multi-classification model may be pre-trained, and the corpus in the first corpus set is classified and identified through the multi-classification model to obtain the classification label of each corpus, where the classification label is used to indicate the business area to which the corpus belongs.
[0024] In some embodiments, after obtaining the classification label of each corpus in the first corpus set, based on the classification label, the corpus in the first corpus set is mapped to different business areas, and then the second corpus set corresponding to each business area is obtained.
[0025] In some embodiments, the classification labels of the corpus can be standardized and stored.
[0026] S103. For each business area, extract key points from each corpus in the second corpus set of the business area to obtain a set of key points for the corpus.
[0027] In some embodiments, key points can be extracted from each corpus in the second corpus set of each business area to obtain at least one key point for each corpus, and the at least one key point for each corpus can be screened to obtain a set of key points for the corpus.
[0028] In some embodiments, a large model is used to perform semantic analysis on the corpus to extract as many valuable key points as possible from the corpus. It can be understood that the extracted valuable key points can be used to simulate abnormal fact cases, that is, the extracted valuable key points may be potential violation points in the operation.
[0029] In some embodiments, among the valuable key points extracted from the corpus, there may be key points with relatively weak industry relevance or high overlap. The key points with relatively weak industry relevance and high overlap can be removed from the valuable key points, and then a set of key points for the corpus is obtained.
[0030] In some embodiments, the key points in the set of key points for the corpus can be sorted out. For example, these key points can be formatted to obtain key points in a standardized format, and the key points in the protected format are stored.
[0031] S104. Simulate abnormal fact cases based on the second corpus set and the set of key points for the corpus to obtain a data set of abnormal fact cases for the business area.
[0032] In some embodiments, the abnormal fact cases can include, but are not limited to: events violating legal documents of the coal industry, events violating regulatory documents of the coal industry, and events violating standard documents of the coal industry.
[0033] In some implementations, after obtaining the set of key points for each corpus, the description information of the abnormal fact case to be simulated can be configured. Optionally, the description information can include, but is not limited to: the occurrence time, location, actor, and action content of the abnormal fact case.
[0034] Furthermore, based on the corpus and the key points of the corpus, combined with the pre-configured description information of the abnormal fact case, the abnormal fact case is simulated to obtain a data set of abnormal fact cases for the business area.
[0035] It can be understood that the data set of abnormal fact cases in the business domain can include not only simulated abnormal fact cases but also abnormal fact cases collected in actual operations.
[0036] In the embodiments of the present application, an operation constraint file of the coal industry can be obtained, text mining can be performed on the operation constraint file to determine a first corpus set of the coal industry, and the first corpus set can be classified to obtain a second corpus set for each business domain. Further, key points of each corpus in the second corpus set are extracted to obtain a key point set of the corpus, and abnormal fact cases in the business domain are simulated based on the second corpus set and the key point set of the corpus, so as to obtain a data set of abnormal fact cases in the business domain. In the present application, through tools such as text mining and data cleaning, standardized entry corpora suitable for simulating abnormal fact cases in relevant laws, regulations, and standards of the coal industry can be sorted out and screened. These standardized entry corpora can lay a data foundation for the batch simulation of abnormal fact cases in the coal industry, making up for the problems of insufficient abnormal fact cases and diverse formats in the industry. Further, based on the key points of the standardized entry corpora, abnormal fact cases can be automatically simulated, providing high-quality training data for the recognition of abnormal fact cases by the large model in the coal industry, which is beneficial to improving the training effect of the large model, and thus beneficial to the inference accuracy of the large model, and can accelerate the popularization in the coal industry.
[0037] Figure 2 It is a schematic flowchart of another method for synthesizing abnormal fact case data for the coal industry provided by the embodiments of the present application. As Figure 2 shown, the method for synthesizing abnormal fact case data for the coal industry may include but is not limited to the following steps: S201, obtain a dictionary set of the coal industry.
[0038] In some embodiments, legal documents, regulatory documents, and standard documents related to the coal industry are obtained as operation constraint files of the coal industry. Further, a dictionary set is established through the operation constraint files Doc ={ D 1, D 2, …, D n}, where D i represents a certain original operation constraint file. For example, it can be a regulatory document, a legal document, or a standard document. Optionally, the original operation constraint file can be a pdf file or a Word file, and the file type of the operation constraint file is not limited in the present application.
[0039] S202, perform text mining on the operation constraint file to obtain an original structured corpus.
[0040] In some embodiments, the text mining model R can be used to extract the original corpus one by one from the original job constraint file D i wherein represents the original structured corpus, represents the original structured corpus, represents the original structured corpus, D i the j th original corpus extracted from the original job constraint file
[0041] S203. For the corpus D i subordinate to the job constraint file according to the corpus and the overall information density of the job constraint file D i determine the first index of the job constraint file D i .
[0042] It can be understood that the corpus is any original corpus in the structured corpus
[0043] In some embodiments, there is a mapping relationship between each corpus and the original job constraint file, and based on the mapping relationship, the job constraint file to which each corpus belongs can be determined
[0044] In some embodiments, for the corpus D i of the job constraint file the parameters of the corpus can be obtained, wherein the parameter is used to represent the frequency of occurrence of the regulatory keywords in the corpus in the job constraint file D i .
[0045] In some embodiments, the overall information density of the job constraint file D i is obtained. Further, after the overall information density is obtained, the Frobenius norm of the overall information density can be determined
[0046] In some embodiments, through word frequency analysis, the frequency of occurrence of each word in the job constraint file D i can be calculated, and based on the frequency of occurrence of each word, the job constraint file D iThe overall information density.
[0047] In some embodiments, a job constraint file can be obtained D i The term frequency-inverse document frequency (TF-IDF) of each word in it is determined, and based on the TF-IDF of each word, the overall information density of the job constraint file D i is determined.
[0048] In some embodiments, the entropy of the content of the job constraint file can be calculated D i Based on the entropy of the content of the job constraint file D i the overall information density of the job constraint file is determined. D i is determined.
[0049] Furthermore, according to the corpus of , and the Frobenius norm of the overall information density of the job constraint file D i the first index of the job constraint file can be determined. D i is determined.
[0050] In some embodiments, based on the following formula (1), the first index of the job constraint file D i is determined:
[0051] wherein, IR ( D i ) represents the first index of the job constraint file D i ; ω is the expert experience weight coefficient; represents the Frobenius norm of the overall information density of the job constraint file D i ; n represents the number of corpora extracted from the job constraint file D i , n and the value of
[0052] In some embodiments, the information retention degree of the job constraint file D i is represented by the first index of the job constraint file D i .
[0053] S204, Obtain the corpus 's term frequency-inverse document frequency, position weight value, and complexity evaluation value.
[0054] S205, Determine the second index of the corpus according to the term frequency-inverse document frequency, position weight value, and complexity evaluation value of the corpus.
[0055] In some embodiments, the term frequency (TF) and inverse document frequency (IDF) of the corpus can be obtained. Further, perform a multiplication operation on the TF and IDF of the corpus to obtain the TF_IDF of the corpus, which can be expressed as .
[0056] In some embodiments, the position of the corpus in the job constraint file can be located. According to the position of the corpus D i in the job constraint file, the corresponding position weight value of the corpus in the job constraint file D i can be calculated. Optionally, a position weight function Pos of the corpus can be pre-constructed, and the corresponding position weight value of the corpus is determined based on the position weight function Pos. That is to say, can represent the position weight value of the corpus of the corpus.
[0057] In some embodiments, the complexity evaluation value of the corpus can be determined based on information entropy. Optionally, a complexity evaluation function Ent of the corpus is pre-constructed, and the corresponding complexity evaluation value of the corpus is determined based on the complexity evaluation function Ent. That is to say, can represent the complexity evaluation value of the corpus of the corpus. It should be noted that the higher the information entropy, the greater the uncertainty of the information, and the higher the complexity of the key points z
[0058] Further, after obtaining the term frequency-inverse document frequency, position weight value, and complexity evaluation value of the corpus , the second index of the corpus can be determined based on the term frequency-inverse document frequency, position weight value, and complexity evaluation value of the corpus of the corpus.
[0059] Optionally, the following formula (2) can be used to determine the second index of the corpus :
[0060] where represents the second index of the corpus . Optionally, the second index can represent the information value of the corpus .
[0061] α, β, γ are weight coefficients respectively, and can be pre-calibrated parameters
[0062] is the term frequency-inverse document frequency of the corpus .
[0063] is the position weight value of the corpus .
[0064] is the complexity evaluation value based on information entropy of the corpus .
[0065] S206, filter the structured corpus according to the first index of the job constraint file D i and the second index of the corpus to obtain the first corpus set
[0066] In some embodiments, according to the second index of the corpus and the first index of the job constraint file D i , determine the retention probability of the corpus .
[0067] In some embodiments, the following formula (3) can be used to determine the retention probability of the corpus :
[0068] where represents the retention probability of the corpus ; σ is the sigmoid function λ is the weight coefficient τ is the information threshold
[0069] After obtaining the retention probability of the corpus , compare the retention probability of the corpus with the set retention threshold to determine whether to retain the corpus 。
[0070] In some embodiments, in response to the retention probability of the corpus being greater than or equal to a set retention threshold, retain the original corpus 。
[0071] In some embodiments, in response to the retention probability of the corpus being less than the set retention threshold, delete the corpus from the original structured corpus 。
[0072] Further, based on the original corpus retained in the structured corpus, obtain a first corpus set.
[0073] It can be understood that based on the noise filtering condition that the retention probability is greater than or equal to the set retention threshold, the original corpus in the original structured corpus can be standardized and filtered, that is, the original corpus belonging to the structured corpus can be retained when it meets the filtering conditions ( ), and further, based on the retained original corpus, construct a first corpus set that conforms to the simulated abnormal fact cases 。
[0074] S207. Classify the first corpus set to obtain a second corpus set for different business fields in the coal industry.
[0075] For the specific introduction of step S207, reference can be made to the relevant records in the embodiments of the present application, which will not be elaborated here.
[0076] S208. For each business field, extract key points from each corpus in the second corpus set of the business field to obtain a key point set of the corpus.
[0077] For the specific introduction of step S208, reference can be made to the relevant records in the embodiments of the present application, which will not be elaborated here.
[0078] S209. According to the second corpus set and the key point set of the corpus, simulate abnormal fact cases to obtain a data set of abnormal fact cases in the business field.
[0079] For the specific introduction of step S209, reference can be made to the relevant records in the embodiments of the present application, which will not be elaborated here.
[0080] In the embodiment of the present application, for the entry corpus in the documents such as regulations and standards that are sorted out in batches, the entry corpus with weak knowledge and high repetition can be eliminated through screening and filtering mechanisms, which is conducive to improving the accuracy and professionalism of the original corpus. Furthermore, through the evaluation of indicators such as word frequency inverse file frequency, position weight, information entropy, etc., the overall information retention of the file and the information value of a single corpus are calculated respectively, and the retention calculation of the corpus is combined with the comprehensive evaluation of the two to eliminate the corpus with weak knowledge, high repetition, and strong noise. This process not only improves the quality and efficiency of the extracted original corpus, but also ensures the effectiveness of the application of the obtained simulated abnormal fact cases, so that it can better reflect the actual needs and capabilities of the industry, and then optimize the accuracy of the large model related to the coal industry in identifying illegal and irregular behaviors.
[0081] Figure 3 A flow chart of another method for synthesizing abnormal fact case data for the coal industry provided in an embodiment of the present application. Figure 3 As shown, the synthesis method of abnormal fact case data for the coal industry may include but is not limited to the following steps: S301, obtaining a dictionary set of the coal industry, and determining a first corpus set of the coal industry based on the operation constraint file of the coal industry in the dictionary set.
[0082] For a detailed description of step S301, please refer to the relevant records in the various embodiments of the present application, which will not be repeated here.
[0083] S302, classify the first corpus set to obtain a second corpus set of different business fields in the coal industry.
[0084] In some embodiments, the set corpus classification set can be expressed as Class ={ C 1, C 2, …, C n}, respectively used to represent specific business areas in the coal industry such as underground mining and open-pit mining, among which C i Indicates a business area.
[0085] In some embodiments, a multi-classification model may be called to perform classification operations on the original corpus in the first corpus set, thereby obtaining at least one classification label for each original corpus.
[0086] For example, the multi-classification model is C, and the original corpus in the first corpus set is classified by the multi-classification model C. By performing classification operations, we can get the original corpus That is, the original corpus Input into the multi-classification model C, and output the original corpus through the multi-classification model C with one or more classification labels. The process of outputting the classification labels of the original corpus through the multi-classification model C can be characterized as and .
[0087] It can be understood that any corpus in the first corpus set may be mapped to multiple specific business areas. For example, any corpus in the first corpus set can be classified into at least one specific business area. That is to say can be classified into at least one specific business area.
[0088] S303. For any corpus in the second corpus set , extract the key points of the corpus through the first agent.
[0089] S304. Obtain the similarity between the key points and the corpus , the complexity evaluation value of the key points, and the compliance deviation degree.
[0090] S305. Perform weighted calculation on the similarity, complexity evaluation value, and compliance deviation degree corresponding to the key points to obtain the scoring value of the key points.
[0091] S306. Filter the key points of the corpus according to the scoring value of the key points to obtain the key point set of the corpus .
[0092] In some embodiments, for any corpus in the second corpus set , the key points of the corpus can be extracted. Optionally, the pre-constructed first agent can be called to extract the key points of the corpus through the first agent.
[0093] Exemplarily, the first agent is Z. Through the first agent Z, the key points of the corpus are extracted to obtain at least one key point of the corpus . Optionally, the corpus z is input into the first agent Z, and through the first agent Z, one or more key points can be extracted from the corpus . The process of extracting the key points of the corpus z through the first agent Z can be characterized as and and .
[0094] In some embodiments, the corpus of each key point z can be obtained. The similarity, the complexity evaluation value, and the compliance deviation degree of the key point are further obtained. Further, weighted calculations are performed on the similarity, the complexity evaluation value, and the compliance deviation degree corresponding to the key point z to obtain the scoring value of the key point. z
[0095] Optionally, the scoring value of the key point z can be determined using the following formula (4):
[0096] It should be noted that the key point z is the key point extracted from the corpus by the first intelligent agent Z, and the number of key points extracted from the corpus is greater than or equal to 2. That is, the following conditions need to be satisfied: and and .
[0097] Among them, 1, λ 2, λ 3 in formula (4) are weight coefficients, which are parameters that can be calibrated in advance. λ
[0098] Sim is a cosine similarity calculation function used to calculate the cosine similarity between the key point z and the corpus from which it is derived. The cosine similarity can characterize the semantic correlation between the key point and the corpus from which it is derived. z and the corpus from which it is derived.
[0099] Ent represents a complexity evaluation function based on information entropy, which is used to evaluate the complexity of the key point z .
[0100] Compl represents a compliance deviation degree calculation function, which is used to evaluate the compliance deviation degree of the key point z . Optionally, the compliance deviation degree calculation function can be a function pre-trained based on compliance knowledge points. Through this function, the compliance deviation degree of the key point z can be scored. It can be understood that the higher the score of the key point z , the greater the compliance deviation degree, indicating that the key point z is more likely to be non-compliant; the lower the score of the key point z , the smaller the compliance deviation degree, indicating that the key point z is less likely to be non-compliant.
[0101] It is understandable that the scoring value of each key point extracted from the corpus can be calculated through formula (4). Further, each key point extracted from the corpus is filtered according to the scoring value, so as to eliminate the key points with weak corpus relevance, blurred clarity, and semantic repetition from all key points, and obtain the key point set of the corpus: wherein, ≥1, and any key point in the key point set of the corpus is a high-value key point for simulating abnormal fact cases. And n ≥1, where any key point in the key point set of the corpus is a high-value key point for simulating abnormal fact cases. is a high-value key point for simulating abnormal fact cases.
[0102] S307. Simulate abnormal fact cases according to the second corpus set and the key point set of the corpus, and obtain a data set of abnormal fact cases in the business domain.
[0103] For the specific introduction of step S307, reference can be made to the relevant records in the embodiments of the present application, which will not be elaborated here.
[0104] In the embodiments of the present application, by obtaining the second corpus set from documents such as coal industry regulations and standards, extracting key points from the corpus of the second corpus set through the first intelligent agent, scoring the key points, and filtering and screening the key points based on the scoring value, key points with strong corpus relevance, high clarity, and non-repeated semantics can be retained, which is beneficial to generating more accurate and effective simulated abnormal fact cases in the subsequent process, thereby being able to increase the data of abnormal fact cases in the coal industry. Moreover, using documents such as regulations and standards as the basis for simulating abnormal fact cases can solve the problem of inconsistent understanding of abnormal fact cases.
[0105] Figure 4 is a schematic flow chart of a method for synthesizing abnormal fact case data for the coal industry provided by the embodiments of the present application. As Figure 4 shown, the method for synthesizing abnormal fact case data for the coal industry may include but is not limited to the following steps: S401. Obtain a dictionary set of the coal industry, and determine the first corpus set of the coal industry based on the operation constraint documents in the dictionary set.
[0106] For the specific introduction of step S401, reference can be made to the relevant records in the embodiments of the present application, which will not be elaborated here.
[0107] S402. Classify the first corpus set to obtain a second corpus set of different business domains in the coal industry.
[0108] For the specific introduction of step S402, reference can be made to the relevant records in the embodiments of the present application, which will not be elaborated here.
[0109] S403. Extract key points from each corpus in the second corpus set of the business domain to obtain a key point set of the corpus.
[0110] For the specific introduction of step S403, reference can be made to the relevant records in the embodiments of the present application, which will not be elaborated here.
[0111] S404. Obtain the preset description information of the abnormal fact case to be simulated.
[0112] In some embodiments, the description information includes the occurrence time, location, actor, and action content of the abnormal fact case.
[0113] Optionally, the description information of the abnormal fact case to be simulated can form a quadruple data. For example, the quadruple data of the abnormal fact case to be simulated can be marked as Ω = ( tim , pos , chara , act ), where tim represents the occurrence time of the abnormal fact case to be simulated; pos represents the occurrence location of the abnormal fact case to be simulated; chara represents the actor of the abnormal fact case to be simulated; act represents the action content of the abnormal fact case to be simulated.
[0114] S405. Invoke the second intelligent agent, and through the second intelligent agent, generate a simulated abnormal fact case of any key point according to the description information, corpus, and any key point extracted from the corpus.
[0115] In some embodiments, the description information, corpus, and any key point extracted from the corpus can be input into the second intelligent agent, and through the second intelligent agent, output a simulated abnormal fact case of any key point according to the description information, corpus, and any key point extracted from the corpus.
[0116] Exemplarily, the second intelligent agent is Generate, any corpus is , and any key point extracted from the corpus is . Input , , and the quadruple data Ω = ( tim , pos , chara , act ) into the second intelligent agent Generate, and the second intelligent agent Generate processes the input data to output a simulated abnormal fact case 。In this application, the process of generating a simulated abnormal fact case based on the output of the second agent Generate can be expressed as 。
[0117] S406, obtaining a data set of abnormal fact cases in the business domain based on the simulated abnormal fact cases at any key points.
[0118] In some embodiments, existing abnormal fact cases in the business domain are determined. Further, the semantic similarity between the simulated abnormal fact cases at any key points and the existing abnormal fact cases is obtained. In response to the semantic similarity being less than the set similarity threshold, the database in the business domain is updated based on the simulated abnormal fact cases to obtain a data set of abnormal fact cases; in response to the semantic similarity being greater than or equal to the set similarity threshold, the simulated abnormal fact cases are deleted.
[0119] In some embodiments, when it is determined that the semantic similarity of the simulated abnormal fact cases is less than the set similarity threshold, the simulated abnormal fact cases can be cached through an array for subsequent parsing processes to read from the array.
[0120] Exemplarily, the following formula (5) is used to calculate the semantic similarity between the simulated abnormal fact cases and the existing abnormal fact cases in the business domain q :
[0121] where represents and q 's semantic similarity; and respectively represent and q 's semantic vectors; θ represents the set similarity threshold.
[0122] Optionally, if the semantic similarity is greater than or equal to the similarity threshold θ then it is determined that the simulated abnormal fact case and the existing abnormal fact case q are duplicates, and the simulated abnormal fact case can be discarded.
[0123] Optionally, if the semantic similarity is less than the similarity threshold θ then it is determined that the simulated abnormal fact case and the existing abnormal fact case q are not duplicates, and the simulated abnormal fact case can be stored in the array 。
[0124] In some embodiments, to ensure the interpretability of the content of the simulated abnormal fact cases, corresponding analysis results need to be generated for the simulated abnormal fact cases. Determine the corpus (i.e., the source corpus) and key points associated with the simulated abnormal fact cases. Combine the corpus and key points associated with the simulated abnormal fact cases, and perform content analysis on the simulated abnormal fact cases to obtain the analysis results of the simulated abnormal fact cases.
[0125] In some embodiments, a third intelligent agent can be invoked to perform content analysis on the simulated abnormal fact cases through the third intelligent agent to obtain the analysis results of the simulated abnormal fact cases. Optionally, the corpus and key points associated with the simulated abnormal fact cases, as well as the simulated abnormal fact cases, are input into the third intelligent agent. The third intelligent agent performs content analysis on the simulated abnormal fact cases according to the corpus and key points associated with the simulated abnormal fact cases to obtain the analysis results of the simulated abnormal fact cases. In the implementation of this application, not only can the problems of lack of standardized data and lack of parsing descriptions be solved, but also the quality and efficiency of simulating abnormal fact cases can be improved, providing high-quality data for the identification of abnormal fact cases in large models related to the coal industry.
[0126] In some embodiments, the simulated abnormal fact cases with a semantic similarity less than the set similarity threshold cached in the array can be read and input into the third intelligent agent for parsing together with the corpus and key points associated with the simulated abnormal fact cases to obtain the analysis results of the simulated abnormal fact cases.
[0127] Exemplarily, the third intelligent agent is Analysis. The simulated abnormal fact cases , the associated corpus and key points are input into the third intelligent agent Analysis, and the analysis results are output through the third intelligent agent Analysis. In this application, the process of performing content analysis on by the third intelligent agent Analysis can be represented as .
[0128] Furthermore, according to the simulated abnormal fact cases, the analysis results, and the corpus associated with the simulated abnormal fact cases, triples are constructed, and the format of the triples is converted to obtain standardized triples. Based on the standardized triples, the database is updated to obtain a data set of abnormal fact cases.
[0129] Exemplarily, the simulated abnormal fact cases , the analysis results , and the associated source corpus form a triple , Further, the triples are standardized using a specific JSON format, and the standardized triples are stored. For example, the standardized triples can be stored in the group .
[0130] In some embodiments, before updating the database based on the simulated abnormal fact cases to obtain a data set of abnormal fact cases, a large model can also be called to automatically verify the simulated abnormal fact cases through the large model. Optionally, the large model can automatically verify the simulated abnormal fact cases in combination with the parsing results of the simulated abnormal fact cases, the associated corpus, and key points.
[0131] In some embodiments, after determining that the simulated abnormal fact cases pass the automatic verification, the simulated abnormal fact cases are sent to the review device. The simulated abnormal fact cases can be displayed on the review device. The review device can monitor the review data of the simulated abnormal fact cases input by the reviewers and provide feedback on the review data of the simulated abnormal fact cases. The execution entity of the present application can receive the review data of the simulated abnormal fact cases fed back by the review device. Further, according to the review data, the scoring value of the simulated abnormal fact cases is determined. In response to the scoring value of the simulated abnormal fact cases being greater than the set scoring threshold, the classification label of the corpus associated with the simulated abnormal fact cases is obtained, and according to the classification label and the standardized triples of the simulated abnormal fact cases, a quadruple of the simulated abnormal fact cases is obtained. Finally, the quadruple of the simulated abnormal fact cases is stored in the database corresponding to the business domain.
[0132] In some embodiments, the quadruple of the simulated abnormal fact cases includes the simulated abnormal fact cases, the parsing results of the simulated abnormal fact cases, the associated corpus, and the classification label of the corpus.
[0133] In some embodiments, the review data of the simulated abnormal fact cases may include, but is not limited to, review values in dimensions such as content reasonableness, legal and regulatory consistency, and coal industry relevance.
[0134] In some embodiments, the review values of each dimension in the review data are weighted and calculated to obtain the scoring value of the simulated abnormal fact cases.
[0135] Exemplarily, a large model related to the coal industry (for example, it can be called a mine large model) is used to automatically verify the simulated abnormal fact cases That is to say, the large model related to the coal industry can be based on the simulated abnormal fact cases Parsing results , Source corpus , Corresponding key points , To verify the simulated abnormal fact cases Conduct a legal consistency assessment and a verification of relevance to the coal industry to ensure the accuracy and reliability of the simulated abnormal fact cases.
[0136] Optionally, configure the consistency threshold in advance θ g and the relevance threshold θ l , and the large model related to the coal industry can output the consistency parameters and relevance parameters of the simulated abnormal fact cases . Compare the consistency parameters of the simulated abnormal fact cases with the consistency threshold θ g , and compare the relevance parameters of the simulated abnormal fact cases with the relevance threshold θ l . After the consistency parameters and relevance parameters of the simulated abnormal fact cases pass the determination of the consistency threshold θ g and the relevance threshold θ l , the large model related to the coal industry can pass the automatic verification of the simulated abnormal fact cases .
[0137] Optionally, after the automatic verification passes, the simulated abnormal fact cases can be manually reviewed. The simulated abnormal fact cases can be sent to the review device, and then domain experts can conduct a detailed inspection of the simulated abnormal fact cases on the review device. It can be understood that the experts comprehensively evaluate the content rationality, legal consistency, and relevance to the coal industry of the simulated abnormal fact cases based on professional knowledge and experience.
[0138] Furthermore, the review device can feedback the review data of the domain experts to the execution entity of this application, and the execution entity of this application can calculate the score value of the simulated abnormal fact cases based on the scoring formula. Optionally, the review device can calculate the score value of the simulated abnormal fact cases based on the scoring formula, obtain the score value of the simulated abnormal fact cases, and feedback the score value of the simulated abnormal fact cases to the execution entity of this application.
[0139] For example, the scoring formula for the simulated abnormal fact cases is Sorce = αa + βr + γu , where α , β , γ are the weight coefficients, a is the review value of rationality, r is the review value of consistency, u is the review value of relevance.
[0140] After determining the scoring value of the simulated abnormal fact case Sorce then, when the scoring value Sorce exceeds a predetermined threshold Sorce min can the artificial verification of the simulated abnormal fact case be passed.
[0141] Optionally, the simulated abnormal fact cases that pass the automatic verification and artificial verification are processed in the format of a quadruple to obtain the quadruple of the simulated abnormal fact case and store it in the database of the business domain.
[0142] On the basis of the above embodiments, after obtaining the data set of abnormal fact cases in different business domains, the abnormal recognition model of the business domain can be trained based on the data set of abnormal fact cases to obtain the target abnormal recognition model of the business domain, and then the abnormal operation behavior can be recognized based on the target abnormal recognition model, so as to achieve the purpose of standardizing the operation behavior in the coal industry.
[0143] In the embodiments of the present application, by verification, the simulated abnormal fact cases that meet the requirements can be retained, which can significantly improve the simulation efficiency and quality of the abnormal fact cases in the coal industry. Moreover, each module works together, from the extraction and screening of the original corpus, to the synthesis and analysis of the abnormal fact cases, and finally through double verification and formatting for storage in the database, forming a standardized and highly reusable abnormal fact case acquisition process.
[0144] Furthermore, through the organic combination of automatic verification and manual review, the rationality of the synthesized simulated abnormal fact cases, the consistency with the corpus, and the relevance to the coal industry are ensured. This automated process not only improves the efficiency of event synthesis, but also enhances the standardization level of abnormal fact cases, making the abnormal fact cases more scientific and consistent, making the training and reasoning of large models related to the coal industry more accurate, and providing technical support for the intelligent transformation of the coal industry.
[0145] To implement the above method for synthesizing abnormal fact case data for the coal industry, the embodiments of the present application also provide an apparatus for synthesizing abnormal fact case data for the coal industry. The apparatus for synthesizing abnormal fact case data for the coal industry includes: a first corpus acquisition module, a second corpus acquisition module, a key point acquisition module, and an event simulation module.
[0146] The first corpus acquisition module is used to acquire the dictionary set of the coal industry and determine the first corpus set of the coal industry based on the operation constraint file of the coal industry in the dictionary set; The second corpus acquisition module is used to classify the first corpus set to obtain the second corpus set of different business domains in the coal industry; A key point acquisition module, configured to extract key points from each corpus in the second corpus set of the business domain for each business domain, to obtain a key point set of the corpus; An event simulation module, configured to simulate abnormal fact cases according to the second corpus set and the key point set of the corpus, to obtain a data set of abnormal fact cases of the business domain.
[0147] In some embodiments, the first corpus acquisition module is further configured to: Perform text mining on the job constraint file to obtain an original structured corpus; For the corpus D i belonging to the job constraint file , according to the corpus and the overall information density of the job constraint file D i , determine a first index of the job constraint file D i ; where the corpus is any original corpus in the structured corpus; Obtain the term frequency-inverse document frequency, position weight value, and complexity evaluation value of the corpus ; According to the term frequency-inverse document frequency, the position weight value, and the complexity evaluation value, determine a second index of the corpus ; According to the first index of the job constraint file D i and the second index of the corpus , filter the structured corpus to obtain the first corpus set.
[0148] In some embodiments, the first corpus acquisition module is further configured to: According to the second index of the corpus and the first index of the job constraint file D i , determine the retention probability of the corpus ; In response to the retention probability of the corpus being greater than or equal to a set retention threshold, retain the original corpus ; In response to the retention probability of the corpus being less than the set retention threshold, delete the corpus from the structured corpus ; Based on the original corpus retained in the structured corpus, obtain the first corpus set.
[0149] In some embodiments, the key point acquisition module is further configured to: For any corpus in the second corpus set , extract the key points of the corpus through the first agent; Obtain the similarity between the key points and the corpus , the complexity evaluation value and compliance deviation degree of the key points; Perform weighted calculation on the similarity, the complexity evaluation value and the compliance deviation degree to obtain the scoring value of the key points; Filter the key points of the corpus according to the scoring value to obtain the key point set of the corpus .
[0150] In some embodiments, the event simulation module is further configured to: Obtain the preset description information of the abnormal fact case to be simulated, where the description information includes the occurrence time, location, actor and action content of the abnormal fact case; Call the second agent, and through the second agent, generate a simulated abnormal fact case of the arbitrary key point according to the description information, the corpus and any key point extracted from the corpus; Based on the simulated abnormal fact case of the arbitrary key point, obtain the data set of the abnormal fact cases in the business domain.
[0151] In some embodiments, the event simulation module is further configured to: Determine the existing abnormal fact cases in the business domain; Obtain the semantic similarity between the simulated abnormal fact case and the existing abnormal fact case; In response to the semantic similarity being less than the set similarity threshold, update the database in the business domain based on the simulated abnormal fact case to obtain the data set of the abnormal fact cases.
[0152] In some embodiments, the event simulation module is further configured to: Perform content analysis on the simulated abnormal fact case, the corpus and key points associated with the simulated abnormal fact case to obtain the analysis result of the simulated abnormal fact case; Construct a triple according to the simulated abnormal fact case, the analysis result and the corpus associated with the simulated abnormal fact case; Perform format conversion on the triple to obtain a standardized triple, and update the database based on the standardized triple to obtain the data set of the abnormal fact cases.
[0153] In some embodiments, the event simulation module is further configured to: Before updating the database based on the standardized triples to obtain the data set of the abnormal fact cases, call the large model; Automatically verify the simulated abnormal fact case through the large model according to the parsing result of the simulated abnormal fact case, the corpus and key points associated with the simulated abnormal fact case.
[0154] In some embodiments, the event simulation module is further configured to: After determining that the simulated abnormal fact case passes the automatic verification, send the simulated abnormal fact case to the review device; Receive the review data of the simulated abnormal fact case fed back by the review device; Determine the scoring value of the simulated abnormal fact case according to the review data; In response to the scoring value of the simulated abnormal fact case being greater than the set scoring threshold, obtain the classification label of the corpus associated with the simulated abnormal fact case; Obtain the quadruple of the simulated abnormal fact case according to the classification label and the standardized triple of the simulated abnormal fact case; Store the quadruple of the simulated abnormal fact case in the database corresponding to the business domain.
[0155] The synthesis device for abnormal fact case data for the coal industry provided by the embodiments of the present application can obtain the operation constraint files of the coal industry, perform text mining on the operation constraint files, determine the first corpus set of the coal industry, and classify the first corpus set to obtain the second corpus set of each business domain. Further, key points of each corpus in the second corpus set are extracted to obtain the key point set of the corpus. Based on the second corpus set and the key point set of the corpus, abnormal fact cases are simulated to obtain the data set of abnormal fact cases in the business domain. In the present application, through tools such as text mining and data cleaning, standardized item corpora suitable for simulating abnormal fact cases in the relevant laws, regulations and standards of the coal industry can be sorted out and screened. These standardized item corpora can lay a data foundation for the batch simulation of abnormal fact cases in the coal industry, and make up for the problems of insufficient abnormal fact cases and diverse formats in the industry. Further, based on the key points of the standardized item corpora, abnormal fact cases can be automatically simulated, providing high-quality data for the identification of abnormal fact cases by the large model in the coal industry, which is beneficial to training the large model and significantly improving the use effect and accuracy of the large model in the coal industry.
[0156] To implement the above embodiments, the present application further provides an electronic device, including: a processor, and a memory communicatively connected to the processor; the memory stores computer-executable instructions; the processor executes the computer-executable instructions stored in the memory to implement the method provided in the foregoing embodiments. To implement the above embodiments, the present application further provides a computer-readable storage medium storing computer-executable instructions, which are used to implement the method provided in the foregoing embodiments when executed by a processor.
[0157] To implement the above embodiments, the present application further provides a computer program product including a computer program, which implements the method provided in the foregoing embodiments when executed by a processor.
[0158] The collection, storage, use, processing, transmission, provision, and application of the user's personal information involved in the present application all comply with the provisions of relevant laws and regulations and do not violate public order and good customs.
[0159] It should be noted that personal information from users should be collected for legal and reasonable purposes and should not be shared or sold outside of these legal uses. In addition, such collection / sharing should be carried out after obtaining the informed consent of the user, including but not limited to notifying the user to read the user agreement / user notice and sign an agreement / authorization including authorizing relevant user information before the user uses the function. In addition, any necessary steps should be taken to protect and secure access to such personal information data and ensure that others with access to the personal information data comply with their privacy policies and procedures.
[0160] The present application anticipates providing embodiments that allow users to selectively block the use or access of personal information data. That is, the present application anticipates providing hardware and / or software to prevent or block access to such personal information data. Once personal information data is no longer needed, the risk can be minimized by restricting data collection and deleting the data. In addition, when applicable, personal identifiers are removed from such personal information to protect the privacy of the user.
[0161] In the descriptions of the foregoing embodiments, the descriptions referring to terms such as "one embodiment", "some embodiments", "example", "specific example", or "some examples" mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present application. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described may be combined in any one or more embodiments or examples in a suitable manner. In addition, without contradiction, those skilled in the art may combine and combine the different embodiments or examples described in this specification and the features of the different embodiments or examples.
[0162] In addition, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one of such features. In the description of the present application, "a plurality of" means at least two, such as two, three, etc., unless otherwise specifically defined.
[0163] Any process or method description in a flowchart or described in other ways herein may be understood to represent a module, segment, or portion of code including one or more executable instructions for implementing a customized logic function or process, and the scope of the preferred embodiments of the present application includes additional implementations, where the functions may be executed in a manner that is not shown or discussed, including in a substantially simultaneous manner or in a reverse order according to the functions involved, which should be understood by those skilled in the art to which the embodiments of the present application pertain.
[0164] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a definitional sequence list of executable instructions for implementing logical functions, and can be specifically implemented in any computer-readable medium for use by an instruction execution system, apparatus, or device (such as a computer-based system, a system including a processor, or other systems that can fetch and execute instructions from the instruction execution system, apparatus, or device), or in conjunction with these instruction execution systems, apparatus, or devices. For the purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by or in conjunction with an instruction execution system, apparatus, or device. More specific examples (a non-exhaustive list) of computer-readable media include the following: an electrical connection portion having one or more wirings (electronic device), a portable computer diskette (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disc read-only memory (CDROM). Additionally, the computer-readable medium can even be paper or other suitable media on which the program can be printed, since the program can be obtained electronically, for example, by optically scanning the paper or other media, followed by editing, interpretation, or otherwise processing as appropriate, and then storing it in a computer memory.
[0165] It should be understood that various parts of the present application can be implemented using hardware, software, firmware, or a combination thereof. In the above-described embodiments, multiple steps or methods can be implemented using software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented using hardware, as in another embodiment, any one or a combination of the following techniques well known in the art can be used: discrete logic circuits having logic gate circuits for implementing logical functions on data signals, application specific integrated circuits having suitable combinational logic gate circuits, programmable gate arrays (PGAs), field programmable gate arrays (FPGAs), etc.
[0166] Those of ordinary skill in the art of this technology can understand that all or part of the steps carried by the methods of the above embodiments can be completed by instructing relevant hardware through a program, and the program can be stored in a computer-readable storage medium. When the program is executed, it includes one or a combination of the steps of the method embodiments.
[0167] In addition, each functional unit in various embodiments of the present application may be integrated into a processing module, may exist physically alone for each unit, or two or more units may be integrated into one module. The above-mentioned integrated module may be implemented in the form of hardware or in the form of a software functional module. When the integrated module is implemented in the form of a software functional module and sold or used as an independent product, it may also be stored in a computer-readable storage medium.
[0168] The above-mentioned storage medium may be a read-only memory, a magnetic disk, an optical disc, etc. Although the embodiments of the present application have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limiting the present application. Those of ordinary skill in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present application.
Claims
1. A synthesis method for abnormal fact case data in the coal industry, characterized in that, The method includes: Obtain a dictionary set of the coal industry, and determine a first corpus set of the coal industry based on the operation constraint files in the dictionary set of the coal industry; Classify the first corpus set to obtain a second corpus set in different business areas of the coal industry; For each business area, extract key points from each corpus in the second corpus set of the business area to obtain a key point set of the corpus; Perform abnormal fact case simulation according to the second corpus set and the key point set of the corpus to obtain a data set of abnormal fact cases in the business area.
2. The method according to claim 1, characterized in that, The determining the first corpus set of the coal industry based on the operation constraint files in the dictionary set includes: Perform text mining on the operation constraint files to obtain an original structured corpus; For the corpora belonging to the job constraint file D i ; determine a first metric of the job constraint file according to the overall information density of the corpora and the job constraint file D i ; wherein the corpora D i are any raw corpora in the structured corpus ; Obtain the corpus of the term frequency-inverse document frequency, position weight value, and complexity evaluation value; Determine a second metric of the corpus according to the term frequency-inverse document frequency, the position weight value, and the complexity evaluation value ; According to the first index of the job constraint file D i and the second index of the corpus filter the structured corpus to obtain the first corpus set.
3. The method according to claim 2, wherein The first index according to the operation constraint file D i and the second index of the corpus are used to filter the structured corpus to obtain a first corpus set, including: Based on the second metric of the said corpus and the first metric of the said job constraint file D i determine the retention probability of the said corpus ; In response to the corpus whose retention probability is greater than or equal to a set retention threshold, retain the original corpus ; In response to the corpus whose retention probability is less than a set retention threshold, delete the corpus from the structured corpus ; Obtain the first corpus set based on the original corpus retained in the structured corpus.
4. The method according to claim 1, characterized in that, The extracting key points from the second corpus set to obtain a key point set corresponding to the business area includes: For any corpus in the second corpus set , extract the key points of the corpus through the first agent; Obtain the similarity between the key points and the corpus , the complexity evaluation value and compliance deviation degree of the key points; Perform weighted calculation on the similarity, the complexity evaluation value, and the compliance deviation degree to obtain a scoring value of the key point; Filter the key points of the corpus according to the scoring value to obtain the key point set of the corpus .
5. The method according to claim 1, characterized in that The performing abnormal fact case simulation according to the second corpus set and the key point set of the corpus to obtain a data set of abnormal fact cases in the business area includes: Obtain preset description information of the abnormal fact case to be simulated, where the description information includes the occurrence time, location, actor, and action content of the abnormal fact case; Invoke a second intelligent agent, and through the second intelligent agent, generate a simulated abnormal fact case of any key point according to the description information, the corpus, and any key point extracted from the corpus; Obtain a data set of abnormal fact cases in the business area based on the simulated abnormal fact case of any key point.
6. The method according to claim 5, characterized in that, The obtaining a data set of abnormal fact cases in the business area based on the simulated abnormal fact case of any key point includes: Determine the existing abnormal fact cases in the business area; Obtain the semantic similarity between the simulated abnormal fact case and the existing abnormal fact cases; In response to the semantic similarity being less than a set similarity threshold, update the database of the business area based on the simulated abnormal fact case to obtain a data set of abnormal fact cases.
7. The method according to claim 6, characterized in that, The updating the database of the business area based on the simulated abnormal fact case to obtain a data set of abnormal fact cases includes: Invoke a third intelligent agent, and through the third intelligent agent, perform content analysis on the simulated abnormal fact case, the corpus associated with the simulated abnormal fact case, and the key points to obtain an analysis result of the simulated abnormal fact case; Construct a triple according to the simulated abnormal fact case, the analysis result, and the corpus associated with the simulated abnormal fact case; Perform format conversion on the triple to obtain a standardized triple, and update the database based on the standardized triple to obtain a data set of abnormal fact cases.
8. The method according to claim 7, characterized in that, Before updating the database based on the standardized triples to obtain the data set of abnormal fact cases, the following steps are further included: Call a large model; Based on the large model, automatically verify the simulated abnormal fact case according to the parsing result of the simulated abnormal fact case, the corpus and key points associated with the simulated abnormal fact case.
9. The method according to claim 8, wherein Updating the database based on the standardized triples to obtain the data set of abnormal fact cases includes: After determining that the simulated abnormal fact case passes the automatic verification, send the simulated abnormal fact case to the review device; Receive the review data of the simulated abnormal fact case feedback by the review device; Determine the score value of the simulated abnormal fact case according to the review data; In response to the score value of the simulated abnormal fact case being greater than the set score threshold, obtain the classification label of the corpus associated with the simulated abnormal fact case; According to the classification label and the standardized triples of the simulated abnormal fact case, obtain the quadruple of the simulated abnormal fact case; Store the quadruple of the simulated abnormal fact case into the database corresponding to the business domain.
10. The method according to claim 7, characterized in that, The method further includes: Cache the simulated abnormal fact cases with semantic similarity less than the set similarity threshold in an array; Read and cache the simulated abnormal fact cases from the array, and input the simulated abnormal fact cases, the corpus and key points associated with the simulated abnormal fact cases into the third intelligent agent for content analysis.
Citation Information
Patent Citations
Case matching-based network security operation aid decision-making method, system and device
CN117350288A
Large model and knowledge graph fused coal industry knowledge base construction and query method
CN118861313A
Professional lexicon construction method for NLP word segmentation in coal industry
CN119443085A
Method and device for determining training data, training method, equipment, medium
CN119782809A
System and methods for generating treebanks for natural language processing by modifying parser operation through introduction of constraints on parse tree structure
US20160259851A1