A risk factor identification method, device and electronic equipment

By using the Bert causality extraction model to identify and assess risk factors in the construction process of digital twin watersheds, the technical problem of rapid risk assessment has been solved, and accurate risk identification and assessment of sensitive information of water conservancy projects has been achieved.

CN120013237BActive Publication Date: 2025-11-21POWERCHINA HUADONG ENG CORP LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510091941.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-21
Publication Date
2025-11-21
Estimated Expiration
2045-01-21

AI Technical Summary

Technical Problem

How to quickly assess the risk factors in the construction of digital twin watersheds, especially the risk assessment of sensitive information such as infrastructure data and cybersecurity involved in water conservancy projects.

Method used

The Bert causal relationship extraction model is used to identify risk factors and their causal relationships in text data. Target sentences are extracted through preprocessing and logical connectives, and the center value is calculated to determine key risk factors.

Benefits of technology

It enables the accurate identification and assessment of risk factors during the construction of digital twin watersheds, improving the efficiency and accuracy of risk assessment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120013237B_ABST
    Figure CN120013237B_ABST
Patent Text Reader

Abstract

The application provides a risk factor identification method and device and electronic equipment, and the method comprises the following steps: obtaining to-be-identified text data for constructing a virtual model of a water conservancy system; processing the to-be-identified text data into first words and first short sentences; inputting the first words and the first short sentences into a pre-trained BERT causal relationship extraction model, outputting the causal relationship between the first words in each first short sentence, and obtaining first sentences with causal relationships; the BERT causal relationship extraction model is pre-trained based on preset key words and sentences of a digital twin river basin construction problem and the causal relationship of the preset key words and sentences; according to a preset logical association word, target sentences related to risks in the first sentences are extracted; the central value of the target sentences is calculated; and according to the central value, the risk factors of the to-be-identified text data are determined. The method can realize accurate identification and evaluation of risks in the virtual model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data processing technology, and in particular to a method, apparatus and electronic device for identifying risk factors. Background Technology

[0002] The application of digital twin watershed technology in water conservancy projects enables project managers and decision-makers to accurately simulate various complex water conservancy processes. The construction of digital twin watersheds involves not only various engineering structural information but also a wealth of sensitive information, such as infrastructure data, cybersecurity, and algorithm accuracy.

[0003] Currently, the technical challenge of quickly assessing the risk factors present during the construction of digital twin watersheds urgently needs to be addressed. Summary of the Invention

[0004] The purpose of this invention is to provide a risk factor identification method, apparatus, and electronic device to alleviate the technical problem of rapidly assessing risk factors in the construction of digital twin watersheds, and to achieve rapid assessment of risk factors in the construction of digital twin watersheds.

[0005] In a first aspect, embodiments of the present invention provide a risk factor identification method applied to a server; the server is equipped with a virtual model of a water conservancy system based on digital twin watershed technology; the method includes: acquiring text data to be identified for constructing the virtual model; preprocessing the text data to be identified to obtain a first word and a first short sentence; inputting the first word and the first short sentence into a pre-trained Bert causality extraction model, outputting the causal relationship between the first word in each of the first short sentences, to obtain a first statement with a causal relationship; the Bert causality extraction model is pre-trained based on preset keywords and phrases of the digital twin watershed construction problem and the causal relationship of the preset keywords and phrases; extracting target statements related to risk in the first statement according to preset logical connectives; calculating the center value of the target statements; and determining the risk factors of the text data to be identified based on the center value.

[0006] In a preferred embodiment of the present invention, the text data to be identified includes: data text data, algorithm text data, and computing power text data for constructing the virtual model; the virtual model includes: the digital resource model, data engine model, data model, water conservancy perception model, water conservancy professional model, visualization model, knowledge platform model, simulation model, water conservancy cloud construction model, and water conservancy information network construction model of the water conservancy system.

[0007] In the preferred embodiment of the present application, before the step of inputting the words and the short sentences into the pre-trained BERT causal relationship extraction model and outputting the causal relationship between the words and the short sentences to obtain the sentences with causal relationship, the method comprises the following steps: S1: obtaining initial text data; the initial text data meets the preset judgment standard of the digital twin flow field construction problem; S2: screening the preset keyword sentences from the initial text data according to the preset keyword; S3: preprocessing the preset keyword sentences to obtain the second words and the second short sentences of the preset keyword sentences and the dependency relationship of the second words; S4: extracting the causal structure of the second words, the second short sentences and the dependency relationship based on the preset feature extraction parameters; S5: training the preset initial BERT causal relationship extraction model through the causal structure until the preset training standard is reached to obtain an intermediate BERT causal relationship extraction model; S6: obtaining the manually annotated causal structure of the second words, the second short sentences and the dependency relationship; S7: training the intermediate BERT causal relationship extraction model through the manually annotated causal structure until the training standard is reached to obtain the trained BERT causal relationship extraction model.

[0008] In the preferred embodiment of the present application, before the step of obtaining the manually annotated causal structure of the second words, the second short sentences and the dependency relationship, the method comprises the following step: inputting the second words, the second short sentences and the dependency relationship into the pre-trained GPT large language model to output the manually annotated causal structure.

[0009] In the preferred embodiment of the present application, after the step of training the intermediate BERT causal relationship extraction model through the manually annotated causal structure until the training standard is reached to obtain the trained BERT causal relationship extraction model, the method further comprises the following steps: B1: inputting the causal structure and the manually annotated causal structure into the trained BERT causal relationship extraction model to respectively output the second sentences and the third sentences with causal relationship; B2: analyzing the correlation of the second sentences and the third sentences by statistical analysis software; B3: when the correlation is greater than a preset threshold, saving the trained BERT causal relationship extraction model; when the correlation is less than or equal to the preset threshold, repeating the steps S1 to S7 and the steps B1 to B2 until the correlation is greater than the preset threshold, and then saving the trained BERT causal relationship extraction model.

[0010] In the preferred embodiment of the present application, the step of calculating the center value of the target sentence comprises the following step: calculating the center value of the target sentence based on the social network analysis method.

[0011] In the preferred embodiment of the present application, according to the central value, the step of determining the risk factors of the to-be-identified text data comprises: sorting the central value to obtain a sorted central value; selecting a target central value greater than a preset threshold from the sorted central value; and determining the risk factors of the to-be-identified text data according to a target sentence corresponding to the target central value.

[0012] In the preferred embodiment of the present application, the logical association words include: cause, lead to, result in, solve, how, can, lack, need, do, then, and, and therefore.

[0013] In a second aspect, an embodiment of the present application provides a risk factor identification device applied to a server; the server is provided with a water conservancy system virtual model based on digital twin river basin technology; the device comprises: a data acquisition module configured to acquire to-be-identified text data for constructing the virtual model; a preprocessing module configured to preprocess the to-be-identified text data to obtain first words and first short sentences of the to-be-identified text data; a cause-effect relationship determination module configured to input the first words and the first short sentences into a pre-trained BERT cause-effect relationship extraction model to output a cause-effect relationship between the first words in each first short sentence, thereby obtaining first sentences with a cause-effect relationship; the BERT cause-effect relationship extraction model is pre-trained based on preset key words and sentences of a digital twin river basin construction problem and a cause-effect relationship of the preset key words and sentences; a sentence extraction module configured to extract target sentences related to risks in the first sentences according to preset logical association words; a central value calculation module configured to calculate central values of the target sentences; and a risk factor determination module configured to determine risk factors of the to-be-identified text data according to the central values.

[0014] In a third aspect, an embodiment of the present application further provides an electronic device, which comprises a processor and a memory; the memory stores computer executable instructions capable of being executed by the processor; and the processor executes the computer executable instructions to implement the method.

[0015] The embodiments of the present application have the following beneficial technical effects:

[0016] The embodiment of the present application provides a risk factor identification method, device and electronic equipment, which are applied to a server; a water conservancy system virtual model based on digital twin river basin technology is arranged on the server; the method comprises the following steps: obtaining to-be-identified text data for constructing the virtual model; pre-processing the to-be-identified text data to obtain first words and first short sentences of the to-be-identified text data; inputting the first words and the first short sentences into a pre-trained BERT causal relationship extraction model to output a causal relationship between the first words in each first short sentence, and obtaining a first sentence with the causal relationship; the BERT causal relationship extraction model is pre-trained based on preset key words and sentences of a digital twin river basin construction problem and a causal relationship of the preset key words and sentences; according to a preset logical association word, a target sentence related to a risk in the first sentence is extracted; a central value of the target sentence is calculated; and a risk factor of the to-be-identified text data is determined according to the central value. The method uses the BERT causal relationship extraction model to identify a risk factor and a causal relationship in the text for constructing the virtual model, finally calculates a central value to determine a key risk factor, so that accurate identification and evaluation of the risk existing in the virtual model are realized. BRIEF DESCRIPTION OF DRAWINGS

[0017] In order to more clearly illustrate the technical solutions in the specific embodiments of the present application or the prior art, the drawings needed in the specific embodiments or prior art description will be briefly introduced. Obviously, the drawings in the following description are some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.

[0018] Figure 1 A flowchart of a risk factor identification method provided by the embodiment of the present application is shown in the figure.

[0019] Figure 2 An analysis result diagram of a logical association word provided by the embodiment of the present application is shown in the figure.

[0020] Figure 3 A flowchart of a construction method of a BERT causal relationship extraction model provided by the embodiment of the present application is shown in the figure.

[0021] Figure 4 A structural diagram of a risk factor identification device provided by the embodiment of the present application is shown in the figure.

[0022] Figure 5 A structural diagram of an electronic device provided by the embodiment of the present application is shown in the figure.

[0023] Icon: 31 - data acquisition module; 32 - preprocessing module; 33 - causal relationship determination module; 34 - sentence extraction module; 35 - central value calculation module; 36 - risk factor determination module; 41 - memory; 42 - processor; 43 - bus; 44 - communication interface. DETAILED DESCRIPTION

[0024] To make the purposes, technical solutions and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, rather than all the embodiments. The components of the embodiments of the present application described and shown in the drawings herein can be arranged and designed in various different configurations.

[0025] The application of digital twin basin technology in water conservancy projects enables engineering managers and decision-makers to accurately simulate various complex water conservancy processes. The construction of a digital twin basin not only involves various engineering structure information, but also includes a large amount of sensitive information, such as infrastructure data, network security, and algorithm accuracy. Therefore, the technical problem of how to quickly evaluate the risk factors existing in the construction process of a digital twin basin needs to be solved.

[0026] Based on this, the embodiments of the present application provide a risk factor identification method, device and electronic equipment. The method uses a BERT causal relationship extraction model to identify risk factors and their causal relationships in the text for constructing a virtual model, and finally calculates the central value to determine the key risk factors, thereby realizing accurate identification and evaluation of risks existing in the virtual model. In order to facilitate understanding, first introduce a risk factor identification method.

[0027] Embodiment 1

[0028] In this embodiment, Figure 1 A flowchart of a risk factor identification method provided by the embodiments of the present application is shown.

[0029] As seen, Figure 1 The method comprises:

[0030] Step S101: acquiring to-be-identified text data for constructing the virtual model.

[0031] In this embodiment, the to-be-identified text data comprises: algorithm text data, algorithm text data and computing power text data for constructing the virtual model; and the virtual model comprises: a digital resource model, a data engine model, a data model, a water conservancy perception model, a water conservancy professional model, a visualization model, a knowledge platform model, a simulation and emulation model, a water conservancy cloud construction model and a water conservancy information network construction model of the water conservancy system.

[0032] Step S102: Preprocessing the above-mentioned to-be-recognized text data to obtain first words and first short sentences of the above-mentioned to-be-recognized text data.

[0033] In actual operation, the above-mentioned S102 includes: segmenting the to-be-recognized text data into words or phrases, and deleting common and meaningless words such as "of", "is", "in", etc.

[0034] Step S103: inputting the above-mentioned first words and the above-mentioned first short sentences into a pre-trained BERT causal relationship extraction model, outputting the causal relationship between the above-mentioned first words in each of the above-mentioned first short sentences, and obtaining first sentences with causal relationship; the above-mentioned BERT causal relationship extraction model is pre-trained based on preset key words and sentences of digital twin river basin construction problems and the causal relationship of the above-mentioned preset key words and sentences.

[0035] Step S104: extracting target sentences related to risks in the above-mentioned first sentences according to preset logical association words.

[0036] Here, the above-mentioned logical association words include: cause, cause, cause, solve, how, can, lack, need, do, then and therefore.

[0037] Further, the present application performs logical analysis through the logical association words "cause, cause, cause, solve, how, can, lack, need, do, then, therefore" to extract target sentences related to risks in the above-mentioned first sentences.

[0038] In order to facilitate understanding, Figure 2 A schematic diagram of an analysis result of a logical association word provided by an embodiment of the present application.

[0039] Step S105: calculating the center value of the above-mentioned target sentences.

[0040] In this embodiment, the step of calculating the center value of the above-mentioned target sentences includes: calculating the center value of the above-mentioned target sentences based on a social network analysis method.

[0041] Specifically, the center value reflects the core of the target sentence. The stronger the center position is, the more attention the concept will receive. The center value of each risk factor is calculated by using the SNA method, and the calculation formula is as follows:

[0042]

[0043] Wherein, CD(P k ) is the number of concepts connected to the concept k in the causal graph composed of the above-mentioned target sentences, d(p i , p k ) is the distance between the concepts P iThe shortest path to each concept in the causal graph, n is the number of concepts in the causal graph.

[0044] Step S106: According to the above-mentioned central value, the risk factors of the above-mentioned to be identified text data are determined.

[0045] In specific implementation, the step of determining the risk factors of the above-mentioned to be identified text data according to the above-mentioned central value comprises: sorting the above-mentioned central value to obtain sorted central values; selecting target central values greater than a preset threshold from the sorted central values; and determining the risk factors of the above-mentioned to be identified text data according to target sentences corresponding to the above-mentioned target central values.

[0046] The embodiment of the present application provides a risk factor identification method applied to a server; a water conservancy system virtual model based on digital twin river basin technology is arranged on the server; the method comprises the following steps: obtaining to-be-identified text data for constructing the virtual model; pre-processing the to-be-identified text data to obtain first words and first short sentences of the to-be-identified text data; inputting the first words and the first short sentences into a pre-trained BERT causal relationship extraction model to output causal relationships between the first words in each first short sentence, thereby obtaining first sentences with causal relationships; the BERT causal relationship extraction model is pre-trained based on preset key words and sentences of a digital twin river basin construction problem and causal relationships of the preset key words and sentences; target sentences related to risks in the first sentences are extracted according to preset logical association words; the central value of the target sentences is calculated; and the risk factors of the to-be-identified text data are determined according to the central value. The method uses the BERT causal relationship extraction model to identify risk factors and their causal relationships in the text for constructing the virtual model, finally calculates the central value to determine the key risk factors, so as to realize accurate identification and evaluation of risks existing in the virtual model.

[0047] Embodiment 2

[0048] On the basis of the above-mentioned embodiments, Figure 3 A flowchart of a BERT causal relationship extraction model construction method provided by the embodiment of the present application is shown.

[0049] The BERT causal relationship extraction model construction method is implemented before the step of inputting the words and the short sentences into the pre-trained BERT causal relationship extraction model to output the causal relationships between the words and the short sentences and obtain the sentences with causal relationships.

[0050] Specifically, the method comprises:

[0051] Step S1: obtaining initial text data; the initial text data meets preset judgment criteria of a digital twin river basin construction problem.

[0052] Here, the above step S1 collects digital twin river basin construction related information through three aspects of academic library, browsing academic platform and expert interview or lecture, and finally completes the collection of the above initial text data in text form.

[0053] Step S2: According to the preset keyword, the above initial text data is screened according to the above preset keyword sentence.

[0054] Here, the keyword is generally a word indicating a cause-effect relationship, such as: because, therefore, lead and other words. Further, the above initial text data is screened according to the above preset keyword sentence, that is, the sentence combining cause and result.

[0055] Step S3: Preprocessing the above preset keyword sentence to obtain the second word and the second short sentence of the above preset keyword sentence and the dependency relationship of the above second word.

[0056] Step S4: Based on the preset feature extraction parameters, extract the cause-effect structure of the above second word, the above second short sentence and the above dependency relationship.

[0057] Step S5: Train the preset initial BERT cause-effect relationship extraction model through the above cause-effect structure until the preset training standard is reached to obtain an intermediate state BERT cause-effect relationship extraction model.

[0058] Step S6: Obtain the artificial annotated cause-effect structure of the above second word, the above second short sentence and the above dependency relationship.

[0059] Step S7: Train the above intermediate state BERT cause-effect relationship extraction model through the above artificial annotated cause-effect structure until the above training standard is reached to obtain the above trained BERT cause-effect relationship extraction model.

[0060] Among them, before the step of obtaining the artificial annotated cause-effect structure of the above second word, the above second short sentence and the above dependency relationship, it includes: inputting the above second word, the above second short sentence and the above dependency relationship into the pre-trained GPT large language model, and outputting the above artificial annotated cause-effect structure.

[0061] Further, after the step of training the above intermediate state BERT cause-effect relationship extraction model through the above artificial annotated cause-effect structure until the above training standard is reached to obtain the above trained BERT cause-effect relationship extraction model, the above method further comprises:

[0062] Step B1: Input the above cause-effect structure and the above artificial annotated cause-effect structure into the above trained BERT cause-effect relationship extraction model to respectively output the second sentence and the third sentence existing the cause-effect relationship.

[0063] Step B2: analyze the correlation of the second sentence and the third sentence by statistical analysis software.

[0064] Step B3: when the correlation is greater than a preset threshold, save the trained BERT causal relation extraction model; when the correlation is less than or equal to the preset threshold, repeat the steps S1 to S7 and the steps B1 to B2 until the correlation is greater than the preset threshold, and then save the trained BERT causal relation extraction model.

[0065] Further, the preset threshold is 0.8.

[0066] The embodiment of the application provides a method for constructing a BERT causal relation extraction model, which comprises: obtaining initial text data; the initial text data meets a preset judgment standard of digital twin flow field construction problems; filtering preset keyword sentences from the initial text data according to a preset keyword; preprocessing the preset keyword sentences to obtain second words and second short sentences of the preset keyword sentences and dependency relationships of the second words; extracting a causal structure of the second words, the second short sentences and the dependency relationships based on a preset feature extraction parameter; training a preset initial BERT causal relation extraction model through the causal structure until a preset training standard is reached to obtain an intermediate state BERT causal relation extraction model; obtaining an artificial labeled causal structure of the second words, the second short sentences and the dependency relationships; training the intermediate state BERT causal relation extraction model through the artificial labeled causal structure until the training standard is reached to obtain the trained BERT causal relation extraction model.

[0067] Embodiment 3

[0068] On the basis of the above-mentioned embodiments, Figure 4 A structural schematic diagram of a risk factor identification device provided by the embodiment of the application is shown.

[0069] The device is applied to a server, and a water conservancy system virtual model based on digital twin flow field technology is arranged on the server.

[0070] As shown in the figure, Figure 4 The device comprises:

[0071] A data acquisition module 31 is configured to acquire to-be-identified text data for constructing the virtual model.

[0072] A preprocessing module 32 is configured to preprocess the to-be-identified text data to obtain first words and first short sentences of the to-be-identified text data.

[0073] The cause-effect relationship determination module 33 is configured to input the first words and the first short sentences into a pre-trained BERT cause-effect relationship extraction model, output a cause-effect relationship between the first words in each of the first short sentences, and obtain first sentences with a cause-effect relationship.

[0074] The sentence extraction module 34 is configured to extract target sentences related to risks in the first sentences according to preset logical association words.

[0075] The central value calculation module 35 is configured to calculate a central value of the target sentences.

[0076] The risk factor determination module 36 is configured to determine a risk factor of the to-be-recognized text data according to the central value.

[0077] The data acquisition module 31, the preprocessing module 32, the cause-effect relationship determination module 33, the sentence extraction module 34, the central value calculation module 35, and the risk factor determination module 36 are sequentially connected.

[0078] The risk factor recognition device provided in the embodiment has the same technical features as the risk factor recognition method provided in the embodiment, and can solve the same technical problems and achieve the same technical effects. For the convenience and brevity of description, the specific working process of the device described above can refer to the corresponding process in the foregoing method embodiment, which will not be described herein.

[0079] Embodiment 4

[0080] The embodiment provides an electronic device, including a processor and a memory, the memory stores computer executable instructions capable of being executed by the processor, and the processor executes the computer executable instructions to implement the steps of the risk factor recognition method.

[0081] The embodiment provides a computer readable storage medium, wherein the computer readable storage medium stores a computer program, and the computer program is executed by a processor to implement the steps of the risk factor recognition method.

[0082] Referring to Figure 5 The electronic device includes a memory 41 and a processor 42, the memory 41 stores a computer program capable of running on the processor 42, and the processor executes the computer program to implement the steps provided by the risk factor recognition method.

[0083] As Figure 5As shown, the device further includes a bus 43 and a communication interface 44, the processor 42, the communication interface 44 and the memory 41 are connected through the bus 43; the processor 42 is configured to execute the executable modules stored in the memory 41, such as computer programs.

[0084] The memory 41 can include a high-speed random access memory (RAM, Random Access Memory), and can also include a non-volatile memory, such as at least one disk memory. The communication connection between the system network element and at least one other network element is realized through at least one communication interface 44 (which can be wired or wireless), and the Internet, a wide area network, a local area network, a metropolitan area network, etc. can be used.

[0085] The bus 43 can be an ISA bus, a PCI bus, an EISA bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 5 Only one bidirectional arrow is used in the figure, but it does not mean that there is only one bus or only one type of bus.

[0086] The memory 41 is configured to store a program, and the processor 42 is configured to execute the program after receiving an execution instruction. The method performed by the risk factor identification device according to any one of the embodiments of the present application can be applied to the processor 42 or implemented by the processor 42. The processor 42 can be an integrated circuit chip with a processing capability of signals. In the implementation process, each step of the above method can be completed by the integrated logic circuit of hardware in the processor 42 or the instruction in the form of software. The processor 42 can be a general processor, including a central processing unit (CPU), a network processor (NP), etc.; or a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. Each method, step and logic block disclosed in the embodiments of the present application can be implemented or executed. The general processor can be a microprocessor or any conventional processor. The steps of the method disclosed in combination with the embodiments of the present application can be directly embodied as a hardware decoding processor for execution, or a combination of hardware and software modules in the decoding processor for execution. The software module can be located in a random access memory, a flash memory, a read-only memory, a programmable read-only memory or an electrically erasable programmable memory, a register or other mature storage medium in the art. The storage medium is located in the memory 41, and the processor 42 reads the information in the memory 41 and combines the hardware to complete the steps of the above method.

[0087] Further, the embodiment of the present application further provides a machine readable storage medium, the machine readable storage medium stores machine executable instructions, when the machine executable instructions are called and executed by the processor 42, the machine executable instructions cause the processor 42 to implement the above risk factor identification method.

[0088] The electronic device and the computer readable storage medium provided by the embodiments of the present application have the same technical features, so they can also solve the same technical problems and achieve the same technical effects.

[0089] In addition, in the description of the embodiments of the present application, unless specifically defined and limited, the terms "mounting", "connection", "connecting" should be understood in a broad sense, for example, can be fixedly connected, can also be detachably connected, or integrally connected; can be mechanically connected, can also be electrically connected; can be directly connected, can also be indirectly connected through an intermediate medium, can be internal communication of two elements. For those skilled in the art, the specific meaning of the above terms in the present application can be understood according to the specific circumstances.

[0090] In the description of the present application, it should be noted that the terms "center", "upper", "lower", "left", "right", "vertical", "horizontal", "inner", "outer" and the like indicate the orientation or positional relationship shown in the drawings, and are only for the convenience of describing the present application and simplifying the description, and do not indicate or imply that the devices or elements referred to must have a particular orientation, be constructed and operated in a particular orientation, and therefore cannot be understood as a limitation on the present application. In addition, the terms "first", "second", "third" are only for the purpose of description, and cannot be understood as indicating or implying relative importance.

Claims

1. A risk factor identification method characterized by, Be applied to a server; The server is provided with a water conservancy system virtual model based on digital twin river basin technology; The method comprises: Obtain the to-be-identified text data for constructing the virtual model; The to-be-identified text data comprises: algorithm text data, algorithm text data and algorithm text data for constructing the virtual model; Pretreat the to-be-identified text data to obtain the first word and the first short sentence of the to-be-identified text data; The first word and the first short sentence are input into the pre-trained BERT causal relationship extraction model, and the causal relationship between the first word in each first short sentence is output, and the first sentence with causal relationship is obtained; The BERT causal relationship extraction model is pre-trained based on the preset key word of the digital twin river basin construction problem and the causal relationship of the preset key word; According to the preset logical association word, the target sentence related to risk in the first sentence is extracted; Based on social network analysis method, the center value of the target sentence is calculated; According to the center value, the risk factors of the to-be-identified text data are determined.

2. The risk factor identification method of claim 1, wherein, The virtual model comprises: digital resource model, data engine model, data model, water conservancy perception model, water conservancy professional model, visualization model, knowledge platform model, simulation model, water conservancy cloud construction model and water conservancy information network construction model.

3. The risk factor identification method of claim 2, wherein, Before the step of inputting the word and the short sentence into the pre-trained BERT causal relationship extraction model, outputting the causal relationship between the word and the short sentence, and obtaining the sentence with causal relationship, the method comprises: Step S1: obtaining initial text data; The initial text data meets the preset judgment standard of digital twin river basin construction problem; Step S2: according to the preset key word, the preset key word is screened from the initial text data; Step S3: pretreat the preset key word to obtain the second word and the second short sentence of the preset key word and the dependency relationship of the second word; Step S4: based on the preset feature extraction parameter, the causal structure of the second word, the second short sentence and the dependency relationship is extracted; Step S5: train the preset initial BERT causal relationship extraction model through the causal structure until the preset training standard is reached, and obtain the intermediate state BERT causal relationship extraction model; Step S6: obtain the artificial annotation causal structure of the second word, the second short sentence and the dependency relationship; Step S7: train the intermediate state BERT causal relationship extraction model through the artificial annotation causal structure until the training standard is reached, and obtain the trained BERT causal relationship extraction model.

4. The risk factor identification method of claim 3, wherein, Before the step of obtaining the artificial annotation causal structure of the second word, the second short sentence and the dependency relationship, comprising: Input the second word, the second short sentence and the dependency relationship into the pre-trained GPT large language model, and output the artificial annotation causal structure.

5. The risk factor identification method of claim 4, wherein, After the step of training the intermediate BERT causal relation extraction model by the artificial annotated causal structure until the training standard is reached to obtain the trained BERT causal relation extraction model, the method further comprises: Step B1: inputting the causal structure and the artificial annotated causal structure into the trained BERT causal relation extraction model to respectively output a second sentence and a third sentence in which a causal relationship exists; Step B2: analyzing the relevance of the second sentence and the third sentence by statistical analysis software; Step B3: when the relevance is greater than a preset threshold, saving the trained BERT causal relation extraction model; when the relevance is less than or equal to the preset threshold, repeating the steps S1 to S7 and the steps B1 to B2 until the relevance is greater than the preset threshold, and saving the trained BERT causal relation extraction model.

6. The risk factor identification method of claim 1, wherein, According to the center value, the step of determining the risk factor of the to-be-identified text data comprises: sorting the center values to obtain sorted center values; selecting a target center value greater than a preset threshold from the sorted center values; determining the risk factor of the to-be-identified text data according to a target sentence corresponding to the target center value.

7. The risk factor identification method of claim 1, wherein, The logical association words include: cause, cause, cause, solve, how, can, lack, need, do, then, and therefore.

8. A risk factor identification apparatus characterized by comprising: The application is applied to a server; the server is provided with a water conservancy system virtual model based on digital twin flow field technology; the device comprises: a data acquisition module configured to acquire to-be-identified text data for constructing the virtual model; the to-be-identified text data comprises: algorithm text data, algorithm text data, and computing power text data for constructing the virtual model; a preprocessing module configured to preprocess the to-be-identified text data to obtain first words and first short sentences of the to-be-identified text data; a causal relation determination module configured to input the first words and the first short sentences into a pre-trained BERT causal relation extraction model to output a causal relationship between the first words in each first short sentence to obtain a first sentence in which a causal relationship exists; the BERT causal relation extraction model is pre-trained based on preset key words and sentences of a digital twin flow field construction problem and a causal relationship of the preset key words and sentences; a sentence extraction module configured to extract a target sentence related to a risk in the first sentence according to a preset logical association word; a center value calculation module configured to calculate a center value of the target sentence based on a social network analysis method; a risk factor determination module configured to determine a risk factor of the to-be-identified text data according to the center value.

9. An electronic device, comprising: The electronic device comprises a processor and a memory, the memory stores computer executable instructions executable by the processor, and the processor executes the computer executable instructions to implement the method of any one of claims 1 to 7.

Citation Information

Patent Citations

  • Method, device and equipment for constructing causal relationship determination model

    CN112329478A

  • Text recognition method, device and equipment

    CN115510871A