Risk factor identification method and device and electronic equipment
By using the Burt causal extraction model to analyze the text data of the virtual model constructed during the construction of the digital twin river basin, identify and evaluate risk factors, the problem of rapid evaluation of risk factors is solved, and the accurate identification and evaluation of risk factors is achieved.
Patent Information
- Application Number
- CN202510091941.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-21
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2045-01-21
AI Technical Summary
The existing technology is difficult to effectively solve this problem by how to quickly evaluate the risk factors existing in the construction of digital twin river basins.
The Burt causality extraction model is used to obtain the text data to be identified for preprocessing by constructing the virtual model, and the target statements related to risks are extracted, and the center value is calculated to determine the risk factors.
It has achieved accurate identification and evaluation of risk factors during the construction of digital twin river basin, and improved the efficiency and accuracy of risk assessment.
Smart Images

Figure CN120013237A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data processing technology, and in particular to a risk factor identification method, device and electronic equipment. Background Art
[0002] The application of digital twin basin technology in water conservancy projects enables project managers and decision makers to accurately simulate various complex water conservancy processes. The construction of digital twin basins not only involves various engineering structure information, but also includes a lot of sensitive information, such as infrastructure data, network security, and algorithm accuracy.
[0003] At present, the technical problem of how to quickly evaluate the risk factors in the construction process of digital twin river basins needs to be solved urgently. Summary of the invention
[0004] The purpose of the present invention is to provide a risk factor identification method, device and electronic equipment to alleviate the technical problems of rapid assessment of risk factors in the process of digital twin river basin construction, and to achieve rapid assessment of risk factors in the process of digital twin river basin construction.
[0005] In the first aspect, an embodiment of the present invention provides a risk factor identification method, which is applied to a server; a virtual model of a water conservancy system based on digital twin watershed technology is set on the server; the method includes: obtaining text data to be identified for constructing the virtual model; preprocessing the text data to be identified to obtain the first word and the first short sentence of the text data to be identified; inputting the first word and the first short sentence into a pre-trained Burt causal relationship extraction model, outputting the causal relationship between the first words in each of the first short sentences, and obtaining the first sentence with a causal relationship; the Burt causal relationship extraction model is pre-trained based on preset key words and sentences of digital twin watershed construction problems and the causal relationship of the preset key words and sentences; according to preset logical conjunctions, extracting the target sentence related to the risk in the first sentence; calculating the central value of the target sentence; and determining the risk factor of the text data to be identified based on the central value.
[0006] In a preferred embodiment of the present invention, the above-mentioned text data to be identified includes: calculation text data, algorithm text data and computing power text data for constructing the above-mentioned virtual model; the above-mentioned virtual model includes: the digital resource model of the above-mentioned water conservancy system, data engine model, data model, water conservancy perception model, water conservancy professional model, visualization model, knowledge platform model, simulation model, water conservancy cloud construction model and water conservancy information network construction model.
[0007] In a preferred embodiment of the present invention, before the step of inputting the above-mentioned words and the above-mentioned short sentences into a pre-trained Burt causal relationship extraction model, outputting the causal relationship between the above-mentioned words and the above-mentioned short sentences, and obtaining sentences with causal relationships, the above-mentioned method includes: step S1: obtaining initial text data; the above-mentioned initial text data meets the preset judgment criteria for the digital twin watershed construction problem; step S2: according to preset keywords, screening the above-mentioned preset key words and sentences from the above-mentioned initial text data; step S3: preprocessing the above-mentioned preset key words and sentences to obtain the second word and the second short sentence of the above-mentioned preset key words and sentences and the dependency relationship of the above-mentioned second word ; Step S4: Based on the preset feature extraction parameters, extract the causal structure of the second word, the second sentence and the dependency relationship; Step S5: Train the preset initial Burt causal relationship extraction model through the causal structure until the preset training standard is reached, and obtain the intermediate Burt causal relationship extraction model; Step S6: Obtain the manually annotated causal structure of the second word, the second sentence and the dependency relationship; Step S7: Train the intermediate Burt causal relationship extraction model through the manually annotated causal structure until the training standard is reached, and obtain the trained Burt causal relationship extraction model.
[0008] In a preferred embodiment of the present invention, before the step of obtaining the manually annotated causal structure of the second word, the second sentence and the dependency relationship, it includes: inputting the second word, the second sentence and the dependency relationship into a pre-trained GPT large language model, and outputting the manually annotated causal structure.
[0009] In a preferred embodiment of the present invention, after the step of training the intermediate state Burt causal relationship extraction model by the manually annotated causal structure until the training standard is reached and the trained Burt causal relationship extraction model is obtained, the method further includes: step B1: inputting the causal structure and the manually annotated causal structure into the trained Burt causal relationship extraction model, and outputting the second sentence and the third sentence with causal relationship respectively; step B2: analyzing the correlation between the second sentence and the third sentence by statistical analysis software; step B3: when the correlation is greater than a preset threshold, saving the trained Burt causal relationship extraction model; when the correlation is less than or equal to the preset threshold, repeating steps S1 to S7 and steps B1 to B2 until the correlation is greater than the preset threshold, saving the trained Burt causal relationship extraction model.
[0010] In a preferred embodiment of the present invention, the step of calculating the central value of the target sentence includes: calculating the central value of the target sentence based on a social network analysis method.
[0011] In a preferred embodiment of the present invention, the step of determining the risk factors of the above-mentioned text data to be identified based on the above-mentioned central values includes: sorting the above-mentioned central values to obtain sorted central values; selecting target central values greater than a preset threshold from the sorted central values; and determining the risk factors of the above-mentioned text data to be identified based on the target sentence corresponding to the above-mentioned target central value.
[0012] In a preferred embodiment of the present invention, the above-mentioned logical conjunctions include: lead, cause, cause, solve, how, can, lack, need, conduct, then, and, therefore.
[0013] In the second aspect, an embodiment of the present invention provides a risk factor identification device, which is applied to a server; a virtual model of a water conservancy system based on digital twin watershed technology is set on the server; the device includes: a data acquisition module, used to obtain text data to be identified for building the virtual model; a preprocessing module, used to preprocess the text data to be identified, and obtain the first word and the first short sentence of the text data to be identified; a causal relationship determination module, used to input the first word and the first short sentence into a pre-trained Burt causal relationship extraction model, output the causal relationship between the first words in each of the first short sentences, and obtain the first sentence with a causal relationship; the Burt causal relationship extraction model is pre-trained based on preset key words and sentences of digital twin watershed construction problems and the causal relationship of the preset key words and sentences; a sentence extraction module, used to extract the target sentence related to the risk in the first sentence according to preset logical conjunctions; a central value calculation module, used to calculate the central value of the target sentence; a risk factor determination module, used to determine the risk factor of the text data to be identified according to the central value.
[0014] In a third aspect, an embodiment of the present invention further provides an electronic device, comprising a processor and a memory, wherein the memory stores computer-executable instructions that can be executed by the processor, and the processor executes the computer-executable instructions to implement the method.
[0015] The embodiments of the present invention have the following beneficial technical effects:
[0016] The embodiment of the present invention provides a risk factor identification method, device and electronic device, which are applied to a server; the server is provided with a virtual model of a water conservancy system based on digital twin watershed technology; the method includes: obtaining text data to be identified for constructing the virtual model; preprocessing the text data to be identified to obtain the first word and the first short sentence of the text data to be identified; inputting the first word and the first short sentence into a pre-trained Burt causal relationship extraction model, outputting the causal relationship between the first words in each of the first short sentences, and obtaining the first sentence with causal relationship; the Burt causal relationship extraction model is pre-trained based on preset key words and sentences of digital twin watershed construction problems and the causal relationship of the preset key words and sentences; extracting the target sentence related to the risk in the first sentence according to the preset logical association words; calculating the central value of the target sentence; and determining the risk factors of the text data to be identified according to the central value. The method uses the Burt causal relationship extraction model to identify the risk factors and their causal relationships in the text for constructing the virtual model, and finally calculates the central value to determine the key risk factors, thereby realizing the accurate identification and evaluation of the risks existing in the virtual model. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] In order to more clearly illustrate the specific implementation methods of the present invention or the technical solutions in the prior art, the drawings required for use in the specific implementation methods or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are some implementation methods of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0018] Figure 1 A schematic diagram of a process for identifying risk factors provided by an embodiment of the present invention;
[0019] Figure 2 A schematic diagram of analysis results of a logical conjunction provided by an embodiment of the present invention;
[0020] Figure 3 A flowchart of a method for constructing a Burt causal relationship extraction model provided by an embodiment of the present invention;
[0021] Figure 4 A schematic diagram of the structure of a risk factor identification device provided by an embodiment of the present invention;
[0022] Figure 5 A schematic diagram of the structure of an electronic device provided by an embodiment of the present invention.
[0023] Icons: 31-data acquisition module; 32-preprocessing module; 33-causal relationship determination module; 34-statement extraction module; 35-central value calculation module; 36-risk factor determination module; 41-memory; 42-processor; 43-bus; 44-communication interface. DETAILED DESCRIPTION
[0024] In order to make the purpose, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Generally, the components of the embodiments of the present invention described and shown in the drawings here can be arranged and designed in various different configurations.
[0025] The application of digital twin basin technology in water conservancy projects enables project managers and decision makers to accurately simulate various complex water conservancy processes. The construction of digital twin basins not only involves various engineering structure information, but also includes a lot of sensitive information, such as infrastructure data, network security, and algorithm accuracy. Therefore, the technical problem of how to quickly evaluate the risk factors in the construction of digital twin basins needs to be solved urgently.
[0026] Based on this, an embodiment of the present invention provides a risk factor identification method, device and electronic device, which uses the Bert causal relationship extraction model to identify risk factors and their causal relationships in the text of the virtual model, and finally calculates the central value to determine the key risk factors, thereby realizing accurate identification and evaluation of the risks existing in the virtual model. For ease of understanding, a risk factor identification method is first introduced.
[0027] Example 1
[0028] In this embodiment, Figure 1 A schematic diagram of a flow chart of a risk factor identification method provided by an embodiment of the present invention.
[0029] Depend on Figure 1 As can be seen, the method includes:
[0030] Step S101: obtaining text data to be recognized for constructing the above virtual model.
[0031] In this embodiment, the above-mentioned text data to be identified includes: calculation text data, algorithm text data and computing power text data for constructing the above-mentioned virtual model; the above-mentioned virtual model includes: the digital resource model of the above-mentioned water conservancy system, data engine model, data model, water conservancy perception model, water conservancy professional model, visualization model, knowledge platform model, simulation model, water conservancy cloud construction model and water conservancy information network construction model.
[0032] Step S102: Preprocess the above-mentioned text data to be recognized to obtain the first words and the first short sentences of the above-mentioned text data to be recognized.
[0033] In actual operation, the above S102 includes: splitting the text data to be recognized into words or phrases, and deleting relatively common and meaningless words, such as "of", "is", "in", etc.
[0034] Step S103: Input the above first words and the above first short sentences into a pre-trained BERT causal relationship extraction model, output the causal relationship between the above first words in each of the above first short sentences, and obtain the first statements with causal relationships; the above BERT causal relationship extraction model is pre-trained based on the preset keyword sentences of the digital twin watershed construction problem and the causal relationships of the above preset keyword sentences.
[0035] Step S104: Extract the target statements related to risks in the above first statements according to the preset logical conjunctions.
[0036] Here, the above logical conjunctions include: lead to, cause, result in, solve, how, whether, lack, need, conduct, then, and therefore.
[0037] Furthermore, this application conducts logical analysis through the logical conjunctions "lead to, cause, result in, solve, how, whether, lack, need, conduct, then, therefore", and extracts the target statements related to risks in the above first statements.
[0038] For ease of understanding, Figure 2 FIG. is a schematic diagram of the analysis result of a logical conjunction provided for an embodiment of the present invention.
[0039] Step S105: Calculate the central value of the above target statements.
[0040] In this embodiment, the step of calculating the central value of the above target statements includes: calculating the central value of the above target statements based on the social network analysis method.
[0041] Specifically, the central value reflects the core of the target statement. The stronger the central position, the more attention the concept will receive. The centrality of each risk factor is calculated using the SNA method, and the calculation formula is as follows:
[0042]
[0043] Among them, CD(P k ) is the number of concepts connected to concept k in the causal graph, and the above causal graph is composed of the above target statements. d(p i , p k ) is the distance from concept P iThe shortest path to each concept in the causal graph, where n is the number of concepts in the causal graph.
[0044] Step S106: Determine the risk factor of the text data to be identified based on the central value.
[0045] In a specific implementation, the steps of determining the risk factors of the above-mentioned text data to be identified based on the above-mentioned central values include: sorting the above-mentioned central values to obtain sorted central values; selecting target central values greater than a preset threshold from the sorted central values; and determining the risk factors of the above-mentioned text data to be identified based on the target sentence corresponding to the above-mentioned target central value.
[0046] The embodiment of the present invention provides a risk factor identification method, which is applied to a server; a virtual model of a water conservancy system based on digital twin watershed technology is set on the server; the method includes: obtaining text data to be identified for constructing the virtual model; preprocessing the text data to be identified to obtain the first word and the first short sentence of the text data to be identified; inputting the first word and the first short sentence into a pre-trained Bert causal relationship extraction model, outputting the causal relationship between the first words in each of the first short sentences, and obtaining the first sentence with a causal relationship; the Bert causal relationship extraction model is pre-trained based on preset key words and sentences of digital twin watershed construction problems and the causal relationship of the preset key words and sentences; extracting the target sentence related to the risk in the first sentence according to the preset logical association words; calculating the central value of the target sentence; and determining the risk factors of the text data to be identified according to the central value. The method uses the Bert causal relationship extraction model to identify the risk factors and their causal relationships in the text for constructing the virtual model, and finally calculates the central value to determine the key risk factors, thereby realizing the accurate identification and evaluation of the risks existing in the virtual model.
[0047] Example 2
[0048] Based on the above embodiments, Figure 3 A flowchart of a method for constructing a Burt causal relationship extraction model provided in an embodiment of the present invention.
[0049] Among them, the method for constructing the Burt causal relationship extraction model is implemented before the step of inputting the above-mentioned words and the above-mentioned short sentences into a pre-trained Burt causal relationship extraction model, outputting the causal relationship between the above-mentioned words and the above-mentioned short sentences, and obtaining sentences with causal relationships.
[0050] Specifically, the method includes:
[0051] Step S1: Obtain initial text data; the above initial text data meets the preset judgment criteria for the digital twin watershed construction problem.
[0052] Here, the above step S1 collects relevant information on digital twin watershed construction through academic libraries, browsing academic platforms, and expert interviews or lectures, and finally completes the collection of the above initial text data in text form.
[0053] Step S2: According to the preset keywords, the preset keyword phrases are screened from the initial text data.
[0054] Here, the keyword is generally a word indicating a cause-effect relationship, such as because, so, and leads to, etc. Further, the preset keyword sentence is screened from the initial text data, that is, a sentence combining cause and effect.
[0055] Step S3: pre-processing the preset key words and sentences to obtain the second word and the second short sentence of the preset key words and sentences and the dependency relationship between the second word and the second word.
[0056] Step S4: extracting the second word, the second sentence and the causal structure of the dependency relationship based on preset feature extraction parameters.
[0057] Step S5: The preset initial Burt causal relationship extraction model is trained through the above causal structure until the preset training standard is reached to obtain the intermediate Burt causal relationship extraction model.
[0058] Step S6: Obtain the manually annotated causal structure of the second word, the second short sentence, and the dependency relationship.
[0059] Step S7: The intermediate state Burt causal relationship extraction model is trained by the manually annotated causal structure until the training standard is reached to obtain the trained Burt causal relationship extraction model.
[0060] Among them, before the step of obtaining the manually annotated causal structure of the second word, the second sentence and the dependency relationship, it includes: inputting the second word, the second sentence and the dependency relationship into a pre-trained GPT large language model, and outputting the manually annotated causal structure.
[0061] Further, after the step of training the intermediate state Burt causal relationship extraction model by manually annotating the causal structure until the training standard is reached and the trained Burt causal relationship extraction model is obtained, the method further includes:
[0062] Step B1: Input the above causal structure and the above manually annotated causal structure into the above trained Bert causal relationship extraction model, and output the second sentence and the third sentence with causal relationship respectively.
[0063] Step B2: Analyze the correlation between the second statement and the third statement by using statistical analysis software.
[0064] Step B3: When the above correlation is greater than the preset threshold, save the above trained Burt causal relationship extraction model; when the above correlation is less than or equal to the preset threshold, repeat the above steps S1 to S7 and the above steps B1 to B2 until the above correlation is greater than the preset threshold, and then save the above trained Burt causal relationship extraction model.
[0065] Furthermore, the above preset threshold is 0.8.
[0066] An embodiment of the present invention provides a method for constructing a Burt causal relationship extraction model, the method comprising: obtaining initial text data; the initial text data meets the preset judgment criteria for the digital twin watershed construction problem; screening the preset key words and sentences from the initial text data according to preset keywords; preprocessing the preset key words and sentences to obtain the second word and the second short sentence of the preset key words and sentences and the dependency relationship of the second word; extracting the causal structure of the second word, the second short sentence and the dependency relationship based on preset feature extraction parameters; training the preset initial Burt causal relationship extraction model through the causal structure until the preset training standard is reached to obtain an intermediate Burt causal relationship extraction model; obtaining the manually annotated causal structure of the second word, the second short sentence and the dependency relationship; training the intermediate Burt causal relationship extraction model through the manually annotated causal structure until the training standard is reached to obtain the trained Burt causal relationship extraction model.
[0067] Example 3
[0068] Based on the above embodiments, Figure 4 A schematic diagram of the structure of a risk factor identification device provided by an embodiment of the present invention.
[0069] Among them, the device is applied to a server; the above-mentioned server is provided with a virtual model of a water conservancy system based on digital twin watershed technology.
[0070] Depend on Figure 4 As can be seen, the device includes:
[0071] The data acquisition module 31 is used to acquire the text data to be recognized for constructing the above virtual model.
[0072] The preprocessing module 32 is used to preprocess the text data to be recognized to obtain the first word and the first short sentence of the text data to be recognized.
[0073] The causal relationship determination module 33 is used to input the above-mentioned first word and the above-mentioned first short sentence into a pre-trained Burt causal relationship extraction model, output the causal relationship between the above-mentioned first words in each of the above-mentioned first short sentences, and obtain the first sentence with a causal relationship; the above-mentioned Burt causal relationship extraction model is pre-trained based on preset key words and phrases of the digital twin watershed construction problem and the causal relationship of the above-mentioned preset key words and phrases.
[0074] The sentence extraction module 34 is used to extract the target sentence related to the risk in the first sentence according to the preset logical association words.
[0075] The central value calculation module 35 is used to calculate the central value of the target sentence.
[0076] The risk factor determination module 36 is used to determine the risk factor of the text data to be identified according to the central value.
[0077] Among them, the data acquisition module 31, the preprocessing module 32, the causal relationship determination module 33, the sentence extraction module 34, the central value calculation module 35 and the risk factor determination module 36 are connected in sequence.
[0078] The risk factor identification device provided in the embodiment of the present invention has the same technical features as the risk factor identification method provided in the above embodiment, so it can also solve the same technical problems and achieve the same technical effects. Those skilled in the art can clearly understand that for the convenience and simplicity of description, the specific working process of the device described above can refer to the corresponding process in the above method embodiment, and will not be repeated here.
[0079] Example 4
[0080] This embodiment provides an electronic device, including a processor and a memory, wherein the memory stores computer executable instructions that can be executed by the processor, and the processor executes the computer executable instructions to implement the steps of the risk factor identification method.
[0081] This embodiment provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, the steps of the risk factor identification method are implemented.
[0082] See also Figure 5 The schematic diagram of the structure of an electronic device shown in the figure comprises: a memory 41 and a processor 42. The memory 41 stores a computer program that can be run on the processor 42. When the processor executes the computer program, the steps provided by the above-mentioned risk factor identification method are implemented.
[0083] like Figure 5As shown, the device further includes: a bus 43 and a communication interface 44, a processor 42, a communication interface 44 and a memory 41 are connected via the bus 43; the processor 42 is used to execute executable modules stored in the memory 41, such as computer programs.
[0084] The memory 41 may include a high-speed random access memory (RAM), and may also include a non-volatile memory, such as at least one disk memory. The communication connection between the system network element and at least one other network element is realized through at least one communication interface 44 (which may be wired or wireless), and the Internet, wide area network, local area network, metropolitan area network, etc. may be used.
[0085] The bus 43 may be an ISA bus, a PCI bus, or an EISA bus. The bus may be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 5 Only one bidirectional arrow is used in the diagram, but this does not mean that there is only one bus or only one type of bus.
[0086] Among them, the memory 41 is used to store the program, and the processor 42 executes the program after receiving the execution instruction. The method performed by the risk factor identification device disclosed in any embodiment of the present invention can be applied to the processor 42, or implemented by the processor 42. The processor 42 may be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above method can be completed by the hardware integrated logic circuit or software instructions in the processor 42. The above processor 42 can be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it can also be a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components. The disclosed methods, steps and logic block diagrams in the embodiments of the present invention can be implemented or executed. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc. The steps of the method disclosed in the embodiment of the present invention can be directly embodied as being executed by a hardware decoding processor, or can be executed by a combination of hardware and software modules in the decoding processor. The software module can be located in a storage medium mature in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, or an electrically erasable programmable memory, a register, etc. The storage medium is located in the memory 41, and the processor 42 reads the information in the memory 41 and completes the steps of the above method in combination with its hardware.
[0087] Furthermore, an embodiment of the present invention also provides a machine-readable storage medium, which stores machine-executable instructions. When the machine-executable instructions are called and executed by the processor 42, the machine-executable instructions prompt the processor 42 to implement the above-mentioned risk factor identification method.
[0088] The electronic device and computer-readable storage medium provided by the embodiments of the present invention have the same technical features, and therefore can solve the same technical problems and achieve the same technical effects.
[0089] In addition, in the description of the embodiments of the present invention, unless otherwise clearly specified and limited, the terms "installed", "connected", and "connected" should be understood in a broad sense, for example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be a direct connection or an indirect connection through an intermediate medium, and it can be the internal communication of two components. For ordinary technicians in this field, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.
[0090] In the description of the present invention, it should be noted that the terms "center", "upper", "lower", "left", "right", "vertical", "horizontal", "inner", "outer", etc., indicating the orientation or positional relationship, are based on the orientation or positional relationship shown in the drawings, and are only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as limiting the present invention. In addition, the terms "first", "second", and "third" are used for descriptive purposes only, and cannot be understood as indicating or implying relative importance.
Claims
1. A risk factor identification method, characterized in that: Applied to a server; a virtual model of a water conservancy system based on digital twin watershed technology is set on the server; the method comprises: Acquire text data to be recognized for constructing the virtual model; Preprocessing the text data to be recognized to obtain a first word and a first short sentence of the text data to be recognized; Inputting the first word and the first short sentence into a pre-trained Burt causal relationship extraction model, outputting the causal relationship between the first words in each of the first short sentences, and obtaining a first sentence with a causal relationship; the Burt causal relationship extraction model is pre-trained based on preset key words and sentences of the digital twin watershed construction problem and the causal relationship of the preset key words and sentences; Extracting target sentences related to the risk in the first sentence according to preset logical association words; Calculating the central value of the target sentence; The risk factor of the text data to be identified is determined according to the central value.
2. The risk factor identification method according to claim 1, characterized in that: The text data to be identified includes: calculation text data, algorithm text data and computing power text data for constructing the virtual model; the virtual model includes: the digital resource model, data engine model, data model, water conservancy perception model, water conservancy professional model, visualization model, knowledge platform model, simulation model, water conservancy cloud construction model and water conservancy information network construction model of the water conservancy system.
3. The risk factor identification method according to claim 2, characterized in that: Before the step of inputting the word and the short sentence into a pre-trained Bert causal relationship extraction model, outputting the causal relationship between the word and the short sentence, and obtaining a sentence with a causal relationship, the method includes: Step S1: Acquire initial text data; the initial text data meets the preset judgment criteria for the digital twin watershed construction problem; Step S2: According to the preset keywords, the preset keyword phrases are screened from the initial text data; Step S3: preprocessing the preset key words and sentences to obtain the dependency relationship between the second word of the preset key words and sentences and the second short sentence and the second word; Step S4: extracting the second word, the second short sentence and the causal structure of the dependency relationship based on preset feature extraction parameters; Step S5: training the preset initial Burt causal relationship extraction model through the causal structure until the preset training standard is reached to obtain the intermediate Burt causal relationship extraction model; Step S6: obtaining the manually annotated causal structure of the second word, the second short sentence and the dependency relationship; Step S7: The intermediate state Burt causal relationship extraction model is trained by the manually annotated causal structure until the training standard is reached to obtain the trained Burt causal relationship extraction model.
4. The risk factor identification method according to claim 3, characterized in that: Before the step of obtaining the second word, the second short sentence, and the manually annotated causal structure of the dependency relationship, the method includes: The second word, the second short sentence and the dependency relationship are input into a pre-trained GPT large language model, and the manually annotated causal structure is output.
5. The risk factor identification method according to claim 4, characterized in that: After the step of training the intermediate state Burt causal relationship extraction model by manually annotating the causal structure until the training standard is reached and the trained Burt causal relationship extraction model is obtained, the method further includes: Step B1: inputting the causal structure and the manually annotated causal structure into the trained Bert causal relationship extraction model, and outputting the second sentence and the third sentence having the causal relationship respectively; Step B2: analyzing the correlation between the second statement and the third statement by statistical analysis software; Step B3: When the correlation is greater than a preset threshold, save the trained Burt causal relationship extraction model; when the correlation is less than or equal to the preset threshold, repeat steps S1 to S7 and steps B1 to B2 until the correlation is greater than the preset threshold, and then save the trained Burt causal relationship extraction model.
6. The risk factor identification method according to claim 1, characterized in that: The step of calculating the central value of the target sentence comprises: Based on the social network analysis method, the central value of the target sentence is calculated.
7. The risk factor identification method according to claim 1, characterized in that: The step of determining the risk factor of the text data to be identified according to the central value includes: Sorting the central values to obtain sorted central values; Delete the target center value that is greater than the preset threshold from the sorted center values; The risk factor of the text data to be recognized is determined according to the target sentence corresponding to the target central value.
8. The risk factor identification method according to claim 1, characterized in that: The logical conjunctions include: lead, cause, cause, solve, how, can, lack, need, conduct, then and therefore.
9. A risk factor identification device, characterized in that: Applied to a server; a virtual model of a water conservancy system based on digital twin watershed technology is provided on the server; the device comprises: A data acquisition module, used to acquire text data to be recognized for constructing the virtual model; A preprocessing module, used for preprocessing the text data to be recognized to obtain a first word and a first short sentence of the text data to be recognized; A causal relationship determination module is used to input the first word and the first short sentence into a pre-trained Burt causal relationship extraction model, output the causal relationship between the first words in each of the first short sentences, and obtain a first sentence with a causal relationship; the Burt causal relationship extraction model is pre-trained based on preset key words and sentences of the digital twin watershed construction problem and the causal relationship of the preset key words and sentences; A sentence extraction module, used to extract a target sentence related to the risk in the first sentence according to preset logical association words; A central value calculation module, used to calculate the central value of the target sentence; The risk factor determination module is used to determine the risk factor of the text data to be identified according to the central value.
10. An electronic device, characterized in that: The electronic device comprises a processor and a memory, wherein the memory stores computer-executable instructions that can be executed by the processor, and the processor executes the computer-executable instructions to implement the method according to any one of claims 1 to 8.
Citation Information
Patent Citations
Key subsystem recognition method and system influencing numerical control equipment reliability
CN110716533A
Method, device and equipment for constructing causal relationship determination model
CN112329478A
Text recognition method, device and equipment
CN115510871A
Digitalized generation method and system for risk management and control of maintenance task of transformer substation, and medium
CN118521294A
Causal relationship strength evaluation method and device, equipment, storage medium and product
CN118569249A