Risk analysis method, device, equipment and computer-readable storage medium

By building a rule labeling system and a model identification system, combining a linear regression model and a decision tree regression model, qualitative and quantitative analysis of long text data is solved, and the problem of low efficiency of web page risk assessment in the existing technology is achieved, and rapid and accurate risk analysis and improved audit efficiency is achieved.

CN112818699BActive Publication Date: 2025-07-25WEBANK (CHINA)
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202110235177.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-03-03
Publication Date
2025-07-25
Estimated Expiration
2041-03-03

AI Technical Summary

Technical Problem

In the prior art, web page risk assessment mainly relies on user reports and qualitative analysis, which is inefficient and difficult to quickly and accurately determine the risks of web pages, and quantitative analysis cannot be achieved.

Method used

By obtaining long text data with a byte length greater than or equal to the preset length threshold, using the pre-constructed rule labeling system and model identification system, the first feature vector and the second feature vector are generated respectively, and the trained first regression model and the second regression model are input to fuse the first risk intensity and the second risk intensity to achieve qualitative and quantitative risk analysis.

Benefits of technology

It improves the accuracy and audit efficiency of web page risk analysis, can quickly and accurately determine the risk level of web pages, and reduces the workload of manual review.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN112818699B_ABST
    Figure CN112818699B_ABST
Patent Text Reader

Abstract

The present application provides a risk analysis method, apparatus, device and computer-readable storage medium. Among them, the method includes: obtaining long text data to be analyzed with a byte length greater than or equal to a preset length threshold; inputting the data to be analyzed into a pre-constructed rule label system to obtain a first feature vector of the data to be analyzed; inputting the data to be analyzed into a pre-constructed model recognition system to obtain a second feature vector of the data to be analyzed; inputting the first feature vector into a trained first regression model to obtain a first risk intensity; inputting the second feature vector into a trained second regression model to obtain a second risk intensity; performing a fusion process on the first risk intensity and the second risk intensity to obtain the risk intensity of the data to be analyzed, thus realizing qualitative and quantitative analysis of the risk of long text data to be analyzed and improving the accuracy of risk analysis.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of machine learning technology, and relates to, but is not limited to, a risk analysis method, device, equipment and computer-readable storage medium. Background Art

[0002] With the rapid development of Internet technology, more and more information is obtained by browsing the Internet. In the related technology for detecting web pages that have published illegal or bad information, one method mainly relies on user reports and is reviewed by reviewers. Obviously, this method is extremely inefficient and the network security of users cannot be guaranteed for web pages. Another method is based on risk assessment for analysis, but risk assessment can only perform qualitative analysis on risks, and quantitative analysis of risks still requires reviewers to review and determine one by one, making it difficult to quickly and accurately determine the risks of web pages. Summary of the Invention

[0003] Embodiments of this application provide a risk analysis method, device, equipment, computer-readable storage medium and computer program product, which can quickly and accurately determine the risks of web pages.

[0004] The technical solution of the embodiments of this application is implemented as follows:

[0005] Embodiments of this application provide a risk analysis method, and the method includes:

[0006] Obtain data to be analyzed, where the data to be analyzed is a long text with a byte length greater than or equal to a preset length threshold;

[0007] Input the data to be analyzed into a pre-constructed rule tagging system to obtain a first feature vector of the data to be analyzed;

[0008] Input the data to be analyzed into a pre-constructed model recognition system to obtain a second feature vector of the data to be analyzed;

[0009] Input the first feature vector into a trained first regression model to obtain a first risk intensity;

[0010] Input the second feature vector into a trained second regression model to obtain a second risk intensity;

[0011] Perform fusion processing on the first risk intensity and the second risk intensity to obtain the risk intensity of the data to be analyzed.

[0012] Embodiments of this application provide a risk analysis device, and the device includes:

[0013] A first acquisition module, configured to acquire data to be analyzed, where the data to be analyzed is a long text with a byte length greater than or equal to a preset length threshold;

[0014] A first input module, configured to input the data to be analyzed into a pre-constructed rule tagging system to obtain a first feature vector of the data to be analyzed;

[0015] A second input module, configured to input the data to be analyzed into a pre-constructed model recognition system to obtain a second feature vector of the data to be analyzed;

[0016] The first input module is further configured to input the first feature vector into a trained first regression model to obtain a first risk intensity;

[0017] The second input module is further configured to input the second feature vector into a trained second regression model to obtain a second risk intensity;

[0018] A fusion module, configured to perform a fusion process on the first risk intensity and the second risk intensity to obtain a risk intensity of the data to be analyzed.

[0019] An embodiment of the present application provides a risk analysis device, including:

[0020] A memory, configured to store executable instructions;

[0021] A processor, configured to implement the method provided by the embodiment of the present application when executing the executable instructions stored in the memory.

[0022] An embodiment of the present application provides a computer-readable storage medium, on which executable instructions are stored, and when the executable instructions are executed by a processor, the method provided by the embodiment of the present application is implemented.

[0023] An embodiment of the present application provides a computer program product, including a computer program, and when the computer program is executed by a processor, the method provided by the embodiment of the present application is implemented.

[0024] The embodiment of the present application has the following beneficial effects:

[0025] In the risk analysis method provided by the embodiments of the present application, first, the data to be analyzed is obtained, and the data to be analyzed is a long text with a byte length greater than or equal to a preset length threshold; the data to be analyzed is input into a pre-constructed rule label system to obtain a first feature vector of the data to be analyzed, and the data to be analyzed is input into a pre-constructed model recognition system to obtain a second feature vector of the data to be analyzed; the first feature vector is input into a trained first regression model to obtain a first risk intensity, and the second feature vector is input into a trained second regression model to obtain a second risk intensity; the first risk intensity and the second risk intensity are fused to obtain the risk intensity of the data to be analyzed, realizing qualitative and quantitative analysis of the risk of the long text data to be analyzed, which can improve the accuracy of risk analysis. Moreover, the reviewer can improve the review efficiency of risk analysis based on the more accurate risk analysis results. Description of the Drawings

[0026] Figure 1A It is a schematic diagram of a network architecture of the risk analysis method provided by the embodiments of the present application;

[0027] Figure 1B It is another schematic diagram of a network architecture of the risk analysis method provided by the embodiments of the present application;

[0028] Figure 2 It is a schematic diagram of the composition structure of the risk analysis device provided by the embodiments of the present application;

[0029] Figure 3 It is a schematic diagram of an implementation process of the risk analysis method provided by the embodiments of the present application;

[0030] Figure 4 It is a schematic diagram of a method for analyzing the body content of an opinion page provided by the embodiments of the present application;

[0031] Figure 5 It is still another schematic diagram of an implementation process of the risk analysis method provided by the embodiments of the present application. Detailed Embodiments

[0032] In order to make the purpose, technical solutions, and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings. The described embodiments should not be regarded as limitations on the present application. All other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of the present application.

[0033] In the following description, reference is made to "some embodiments" which describe a subset of all possible embodiments. However, it can be understood that "some embodiments" can be the same subset or different subsets of all possible embodiments and can be combined with each other without conflict.

[0034] In the following description, the terms "first", "second", and "third" only distinguish similar objects and do not represent a specific order for the objects. Understandably, "first", "second", and "third" can be interchanged in a specific order or sequence when permitted, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein.

[0035] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which this application belongs. The terms used herein are only for the purpose of describing the embodiments of this application and are not intended to limit this application.

[0036] The following describes an exemplary application of the device for implementing the embodiments of the present application. The device provided by the embodiments of the present application can be implemented as a terminal device. Hereinafter, the exemplary applications covered by the terminal device when the device is implemented as a terminal device will be described.

[0037] Figure 1A It is a schematic diagram of a network architecture for a risk analysis method provided by an embodiment of the present application. As Figure 1A shown, the network architecture at least includes a risk analysis device 100, a network 200, a terminal 300, and a server 400. To support an exemplary application, the risk analysis device 100 is a device for performing risk analysis, which can be a server or a terminal device such as a desktop computer or a laptop computer. The server 400 stores a large amount of public opinion web page data. The terminal 300 is a terminal for reviewers to perform review operations, which can be a device such as a laptop computer, a tablet computer, a desktop computer, or a smart terminal. The risk analysis device 100 is connected to the terminal 300 and the server 400 through the network 200. The network 200 can be a wide area network, a local area network, or a combination of the two, and uses wireless or wired links to achieve data transmission.

[0038] The risk analysis device 100 receives a risk analysis request sent by the terminal 300, and based on this risk analysis request, sends a data acquisition request for obtaining data to be analyzed to the server 400. Among them, the data to be analyzed is a long text with a byte length greater than or equal to a preset length threshold. The server 400 responds to this data acquisition request, and obtains the body content of a public opinion web page from multiple websites stored in itself as the data to be analyzed and sends it to the risk analysis device 100. This data to be analyzed is generally a long text. The long text data to be analyzed is input into a pre-established rule tagging system, model recognition system, first regression model, and second regression model to obtain the risk intensity of the data to be analyzed. Here, the rule tagging system and model recognition system are established based on training data of short texts with a byte length less than the preset length threshold. In this way, quantitative analysis of the public opinion risk of this long text is realized. Then, the risk analysis device 100 sends the risk intensity to the terminal 300 and outputs it, so that the review personnel can review the analysis result.

[0039] Figure 1B Another schematic diagram of the network architecture of the risk analysis method provided by the embodiment of the present application is as follows Figure 1B As shown, this network architecture at least includes a risk analysis device 100, a network 200, and a terminal 300. To support an exemplary application, the risk analysis device 100 can be a server, on which a large amount of public opinion web page data is stored. The terminal 300 is a terminal for the review personnel to perform review operations, and can be devices such as a laptop computer, a tablet computer, a desktop computer, and a smart terminal. The risk analysis device 100 is connected to the terminal 300 through the network 200. The network 200 can be a wide area network or a local area network, or a combination of the two, and uses wireless or wired links to implement data transmission.

[0040] The risk analysis device 100 receives a risk analysis request sent by the terminal 300, and based on this risk analysis request, obtains the body content of a public opinion web page from multiple websites stored in itself as the data to be analyzed. This data to be analyzed is generally a long text. The long text data to be analyzed is input into a pre-established rule tagging system, model recognition system, first regression model, and second regression model to obtain the risk intensity of the data to be analyzed, realizing quantitative analysis of the public opinion risk of this long text. Then, the risk analysis device 100 sends the risk intensity to the terminal 300 and outputs it, so that the review personnel can review the analysis result.

[0041] Figure 1A The difference between the network architecture shown and Figure 1B the network architecture shown is that Figure 1A in it, the risk analysis device 100 obtains the data to be analyzed from another server 400, Figure 1B and in it, the risk analysis device 100 obtains the data to be analyzed from itself. Figure 1BAfter the risk analysis device 100 in the shown network architecture obtains the data to be analyzed, the risk analysis method it executes is the same as that of Figure 1A the risk analysis device 100 in the shown network architecture.

[0042] In some other embodiments, Figure 1A the shown network architecture or Figure 1B the risk analysis device 100 in the shown network architecture can also trigger risk analysis based on other conditions. For example, it can automatically trigger risk analysis based on a pre-set timer. At this time, the terminal 300 does not need to send a risk analysis request to the risk analysis device 100.

[0043] In still some other embodiments, the risk analysis device 100 and the terminal 300 in the above network architecture can be the same device. The analysis result of the data to be analyzed is directly output on the display interface of the risk analysis device 100, and the reviewer can view the risk intensity through the risk analysis device 100.

[0044] The device provided by the embodiments of the present application can be implemented in a hardware or a combination of software and hardware manner. The following describes various exemplary implementations of the device provided by the embodiments of the present application.

[0045] According to Figure 2 the exemplary structure of the risk analysis device shown here, the risk analysis device is shown by taking the risk analysis device 100 as an example. It can be foreseen that there are other exemplary structures of the risk analysis device. Therefore, the structure described here should not be regarded as a limitation. For example, some components described below can be omitted, or components not described below can be added to meet the special needs of certain applications.

[0046] Figure 2 The shown risk analysis device 100 includes: at least one processor 110, a memory 140, at least one network interface 120, and a user interface 130. Each component in the risk analysis device 100 is coupled together through a bus system 150. It can be understood that the bus system 150 is used to realize the connection and communication between these components. In addition to the data bus, the bus system 150 also includes a power bus, a control bus, and a status signal bus. However, for the sake of clear illustration, in Figure 2 all kinds of buses are labeled as the bus system 150.

[0047] The user interface 130 can include a display, a keyboard, a mouse, a touchpad, and a touch screen, etc.

[0048] The memory 140 can be a volatile memory or a non-volatile memory, or can include both volatile and non-volatile memories. Among them, the non-volatile memory can be a read-only memory (ROM, Read Only Memory). The volatile memory can be a random access memory (RAM, Random Access Memory). The memory 140 described in the embodiments of the present application is intended to include any suitable type of memory.

[0049] The memory 140 in the embodiments of the present application can store data to support the operation of the risk analysis device 100. Examples of such data include: any computer programs for operating on the risk analysis device 100, such as an operating system and application programs. Among them, the operating system contains various system programs, such as a framework layer, a core library layer, a driver layer, etc., for implementing various basic services and processing hardware-based tasks. The application programs can include various application programs.

[0050] As an example of the method provided in the embodiments of the present application being implemented in software, the method provided in the embodiments of the present application can be directly embodied as a combination of software modules executed by the processor 110. The software modules can be located in a storage medium, and the storage medium is located in the memory 140. The processor 110 reads the executable instructions included in the software modules in the memory 140 and combines the necessary hardware (for example, including the processor 110 and other components connected to the bus 150) to complete the method provided in the embodiments of the present application.

[0051] As an example, the processor 110 can be an integrated circuit chip with signal processing capabilities, such as a general-purpose processor, a digital signal processor (DSP, Digital Signal Processor), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among them, the general-purpose processor can be a microprocessor or any conventional processor, etc.

[0052] The risk analysis method provided in the embodiments of the present application will be described in combination with the exemplary applications and implementations of the terminal provided in the embodiments of the present application.

[0053] Figure 3 It is a schematic diagram of an implementation process of the risk analysis method provided in the embodiments of the present application, which is applied to Figure 1A or Figure 1B the risk analysis device with the network architecture shown, and will be described in combination with Figure 3 the steps shown.

[0054] Step S301, obtain the data to be analyzed.

[0055] Due to the huge amount of fraud information in the entire network, the embodiments of the present application provide a risk analysis method capable of quantitatively analyzing long texts such as web pages, and this method is applied to a risk analysis device.

[0056] When the risk analysis device performs anti-fraud analysis, it first obtains the data to be analyzed. The data to be analyzed can be that the risk analysis device obtains a web page from the network in real time and uses the body text of the web page as the data to be analyzed, or the data to be analyzed is the data stored in the storage space of the risk analysis device itself.

[0057] It should be noted that the risk analysis method provided by the embodiments of the present application can be used for quantitative risk analysis of long texts. Therefore, the data to be analyzed can be a long text with a byte length greater than or equal to a preset length threshold. A text with a byte length less than the preset length threshold is called a short text. Here, the preset length threshold can be set by the designer in advance. For example, it is set to 100 bytes. A text with a length less than 100 bytes is a short text, and a text with a length greater than or equal to 100 bytes is a long text.

[0058] For example, in an auto anti-fraud project, the data to be analyzed is collected from a web page. In the embodiments of the present application, the collected web page body is regarded as the long text data to be analyzed.

[0059] Step S302: Input the data to be analyzed into a pre-constructed rule label system to obtain a first feature vector of the data to be analyzed.

[0060] After obtaining the data to be analyzed, input the data to be analyzed into a pre-constructed rule label system to obtain a first feature vector. The first feature vector is determined according to the qualitative analysis result of the data to be analyzed. The first feature vector includes risk keywords and a first eigenvalue of the risk keyword. The first eigenvalue is used to represent whether the risk keyword exists in the data to be analyzed. When the current risk keyword exists in the data to be analyzed, the first eigenvalue is 1; when it does not exist, the first eigenvalue is 0.

[0061] For example, the predefined risk keyword set includes {"down payment", "household register", "vehicle model", "credit investigation", "overdue"}, where "down payment", "household register", "vehicle model", "credit investigation", and "overdue" are 5 different types of risk keywords.

[0062] Step S303: Input the data to be analyzed into a pre-constructed model recognition system to obtain a second feature vector of the data to be analyzed.

[0063] After obtaining the data to be analyzed, input the data to be analyzed into a pre-constructed model recognition system to obtain a second feature vector, where the second feature vector is determined according to the quantitative analysis result of the data to be analyzed. The second feature vector includes risk keywords and the second eigenvalue of the risk keyword. Here, the second eigenvalue is used to represent the recall rate of the risk keyword in the data to be analyzed, that is, the success rate of retrieving the risk keyword from the data to be analyzed. The text description includes risk keywords "down payment", "household register", and "credit investigation". Based on multi-category recall, the second eigenvalues of the three risk keywords "down payment", "household register", and "credit investigation" are 0.8, 0.9, and 0.93 respectively, so as to obtain the second feature vector as {"down payment": 0.8, "household register": 0.9, "credit investigation": 0.93}.

[0064] It should be noted that the execution order of step S302 and step S303 is not limited. Step S302 can be executed first, step S303 can be executed first, or step S302 and step S303 can be executed simultaneously.

[0065] Step S304, input the first feature vector into the trained first regression model to obtain the first risk intensity.

[0066] Here, the first feature vector is obtained by inputting the data to be analyzed into a pre-constructed rule label system. Input the first feature vector into the trained first regression model for fitting to obtain the first risk intensity. Among them, the first feature vector is a feature vector determined by qualitative analysis of the data to be analyzed. In this way, the first risk intensity of the data to be analyzed is obtained based on qualitative analysis. The first regression model can be a model trained based on linear regression or a model trained based on decision tree regression.

[0067] Step S305, input the second feature vector into the trained second regression model to obtain the second risk intensity.

[0068] Here, the second feature vector is obtained by inputting the data to be analyzed into a pre-constructed model recognition system. Input the second feature vector into the trained second regression model for fitting to obtain the second risk intensity. Among them, the second feature vector is a feature vector determined by quantitative analysis of the data to be analyzed. In this way, the second risk intensity of the data to be analyzed is obtained based on quantitative analysis. The second regression model can be a model trained based on linear regression or a model trained based on decision tree regression.

[0069] It should be noted that the execution order of step S304 and step S305 is not limited. Step S304 can be executed first, step S305 can be executed first, or step S304 and step S305 can be executed simultaneously.

[0070] Step S306: Perform a fusion process on the first risk intensity and the second risk intensity to obtain the risk intensity of the data to be analyzed.

[0071] In the embodiment of the present application, the first risk intensity obtained in step S304 and the second risk intensity obtained in step S305 are fused to obtain the final risk intensity of the data to be analyzed. When fusing, it can be fused based on a pre-determined weight. For example, the first risk intensity is 0.8, the fusion weight is 0.5, the second risk intensity is 0.6, and the fusion weight is 0.5. Thus, the final risk intensity is 0.5 * 0.8 + 0.5 * 0.6 = 0.7, that is, the risk intensity of the data to be analyzed is 0.7.

[0072] Combined with a pre-established risk level {high risk, medium risk, low risk, no risk}, where the risk intensity corresponding to high risk is 1.0, the risk intensity corresponding to medium risk is 0.7, the risk intensity corresponding to low risk is 0.4, and the risk intensity corresponding to no risk is 0.0. In this way, it can be determined that the risk level of the data to be analyzed is medium risk, that is, the data to be analyzed is medium-risk data.

[0073] The risk analysis method provided by the embodiment of the present application includes obtaining data to be analyzed, where the data to be analyzed is a long text with a byte length greater than or equal to a preset length threshold; inputting the data to be analyzed into a pre-constructed rule label system to obtain the first feature vector of the data to be analyzed; inputting the data to be analyzed into a pre-constructed model recognition system to obtain the second feature vector of the data to be analyzed; inputting the first feature vector into a trained first regression model to obtain the first risk intensity; inputting the second feature vector into a trained second regression model to obtain the second risk intensity; performing a fusion process on the first risk intensity and the second risk intensity to obtain the risk intensity of the data to be analyzed. In this way, qualitative and quantitative analysis of the risk of the long text data to be analyzed is realized, which can improve the accuracy of risk analysis. Moreover, the reviewer can improve the review efficiency of risk analysis based on the risk analysis results with higher accuracy.

[0074] In some embodiments, before Figure 3 in the steps shown in embodiment, the risk analysis device first constructs a rule label system. In one implementation, "constructing a rule label system" can be achieved through the following steps:

[0075] Step S11: Obtain a set of short text keywords.

[0076] Here, the set of short text keywords includes at least one risk keyword.

[0077] When building a rule label system, according to the actual business requirements, analyze the business data to obtain the corresponding set of short text keywords, ShortClasses. In actual implementation, the set of short text keywords can be defined by designers or defined by the risk analysis device based on machine learning.

[0078] For example, in the anti-fraud project of auto finance, the obtained set of short text keywords, ShortClasses = {"low down payment", "household register", "vehicle model", "credit investigation", "overdue"}.

[0079] Step S12, construct a rule-based multi-label system based on the at least one risk keyword to obtain a rule label system.

[0080] In some embodiments, the above step S12 can be implemented as the following steps:

[0081] Step S121, obtain long text training data.

[0082] When a web page content has fraud risk, this web page content generally includes more than one text description fragment with fraud risk.

[0083] Step S122, determine the short text training data corresponding to the long text training data based on the long text training data.

[0084] In actual implementation, in addition to context model recall, other modes or other methods can also be used to determine the short text training data from the long text training data.

[0085] Step S123, determine at least one retrieval term corresponding to each risk keyword according to the short text training data.

[0086] Here, the retrieval term of the risk keyword is a word whose semantic similarity to the risk keyword is greater than a preset threshold.

[0087] To determine the retrieval terms of each risk keyword, based on semantic analysis, words with semantic similarity greater than the preset threshold can be screened out from the words included in the short text training data as the retrieval terms corresponding to the risk keyword.

[0088] Step S124, construct a retrieval dictionary for each risk keyword based on the at least one retrieval term corresponding to each risk keyword.

[0089] Step S125, construct a rule-based multi-label system according to the short text training data and the retrieval dictionary of each risk keyword to obtain a rule label system.

[0090] Input the short text training data and the retrieval dictionaries of each risk keyword into the neural network model for training, construct a rule-based multi-label system, and obtain a rule label system.

[0091] Input the short text into the rule label system, and its output result is the first feature vector.

[0092] In the embodiments of the present application, by obtaining long text training data, determining short text training data corresponding to the long text training data based on the long text training data, and then combining a short text keyword set including at least one risk keyword, according to at least one risk keyword, construct a rule-based multi-label system to obtain a rule label system. By using the long text training data as a combination of short text training data, a rule label system that can be used for long text data to be analyzed is constructed using short text keywords, so as to qualitatively analyze the risks of long text data to be analyzed, and the accuracy of risk analysis can be improved.

[0093] In some embodiments, after step S125 and before step S302, the risk analysis device also needs to pre-train a first regression model. In one implementation, training the first regression model can be achieved through the following steps:

[0094] Step S126, obtain a long text level set.

[0095] Here, the long text level set includes at least one risk level.

[0096] According to the actual business requirements, analyze the business data to determine the long text level set LongClasses. In actual implementation, the long text level set can be defined by designers or defined by the risk analysis device based on machine learning.

[0097] For example, in the automotive finance anti-fraud project, the obtained long text level set LongClasses = {high risk, medium risk, low risk, no risk}. At the same time, the risk intensity corresponding to each risk level can be defined. For example, the risk intensity corresponding to high risk is 1.0, the risk intensity corresponding to medium risk is 0.7, the risk intensity corresponding to low risk is 0.4, and the risk intensity corresponding to no risk is 0.0.

[0098] Step S127, obtain the risk level of the long text training data.

[0099] Here, the risk level of the training data can be set by the user based on the risk of the long text training data. Further, the user can set the risk intensity of the long text training data.

[0100] Step S128: Train a first regression model according to the risk level of the long text training data and the rule label system to obtain a trained first regression model.

[0101] Input the long text training data into the rule label system to obtain a first training feature vector. In another implementation, the risk level of the long text training data and the first training feature vector can be input into the first regression model for training to obtain a trained first regression model. In another implementation, the risk intensity of the long text training data and the first training feature vector can be input into the first regression model for training to obtain a trained first regression model.

[0102] Here, the first regression model can be a model trained based on linear regression or a model trained based on decision tree regression.

[0103] After obtaining the first regression model, implementing step S302 to obtain the first feature vector can be achieved by: inputting the data to be analyzed into the rule label system, searching for the search terms included in the retrieval dictionary of each risk keyword in the data to be analyzed to obtain the search results of each risk keyword; based on each risk keyword and the search results of each risk keyword, determine the first feature vector of the data to be analyzed.

[0104] Search for each search term included in the retrieval dictionary of each short text type in the data to be analyzed; when one of the search terms included in the retrieval dictionary of the current short text type is found in the data to be analyzed, determine that the matching result of the current short text type is a successful match, that is, the first feature value is 1; when none of the search terms included in the retrieval dictionary of the current short text type is found in the data to be analyzed, determine that the matching result of the current short text type is a failed match, that is, the first feature value is 0.

[0105] In the embodiments of the present application, first train the first regression model, and then use the trained first regression model to fit the first feature vector of the data to be analyzed, which can achieve quantitative analysis of the risk of the long text data to be analyzed, thereby improving the accuracy of risk analysis.

[0106] In some embodiments, before Figure 3 the steps of the illustrated embodiment S302, the risk analysis device also needs to pre-construct a model recognition system. In one implementation, "constructing a model recognition system" can be achieved through the following steps:

[0107] Step S21: Obtain a set of short text keywords.

[0108] Here, the set of short text keywords includes at least one risk keyword.

[0109] Only one of step S21 and step S11 needs to be executed. When the risk analysis device constructs the rule label system first, only step S11 can be executed without repeating step S21; when the risk analysis device constructs the model recognition system first, only step S21 can be executed without repeating step S11.

[0110] Step S22: Construct a multi-label recognition system based on the model according to the at least one risk keyword to obtain a model recognition system.

[0111] In some embodiments, the above step S22 can be implemented as the following steps:

[0112] Step S221: Obtain long text training data.

[0113] Step S222: Determine short text training data corresponding to the long text training data based on the long text training data.

[0114] For the implementation manners of step S221 and step S222, refer to the corresponding descriptions in step S121 and step S122 respectively.

[0115] Step S223: Determine the labels of the short text training data according to the at least one risk keyword.

[0116] The risk keywords included in multiple short text training data can be different, and the label of the short text training data is the risk keyword included in the short text training data.

[0117] Step S224: Construct a multi-label recognition system based on the model according to the short text training data and the labels of the short text training data to obtain a model recognition system.

[0118] Input the short text training data and the labels of the short text training data into a neural network model for training to construct a multi-label recognition system based on the model to obtain a model recognition system. Since the short text keyword set includes multiple risk keywords and a short text training data can contain multiple risk keywords, the model recognition system is a multi-label classification system.

[0119] Input the short text into the multi-label classification system, and the output result is the second feature vector.

[0120] In the embodiments of the present application, by obtaining long text training data, determining short text training data corresponding to the long text training data based on the long text training data, and then combining a short text keyword set including at least one risk keyword obtained in advance, a model-based multi-label recognition system is constructed according to at least one risk keyword to obtain a model recognition system. By using the long text training data as a combination of short text training data, a model recognition system that can be used for long text data to be analyzed is realized by using short text keywords, so as to realize quantitative analysis of the risks of long text data to be analyzed, and the accuracy of risk analysis can be improved.

[0121] In some embodiments, after the above step S224 and before step S302, the risk analysis device also needs to pre-train a second regression model. In one implementation, training the second regression model can be achieved through the following steps:

[0122] Step S225, obtain a long text level set.

[0123] Here, the long text level set includes at least one risk level.

[0124] Step S226, obtain the risk level of the long text training data.

[0125] For the implementation manners of step S225 and step S226, refer to the corresponding descriptions in step S126 and step S127 respectively.

[0126] Step S227, train a second regression model according to the risk level of the long text training data and the model recognition system to obtain a trained second regression model.

[0127] Input the long text training data into the model recognition system to obtain a second training feature vector. In another implementation, the risk level of the long text training data and the second training feature vector can be input into the second regression model for training to obtain a trained second regression model. In another implementation, the risk intensity of the long text training data and the second training feature vector can be input into the second regression model for training to obtain a trained second regression model.

[0128] Here, the second regression model can be a model trained based on linear regression or a model trained based on decision tree regression.

[0129] In some embodiments, after obtaining the risk intensity of the data to be analyzed in the above step S306, the risk analysis device can further send the risk intensity to the review terminal so that the reviewers can view the analysis results and review the analysis results.

[0130] In the embodiments of the present application, first, a second regression model is trained, and then the trained second regression model is used to fit the second feature vector of the data to be analyzed, so as to realize the quantitative analysis of the risk of the long text data to be analyzed.

[0131] Based on the foregoing embodiments, the embodiments of the present application further provide a risk analysis method, and the risk analysis method includes the following steps:

[0132] Step S401: Obtain a short text keyword set.

[0133] Here, the short text keyword set includes at least one risk keyword, and the short text is a text with a byte length less than a preset length threshold.

[0134] Step S402: Obtain long text training data.

[0135] Here, the long text is a text with a byte length greater than or equal to the preset length threshold.

[0136] Step S403: Determine the short text training data corresponding to the long text training data based on the long text training data.

[0137] Step S404: Determine at least one retrieval term corresponding to each risk keyword according to the short text training data.

[0138] Here, the retrieval term of the risk keyword is a term with a semantic similarity greater than a preset threshold to the risk keyword.

[0139] Step S405: Construct a retrieval dictionary for each risk keyword based on at least one retrieval term corresponding to each risk keyword.

[0140] Step S406: Construct a rule-based multi-label system according to the short text training data and the retrieval dictionary of each risk keyword to obtain a rule label system.

[0141] Step S407: Obtain a long text level set.

[0142] Here, the long text level set includes at least one risk level.

[0143] Step S408: Obtain the risk level of the long text training data.

[0144] Step S409: Train a first regression model according to the risk level of the long text training data and the rule label system to obtain a trained first regression model.

[0145] Step S410: Determine the label of the short text training data according to the at least one risk keyword.

[0146] Step S411: Based on the short text training data and the labels of the short text training data, construct a multi-label recognition system based on a model to obtain a model recognition system.

[0147] It should be noted that step S411 can be executed after step S403 and before any one of steps S404 to S410.

[0148] Step S412: Train a second regression model according to the risk level of the long text training data and the model recognition system to obtain a trained second regression model.

[0149] Step S413: Receive a risk analysis request.

[0150] Here, the risk analysis request carries the web page to be analyzed.

[0151] Step S414: Obtain data to be analyzed based on the web page to be analyzed.

[0152] Step S415: Input the data to be analyzed into a pre-constructed rule label system to obtain the first feature vector of the data to be analyzed.

[0153] In some embodiments, step S415 can be implemented as: input the data to be analyzed into the rule label system, search for the search terms included in the retrieval dictionary of the risk keyword short text types in the data to be analyzed for matching, and obtain the matching search results of each risk keyword short text type; based on each risk keyword short text type and the matching search results of each risk keyword short text type, determine the first feature vector of the data to be analyzed.

[0154] Step S416: Input the data to be analyzed into a pre-constructed model recognition system to obtain the second feature vector of the data to be analyzed.

[0155] Step S417: Input the first feature vector into the trained first regression model to obtain the first risk intensity.

[0156] Step S418: Input the second feature vector into the trained second regression model to obtain the second risk intensity.

[0157] Step S419: Perform a fusion process on the first risk intensity and the second risk intensity to obtain the risk intensity of the data to be analyzed.

[0158] Step S420: Sort the risk intensities of multiple data to be analyzed according to a preset sorting rule to obtain a risk analysis result.

[0159] Sort the risk intensities of multiple data to be analyzed, which is convenient for reviewers to review the risk analysis results and can improve the review efficiency.

[0160] Step S421: Output the risk analysis result.

[0161] In some embodiments, after the risk analysis device sends the data, the display terminal can sort and display the data to be analyzed in long text according to the public opinion risk value, and can also search and display according to the semantic tags of the short text corresponding to the data to be analyzed.

[0162] In the risk analysis method provided in the embodiments of the present application, first, a rule label system and a model recognition system are constructed based on long text training data, and a first regression model and a second regression model are trained. After obtaining the data to be analyzed, the data to be analyzed is input into the pre-constructed rule label system to obtain the first feature vector of the data to be analyzed; the data to be analyzed is input into the pre-constructed model recognition system to obtain the second feature vector of the data to be analyzed; then the first feature vector is input into the trained first regression model to obtain the first risk intensity; the second feature vector is input into the trained second regression model to obtain the second risk intensity; the first risk intensity and the second risk intensity are fused to obtain the risk intensity of the data to be analyzed, realizing qualitative and quantitative analysis of the risk of the long text data to be analyzed, which can improve the accuracy of risk analysis. Moreover, the reviewer reviews based on the risk analysis result with higher accuracy, which can improve the review efficiency of risk analysis.

[0163] Next, an exemplary application of the embodiments of the present application in an actual application scenario will be described.

[0164] Currently, there are many fraudulent advertisements on the Internet.

[0165] Due to the huge number of public opinions, a fraud search system is needed to recommend risk public opinions to reviewers according to the fraud risk intensity every day.

[0166] The fraud search system should have two functions: 1) It can identify fraud risk public opinions, such as identifying the risk public opinions in the web page text. 2) It can sort the identified risk public opinions in reverse order according to the risk intensity. For example, the risk intensity is between 0 and 1, and the greater the risk intensity, the greater the fraud risk of the public opinion (0.0 is no risk, 0.6 is medium-high risk, 1.0 is high risk), and the higher the ranking, so that the reviewer can process it in time.

[0167] In related technologies, when modeling the multi-classification of public opinion web page texts, most directly perform regression modeling on long texts, making it difficult to accurately determine the risk intensity of long-text public opinion, resulting in a low accuracy of the risk intensity of public opinion. It may recommend some public opinions with non-fraud risks to reviewers, leading to low review efficiency for reviewers. It may also fail to timely feedback some public opinions with fraud risks to reviewers, resulting in missed reviews and missed detections.

[0168] Figure 4 FIG. is a schematic process diagram for analyzing the main content of a public opinion web page provided by an embodiment of the present application. As Figure 4 shown, in the analysis process provided by the embodiment of the present application, it includes three models: a rule-based tag system 42, a model-based multi-label classification system 43, and a linear regression model 44. Using these three models to qualitatively and quantitatively analyze the main content 41, the risk intensity 45 is obtained to achieve the purpose of quantitatively analyzing the public opinion risk of long texts.

[0169] Since there are a large number of public opinion situations for long texts, which is not conducive to manual annotation. There are many categories of short-text public opinions, which can greatly reduce the labor cost, but the risk intensity of each category is very limited. In the embodiment of the present application, the public opinion of long texts is regarded as a combination of the public opinion of short texts. The risk of short-text public opinion is qualitatively and quantitatively analyzed through a rule-based tag system and a tag system based on multi-label classification (qualitative analysis is to analyze which tag the short text belongs to, and quantitative analysis is the probability of the tag). Then, based on the risk quantification of the two systems for short texts, regression models (which can be linear regression or decision tree regression) are respectively constructed. Then, based on the results of the two regression models, fusion is performed to obtain the quantitative analysis result of the public opinion risk of long texts.

[0170] In the embodiment of the present application, after obtaining the quantitative analysis result of the public opinion risk, the result can be displayed in various ways in the fraud search system. For example, the long texts can be sorted and displayed according to the public opinion risk value, or searched and displayed through the semantic tags of the short texts.

[0171] Figure 5 FIG. is another schematic implementation process diagram of the risk analysis method provided by the embodiment of the present application. As Figure 5 shown, the method includes the following steps:

[0172] Step S501, obtain a long-text category set (corresponding to the long-text level set in the above text).

[0173] Here, the long-text category set can be obtained based on the input operation of the reviewer. The reviewer designs the long-text category set LongClasses according to the business requirements of the project and inputs the set LongClasses into the fraud search system.

[0174] For example, in the automotive anti-fraud project, public opinion data is collected from web pages. Taking the body text of a collected web page as a long text, through reading and analyzing the corpus, the auditor summarizes the long text category set LongClasses corresponding to the body text of this web page as LongClasses = {High risk, Medium risk, Low risk, No risk}.

[0175] Step S502: Build a long text - long text category corpus based on the existing data of the project.

[0176] Continuously increase the long text - long text category corpus through manual annotation and regular expression matching.

[0177] Step S503: Design a short text category set (corresponding to the short text keyword set in the above text) according to the business requirements of the project.

[0178] Manually summarize the short text category set ShortClasses by reading and analyzing the corpus.

[0179] Short text category set ShortClasses = {Low down payment, Household register, Vehicle model, Credit investigation, Overdue}.

[0180] Step S504: Build a short text - short text category corpus based on the existing data of the project.

[0181] Continuously increase the short text - short text category corpus through manual annotation and regular expression matching.

[0182] Step S505: Build a rule-based multi-label system according to the category system of short texts.

[0183] Step S5051: Assume that there are multiple categories for the short text. For each category, manually organize the regular expressions patterns that can cover (or cover most of) the statements of this category. This step can generate a regular expression dictionary label-patterns for multiple categories.

[0184] Step S5052: In the prediction stage, for the prediction text, if the regular expression set label.patterns of this category can match a result in the prediction text, then it is considered that the text contains the semantics of this category.

[0185] In this way, we have built a multi-label recognition system for the prediction text.

[0186] Step S506: Determine the first risk intensity.

[0187] Integrate the long text - long text category corpus with the multi - label recognition system in step S505 to obtain the manually annotated labels of the long text and the output of the rule - based multi - label recognition system (regarded as the feature vector of the long text), thereby obtaining a lot of feature vectors of long texts and the corresponding category sets of long texts.

[0188] Step S5061: For the category set of long texts {high - risk, medium - risk, low - risk, risk - free}, convert it to a risk intensity from 0 to 1.

[0189] Set high - risk: 1.0, medium - risk: 0.7, low - risk: 0.4, risk - free: 0.0.

[0190] Step S5062: According to the feature vector and the risk intensity of the long text, train a linear regression model.

[0191] Step S5063: Based on the trained linear regression model, a risk intensity fitting can be made for the output (feature vector) of the rule - based multi - label recognition system to obtain the first risk intensity.

[0192] Step S507: According to the short text - short text category corpus, train a short - text multi - label classification model.

[0193] Step S5071: Short - text multi - category recall.

[0194] Train a distance - to - context model to recall short texts from long texts. Since the short - text category set is multi - category, this is a multi - category recall here.

[0195] Step S5072: According to the short - text multi - label classification, construct a model - based multi - label recognition system.

[0196] There are multiple short - text category sets, and a short text can contain multiple categories, so it is a multi - label classification system.

[0197] Step S508: Determine the second risk intensity.

[0198] Integrate the long text - long text category corpus with the output results of the model - based multi - label recognition system in step S507 to obtain the manually annotated labels of the long text and the output of the model - based multi - label recognition system (regarded as the feature vector of the long text), thereby obtaining a lot of feature vectors of long texts and the corresponding category sets of long texts.

[0199] Step S5081: Convert the category set of long texts {high - risk, medium - risk, low - risk, risk - free} to a risk intensity from 0 to 1.

[0200] Set high - risk: 1.0, medium - risk: 0.7, low - risk: 0.4, risk - free: 0.0.

[0201] Step S5082: Train a linear regression model based on the feature vector and the risk intensity of the long text.

[0202] Step S5083: Based on the trained linear regression model, the risk intensity can be fitted for the output (feature vector) of the multi-label recognition system based on the model, and a second risk intensity is obtained.

[0203] Step S509: Integrate the results of the first risk intensity and the second risk intensity to obtain the final result of the risk intensity.

[0204] For example, the final risk intensity = 0.5 * (the first risk intensity) + 0.5 * (the second risk intensity).

[0205] The situation of automotive public opinion fraud is complex, but based on business considerations, it is found that the aspects involved in fraud are limited. For example, {low down payment, household registration, vehicle model, credit investigation, overdue payment}. Therefore, a limited short text multi-category set can be designed.

[0206] Micro-level analysis. Public opinions with fraud risks basically express one or several aspects in the short text multi-category set. We use the analysis at this micro-level of multiple categories of short texts as the basis for the fraud risk of long texts. Macro-level analysis. Based on the basis at the micro-level and the multi-categories of long texts, a model is built to obtain the risk intensity.

[0207] Due to a large amount of risk public opinions, a search system with retrieval and recommendation functions is required. This system can display to the reviewers in reverse order according to the risk intensity value. In terms of the model effect, it is more detailed and accurate, and based on modeling, making the model more interpretable. In terms of the review by reviewers, it is easy for reviewers to review, improving the review efficiency.

[0208] Next, continue to describe the exemplary structure of the implementation of the risk analysis device provided in the embodiments of the present application as software modules. In some embodiments, as Figure 2 shown, the software modules in the risk analysis device 60 stored in the memory 140 may include:

[0209] The first acquisition module 61 is used to acquire the data to be analyzed, and the data to be analyzed is a long text with a byte length greater than or equal to a preset length threshold;

[0210] The first input module 62 is used to input the data to be analyzed into a pre-constructed rule label system to obtain the first feature vector of the data to be analyzed;

[0211] The second input module 63 is used to input the data to be analyzed into a pre-constructed model recognition system to obtain the second feature vector of the data to be analyzed;

[0212] The first input module 62 is further configured to input the first feature vector into the trained first regression model to obtain a first risk intensity;

[0213] The second input module 63 is further configured to input the second feature vector into the trained second regression model to obtain a second risk intensity;

[0214] The fusion module 64 is configured to perform a fusion process on the first risk intensity and the second risk intensity to obtain the risk intensity of the data to be analyzed.

[0215] In some embodiments, the risk analysis device 60 further includes:

[0216] A second acquisition module, configured to acquire a short text keyword set, where the short text keyword set includes at least one risk keyword, and the short text is a text with a byte length less than a preset length threshold;

[0217] A first construction module, configured to construct a rule-based multi-label system based on the at least one risk keyword to obtain a rule label system.

[0218] In some embodiments, the first construction module is further configured to:

[0219] Acquire long text training data;

[0220] Determine short text training data corresponding to the long text training data based on the long text training data;

[0221] According to the short text training data, determine at least one retrieval term corresponding to each risk keyword, where the retrieval term of the risk keyword is a term with a semantic similarity greater than a preset threshold to the risk keyword;

[0222] Construct a retrieval dictionary for each risk keyword based on the at least one retrieval term corresponding to each risk keyword;

[0223] Construct a rule-based multi-label system according to the short text training data and the retrieval dictionary of each risk keyword to obtain a rule label system.

[0224] In some embodiments, the risk analysis device 60 further includes:

[0225] A third acquisition module, configured to acquire a long text level set, where the long text level set includes at least one risk level;

[0226] A fourth acquisition module, configured to acquire the risk level of the long text training data;

[0227] A training module, configured to train a first regression model according to the risk level of the long text training data and the rule label system, so as to obtain a trained first regression model.

[0228] In some embodiments, the first input module 62 is further configured to:

[0229] Input the data to be analyzed into the rule label system, and search for the search terms included in the retrieval dictionary of each risk keyword in the data to be analyzed, so as to obtain the search results of each risk keyword;

[0230] Based on each risk keyword and the search results of each risk keyword, determine the first feature vector of the data to be analyzed.

[0231] In some embodiments, the risk analysis device 60 further includes:

[0232] A fifth acquisition module, configured to acquire a short text keyword set, where the short text keyword set includes at least one risk keyword;

[0233] A second construction module, configured to construct a model-based multi-label recognition system according to the at least one risk keyword, so as to obtain a model recognition system.

[0234] In some embodiments, the second construction module is further configured to:

[0235] Acquire long text training data;

[0236] Based on the long text training data, determine the short text training data corresponding to the long text training data;

[0237] According to the at least one risk keyword, determine the label of the short text training data;

[0238] According to the short text training data and the label of the short text training data, construct a model-based multi-label recognition system, so as to obtain a model recognition system.

[0239] In some embodiments, the risk analysis device 60 further includes:

[0240] A receiving module, configured to receive a risk analysis request, where the risk analysis request carries a web page to be analyzed;

[0241] In some embodiments, the risk analysis device 60 further includes:

[0242] A sorting module, configured to sort the risk intensities of multiple data to be analyzed according to a preset sorting rule, so as to obtain a risk analysis result;

[0243] An output module, configured to output the risk analysis result.

[0244] It should be noted here that the description of the above embodiments of the risk analysis device is similar to the above method description and has the same beneficial effects as the method embodiments. For the technical details not disclosed in the embodiments of the risk analysis device of the present application, those skilled in the art can refer to the description of the method embodiments of the present application for understanding.

[0245] The embodiments of the present application provide a storage medium storing executable instructions, where the executable instructions, when executed by a processor, will cause the processor to execute the method provided by the embodiments of the present application. For example, as Figure 3 shown in the method.

[0246] In some embodiments, the storage medium may be a memory such as FRAM, ROM, PROM, EPROM, EEPROM, flash memory, magnetic surface memory, optical disc, or CD-ROM; or it may be various devices including one or any combination of the above memories.

[0247] In some embodiments, the executable instructions may be in the form of a program, software, software module, script, or code, written in any form of programming language (including compiled or interpreted languages, or declarative or procedural languages), and may be deployed in any form, including being deployed as an independent program or being deployed as a module, component, subroutine, or other unit suitable for use in a computing environment.

[0248] As an example, the executable instructions may or may not correspond to a file in the file system, may be stored as a part of a file storing other programs or data. For example, they may be stored in one or more scripts in a Hyper Text Markup Language (HTML) document, stored in a single file dedicated to the program being discussed, or stored in multiple cooperating files (for example, files storing one or more modules, subroutines, or code portions).

[0249] As an example, the executable instructions may be deployed to be executed on one computing device, or on multiple computing devices located at one location, or on multiple computing devices distributed at multiple locations and interconnected through a communication network.

[0250] The above is only the embodiments of the present application and is not intended to limit the protection scope of the present application. Any modifications, equivalent replacements, and improvements made within the spirit and scope of the present application are all included in the protection scope of the present application.

Claims

1. A risk analysis method, characterized in that, The method includes: Obtaining the data to be analyzed, where the data to be analyzed is a long text with a byte length greater than or equal to a preset length threshold; Inputting the data to be analyzed into a pre-constructed rule tagging system to obtain a first feature vector of the data to be analyzed; Inputting the data to be analyzed into a pre-constructed model recognition system to obtain a second feature vector of the data to be analyzed; Inputting the first feature vector into a trained first regression model to obtain a first risk intensity; Inputting the second feature vector into a trained second regression model to obtain a second risk intensity; Performing a fusion process on the first risk intensity and the second risk intensity to obtain the risk intensity of the data to be analyzed; The method further includes: Obtaining a short text keyword set, where the short text keyword set includes at least one risk keyword, and the short text is a text with a byte length less than the preset length threshold; Obtaining long text training data; Determining corresponding short text training data for the long text training data based on the long text training data; According to the short text training data, determining at least one retrieval term corresponding to each risk keyword, where the retrieval term corresponding to the risk keyword is a term with a semantic similarity greater than a preset threshold to the risk keyword; Constructing a retrieval dictionary for each risk keyword based on at least one retrieval term corresponding to each risk keyword; Constructing a rule-based multi-label system according to the short text training data and the retrieval dictionary of each risk keyword to obtain a rule tagging system; The method further includes: Obtaining a short text keyword set, where the short text keyword set includes at least one risk keyword; Determining the labels of the short text training data according to the at least one risk keyword; Constructing a model-based multi-label recognition system according to the short text training data and the labels of the short text training data to obtain a model recognition system.

2. The method according to claim 1, wherein The method further includes: Obtaining a long text level set, where the long text level set includes at least one risk level; Obtaining the risk level of the long text training data; Training a first regression model according to the risk level of the long text training data and the rule tagging system to obtain a trained first regression model.

3. The method according to claim 1, wherein The step of inputting the data to be analyzed into a pre-constructed rule tagging system to obtain a first feature vector of the data to be analyzed includes: Inputting the data to be analyzed into the rule tagging system, and searching for retrieval terms included in the retrieval dictionary of each risk keyword in the data to be analyzed to obtain search results for each risk keyword; Determining the first feature vector of the data to be analyzed based on each risk keyword and the search results of each risk keyword.

4. The method according to claim 1, wherein The method further includes: Sorting the risk intensities of multiple data to be analyzed according to a preset sorting rule to obtain a risk analysis result; Outputting the risk analysis result.

5. A risk analysis device, characterized in that, The device includes: A first acquisition module, configured to acquire data to be analyzed, where the data to be analyzed is a long text with a byte length greater than or equal to a preset length threshold; A first input module, configured to input the data to be analyzed into a pre-constructed rule tagging system to obtain a first feature vector of the data to be analyzed; A second input module, configured to input the data to be analyzed into a pre-constructed model recognition system to obtain a second feature vector of the data to be analyzed; The first input module is further configured to input the first feature vector into a trained first regression model to obtain a first risk intensity; The second input module is further configured to input the second feature vector into a trained second regression model to obtain a second risk intensity; A fusion module, configured to perform a fusion process on the first risk intensity and the second risk intensity to obtain a risk intensity of the data to be analyzed; A second acquisition module, configured to acquire a short text keyword set, where the short text keyword set includes at least one risk keyword, and the short text is a text with a byte length less than a preset length threshold; A first construction module, configured to acquire long text training data; determine short text training data corresponding to the long text training data based on the long text training data; determine at least one retrieval term corresponding to each risk keyword according to the short text training data, where the retrieval term corresponding to the risk keyword is a term with a semantic similarity greater than a preset threshold to the risk keyword; construct a retrieval dictionary for each risk keyword based on the at least one retrieval term corresponding to each risk keyword; construct a rule-based multi-label system based on the short text training data and the retrieval dictionary for each risk keyword to obtain a rule tagging system; A fifth acquisition module, configured to acquire a short text keyword set, where the short text keyword set includes at least one risk keyword; A second construction module, configured to determine labels of the short text training data according to the at least one risk keyword; construct a model-based multi-label recognition system based on the short text training data and the labels of the short text training data to obtain a model recognition system.

6. A risk analysis device, characterized in that, The device includes: A memory, configured to store executable instructions; A processor, configured to implement the method according to any one of claims 1 to 4 when executing the executable instructions stored in the memory.

7. A computer-readable storage medium, characterized in that, An executable instruction is stored on the computer-readable storage medium, and when being executed by a processor, the method according to any one of claims 1 to 4 is implemented.

8. A computer program product, comprising a computer program, characterized in that, The computer program, when being executed by a processor, implements the method according to any one of claims 1 to 4.

Citation Information

Patent Citations

  • Method, apparatus and device for predicting trend of IT system performance risk

    CN109062769A

  • Contract term risk intelligent identification method and device, electronic equipment and storage medium

    CN112232088A