Risk control policy analysis method and device, electronic equipment and storage medium
By constructing a test sample set and using generative adversarial networks to generate variant samples, the problem of missed detections in risk control strategies was solved, enabling automated testing and optimization of risk control strategies, improving the accuracy and efficiency of risk control strategies, and ensuring content security.
Patent Information
- Application Number
- CN202310926762.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-07-26
- Publication Date
- 2025-12-05
- Estimated Expiration
- 2043-07-26
AI Technical Summary
Existing risk control strategies have the risk of missing detections when processing internet content, and are difficult to deal with diverse content formats and security issues. Traditional manual processing methods are inefficient and highly subjective, and cannot effectively address content security crises.
A risk control strategy analysis method based on chaos theory is adopted. By obtaining instances of missed detections by the target risk control strategy, a test sample set with target characteristics is constructed. Generative adversarial networks are used to generate variant samples for testing to evaluate the performance of the risk control strategy.
It improves the accuracy and efficiency of risk control strategies, can automatically identify missed detection risks, provide a data foundation to optimize strategies, prevent missed detections, and enhance content security.
Smart Images

Figure CN117076746B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present disclosure relates to the field of data processing, especially the field of artificial intelligence, big data, etc. BACKGROUND
[0002] The Internet is full of various contents, some of which are made by self-media and some of which are made by units. In order to ensure the healthy development of the content ecology and prevent the leakage of contents with security problems, the Internet platform not only manually processes the contents, but also develops a machine processing mode to assist the manual processing, so as to cope with the content security problem. Machine processing generally adopts a corresponding risk control strategy to achieve it. Machine processing can also be referred to as machine review, and one of the advantages of machine review is that it can cope with the content processing work of different media. Some repetitive and computationally intensive work is handed over to the machine for screening, and a large number of definitely illegal contents are first removed, and the remaining contents are then processed by manual processing. However, due to the diversity of risk control security problems, the risk control strategy of machine review also has the risk of missing detection. SUMMARY
[0003] The present disclosure provides a risk control strategy analysis method and device, an electronic device, and a storage medium.
[0004] According to an aspect of the present disclosure, a risk control strategy analysis method is provided, comprising:
[0005] obtaining a target instance missed by a target risk control strategy;
[0006] constructing a test sample set with a target feature based on the target feature causing the target instance to be missed;
[0007] testing the target risk control strategy based on the test sample set to obtain an evaluation result of the target risk control strategy.
[0008] According to another aspect of the present disclosure, a risk control strategy analysis device is provided, comprising:
[0009] an obtaining module configured to obtain a target instance missed by a target risk control strategy;
[0010] a constructing module configured to construct a test sample set with a target feature based on the target feature causing the target instance to be missed;
[0011] a testing module configured to test the target risk control strategy based on the test sample set to obtain an evaluation result of the target risk control strategy.
[0012] According to another aspect of the present disclosure, an electronic device is provided, comprising:
[0013] at least one processor; and
[0014] a memory in communication with the at least one processor; wherein
[0015] The memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform the method of any embodiment of the present disclosure.
[0016] According to another aspect of the present disclosure, there is provided a non-transitory computer readable storage medium storing computer instructions, wherein the computer instructions are used to cause the computer to perform the method according to any embodiment of the present disclosure.
[0017] According to another aspect of the present disclosure, there is provided a computer program product comprising a computer program which, when executed by a processor, implements the method according to any embodiment of the present disclosure.
[0018] The embodiments of the present disclosure can automatically mine a target risk control strategy with risks, and construct a test sample set by causing a target feature of missed detection to test and analyze the strategy, thereby providing a data basis for optimizing the strategy.
[0019] It should be understood that the content described in this part is not intended to identify key or important features of the embodiments of the present disclosure, nor to limit the scope of the present disclosure. Other features of the present disclosure will become apparent from the following description. BRIEF DESCRIPTION OF DRAWINGS
[0020] The accompanying drawings are used to better understand the present scheme and do not limit the present disclosure. Among them:
[0021] Figure 1 is a flowchart of an analysis method of a risk control strategy according to an embodiment of the present disclosure;
[0022] Figure 2 is a flowchart of an analysis method of a risk control strategy according to an embodiment of the present disclosure;
[0023] Figure 3 is a schematic diagram of three requirements to be determined in the analysis method of the risk control strategy according to an embodiment of the present disclosure;
[0024] Figure 4 is a schematic diagram of outputting an evaluation result according to an embodiment of the present disclosure;
[0025] Figure 5 is a schematic diagram of a rotation type of variation according to an embodiment of the present disclosure;
[0026] Figure 6 is a flowchart of an analysis method of a risk control strategy according to an embodiment of the present disclosure;
[0027] Figure 7This is a schematic diagram illustrating three main aspects of the analysis of a risk control strategy according to an embodiment of this disclosure;
[0028] Figure 8 This is a schematic diagram of the structure of an analysis device for a risk control strategy according to an embodiment of the present disclosure;
[0029] Figure 9 This is a block diagram of an electronic device used to implement the risk control strategy analysis method of the embodiments of this disclosure. Detailed Implementation
[0030] The exemplary embodiments of this disclosure are described below with reference to the accompanying drawings, including various details of the embodiments to aid understanding, and should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope of this disclosure. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.
[0031] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature. In the description of this disclosure, "multiple" means two or more, unless otherwise explicitly specified.
[0032] With the increasing frequency and volume of internet content, and the upgrading of markets and industries, the media and forms of information dissemination are becoming more and more diversified. Traditional manual processing methods, due to their low efficiency, high subjectivity, and inconsistent evaluation standards, are no longer able to effectively address content security crises. This has led to the development of machine processing based on risk control strategies.
[0033] One major advantage of machine processing is its ability to handle content processing across different media dimensions, first eliminating a large amount of clearly violating content, and then handing the rest over to human processing. However, due to the diversity of risk control and security issues, risk control strategies also carry the risk of missed detections.
[0034] In view of this, this disclosure provides an analytical method for risk control strategies based on chaotic thinking. This facilitates risk mining of risk control strategies with the risk of missed detection, thereby providing strong data support for improving the performance of risk control strategies. Moreover, it also helps to identify problems in risk control strategies in advance and prevent problems before they occur.
[0035] like Figure 1 The diagram shown is a flowchart illustrating the risk control strategy analysis method provided in this embodiment of the disclosure, including:
[0036] S101, Obtain the target instance that was missed by the target risk control strategy.
[0037] S102, construct a test sample set with the target feature based on the target feature causing the target instance to be missed.
[0038] That is, because the target instance causes the target risk control policy to be missed due to the target feature.
[0039] In implementation, the target instance can be analyzed by a professional to determine the target feature causing the target instance to be missed. And the target feature is input through a human-computer interaction page, so as to facilitate subsequent construction of a test sample set based on the target feature.
[0040] In another embodiment, a missed feature library can be established in advance. The features in the missed feature library are identified in the target instance. The identified features are determined as target features respectively. For example, the target instance can be classified by a neural network model, and the target instance is classified into a category not containing features in the missed feature library, and in the case of containing features in the missed feature library, which specific category of features is contained.
[0041] Since the processing manner of each target feature is the same, the analysis method of the risk control policy provided by the embodiments of the disclosure will be described from the perspective of one target feature.
[0042] S103, test the target risk control policy based on the test sample set to obtain an evaluation result of the target risk control policy.
[0043] Therefore, in the embodiments of the disclosure, the missed target instance is automatically mined, and the corresponding test sample set is constructed based on the target feature causing the missed detection, so as to complete the test of the target risk control policy through the method of chaotic test. The test and evaluation result of the target risk control policy can help to intuitively understand the specific performance of the target risk control policy when facing the target feature. It provides a strong data basis for optimizing the target risk control policy.
[0044] In some embodiments, in order to be able to determine the target feature accurately and thereby construct the test sample set, the step S102 of constructing the test sample set with the target feature based on the target feature causing the target instance to be missed can be implemented as shown in Figure 2
[0045] S201, obtaining the instance type and the missed detection reason of the target instance.
[0046] Among them, the form of publishing content on the Internet is various, which can include text, picture, video, audio, etc. Therefore, the instance type of the target instance can be any of the above forms.
[0047] Wherein, the reason for missing detection of the target instance can be labeled by a processing personnel at a manual processing link, so as to facilitate obtaining the reason for missing detection.
[0048] S202, obtaining a candidate sample set based on the instance type of the target instance, and determining a target feature based on the reason for missing detection.
[0049] In implementation, a plurality of samples with the same instance type as the target instance can be obtained from a public data set for constructing the candidate sample set. In implementation, a sample with the target feature can also be searched from the public data set based on the target feature, so as to construct the candidate sample set. In this way, more test samples can be obtained for the target feature, so as to facilitate sufficient testing of the target risk control strategy, thereby improving the confidence of the test result.
[0050] For example, in the case that the instance type of the target instance is a picture, corresponding samples are obtained from a public data set related to the picture, and the obtained samples and the picture are used together to construct the candidate sample set.
[0051] In order to obtain more test samples for sufficient and highly credible testing of the target risk control strategy, in S203, the candidate sample set is processed based on the target variation type corresponding to the target feature, to obtain a variation sample set.
[0052] It should be understood that the variation processing in the embodiments of the present disclosure refers to generating a variation sample similar to a candidate sample in the candidate sample set and having the target feature as much as possible based on the candidate sample. In this way, the test samples can be expanded through variation processing, so as to construct a test sample set based on the candidate sample set and the variation sample set in S204.
[0053] In summary, in the embodiments of the present disclosure, by obtaining a candidate sample set and constructing a variation sample with the target feature based on the candidate sample set, a large number of test samples can be constructed for the reason for missing detection (i.e., the target feature), so as to facilitate sufficient analysis of the processing effect of the target risk control strategy on the test samples containing the target feature through chaotic testing. In order to accurately understand the missing detection of the target risk control strategy, a powerful data basis is provided for improving the target risk control strategy.
[0054] In implementation, the variation sample can be generated based on a generator and a discriminator of a GAN (Generative Adversarial Networks). The generator is used to learn the way of constructing a sample with the target feature from the candidate sample set, and the discriminator is used to determine whether the generated variation sample has the target feature. In this way, more test samples can be constructed through the GAN network, so as to complete the testing of the target risk control strategy with high quality.
[0055] In some embodiments, as shown in FIG. 1, the target feature is a target image, and the target image is a picture, a video, or a combination thereof. Figure 2 In some embodiments, as shown in FIG. 1, the target feature is a target image, and the target image is a picture, a video, or a combination thereof.
[0056] In some embodiments, as shown in FIG. 1, the target feature is a target image, and the target image is a picture, a video, or a combination thereof.
[0057] For example, as shown in Table 1, different instance types can support different types of variations. It should be noted that Table 1 is only used to illustrate the embodiments of the present disclosure and does not limit the embodiments of the present disclosure.
[0058] Table 1
[0059]
[0060]
[0061] For example, as shown in Table 1, different instance types can support different types of variations. It should be noted that Table 1 is only used to illustrate the embodiments of the present disclosure and does not limit the embodiments of the present disclosure.
[0062] For example, as shown in Table 1, different instance types can support different types of variations. It should be noted that Table 1 is only used to illustrate the embodiments of the present disclosure and does not limit the embodiments of the present disclosure.
[0063] For example, as shown in Table 1, different instance types can support different types of variations. It should be noted that Table 1 is only used to illustrate the embodiments of the present disclosure and does not limit the embodiments of the present disclosure.
[0064] For example, as shown in Table 1, different instance types can support different types of variations. It should be noted that Table 1 is only used to illustrate the embodiments of the present disclosure and does not limit the embodiments of the present disclosure.
[0065] For example, as shown in Table 1, different instance types can support different types of variations. It should be noted that Table 1 is only used to illustrate the embodiments of the present disclosure and does not limit the embodiments of the present disclosure.
[0066] For example, as shown in Table 1, different instance types can support different types of variations. It should be noted that Table 1 is only used to illustrate the embodiments of the present disclosure and does not limit the embodiments of the present disclosure.
[0067] For example, as shown in Table 1, different instance types can support different types of variations. It should be noted that Table 1 is only used to illustrate the embodiments of the present disclosure and does not limit the embodiments of the present disclosure.
[0068] For example, as shown in Table 1, different instance types can support different types of variations. It should be noted that Table 1 is only used to illustrate the embodiments of the present disclosure and does not limit the embodiments of the present disclosure.
[0069] The black edge of the face corresponding to the picture can be understood as blurring the face, and different blurring degrees are different variation degrees.
[0070] For the video in Table 1, the variation of the video corresponds to the variation of the playing speed of the video, that is, the playing speed of the video. Different playing speeds can be understood as different variation degrees. Among them, the playing speed can refer to the playing speed of the picture, or the playing speed of the audio of the video. Either one can be selected, or both can be selected.
[0071] The audio track insertion corresponding to the video refers to the fusion of other sounds into the audio of the video. Different fusion sounds and different fusion positions can be understood as different variation degrees. Among them, the fusion position of the sound refers to the fusion of the selected sound at different positions on the audio timeline.
[0072] For the text in Table 1, different target words can be inserted, and the same target word can be continuously inserted into the same position or scattered and inserted into different positions. Different inserted target words or different inserted positions can be understood as different variation degrees. Among them, the target word, for example, is a word that is not allowed to be published on the network, so as to test the machine review effect of the target risk control strategy on the target word.
[0073] For the audio in Table 1, it can be speeded up, that is, the playing speed of the audio is adjusted, and different playing speeds are different variation degrees.
[0074] For the audio, it can also be audio track insertion for variation, that is, the fusion of other sounds into the audio. Different fusion sounds and different fusion positions of the sound can be understood as different variation degrees. Among them, the fusion position of the sound refers to the fusion of the selected sound at different positions on the audio timeline.
[0075] S2032, based on at least one variation degree, the test sample set is subjected to variation processing of the target variation type, and a variation sample set is obtained.
[0076] Therefore, in the embodiment of the disclosure, for the same variation type, different variation degrees can be used to expand the test sample set. The target risk control strategy is fully tested by variation. In addition, different variation degrees also facilitate the understanding of the risk control effect of the target risk control strategy on different variation degrees, and provide a data basis for optimizing the target risk control strategy.
[0077] In some embodiments, as shown in Figure 2 The instance type and the missed detection reason of the target instance can be implemented as:
[0078] S2011, obtaining the annotation result of the target instance.
[0079] S2012, obtain the instance type of the target instance from the type field of the labeling result, and obtain the missed detection reason from the reason type field of the labeling result.
[0080] In implementation, a full set of missed detection reasons can be provided as a classification standard for the target instance to be classified. A class with a confidence level higher than a confidence threshold is selected as the missed detection reason of the target instance.
[0081] In some other embodiments, the target instance can also be labeled by a processing person to obtain the missed detection reason of the target instance.
[0082] The instance type of the target instance can be determined based on the file format of the target instance, which is not limited in the embodiments of the present disclosure.
[0083] In the embodiments of the present disclosure, the instance type and the missed detection reason of the target instance can be accurately determined through the labeling result, so as to facilitate the production of the test sample set.
[0084] In some embodiments, obtaining the target instance missed by the target risk control policy can be implemented as:
[0085] Step A1, the candidate instance set transferred to the human review link is pushed to the target object for processing.
[0086] That is, a large number of instances to be screened are first processed in the machine review link using the corresponding risk control policy, and the instances passed by the machine review are output to the human review link for further manual processing.
[0087] Since the form of content dissemination has diversity, different risk control policies will be used for different forms of instances. For example:
[0088] 1. Risk control policy for text content: the most basic processing work is to match the word library for classification processing. Unlike manual work, the risk control policy required by machine review can be implemented through AI (Artificial Intelligence), and the screening of text content can be completed through a preset word library, so as to screen out unqualified text content.
[0089] 2. Risk control policy for picture content: the underlying technology is picture recognition, and based on this logic, each scene of picture content can be identified and applied to content processing.
[0090] 3. Risk control policy for video content: video content is composed of audio content and video picture, and the machine processing of video picture can capture picture frames and upload them to the machine review link for processing. The reusable part is the judgment of whether the scene, person, and object are illegal by the picture recognition channel.
[0091] 4. Risk control strategy for audio content: The technical basis for audio recognition is to establish a pronunciation template, which is based on an acoustic model. By matching the pronunciation template, the language and corresponding semantic can be recognized, and the output can be understood by the computer.
[0092] In the process design, the online content information flow will first pass through the machine review risk control strategy. Through machine processing, the platform can help to exclude a large number of exact violation content in advance, and the remaining ones are transferred to the manual processing link for manual evaluation. In this link, there may be cases (instances) that pass the machine review strategy but are rejected by human review, i.e. missed instances. Therefore, such instances can be collected to facilitate in-depth analysis of the reasons for the non-recall of cases, and to lock the problematic risk control strategies that have a risk of missing detection, i.e. to mine target risk control strategies that need to be evaluated and tested.
[0093] Step A2, obtaining the processing result of the target object on the candidate instance set.
[0094] Among them, the processing result can be divided into two categories, one is to confirm that there is no problem, and the other is to confirm that there is a problem, i.e. instances missed by the risk control strategy.
[0095] Step A3, based on the processing result of the candidate instance set, the target instance is selected from the candidate instance set.
[0096] In the embodiments of the present disclosure, through the manual processing link, the missed target instance can be flexibly and accurately screened out. In order to provide data basis for testing the target risk control strategy.
[0097] In some embodiments, based on the processing result, the candidate instance whose missed detection situation meets the preset requirement is selected from the candidate instance set as the target instance.
[0098] For example, the preset requirement required to be met by the missed detection situation can be determined according to the actual situation. In one possible implementation, the preset requirement can be designed manually based on the missed instance. In order to flexibly screen out the target instance.
[0099] In another possible implementation, the risk type, instance source, and product line of each candidate instance in the candidate instance set can be obtained. Among them, the risk type can be determined according to the reason why the candidate instance cannot be published. Different risk types have corresponding risk levels. The instance source can refer to the APP (Application, application program), terminal device, or any device that can represent the instance source.
[0100] In implementation, the preset requirement can include at least one of the following: belonging to the target risk type, belonging to the target risk level, belonging to the target instance source, and belonging to the target product line.
[0101] The embodiment of the present disclosure filters out target instances based on processing results, and requires not only that the instances be missed instances, but also that the missed detection of the missed instances meet preset requirements, thereby accurately filtering out target instances that need to be tested to complete the testing of the target risk control strategy.
[0102] In the embodiment of the present disclosure, a large number of missed instances can be collected for different risk control strategies, thereby obtaining a target case set. Three elements are obtained based on the target case set. As shown in the following table, Figure 3 The three elements include instance type, variation type, and corresponding problematic target risk control strategy.
[0103] After obtaining the corresponding three elements, a risk control chaos experiment can be performed for each target instance of the target risk control strategy to evaluate the target risk control strategy.
[0104] Correspondingly, for each target instance, the target risk control strategy is tested based on the test sample set to obtain the evaluation result of the target risk control strategy, which can be implemented as:
[0105] Step B1, obtaining the machine review result of the target risk control strategy on the test sample set.
[0106] Step B2, obtaining the standard processing result of the test sample set.
[0107] The standard processing result is the correct processing result of the test sample set. The implementation order of step B1 and step B2 is not limited.
[0108] Step B3, determining the evaluation result based on the difference between the machine review result and the standard processing result.
[0109] In the embodiment of the present disclosure, the machine review result of the target risk control strategy is judged based on the labeled processing result, and the target risk control strategy can be accurately evaluated based on the difference between the two.
[0110] For example, based on the three elements of the instance type of the target instance, the target variation type, and the target risk control strategy, the risk control chaos experiment can be implemented as randomly extracting some original pictures or original videos from an online black library, performing specific variations on them, and batch requesting the target risk control strategy, and finally outputting a target risk control strategy evaluation report of the target risk control strategy. The purpose of the experiment is to evaluate the accuracy and recall rate of the target risk control strategy for a certain type of variation case under different variation degrees, and to try to increase the confidence of the experimental results by increasing the test sample set through variation, so as to mine related missed call risk problems and prevent similar online real cases from being missed.
[0111] Therefore, the evaluation result is determined based on the difference between the machine review result and the standard processing result, which can be implemented as:
[0112] Step B31, in the case of including test samples of different degrees of variation in the test sample set, determining the machine review performance of different degrees of variation based on the difference between the machine review result and the standard processing result, the machine review performance including recall rate and / or accuracy rate.
[0113] Step B32, generating an evaluation result based on the machine review performance of different degrees of variation.
[0114] In this way, for the same target variation type of the same risk control strategy, the machine review performance corresponding to different degrees of variation in the target variation type can be analyzed, so as to mine the optimization mode of the target risk control strategy through different degrees of variation.
[0115] In implementation, the recall rate and accuracy rate of the target risk control strategy under different degrees of variation can be shown. As shown in Figure 4 , the abscissa represents the degree of variation, and the ordinate represents the recall rate or accuracy rate. Thus, based on the trend graph shown in Figure 4 , it can be analyzed that under what degree of variation the target risk control strategy seriously misses detection and cannot meet the target requirements. Thus, the target risk control strategy can be optimized to improve the recall rate and accuracy rate under this degree of variation.
[0116] In some embodiments, some degrees of variation are closely related to the threshold of the target risk control strategy. By analyzing the machine review performance of the target risk control strategy under different degrees of variation, the threshold can be optimized. For example, in the case of occlusion of a portrait, the machine review performance of different degrees of occlusion is analyzed. It can be analyzed that when the occlusion is m%, the recall rate and accuracy are poor, and thus the occlusion threshold can be set to be higher than m%.
[0117] In some embodiments, not only the performance of different degrees of variation under the same target risk control strategy in the same target variation type can be analyzed, but also the performance of different degrees of variation under the same target risk control strategy in the same target variation type can be analyzed. In the case that the target risk control strategy includes target features leading to missed detection instances, each target feature corresponds to a respective target variation type. For each target feature, the method of the embodiments of the present disclosure is used to construct a test sample set for testing. So as to test different target variation types corresponding to different target features. On this basis, the evaluation result of each target variation type of the test sample set to the target risk control strategy can be obtained; the evaluation results of the test sample sets of different target variation types to the target risk control strategy are compared to obtain the missed detection degree of the target risk control strategy to different target variation types.
[0118] For example, as shown in Table 2, the target variation types include rotation, mirroring, and splicing. The performance of the target risk control strategy under different variation types is analyzed. The comparative chart can be used to show the target risk control strategy under different instance types.
[0119] Table 2
[0120]
[0121]
[0122] For the target risk control strategy, target instances caused by different missed detection reasons can be obtained. For example, target instance 1 is caused by rotation missed detection, and target instance 2 is caused by occlusion missed detection. Figure 5 As shown, the test includes different occlusion parts and different occlusion degrees, thereby constructing a test sample set 1 for testing the machine review performance of the target risk control strategy on the rotation case. A test sample set 2 is constructed for testing the machine review performance of the target risk control strategy on the occlusion case. Thus, by comparing the machine review performance in the two cases, the direction that needs to be optimized can be mined. As shown in Figure 4 the trend chart, an intuitive and interpretable and credible data basis is provided for optimizing the target risk control strategy.
[0123] Thus, in the embodiments of the present disclosure, by comparing the missed detection cases of different variation types by the target risk control strategy, the overall effect of the target risk control strategy can be provided, so as to mine the optimization direction and optimization target of the target risk control strategy.
[0124] As shown in Figure 6 , the overall flowchart of the analysis method of the risk control strategy provided by the embodiments of the present disclosure is shown. As shown in Figures 4-6 , the instances missed by the machine review can be found, and then the target instances that need to be tested are comprehensively analyzed and mined based on the risk type, instance source, product line, etc. The three elements of the target instance, the instance type, and the target risk control strategy that need to be tested can be determined based on the test requirements. The target variation type and the target variation degree can be determined based on the test requirements. The target variation type, the target variation degree, and the instance type can obtain a candidate sample set to construct a test sample set. Then, the target risk control strategy is tested based on the test sample set, and thus the evaluation result is obtained. The evaluation result can be output and displayed, and can be referred to for display, so as to mine the corresponding optimization strategy. The evaluation result can be analyzed to determine the optimization scheme.
[0125] Figure 7 In summary, the analysis method of the risk control side provided by the embodiments of the present disclosure can mainly include three aspects, as shown in
[0126] 1) Machine review problem strategy mining, that is, positioning the target risk control strategy with problems
[0127] The specific way of mining the target risk control strategy is to input the online sample into the machine review strategy, so as to be processed by at least one risk control strategy included in the machine review strategy. The samples passed by the machine review are transferred to the human review link, and the cases missed by the machine review are evaluated by the human review.
[0128] The human review link can display a labeling template, and the labeling template is used to label the missed detection reason, instance type and risk control policy that misses the instance. Thus, the problematic target risk control policy can be mined based on the labeling.
[0129] For the target risk control policy, the target feature that causes the missed detection can be determined based on the cause of the missed detection. For example, the target feature is occlusion because of occlusion. Thus, the variation type can be determined. By tracing the product line range of the missed detection instance, the personnel of the related product line can jointly develop a test scheme.
[0130] Of course, a candidate sample set as many as possible can also be autonomously mined directly according to the target risk control policy and the target variation type, so as to facilitate the risk control chaos experiment test.
[0131] 2) Risk control chaos experiment test
[0132] A corresponding risk control chaos experiment test platform can be developed, and the test requirement, the variation type, the variation degree and the candidate sample set included in the test requirement are input into the platform. The platform can process the candidate sample set according to the test requirement to construct a test sample set. Then, the test sample set is used to complete the test on the target risk control policy, and finally, an evaluation result report is generated.
[0133] The test requirement can be transmitted to the risk control chaos experiment test platform through a defined interface format. The content required to be output can be customized through an interface, for example, an evaluation result is output in a trend chart as shown in FIG. 2, an evaluation result is output in an example as shown in Table 2, and the like. The platform can be required to randomly extract samples from the test sample set for testing. The platform can test based on the test scheme, and finally, an evaluation result is obtained. Figure 4
[0134] 3) Strategy optimization
[0135] The evaluation result report can be output and displayed in this link, so as to intuitively express the test result of the target risk control policy, so as to further mine and perfect the target risk control policy according to the displayed result. The target risk control policy is iteratively optimized through a specified optimization strategy, and then the optimized target risk control policy can continue to execute the analysis method provided in the embodiments of the present disclosure, so as to iteratively optimize the strategy.
[0136] For the evaluation result output by the risk control chaos experiment, a classification analysis of the missed case can be performed, a variation type with a serious missed call of the target risk control policy is mined, and an effective and feasible technical solution is found and developed, the strategy iteration is promoted, the case call is improved, and the occurrence of content security problems is prevented.
[0137] In some possible embodiments, the constructed test sample set can be used to train the target risk control strategy. By continuously iterating and optimizing the target risk control strategy and analyzing the data difference before and after the optimization, the recall effect of the strategy can be continuously improved.
[0138] Based on the same technical concept, the disclosure also provides an analysis device 800 for a risk control strategy, which includes, as shown in Figure 8
[0139] The acquisition module 801 is configured to acquire a target instance missed by the target risk control strategy.
[0140] The construction module 802 is configured to construct a test sample set with a target feature based on the target feature causing the target instance to be missed.
[0141] The test module 803 is configured to test the target risk control strategy based on the test sample set, and obtain an evaluation result of the target risk control strategy.
[0142] In some embodiments, the construction module includes:
[0143] The first acquisition unit is configured to acquire an instance type and a missed reason of the target instance.
[0144] The determination unit is configured to acquire a candidate sample set based on the instance type of the target instance, and determine the target feature based on the missed reason.
[0145] The variation unit is configured to perform variation processing on the candidate sample set based on a target variation type corresponding to the target feature, and obtain a variation sample set.
[0146] The construction unit is configured to construct the test sample set based on the candidate sample set and the variation sample set.
[0147] In some embodiments, the variation unit is specifically configured to:
[0148] Determine at least one variation degree based on a variation degree range corresponding to the target variation type.
[0149] Perform variation processing on the test sample set based on the at least one variation degree to obtain the variation sample set.
[0150] In some embodiments, the acquisition unit is specifically configured to:
[0151] Acquire a label result of the target instance.
[0152] Acquire the instance type of the target instance from a type field of the label result, and acquire the missed reason from a reason type field of the label result.
[0153] In some embodiments, the acquisition module includes:
[0154] a pushing unit configured to push the set of candidate instances to the human review stage to a target object for processing;
[0155] a processing unit configured to obtain a processing result of the target object on the set of candidate instances;
[0156] a screening unit configured to screen a target instance from the set of candidate instances based on the processing result of the set of candidate instances.
[0157] In some embodiments, the screening unit is specifically configured to:
[0158] screen a candidate instance from the set of candidate instances based on the processing result, as the target instance, in a case where the missed detection situation of the candidate instance meets a preset requirement;
[0159] The preset requirement includes at least one of the following: belonging to a target risk type, belonging to a target risk level, belonging to a target instance source, and belonging to a target product line.
[0160] In some embodiments, the test module includes:
[0161] a second obtaining unit configured to obtain a machine review result of the target risk control policy on the test sample set, and obtain a standard processing result of the test sample set;
[0162] an evaluation unit configured to determine an evaluation result based on a difference between the machine review result and the standard processing result.
[0163] In some embodiments, the evaluation unit is specifically configured to:
[0164] In a case where the test sample set includes test samples of multiple variation degrees, determine machine review performance of different variation degrees based on the difference between the machine review result and the standard processing result, the machine review performance including recall rate and / or accuracy rate;
[0165] generate the evaluation result based on the machine review performance of different variation degrees.
[0166] In some embodiments, in a case where the target risk control policy includes multiple target features resulting in missed detection instances, each target feature corresponds to a respective target variation type;
[0167] The obtaining module is further configured to obtain an evaluation result of each test sample set of each target variation type on the target risk control policy;
[0168] The test module is further configured to compare the evaluation results of the test sample sets of different target variation types on the target risk control policy, to obtain a missed detection degree of the target risk control policy on different target variation types.
[0169] The specific functions and examples of the modules and units of the apparatuses in the embodiments of the present disclosure are described in the related description of the corresponding steps in the method embodiments, which will not be described here.
[0170] In the technical solutions of the present disclosure, the acquisition, storage and application of user personal information comply with relevant laws and regulations and do not violate public order and good customs.
[0171] According to the embodiments of the present disclosure, the present disclosure further provides an electronic device, a readable storage medium and a computer program product.
[0172] Figure 9 A schematic block diagram of an example electronic device 900 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptops, desktops, tablets, personal digital assistants, servers, blade servers, mainframes, and other appropriate computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular telephones, smartphones, wearable devices, and other similar computing devices. The components shown here, their connections and relationships, and their functions, are meant to be examples only, and are not meant to limit implementations of the present disclosure described and / or claimed in this document.
[0173] As shown in Figure 9 The electronic device 900 includes a computing unit 901 that can perform various appropriate actions and processes in accordance with a computer program stored in a read-only memory (ROM) 902 or a computer program loaded from a storage unit 908 into a random access memory (RAM) 903. Various programs and data required for the operation of the electronic device 900 can also be stored in the RAM 903. The computing unit 901, the ROM 902, and the RAM 903 are connected to each other through a bus 904. An input / output (I / O) interface 905 is also connected to the bus 904.
[0174] Various components in the electronic device 900 are connected to the I / O interface 905, including an input unit 906 such as a keyboard, a mouse, etc., an output unit 907 such as various types of displays, a speaker, etc., a storage unit 908 such as a magnetic disk, an optical disk, etc., and a communication unit 909 such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 909 allows the electronic device 900 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.
[0175] The computing unit 901 can be various general and / or special purpose processing components with processing and computing capabilities. Some examples of the computing unit 901 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various specialized artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The computing unit 901 performs various methods and processes described above, such as the analysis method of the risk control policy. For example, in some embodiments, the analysis method of the risk control policy can be implemented as a computer software program tangibly embodied in a machine-readable medium, such as the storage unit 908. In some embodiments, part or all of the computer program can be loaded and / or installed onto the device 900 via the ROM 902 and / or the communication unit 909. When the computer program is loaded onto the RAM 903 and executed by the computing unit 901, one or more steps of the analysis method of the risk control policy described above can be performed. Alternatively, in other embodiments, the computing unit 901 can be configured to perform the analysis method of the risk control policy by any other appropriate means, such as by means of firmware.
[0176] Various implementations of the systems and techniques described above can be realized in digital electronic circuitry, integrated circuitry, a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), a system on a chip (SOC), a complex programmable logic device (CPLD), computer hardware, firmware, software, and / or combinations thereof. These various implementations can include implementation in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which can be special or general purpose, coupled to receive data and instructions from, and to transmit data and instructions to, a storage system, at least one input device, and at least one output device.
[0177] Program code for carrying out methods of the present disclosure can be written in any combination of one or more programming languages. The program code can be provided to a processor or controller of a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the program code, when executed by the processor or controller, produces a means for implementing the functions / acts specified in the flowcharts and / or block diagrams. The program code can be executed entirely on a machine, partially on a machine, partially on a machine and partially on a remote machine or entirely on a remote machine or server.
[0178] In the context of this disclosure, a machine-readable medium can be a tangible medium that contains or stores a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include but is not limited to an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0179] To provide for interaction with a user, the systems and techniques described here can be implemented on a computer having a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can be used to provide for interaction with a user as well; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form, including acoustic, speech, or tactile input.
[0180] The systems and techniques described here can be implemented in a computing system that includes a back end component (e.g., as a data server), or that includes a middleware component (e.g., an application server), or that includes a front end component (e.g., a user computer having a graphical user interface or a Web browser through which a user can interact with an implementation of the systems and techniques described here), or any combination of such back end, middleware, or front end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), and the Internet.
[0181] The computer system can include clients and servers. A client and server are generally remote from each other and typically interact through a communication network. The relationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other. The server can be a cloud server, a server of a distributed system, or a server combined with a blockchain.
[0182] It should be understood that the various forms of flow shown above can be re-ordered, steps added or removed, etc. For example, the steps recited in the present disclosure can be performed in parallel, in series, in a different order, etc., so long as the desired results of the technology disclosed in the present disclosure are achieved, which is not limited herein.
[0183] The above detailed description does not constitute a limitation of the protection scope of the present disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent replacements, and improvements made within the principles of the present disclosure shall be included in the protection scope of the present disclosure.
Claims
1. An analysis method of a risk control strategy, comprising: obtaining a target instance of a missed detection of a target risk control strategy; based on a target feature causing the target instance to be missed, constructing a test sample set with the target feature, comprising: obtaining an instance type of the target instance, a missed detection reason; based on the instance type of the target instance, obtaining a candidate sample set, and determining the target feature based on the missed detection reason; based on a target variation type corresponding to the target feature, performing variation processing on the candidate sample set to obtain a variation sample set; based on the candidate sample set and the variation sample set, constructing the test sample set; based on the test sample set, testing the target risk control strategy to obtain an evaluation result of the target risk control strategy.
2. The method of claim 1, wherein, The variation processing on the candidate sample set based on the target variation type corresponding to the target feature to obtain a variation sample set, comprising: based on a variation degree range corresponding to the target variation type, determining at least one variation degree; based on the at least one variation degree, performing variation processing of the target variation type on the test sample set to obtain the variation sample set.
3. The method of claim 1, wherein, The obtaining of the instance type of the target instance and the missed detection reason, comprising: obtaining a labeling result of the target instance; obtaining the instance type of the target instance from a type field of the labeling result, and obtaining the missed detection reason from a reason type field of the labeling result.
4. The method of any one of claims 1-3, wherein, The obtaining of the target instance of a missed detection of a target risk control strategy, comprising: pushing a candidate instance set flowing to a human review link to a target object for manual processing; obtaining a processing result of the target object on the candidate instance set; the processing result is used to indicate whether the candidate instance has a problem or not; based on the processing result of the candidate instance set, screening the target instance from the candidate instance set.
5. The method of claim 4, wherein, The screening of the target instance from the candidate instance set based on the processing result of the candidate instance set, comprising: based on the processing result, screening a candidate instance whose missed detection situation meets a preset requirement from the candidate instance set as the target instance; wherein, the preset requirement comprises at least one of the following: belonging to a target risk type, belonging to a target risk level, belonging to a target instance source, and belonging to a target product line.
6. The method of claim 1, wherein, The testing of the target risk control strategy based on the test sample set to obtain an evaluation result of the target risk control strategy, comprising: obtaining a machine review result of the target risk control strategy on the test sample set; and obtaining a standard processing result of the test sample set; based on the difference between the machine review result and the standard processing result, determining the evaluation result.
7. The method of claim 6, wherein, The determination of the evaluation result based on the difference between the machine review result and the standard processing result, comprising: in the case that the test sample set includes test samples of multiple variation degrees, based on the difference between the machine review result and the standard processing result, determining machine review performance of different variation degrees, the machine review performance including recall rate and / or accuracy rate; based on the machine review performance of different variation degrees, generating the evaluation result.
8. The method of claim 1, in a case where the target risk control policy includes multiple target features resulting in a missed detection instance, each target feature corresponding to a respective target variation type, further comprising: obtaining evaluation results of the target risk control policy on a test sample set of each target variation type; and comparing the evaluation results of the target risk control policy on the test sample sets of different target variation types to obtain a missed detection degree of the target risk control policy on different target variation types.
9. An analysis device for a risk control policy, comprising: an obtaining module configured to obtain a target instance of a missed detection of a target risk control policy; a constructing module comprising: a first obtaining unit configured to obtain an instance type and a missed detection reason of the target instance; a determining unit configured to obtain a candidate sample set based on the instance type of the target instance and determine a target feature based on the missed detection reason; a variation unit configured to perform variation processing on the candidate sample set based on a target variation type corresponding to the target feature to obtain a variation sample set; a constructing unit configured to construct a test sample set based on the candidate sample set and the variation sample set; and a testing module configured to test the target risk control policy based on the test sample set to obtain an evaluation result of the target risk control policy. The variation unit is specifically configured to: determine at least one variation degree based on a variation degree range corresponding to the target variation type; and perform variation processing on the test sample set based on the at least one variation degree to obtain the variation sample set. The obtaining unit is specifically configured to: obtain a labeling result of the target instance; and obtain the instance type of the target instance from a type field of the labeling result and obtain the missed detection reason from a reason type field of the labeling result. The obtaining module comprises: a pushing unit configured to push a candidate instance set flowing to a human review link to a target object for manual processing; a processing unit configured to obtain a processing result of the target object on the candidate instance set, the processing result being used to indicate whether a candidate instance has a problem or not; and a screening unit configured to screen the target instance from the candidate instance set based on the processing result of the candidate instance set. The screening unit is specifically configured to: screen a candidate instance whose missed detection condition meets a preset requirement from the candidate instance set as the target instance based on the processing result; and wherein the preset requirement comprises at least one of the following: belonging to a target risk type, belonging to a target risk level, belonging to a target instance source, and belonging to a target product line. The testing module comprises: a second obtaining unit configured to obtain a machine review result of the target risk control policy on the test sample set and obtain a standard processing result of the test sample set; and an evaluation unit configured to determine the evaluation result based on a difference between the machine review result and the standard processing result. The evaluation unit is specifically configured to: 10. The apparatus of claim 9, wherein, 11. The apparatus of claim 9, wherein, 12. The apparatus of any of claims 9-11, wherein, 13. The apparatus of claim 12, wherein, 14. The apparatus of claim 9, wherein, 15. The apparatus of claim 14, wherein, In a case where the test sample set includes test samples of multiple variation degrees, machine review performance of different variation degrees, including recall rate and / or accuracy rate, is determined based on differences between the machine review result and the standard processing result; The evaluation result is generated based on the machine review performance of different variation degrees.
16. The apparatus of claim 9, in a case where the target risk control policy includes multiple target features resulting in missed detection instances, each target feature corresponds to a respective target variation type; The obtaining module is further configured to obtain evaluation results of the test sample set of each target variation type on the target risk control policy; The test module is further configured to compare the evaluation results of the test sample set of different target variation types on the target risk control policy to obtain missed detection degrees of the target risk control policy on different target variation types.
17. An electronic device, comprising: at least one processor; and a memory connected to the at least one processor in communication; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1-8.
18. A non-transitory computer readable storage medium having stored thereon computer instructions, wherein, The computer instructions are used to enable the computer to perform the method of any one of claims 1-8.
19. A computer program product comprising a computer program which, when executed by a processor, implements the method of any one of claims 1-8.
Citation Information
Patent Citations
Method and device for improving risk control strategy effect
CN113538071A
Credit risk classification model test method and device, equipment and medium
CN116089869A