Illegal prediction method and related device
By identifying and processing violations of the to be processed data and making violation predictions based on historical audit data, the problems of incomplete evaluation of violation risk and omissions in the existing technology are solved, and more accurate identification of violation content and risk prediction are achieved.
Patent Information
- Application Number
- CN202411998550.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-31
- Publication Date
- 2025-05-02
AI Technical Summary
When identifying and managing violation content, the existing technology cannot fully evaluate the risk of violations by relying on a single model, and the identification method of a single violation element fails to comprehensively consider the combined effect of multiple risk elements, resulting in omissions in identification.
By performing violation identification and processing on the to be processed data, obtaining historical audit data based on risk sets composed of violation elements and/or users who publish data to make violation predictions.
It improves the accuracy and comprehensiveness of the identification of violation content, improves the accuracy of violation risk prediction, and improves processing speed and efficiency without identifying the risk set and user's no audit record.
Smart Images

Figure CN119918004A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer technology, and in particular to a violation prediction method, device, electronic device, computer-readable medium and computer program product. Background Art
[0002] With the explosive growth of Internet content, the identification and management of illegal content has become an important issue. Existing technologies usually rely on a single model to identify illegal content, which may not be able to fully assess the risk of violations. In addition, the single illegal element identification method does not consider the combined effect of multiple risk elements.
[0003] Currently, the model side often identifies a single category, and the audit judgment of whether there is a violation is based on a comprehensive judgment of multiple risk elements. If the auditor lists the various violation elements identified by himself, omissions are inevitable. Summary of the invention
[0004] Multiple aspects of the present application provide a violation prediction method, apparatus, electronic device, computer-readable medium, and computer program product.
[0005] In one aspect of the present application, a violation prediction method is provided, wherein the method comprises:
[0006] By performing violation identification processing on the data to be processed, one or more identified violation elements are obtained;
[0007] Acquire historical audit data corresponding to the data to be processed, where the historical audit data is obtained based on the risk set to be predicted consisting of the one or more violation elements and / or the user who published the data to be processed;
[0008] Based on the historical audit data, violation prediction is performed on the pending number.
[0009] In one aspect of the present application, a device for violation prediction is provided, wherein the device comprises:
[0010] A device for obtaining one or more identified violation elements by performing violation identification processing on the data to be processed;
[0011] A device for obtaining historical audit data corresponding to the data to be processed, wherein the historical audit data is obtained based on the risk set to be predicted composed of the one or more violation elements and / or the user who published the data to be processed;
[0012] A device for predicting violations of the pending number based on the historical audit data.
[0013] Another aspect of the present application provides an electronic device, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method of an embodiment of the present application.
[0014] In another aspect of the present application, a computer-readable storage medium is provided, on which computer program instructions are stored. The computer program instructions can be executed by a processor to implement the method of the embodiment of the present application.
[0015] Another aspect of the present application provides a computer program product, including a computer program, which implements the method of the embodiment of the present application when executed by a processor.
[0016] In the solution provided in the embodiment of the present application, the corresponding historical audit data is obtained by obtaining a risk set based on the violation elements and / or the user who publishes the data to be processed, and violation prediction is performed based on the historical audit data. This can comprehensively consider multiple violation elements and user behaviors, improve the accuracy and comprehensiveness of violating content identification, and improve the accuracy of violation risk prediction. The embodiment of the present application performs violation prediction based on historical audit data, and will only perform review processing when the risk set has not been identified in the previous review process and the data published by the corresponding user has not entered the review process, thereby improving processing speed and efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] In order to more clearly illustrate the technical solutions in the embodiments of the present application, a brief introduction will be given below to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.
[0018] Other features, objects and advantages of the present application will become more apparent by reading the detailed description of non-limiting embodiments made with reference to the following drawings:
[0019] Figure 1 A schematic diagram of a flow chart of a violation prediction method provided in an embodiment of the present application is shown;
[0020] FIG. 2( a ) shows a schematic diagram of an exemplary process of matching based on a risk set and a user identifier according to an embodiment of the present application;
[0021] FIG2( b ) shows a schematic diagram of an exemplary risk score calculation rule according to an embodiment of the present application;
[0022] Figure 3 A schematic diagram of the structure of a device for violation prediction provided by an embodiment of the present application is shown;
[0023] Figure 4 A schematic diagram of the structure of a device suitable for implementing the solution in the embodiment of the present application is shown.
[0024] The same or similar reference numerals in the drawings represent the same or similar components. DETAILED DESCRIPTION
[0025] In order to make the purpose, technical solution and advantages of the embodiments of the present application clearer, the technical solution in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.
[0026] In a typical configuration of the present application, the terminal and the equipment of the service network each include one or more processors (CPU), input / output interface, network interface and memory.
[0027] The memory may include non-permanent storage in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. The memory is an example of a computer-readable medium.
[0028] Computer readable media include permanent and non-permanent, removable and non-removable media, and can be implemented by any method or technology to store information. Information can be computer program instructions, data structures, modules of programs or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, read-only compact disk (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic tape disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device.
[0029] Figure 1 A schematic flow chart of a violation prediction method provided in an embodiment of the present application is shown. The method at least includes step S101, step S102 and step S103.
[0030] In actual scenarios, the execution subject of the method can be a network device, or an application running on a network device, wherein the network device includes but is not limited to a network host, a single network server, a plurality of network server sets, or a collection of computers based on cloud computing, and can be used to implement some processing functions when setting an alarm. Here, the cloud is composed of a large number of hosts or network servers based on cloud computing, wherein cloud computing is a type of distributed computing, a virtual computer composed of a group of loosely coupled computer sets.
[0031] Reference Figure 1 In step S101, one or more illegal elements are identified by performing illegal identification processing on the data to be processed.
[0032] The data to be processed may include various types of data, such as text, image, or a combination of text and image.
[0033] For image data, the illegal elements include but are not limited to identified illegal political elements, illegal scene elements, illegal character elements, etc. For text data, the illegal elements include but are not limited to identified sensitive words.
[0034] The violation identification process is used to automatically identify the violation elements contained in the data to be processed.
[0035] According to one implementation, the method uses multiple models to identify illegal elements for data in different forms, and step S101 further includes step S1011 and step S1012.
[0036] In step S1011, the data type corresponding to the number to be processed is determined.
[0037] The method may classify the data to be processed in a variety of ways to obtain corresponding data types. For example, three data types, namely, pictures, texts, and pictures and texts, may be obtained by classification based on different data forms. Alternatively, corresponding data types may be obtained by more detailed classification. For example, for image types, they may be further divided into subcategories such as faces, cartoon characters, advertisements, and buildings, so as to identify illegal elements through corresponding risk identification models.
[0038] In step S1012, based on the determined data type, a corresponding risk identification model is used to perform violation identification processing on the data to be processed.
[0039] For example, assuming that the data type may include pictures / text / a combination of pictures and text, different models are used to identify illegal elements. For picture-type data, a model based on a detection or classification network structure is used to identify illegal elements contained in the picture. For text-type data, a model based on a BERT structure is used to identify sensitive words in the text as illegal elements.
[0040] Continue to refer to Figure 1 To illustrate, in step S102, historical audit data corresponding to the data to be processed is obtained.
[0041] The historical audit data is obtained based on the risk set to be predicted consisting of the one or more violation elements and / or the user who publishes the data to be processed.
[0042] The historical audit data includes audit result information within a predetermined period of time in the past. Optionally, the audit result information includes historical pass frequency information and historical rejection frequency information. The historical pass frequency information refers to the frequency of audit passes within a predetermined period of time in the past. The historical rejection frequency information refers to the frequency of audit rejections within a predetermined period of time in the past.
[0043] Optionally, the historical audit data corresponds to different business types. The method determines the business type information corresponding to the data to be processed, and then obtains the historical audit data corresponding to the data to be processed under the business type.
[0044] The method for obtaining the historical audit data corresponding to the data to be processed includes but is not limited to:
[0045] 1) obtaining corresponding historical audit data based on the risk set consisting of the one or more violation elements;
[0046] The risk set includes one or more violation elements corresponding to a piece of data. For N pieces of data, N or fewer risk sets may be generated because multiple pieces of data may share the same risk set. For example, if multiple pieces of data contain the same violation element, they will be included in the same risk set. In addition, there is no order distinction between elements in the risk set. For example, the risk sets represented by {'element1', 'element2'} and {'element2', 'element1'} will be regarded as the same risk set. If no violation element is identified, the risk set is empty.
[0047] The method of this embodiment further includes step S104, and step S102 includes step S1021 and step S1022.
[0048] In step S104, historical risk set data is recorded and stored.
[0049] The historical risk set data includes historical audit data corresponding to multiple historical violation risk sets. The multiple historical violation risk sets are obtained based on the audit results of multiple data within a predetermined historical time period. The historical pass times and rejection times corresponding to each historical violation risk set are collected to obtain the historical audit data corresponding to the multiple historical violation risk sets.
[0050] In step S1021, one or more violation elements identified from the data to be processed are combined into a risk set to be predicted corresponding to the data to be processed.
[0051] In step S1022, a match is performed in pre-stored historical risk set data based on the risk set to be predicted to obtain historical audit data corresponding to the risk set to be predicted.
[0052] Specifically, the method includes the following steps of matching based on the risk set to be predicted: matching the risk set to be predicted with any historical violation risk set, comparing the number of violation elements in the risk set to be predicted with the number of violation elements in any historical violation risk set; if the number of violation elements in the risk set to be predicted is equal to the number of violation elements in any historical violation risk set, further checking whether the first violation element in the risk set to be predicted exists in any historical violation risk set; if the number of violation elements in the risk set to be predicted is not equal to the number of violation elements in any historical violation risk set, turning to the next historical violation risk set for matching; if the first violation element in the risk set to be predicted exists in any historical violation risk set, continuing to check one by one whether the remaining violation elements in the risk set to be predicted exist in any historical violation risk set; if any violation element in the risk set to be predicted does not exist in any historical violation risk set, turning to the next historical violation risk set for matching until a target historical violation risk set containing all violation elements in the risk set to be predicted is found, and the historical audit data corresponding to the target historical violation risk set is used as the historical audit data corresponding to the risk set to be predicted.
[0053] This search and matching method can greatly improve the matching speed when searching for the same risk set in the historical review data. For example, assuming that the total number of historical risk sets is K, this method can reduce the matching time by at least K times.
[0054] 2) Obtain corresponding historical audit data based on the user who published the data to be processed;
[0055] According to an embodiment, the method of this embodiment further includes step S105, and step S102 includes step S1023 and step S1024.
[0056] In step S105, the historical audit data of each user is recorded and stored.
[0057] In step S1023, the user identification information of the user who publishes the data to be processed is obtained.
[0058] The user identification information includes various information that can uniquely identify the user, such as user name, email address or phone number, etc.
[0059] In step S1024, based on the user identification information, the historical audit data corresponding to the user is obtained. Specifically, based on the user identification information, a query is performed in the stored historical audit data to obtain the historical audit data matching the user identification information.
[0060] 3) Obtain corresponding historical audit data based on the identified illegal elements and the users who published the data to be processed;
[0061] Specifically, the method obtains historical audit data based on the identified illegal elements, obtains historical audit data based on the user who publishes the data to be processed, and merges the obtained data, thereby using the merged data as the historical audit data corresponding to the data to be processed.
[0062] According to one embodiment, the method further includes step S106.
[0063] In step S106, the historical audit data of the historical risk set data and / or the historical audit data of each user are updated regularly.
[0064] Optionally, the method clears the stored data that is not within the predetermined time range each time the historical risk set data and / or the historical audit data of each user is updated. For example, only the historical audit data of the past 7 days is stored, and the data before 7 days is cleared.
[0065] The method of the embodiment of the present application is described below with reference to an example.
[0066] According to the first example of the present application, referring to the schematic diagram shown in Figure 2(a), the data to be processed is subjected to violation identification, and the four identified violation elements constitute a risk element set, which is represented as (e1, e2, e3, e4). In addition, this example obtains the identification information (user id) of the user who published the data to be processed.
[0067] The historical audit data in this example includes data collected based on the risk element set (e1, e2, e3, e4) and data collected based on the user ID.
[0068] The process of collecting data based on user ID includes: matching historical user data based on user ID to obtain historical behavior data of the corresponding user. The historical user data includes the violation information of multiple historical user IDs that have been stored. For each user ID, the violation information of multiple historical user IDs is obtained by collecting the historical pass times and rejection times of the data published by the user ID. For example, as shown in FIG2(a), the historical pass times of the user corresponding to the user ID are recorded as p11, and the historical rejection times are recorded as r11.
[0069] Furthermore, the system of this example updates data regularly (eg, every day), and the updated content includes user ID, the number of approvals and rejections of published data, so as to reflect the latest user behaviors and risk conditions.
[0070] Among them, the process of collecting data based on the risk element set (e1, e2, e3, e4) includes: matching the historical risk set data based on the current risk set to obtain the corresponding historical audit data. Among them, the historical risk set data includes historical audit data corresponding to multiple stored historical violation risk sets. By collecting the historical pass times and rejection times corresponding to each historical violation risk set, the historical audit data corresponding to multiple historical violation risk sets are obtained. For example, as shown in the figure, the historical audit data corresponding to the historical violation risk set R1 includes a violation element e1, and the historical pass times of the violation element e1 are recorded as p21, and the historical rejection times are recorded as r21.
[0071] Furthermore, the method of this example periodically updates the historical audit data, for example, retaining the historical audit data from the current time to M days ago.
[0072] Continue to refer to Figure 1 To illustrate, in step S103, the violation prediction of the pending number is performed based on the historical audit data.
[0073] According to one embodiment, in step S103, based on the historical audit data, the risk score corresponding to the data to be processed is calculated according to a predetermined score calculation rule, and whether the data to be processed violates the rules is determined according to the calculated risk score.
[0074] Specifically, if the risk score exceeds the preset risk threshold, the data to be processed is judged to be in violation of the rules, and a "rejected" audit result is obtained; if the risk score does not exceed the preset risk threshold, the data to be processed is judged to be not in violation of the rules, and a "passed" audit result is obtained.
[0075] According to one embodiment, the risk score is calculated based on the risk set violation score and the user violation score.
[0076] Continuing to explain the first example above, referring to FIG2(b), the risk score calculation rule of this example is as follows:
[0077] The risk score is calculated based on the formula (1) shown below:
[0078] Risk score = w × risk set violation score + (1-w) × user violation score (1)
[0079] Wherein, w represents the weight coefficient, which is used to adjust the proportion of the risk set violation score and the user violation score in the total risk score. The weight coefficient w can be obtained based on linear fitting or can be determined based on experience. The higher the risk score value calculated based on formula (1), the greater the probability of violation.
[0080] If there is no risk set violation data, the risk set violation score is 1. If there is risk set violation data, the risk set violation score is calculated based on the following formula:
[0081] Risk set violation score = historical number of rejections for the same violation risk set / (historical number of approvals for the same violation risk set + historical number of rejections for the same violation risk set) (2)
[0082] If there is no audit data for the same user, the user violation score is 1. If there is audit data for the same user, the user violation score is calculated based on the following formula:
[0083] User violation score = historical rejection times of the same user / (historical approval times of the same user + historical rejection times of the same user) (3)
[0084] The risk score calculated based on formula (1) is compared with the preset violation threshold: if the risk score is greater than the violation threshold, it is judged as a violation; if the risk score is not greater than the violation threshold, it is judged as not a violation.
[0085] The risk score calculation formula based on historical data in this example combines the risk set violation score and the user violation score to more accurately predict the probability of violation.
[0086] According to one embodiment, if the historical audit data corresponding to the risk set to be predicted or the publishing user is not obtained, that is, the risk set to be predicted is not identified in the previous audit process and the data published by the corresponding user has not entered the audit process, the violation elements contained in the risk set to be predicted are audited and processed, and the historical audit data corresponding to the risk set to be predicted or the publishing user is stored based on the audit results. When the risk set to be predicted or the publishing user is successfully matched later, the historical audit data corresponding to the predicted risk set or the publishing user can be obtained from the stored historical audit data.
[0087] According to the method of the embodiment of the present application, by obtaining corresponding historical audit data based on the risk set composed of violation elements and / or the user who published the data to be processed, and making violation predictions based on the historical audit data, it is possible to comprehensively consider multiple violation elements and user behaviors, improve the accuracy and comprehensiveness of illegal content identification, and improve the accuracy of violation risk prediction; the embodiment of the present application makes violation predictions based on historical audit data, and will only perform review processing when the risk set has not been identified in the previous review process and the data published by the corresponding user has not entered the review process, thereby improving processing speed and efficiency.
[0088] Figure 3 A schematic diagram of the structure of a device for violation prediction provided in an embodiment of the present application is shown.
[0089] The device includes: a device for obtaining one or more identified violation elements by performing violation identification processing on the data to be processed (hereinafter referred to as "violation element identification device 101"), a device for obtaining historical audit data corresponding to the data to be processed (hereinafter referred to as "historical data acquisition device 102"), and a device for predicting violations of the data to be processed based on the historical audit data (hereinafter referred to as "violation prediction device 103").
[0090] Reference Figure 3 The illegal element identification device 101 performs illegal identification processing on the data to be processed to obtain one or more identified illegal elements.
[0091] The data to be processed may include various types of data, such as text, image, or a combination of text and image.
[0092] For image data, the illegal elements include but are not limited to identified illegal political elements, illegal scene elements, illegal character elements, etc. For text data, the illegal elements include but are not limited to identified sensitive words.
[0093] The violation identification process is used to automatically identify the violation elements contained in the data to be processed.
[0094] According to one implementation, the illegal element identification device 101 uses multiple models to identify illegal elements for data in different forms, and the illegal element identification device 101 further includes a type determination device and an identification processing device.
[0095] The type determination device determines the data type corresponding to the number to be processed.
[0096] The data to be processed may be classified in a variety of ways to obtain the corresponding data types. For example, the data may be classified based on different data forms to obtain three data types: pictures, texts, and pictures and texts combined. Alternatively, the corresponding data types may be obtained by performing more detailed classification. For example, the image type may be further classified into subcategories such as faces, cartoon characters, advertisements, and buildings, so as to identify the illegal elements through the corresponding risk identification model.
[0097] The identification and processing device uses a corresponding risk identification model to perform violation identification processing on the data to be processed based on the determined data type.
[0098] For example, assuming that the data type may include pictures / text / a combination of pictures and text, different models are used to identify illegal elements. For picture type data, the recognition and processing device uses a model based on a detection or classification network structure to identify illegal elements contained in the picture. For text type data, a model based on a BERT structure is used to identify sensitive words in the text as illegal elements.
[0099] Continue to refer to Figure 3 To explain, the historical data acquisition device 102 acquires historical review data corresponding to the data to be processed.
[0100] The historical audit data is obtained based on the risk set to be predicted consisting of the one or more violation elements and / or the user who publishes the data to be processed.
[0101] The historical audit data includes audit result information within a predetermined period of time in the past. Optionally, the audit result information includes historical pass frequency information and historical rejection frequency information. The historical pass frequency information refers to the frequency of audit passes within a predetermined period of time in the past. The historical rejection frequency information refers to the frequency of audit rejections within a predetermined period of time in the past.
[0102] Optionally, the historical audit data corresponds to different business types. The historical data acquisition device 102 determines the business type information corresponding to the data to be processed, and then acquires the historical audit data corresponding to the data to be processed under the business type.
[0103] The method in which the historical data acquisition device 102 acquires the historical audit data corresponding to the data to be processed includes but is not limited to:
[0104] 1) obtaining corresponding historical audit data based on the risk set consisting of the one or more violation elements;
[0105] The risk set includes one or more violation elements corresponding to a piece of data. For N pieces of data, N or fewer risk sets may be generated because multiple pieces of data may share the same risk set. For example, if multiple pieces of data contain the same violation element, they will be included in the same risk set. In addition, there is no order distinction between elements in the risk set. For example, the risk sets represented by {'element1', 'element2'} and {'element2', 'element1'} will be regarded as the same risk set. If no violation element is identified, the risk set is empty.
[0106] The device of this embodiment further includes a historical risk storage device, and the historical data acquisition device 102 includes a risk set composition device and a risk set data acquisition device.
[0107] The historical risk storage device records and stores historical risk set data.
[0108] The historical risk set data includes historical audit data corresponding to multiple historical violation risk sets. The multiple historical violation risk sets are obtained based on the audit results of multiple data within a predetermined historical time period. The historical pass times and rejection times corresponding to each historical violation risk set are collected to obtain the historical audit data corresponding to the multiple historical violation risk sets.
[0109] The risk set forming device forms a risk set to be predicted corresponding to the data to be processed by using one or more violation elements identified from the data to be processed.
[0110] The risk set data acquisition device matches the pre-stored historical risk set data based on the risk set to be predicted to acquire the historical review data corresponding to the risk set to be predicted.
[0111] Specifically, the process of the risk set data acquisition device matching based on the risk set to be predicted includes: matching the risk set to be predicted with any historical violation risk set, comparing the number of violation elements in the risk set to be predicted with the number of violation elements in any historical violation risk set; if the number of violation elements in the risk set to be predicted is equal to the number of violation elements in any historical violation risk set, further checking whether the first violation element in the risk set to be predicted exists in any historical violation risk set; if the number of violation elements in the risk set to be predicted is not equal to the number of violation elements in any historical violation risk set, turning to the next historical violation risk set for matching; if the first violation element in the risk set to be predicted exists in any historical violation risk set, continuing to check one by one whether the remaining violation elements in the risk set to be predicted exist in any historical violation risk set; if any violation element in the risk set to be predicted does not exist in any historical violation risk set, turning to the next historical violation risk set for matching until the target historical violation risk set containing all the violation elements of the risk set to be predicted is found, and the historical audit data corresponding to the target historical violation risk set is used as the historical audit data corresponding to the risk set to be predicted.
[0112] This search and matching method can greatly improve the matching speed when searching for the same risk set in the historical review data. For example, assuming that the total number of historical risk sets is K, this method can reduce the matching time by at least K times.
[0113] 2) Obtain corresponding historical audit data based on the user who published the data to be processed;
[0114] According to one embodiment, the device of this embodiment further includes a user data storage device, and the historical data acquisition device 102 includes an identification acquisition device and a user data acquisition device.
[0115] The user data storage device records and stores the historical audit data of each user.
[0116] The identification acquisition device acquires the user identification information of the user who publishes the data to be processed.
[0117] The user identification information includes various information that can uniquely identify the user, such as user name, email address or phone number, etc.
[0118] The user data acquisition device acquires the historical audit data corresponding to the user based on the user identification information. Specifically, based on the user identification information, the stored historical audit data is searched to obtain the historical audit data matching the user identification information.
[0119] 3) Obtain corresponding historical audit data based on the identified illegal elements and the users who published the data to be processed;
[0120] Specifically, the historical data acquisition device 102 acquires historical audit data based on the identified illegal elements, acquires historical audit data based on the user who publishes the data to be processed, and merges the obtained data, thereby using the merged data as the historical audit data corresponding to the data to be processed.
[0121] According to one embodiment, the method further comprises a data updating device.
[0122] The data updating device periodically updates the historical audit data of the historical risk set data and / or the historical audit data of each user.
[0123] Optionally, the data updating device clears the stored data that is not within the predetermined time range each time the historical risk set data and / or the historical audit data of each user are updated. For example, only the historical audit data of the past 7 days is stored, and the data before 7 days is cleared.
[0124] Continue to refer to Figure 3 To explain, the violation prediction device 103 predicts violations of the number to be processed based on the historical audit data.
[0125] According to one embodiment, the violation prediction device 103 calculates the risk score corresponding to the data to be processed based on the historical audit data according to a predetermined score calculation rule, and determines whether the data to be processed violates the rules based on the calculated risk score.
[0126] Specifically, if the risk score exceeds the preset risk threshold, the data to be processed is judged to be in violation of the rules, and a "rejected" audit result is obtained; if the risk score does not exceed the preset risk threshold, the data to be processed is judged to be not in violation of the rules, and a "passed" audit result is obtained.
[0127] According to one embodiment, the risk score is calculated based on the risk set violation score and the user violation score.
[0128] According to one embodiment, if the historical audit data corresponding to the risk set to be predicted or the publishing user is not obtained, that is, the risk set to be predicted is not identified in the previous audit process and the data published by the corresponding user has not entered the audit process, the violation elements contained in the risk set to be predicted are audited and processed, and the historical audit data corresponding to the risk set to be predicted or the publishing user is stored based on the audit results. When the risk set to be predicted or the publishing user is successfully matched later, the historical audit data corresponding to the predicted risk set or the publishing user can be obtained from the stored historical audit data.
[0129] According to the device of the embodiment of the present application, by acquiring corresponding historical audit data based on a risk set composed of violation elements and / or the user who published the data to be processed, and making violation predictions based on the historical audit data, it is possible to comprehensively consider multiple violation elements and user behaviors, improve the accuracy and comprehensiveness of violating content identification, and improve the accuracy of violation risk prediction; the embodiment of the present application makes violation predictions based on historical audit data, and will only perform review processing when the risk set has not been identified in the previous review process and the data published by the corresponding user has not entered the review process, thereby improving processing speed and efficiency.
[0130] Based on the same inventive concept, an electronic device is also provided in an embodiment of the present application, and the method corresponding to the electronic device may be the violation prediction method in the aforementioned embodiment, and its principle of solving the problem is similar to that of the method. The electronic device provided in an embodiment of the present application includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the methods and / or technical solutions of the aforementioned multiple embodiments of the present application.
[0131] The electronic device may be a user device, or a device formed by integrating a user device and a network device through a network, or may be an application running on the above device. The user device includes but is not limited to various terminal devices such as computers, mobile phones, tablet computers, smart watches, and bracelets. The network device includes but is not limited to network hosts, single network servers, multiple network server sets, or cloud computing-based computer sets, which can be used to implement some processing functions when setting an alarm. Here, the cloud is composed of a large number of hosts or network servers based on cloud computing, where cloud computing is a type of distributed computing, a virtual computer composed of a group of loosely coupled computer sets.
[0132] Figure 4The structure of a device suitable for implementing the method and / or technical solution in the embodiment of the present application is shown, and the device 1200 includes a central processing unit (CPU, Central Processing Unit) 1201, which can perform various appropriate actions and processes according to the program stored in the read-only memory (ROM, Read Only Memory) 1202 or the program loaded from the storage part 1208 to the random access memory (RAM, Random Access Memory) 1203. In RAM1203, various programs and data required for system operation are also stored. CPU 1201, ROM 1202 and RAM 1203 are connected to each other through bus 1204. Input / output (I / O, Input / Output) interface 1205 is also connected to bus 1204.
[0133] The following components are connected to the I / O interface 1205: an input section 1206 including a keyboard, a mouse, a touch screen, a microphone, an infrared sensor, etc.; an output section 1207 including a cathode ray tube (CRT), a liquid crystal display (LCD), an LED display, an OLED display, etc., and a speaker, etc.; a storage section 1208 including one or more computer-readable media such as a hard disk, an optical disk, a magnetic disk, a semiconductor memory, etc.; and a communication section 1209 including a network interface card such as a LAN (Local Area Network) card, a modem, etc. The communication section 1209 performs communication processing via a network such as the Internet.
[0134] In particular, the methods and / or embodiments in the embodiments of the present application may be implemented as computer software programs. For example, the embodiments disclosed in the present application include a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes program code for executing the method shown in the flowchart. When the computer program is executed by the central processing unit (CPU) 1201, the above functions defined in the method of the present application are executed.
[0135] Another embodiment of the present application further provides a computer-readable storage medium having computer program instructions stored thereon, wherein the computer program instructions can be executed by a processor to implement the methods and / or technical solutions of any one or more embodiments of the present application described above.
[0136] Specifically, the present embodiment may adopt any combination of one or more computer-readable media. The computer-readable medium may be a computer-readable signal medium or a computer-readable storage medium. The computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or device, or any combination thereof. More specific examples (non-exhaustive list) of computer-readable storage media include: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In this document, a computer-readable storage medium may be any tangible medium containing or storing a program that may be used by or in combination with an instruction execution system, device, or device.
[0137] Computer readable signal media may include a data signal propagated in baseband or as part of a carrier wave, which carries a computer readable program code. Such propagated data signals may take a variety of forms, including, but not limited to, electromagnetic signals, optical signals, or any suitable combination of the above. Computer readable signal media may also be any computer readable medium other than a computer readable storage medium, which may send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device.
[0138] Program code embodied on a computer readable medium may be transmitted using any appropriate medium, including but not limited to wireless, wireline, optical fiber cable, RF, etc., or any suitable combination of the foregoing.
[0139] Computer program code for performing the operations of the present application may be written in one or more programming languages or a combination thereof, including object-oriented programming languages, such as Java, Smalltalk, C++, and conventional procedural programming languages, such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a separate software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0140] The flow chart or block diagram in the accompanying drawings shows the possible architecture, function and operation of the equipment, method and computer program product according to various embodiments of the present application. In this regard, each square box in the flow chart or block diagram can represent a module, a program segment or a part of a code, and the module, the program segment or a part of the code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some implementations as replacements, the functions marked in the square box can also occur in a sequence different from that marked in the accompanying drawings. For example, two square boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each square box in the block diagram and / or flow chart, and the combination of the square boxes in the block diagram and / or flow chart can be implemented with a dedicated system for hardware that performs a specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
[0141] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0142] In the several embodiments provided in the present application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation. For example, multiple units or page components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interfaces, devices or units, which can be electrical, mechanical or other forms.
[0143] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0144] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of hardware plus software functional units.
[0145] The above-mentioned integrated unit implemented in the form of a software functional unit can be stored in a computer-readable storage medium. The above-mentioned software functional unit is stored in a storage medium, including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) or a processor to perform some steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (Read-Only Memory, ROM), random access memory (Random Access Memory, RAM), disk or optical disk and other media that can store program codes.
[0146] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit it. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application.
[0147] In addition, it is clear that the word "comprising" does not exclude other units or steps, and the singular does not exclude the plural. Multiple units or devices stated in a device claim can also be implemented by one unit or device through software or hardware. The words first, second, etc. are used to indicate names, and do not indicate any particular order.
Claims
1. A violation prediction method, wherein: The method comprises: By performing violation identification processing on the data to be processed, one or more identified violation elements are obtained; Acquire historical audit data corresponding to the data to be processed, where the historical audit data is obtained based on the risk set to be predicted consisting of the one or more violation elements and / or the user who published the data to be processed; Based on the historical audit data, violation prediction is performed on the pending number.
2. The method according to claim 1, wherein: The method further comprises: Record and store historical risk set data, wherein the historical risk set data includes historical audit data corresponding to multiple historical violation risk sets; Wherein, the acquisition of historical audit data corresponding to the data to be processed includes: One or more violation elements identified from the data to be processed form a risk set to be predicted corresponding to the data to be processed; Based on the risk set to be predicted, matching is performed in pre-stored historical risk set data to obtain historical review data corresponding to the risk set to be predicted.
3. The method according to claim 2, wherein the matching in pre-stored historical risk set data based on the risk set comprises: Matching the risk set to be predicted with any historical violation risk set, and comparing the number of violation elements in the risk set to be predicted with the number of violation elements in any historical violation risk set; If the number of violation elements in the risk set to be predicted is equal to the number of violation elements in any of the historical violation risk sets, then further check whether the first violation element in the risk set to be predicted exists in any of the historical violation risk sets; if the number of violation elements in the risk set to be predicted is not equal to the number of violation elements in any of the historical violation risk sets, then turn to the next historical violation risk set for matching; If the first violation element in the risk set to be predicted exists in any of the historical violation risk sets, continue to check one by one whether the remaining violation elements in the risk set to be predicted exist in any of the historical violation risk sets; if any violation element in the risk set to be predicted does not exist in any of the historical violation risk sets, turn to the next historical violation risk set for matching until the target historical violation risk set containing all the violation elements of the risk set to be predicted is found, and the historical audit data corresponding to the target historical violation risk set is used as the historical audit data corresponding to the risk set to be predicted.
4. The method according to claim 1, wherein: The method further comprises: Record and store historical audit data of each user; Wherein, the acquisition of historical audit data corresponding to the data to be processed includes: Obtain user identification information of the user who published the data to be processed; Based on the user identification information, historical audit data corresponding to the user is obtained.
5. The method according to claim 1, wherein: The method adopts multiple models to identify illegal elements for data of different forms. The one or more illegal elements identified by performing illegal identification processing on the data to be processed include: Determine the data type corresponding to the number to be processed; Based on the determined data type, the corresponding risk identification model is used to identify violations of the data to be processed.
6. The method according to claim 1, wherein: The predicting of violations of the pending number based on the historical audit data includes: Based on the historical audit data, the risk score corresponding to the data to be processed is calculated according to a predetermined score calculation rule; The calculated risk score is used to determine whether the data to be processed is in violation of regulations.
7. The method according to any one of claims 1 to 6, wherein: The method further comprises: The historical audit data of the historical risk set data and / or the historical audit data of each user are updated regularly.
8. The method according to any one of claims 1 to 6, wherein: The method further comprises: If the historical audit data corresponding to the risk set to be predicted or the publishing user is not obtained, the violation elements included in the risk set to be predicted are audited, and the historical audit data corresponding to the risk set to be predicted or the publishing user is stored based on the audit result.
9. A device for violation prediction, wherein: The device comprises: A device for obtaining one or more identified violation elements by performing violation identification processing on the data to be processed; A device for obtaining historical audit data corresponding to the data to be processed, wherein the historical audit data is obtained based on the risk set to be predicted composed of the one or more violation elements and / or the user who published the data to be processed; A device for predicting violations of the pending number based on the historical audit data.
10. An electronic device, comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 8.
11. A computer readable medium having computer program instructions stored thereon, wherein the computer program instructions can be executed by a processor to implement the method according to any one of claims 1 to 8.
12. A computer program product, comprising a computer program, wherein when the computer program is executed by a processor, the method according to any one of claims 1 to 8 is implemented.