Information identification method, system and device and storage medium
Through dynamic adjustment of feature engineering and output thresholds, an identification model is constructed, which solves the problem of time-consuming and labor-intensive construction of anti-fraud model and insufficient real-time response capabilities in the existing technology, and achieves efficient and accurate information identification and anti-fraud capabilities.
Patent Information
- Application Number
- CN202510263506.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-06
- Publication Date
- 2025-07-08
AI Technical Summary
In the prior art, manual rule modeling and machine learning training are time-consuming and labor-intensive in the construction of anti-fraud models, difficult to keep up with the speed of changes in new fraud methods, and lack the ability to respond quickly to real-time data.
By obtaining call information, using feature engineering to determine fraud feature indicators, initialize feature thresholds, and update feature thresholds based on output thresholds, build an identification model to achieve evaluation of target call information.
It improves the efficiency and accuracy of information identification, can automatically adapt to changes in fraud methods, reduces related costs, and improves the effectiveness of anti-fraud work.
Smart Images

Figure CN120278728A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technology, and in particular to an information recognition method, system, device and storage medium. Background Art
[0002] Currently, the substantial increase in the quantity of various information has made information recognition and information screening indispensable technical means to improve the efficiency of various types of work. Among them, the implementation of telecommunications anti-fraud work relies on information recognition, and accurate recognition of fraud numbers is very important for improving the effectiveness of anti-fraud work. The automatically generated anti-fraud model can not only improve the ability to detect and prevent fraud, but also effectively reduce related costs, providing strong support for maintaining network security and social stability.
[0003] In related technologies, the anti-fraud model is constructed by means of manual rule modeling and machine learning training. Manual rule modeling usually relies on manual rule formulation and expert experience. This method is not only time-consuming and laborious, but also difficult to keep up with the changing speed of new fraud means. Machine learning training often trains based on static data sets, and it is necessary to regularly select samples and data sets, lacking the ability to quickly respond to real-time data and the adaptive ability, resulting in poor performance when facing emerging threats. Summary of the Invention
[0004] An object of the present invention is to solve at least to some extent one of the technical problems existing in the prior art.
[0005] To this end, an object of the present invention is to provide an efficient information recognition method, system, device and storage medium.
[0006] In order to achieve the above technical object, on the one hand, an embodiment of the present invention provides an information recognition method, including the following steps: obtaining call information; determining a number of features based on the call information according to feature engineering, and initializing a feature threshold for each of the features; the feature engineering is used to characterize the feature indicators for identifying fraud; updating the feature threshold according to an output quantity threshold, and constructing an identification model according to the updated feature threshold; the output quantity threshold is the number of warning numbers expected to be obtained based on the identification model; evaluating the target call information according to the identification model to determine the target warning number. In the embodiment of the present application, features related to information recognition are determined through feature engineering, and an identification model is constructed according to the output quantity threshold for information recognition, which is beneficial to improving the efficiency and accuracy of information recognition.
[0007] In some embodiments, for the information recognition method of the embodiment of the present invention, the initializing the feature threshold for each of the features includes:
[0008] Traverse the work order numbers in the call information to determine the threshold range of each feature;
[0009] Initialize the feature threshold of each feature according to the threshold range.
[0010] In some embodiments, in one embodiment of the present invention, the updating the feature threshold according to the output volume threshold includes:
[0011] If the output volume of the current recognition model meets the output volume threshold, determine that the updated feature threshold is the current feature threshold;
[0012] Or, if the output volume of the current recognition model does not meet the output volume threshold and the output volume no longer decreases with the update of the feature threshold, determine that the updated feature threshold is the current feature threshold;
[0013] Or, if the output volume of the current recognition model does not meet the output volume threshold, update the feature threshold.
[0014] In some embodiments, in one embodiment of the present invention, the updating the feature threshold includes:
[0015] According to the recognition model, determine the feature output volume corresponding to each feature threshold to obtain an output volume set;
[0016] According to the output volume set, determine the first feature output volume with the smallest value and the first feature corresponding to the first feature output volume;
[0017] Update the first feature threshold corresponding to the first feature.
[0018] In some embodiments, in one embodiment of the present invention, the updating the feature threshold includes:
[0019] Based on the feature threshold before update, the number of warning numbers obtained through the recognition model is the first number;
[0020] Based on the feature threshold after update, the number of the warning numbers obtained through the recognition model is the second number; the second number is less than or equal to the first number.
[0021] In some embodiments, in one embodiment of the present invention, the output volume threshold is determined through the following steps:
[0022] Determine the output volume threshold according to the anti-fraud ranking of the operator;
[0023] Or, determine the output volume threshold according to the holiday characteristics.
[0024] In some embodiments, the call information is determined through the following steps:
[0025] Determine sample data;
[0026] Determine the call information according to the dispatch time in the sample data.
[0027] On the other hand, an embodiment of the present invention provides an information recognition system, including:
[0028] A first module for obtaining call information;
[0029] A second module for determining several features based on the call information according to feature engineering and initializing the feature threshold of each feature; the feature engineering is used to characterize the feature indicators for identifying fraud;
[0030] A third module for updating the feature threshold according to the output quantity threshold and constructing an identification model according to the updated feature threshold; the output quantity threshold is the number of warning numbers expected to be obtained based on the identification model;
[0031] A fourth module for evaluating the target call information according to the identification model and determining the target warning number.
[0032] On the other hand, an embodiment of the present invention provides an information recognition device, including:
[0033] At least one processor;
[0034] At least one memory for storing at least one program;
[0035] When the at least one program is executed by the at least one processor, the at least one processor implements the above information recognition method.
[0036] On the other hand, an embodiment of the present invention provides a storage medium, in which a processor-executable program is stored, and the processor-executable program is used to implement the above information recognition method when executed by a processor.
[0037] The embodiments of the present application at least include the following beneficial effects: The method provided by the embodiments of the present invention includes: obtaining call information; determining several features based on the call information according to feature engineering, and initializing the feature threshold of each feature; the feature engineering is used to characterize the feature indicators for identifying fraud; updating the feature threshold according to the output quantity threshold, and constructing an identification model according to the updated feature threshold; the output quantity threshold is the number of warning numbers expected to be obtained based on the identification model; evaluating the target call information according to the identification model to determine the target warning number. The embodiments of the present application determine the features related to information identification through feature engineering, and construct an identification model according to the output quantity threshold for information identification, which is beneficial to improving the efficiency and accuracy of information identification. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following introduces the accompanying drawings of the relevant technical solutions in the embodiments of the present invention or the prior art. It should be understood that the accompanying drawings in the following introduction are only for conveniently and clearly expressing some embodiments of the technical solutions in the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative efforts.
[0039] Figure 1 It is a schematic flowchart of an embodiment of the information identification method provided by the present invention;
[0040] Figure 2 It is a schematic flowchart of another embodiment of the information identification method provided by the present invention;
[0041] Figure 3 It is a schematic interface diagram of an embodiment of the sample determination process provided by the present invention;
[0042] Figure 4 It is a schematic interface diagram of an embodiment of the identification result provided by the present invention;
[0043] Figure 5 It is a schematic interface diagram of another embodiment of the identification result provided by the present invention;
[0044] Figure 6 It is a schematic structural diagram of an embodiment of the information identification system provided by the present invention;
[0045] Figure 7 It is a schematic structural diagram of an embodiment of the information identification device provided by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0046] Embodiments of the present invention will be described in detail below. Examples of the embodiments are shown in the accompanying drawings, where like or similar reference numerals denote like or similar elements or elements having like or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are only for explaining the present invention and should not be construed as limiting the present invention. For the step numbers in the following embodiments, they are only set for the convenience of elaboration and explanation, and no limitation is imposed on the order between steps. The execution order of each step in the embodiments can be adaptively adjusted according to the understanding of those skilled in the art.
[0047] In current society, fraud cases occur frequently, and traditional anti-fraud measures have problems of low efficiency and high cost. With the rapid development of Internet and mobile communication technologies, fraud behaviors have become increasingly concealed and complex. Traditional anti-fraud measures relying on manual analysis and empirical judgment are difficult to cope with the growing fraud risks. Therefore, it has become an urgent need to develop an anti-fraud model that is efficient, accurate, and can automatically adapt to constantly changing fraud methods. Such an automatically generated anti-fraud model can not only improve the ability to detect and prevent fraud, but also effectively reduce related costs, providing strong support for maintaining network security and social stability.
[0048] There are two ways for traditional technologies to handle fraud work order numbers, namely manual rule modeling and machine learning training. Manual rule modeling usually relies on manual rule formulation and expert experience. This method is not only time-consuming and laborious, but also difficult to keep up with the changing speed of new fraud methods. In addition, machine learning training often trains based on static data sets, requires regular sample and data set selection, lacks the ability to quickly respond to real-time data and adaptive ability, resulting in poor performance when facing emerging threats.
[0049] The information recognition method and system proposed according to the embodiments of the present invention will be described in detail below with reference to the accompanying drawings. First, the information recognition method proposed according to the embodiments of the present invention will be described with reference to the accompanying drawings.
[0050] Referring to Figure 1 , an information recognition method is provided in an embodiment of the present invention. The information recognition method in the embodiment of the present invention can be applied to a terminal, can also be applied to a server, or can be software running on a terminal or a server, etc. The terminal can be a tablet computer, a notebook computer, a desktop computer, etc., but is not limited thereto. The server can be an independent physical server, can also be a server cluster or a distributed system composed of multiple physical servers, or can be a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. The information recognition method in the embodiment of the present invention mainly includes the following steps:
[0051] S100: Obtain call information;
[0052] S200: Determine a number of features based on the call information according to feature engineering, and initialize a feature threshold for each of the features; the feature engineering is used to characterize the feature indicators for identifying fraud;
[0053] S300: Update the feature threshold according to the output quantity threshold, and construct an identification model according to the updated feature threshold; the output quantity threshold is the number of warning numbers expected to be obtained based on the identification model;
[0054] S400: Evaluate the target call information according to the identification model to determine the target warning number.
[0055] In some possible implementation manners, the call information in this application is the sample information for training, the feature threshold is the threshold for screening the call information, and the feature threshold adjusts the output quantity of the identification model to a certain extent.
[0056] Optionally, in an embodiment of the present invention, the initializing the feature threshold for each of the features includes:
[0057] Traverse the work order numbers in the call information to determine the threshold range for each of the features;
[0058] Initialize the feature threshold for each of the features according to the threshold range.
[0059] In some possible implementation manners, the initialization of the feature threshold is such that the output quantity of the identification model is the largest, that is, so that all warning numbers are identified.
[0060] Optionally, in an embodiment of the present invention, the updating the feature threshold according to the output quantity threshold includes:
[0061] If the output quantity of the current identification model meets the output quantity threshold, determine that the updated feature threshold is the current feature threshold;
[0062] Or, if the output quantity of the current identification model does not meet the output quantity threshold, and the output quantity no longer decreases as the feature threshold is updated, determine that the updated feature threshold is the current feature threshold;
[0063] Or, if the output quantity of the current identification model does not meet the output quantity threshold, update the feature threshold.
[0064] Optionally, in an embodiment of the present invention, the updating the feature threshold includes:
[0065] According to the recognition model, determine the feature output quantity corresponding to each feature threshold to obtain an output quantity set;
[0066] According to the output quantity set, determine the first feature output quantity with the smallest value and the first feature corresponding to the first feature output quantity;
[0067] Update the first feature threshold corresponding to the first feature.
[0068] Optionally, in an embodiment of the present invention, the updating the feature threshold includes:
[0069] Based on the feature threshold before update, the number of warning numbers obtained through the recognition model is the first number;
[0070] Based on the feature threshold after update, the number of warning numbers obtained through the recognition model is the second number; the second number is less than or equal to the first number.
[0071] In some possible implementation manners, the output quantity corresponding to the updated first feature threshold is less than the output quantity corresponding to the first feature threshold before update.
[0072] Optionally, in an embodiment of the present invention, the output quantity threshold is determined through the following steps:
[0073] Determine the output quantity threshold according to the anti-fraud ranking of the operator;
[0074] Or determine the output quantity threshold according to the holiday characteristics.
[0075] Optionally, in an embodiment of the present invention, the call information is determined through the following steps:
[0076] Determine sample data;
[0077] Determine the call information according to the dispatch time in the sample data.
[0078] Next, with reference to Figure 2 shown below, a specific embodiment is used to introduce the information recognition method of the present application in detail:
[0079] Step S001, data acquisition: The data that needs to be acquired in this proposal is the call signaling bill data of the operator, and the format is as
[0080] the call signaling and bill data shown in Table 1:
[0081]
[0082] Table 1
[0083] Step S002, feature engineering:
[0084] Feature engineering determines the upper limit of the model, and the rule algorithm only approaches this upper limit infinitely. Therefore, the construction of feature engineering is extremely crucial. The following are hundreds of fraud feature indicators summarized based on the current fraud behavior characteristics, which can be calculated based on the data obtained in step S001 and can strongly support the construction of the fraud model. Table 2 shows an embodiment of feature engineering.
[0085]
[0086]
[0087]
[0088]
[0089]
[0090]
[0091]
[0092] Table 2
[0093] Step S003, sample selection: The user customizes and selects the sample data to be trained according to the sample dispatch time, such as the sample data table shown in Table 3. In some embodiments, as Figure 3 shown, the time node of the sample data is set through the visualization interface, and the sample data is determined according to the set time node, that is, the call information in this application.
[0094]
[0095] Table 3
[0096] Step S004, algorithm implementation:
[0097] The model algorithm serves the business of telecommunications operators. The automated anti-fraud model is to meet the assessment of the Ministry of Industry and Information Technology. Currently, among communication operators, they are ranked according to the case rate of work orders to promote major operators to do a good job in anti-fraud. Therefore, the strategy of the automated model needs to support the dynamic balance of tightness according to the ranking. For example, when the anti-fraud index ranking is poor, the model needs to be able to shut down more case-related work orders, even if it causes a large number of misidentifications. When the anti-fraud index ranking is good, the model needs to balance complaints and anti-fraud rankings. At this time, it can shut down fewer numbers and be able to shut down some of the case-related work orders. Even if a small number of case-related work order numbers are not identified, it will not affect the overall anti-fraud ranking. Only in this way can the anti-fraud work and customer satisfaction be balanced.
[0098] First, clarify the anti-fraud assessment indicators. The current assessment indicator is mainly the case involvement rate = the number of cases involved / the number of operator numbers in the province. The assessment method is ranking. That is, if the case involvement rate indicator of the relevant department ranks in the top 30% among multiple telecom operators, it is excellent; the middle 30% - 70% is good; and the bottom 30% is poor.
[0099] The implementation of the algorithm is closely related to the anti-fraud ranking. Because the recall rate and output volume of the model always maintain a dynamic balance. Covering more case-related work order numbers often means a large output volume. A small output volume often means that only a small part of the case-related work order numbers can be covered.
[0100] Combined with the actual business, the implementation steps of this application are as follows:
[0101] 1. According to the current anti-fraud work of telecom operators, the severity level of the anti-fraud situation can be determined based on the current anti-fraud ranking, which is divided into excellent (top 30%), good (ranking in the middle 30% - 70%), and poor (ranking in the bottom 30%). Different rankings adopt different strategies. Of course, more generally, the strategy can be adjusted according to the current festival attributes (such as holidays, when people's awareness of prevention is weak, so the intensity is strengthened. Because there will be abnormal behaviors such as returning home for migration and increased communication with relatives and friends, in order to reduce the misjudgment of normal users, the intensity is appropriately reduced at this time), the entrance examination season and other attributes.
[0102] 2. Determine the output volume threshold of the model:
[0103] When the ranking is poor, the goal is to cover more case-related work order numbers. Even if there are more misjudgments in the numbers output by the model, it doesn't matter. At this time, select the maximum daily shutdown volume that the model can bear. According to the number shutdown experience of relevant operators, 0.015% of the total number of users of the telecom operator is the maximum shutdown volume that the model can bear. That is, assuming a provincial telecom operator with a user volume of 20 million, the maximum shutdown volume that the model can bear per day is 3000 numbers.
[0104] When the ranking is excellent, the goal is that the model needs to balance complaints and anti-fraud ranking. At this time, some numbers can be shut down less, and some of the case-related work orders can be shut down. Even if a small part of the case-related work order numbers are not recognized, it will not affect the overall anti-fraud ranking. At this time, according to the number shutdown experience of relevant operators, 0.005% of the total number of users of the telecom operator is the recommended shutdown volume that the model can bear. That is, assuming a telecom operator with a user volume of 20 million, the maximum recommended shutdown volume that the model can bear per day is 1000 numbers.
[0105] When the ranking is "good", according to the experience of shutting down numbers by provincial operators, 0.010% of the total number of users of telecom operators across the province is the recommended shutdown volume of the model that can be tolerated. That is, assuming a provincial telecom operator with a user base of 20 million, the maximum number of numbers that the model can recommend for shutdown in a single day is 2,000 numbers.
[0106] 3. First-round hit calculation and threshold update: In the first round, according to the feature thresholds in the feature engineering of step S002, combined with logical operators (including but not limited to greater than, less than, equal to, greater than or equal to, less than or equal to, AND, YES, NO, etc.), traverse and analyze the work order numbers to obtain the threshold range of each feature (for example, the range of the cumulative call volume on the same day for all fraud numbers in December 2024 is 80 - 100 times, the range of the cumulative number of called parties on the same day is 68 - 95 times, and the range of the cumulative number of cities where the called parties are located on the same day is 6 - 10 times), and obtain the initialization rule, with the goal that each input work order number must be hit. Calculate the hit situation under each feature threshold, select the feature with the smallest output volume for threshold update, and record the changes before and after the update. This process will be repeated until the threshold of the output volume is met, or in the case where the output volume is not met, the output volume no longer decreases.
[0107] 4. Dynamically adjust the number of work order numbers that are hit: When the threshold of the output volume is not met and the output volume no longer decreases, gradually reduce the number of work order numbers that are hit, and continue to execute the previous step until the threshold of the output volume is met.
[0108] 5. Model evaluation and result output: According to the above steps, a rule model is finally obtained, and the evaluation result and warning numbers are output. As Figure 4 shown. It is also possible to display each warning number through a visualization interface, as shown in Figure 5 shown.
[0109] The following takes an embodiment as an example for detailed description:
[0110] Currently, the selected type of work order related to the case is high-frequency outbound fraud, with 3 fraud numbers and another 4 normal numbers. In the first round, it is required that each input work order number must be hit. Referring to Table 4, the system first extracts the features of the cumulative call volume on the same day, the cumulative number of called parties on the same day, and the cumulative number of cities where the called parties are located on the same day for the 3 fraud numbers, and then obtains the minimum values of these three features, which are 80, 68, and 6 respectively; at this time, the initialization rule is that when the cumulative call volume >= 80 and the cumulative number of called parties >= 68 and the cumulative number of cities where the called parties are located >= 6, all fraud numbers are hit. At this time, the output number of numbers is 4.
[0111] Assume that at this time, the model output thresholds corresponding to excellent, good, and poor anti-fraud rankings are calculated to be 2, 3, and 4 respectively according to the above method.
[0112] If the anti-fraud ranking is poor, the maximum threshold output by the model is 4. At this time, the initialization rules can just be satisfied, and the model rules are generated.
[0113] If the anti-fraud ranking is good, the maximum threshold output by the model is 3. At this time, the initialization rules do not meet the requirements. The next step is to select the feature with the smallest output volume for update. At this time, the output volumes of the cumulative call volume on the same day >= 80, the cumulative number of called parties on the same day >= 68, and the cumulative number of called party home cities on the same day >= 6 in the initialization rules are 7, 6, and 5 respectively. Then the cumulative number of called party home cities on the same day is the feature with the smallest output volume. The system selects to update its threshold and traverses downward (that is, the feature of "cumulative number of called party home cities on the same day" is incremented by 1 in turn). When the rules are adjusted to the cumulative call volume on the same day >= 80, the cumulative number of called parties on the same day >= 68, and the cumulative number of called party home cities on the same day >= 7, the output volume is 2, which meets the model output threshold, but can only cover 2 of the 3 involved work order numbers. That is, at this time, it is no longer possible to hit the fraud number E, and the situation where the output volume cannot decrease occurs while ensuring that all fraud numbers are covered. At this time, step 4, dynamically adjusting the number of hit work order numbers, is triggered: when the threshold of the output volume is not met and the output volume no longer decreases, gradually reduce the number of hit work order numbers, and continue to execute the second step until the threshold of the output volume is met.
[0114] That is, the number of hit work order numbers is reduced from 3 to 2. At the same time, the program records the fraud number E that was not hit in the previous step and uses it as the work order number to be preferentially excluded. Then, according to the method in step 3, when the cumulative call volume on the same day >= 80, the cumulative number of called parties on the same day >= 68, and the cumulative number of called party home cities on the same day >= 7, the model rules with an output volume of 2 can be obtained. At this time, it meets the requirement that the maximum model output threshold is 3 and can cover 2 involved work orders. At this time, the traversal ends.
[0115] If the anti-fraud ranking is excellent, the maximum threshold output by the model is 2. At this time, the output quantity of the initialization rules is 4 and does not meet the requirements. According to the algorithm, select the feature with the smallest output volume for update. At this time, it is the same as the previous scenario where the "anti-fraud ranking is good", which will not be elaborated here. The final output number is 2, and the traversal ends.
[0116]
[0117] Table 4
[0118] Step S005, model generation: After meeting the high-efficiency evaluation, the model is automatically generated, and the system automatically generates the numbers that meet the model output.
[0119] The method provided by this application has stronger interpretability. The existing automated anti-fraud models mainly rely on machine learning algorithm models for updates, without clear model rules, resulting in non-interpretability. After the model is shut down and users file complaints, the customer service staff of telecom operators cannot explain to users the reasons for the shutdown, leading to passivity in work.
[0120] The method provided by this application has a more perfect feature engineering. This solution deeply explores the communication behavior characteristics of users and constructs hundreds of feature engineering indicators, which can support the accuracy of the automated model and the fraud coverage rate.
[0121] The method provided by this application has the ability of differentiated and refined governance. According to the different anti-fraud assessment pressures and anti-fraud situations faced by telecom operators, this solution carefully designs reasonable automated model strategies, which can differentially respond to fraud numbers in different periods and achieve a balance between anti-fraud work and user perception.
[0122] In summary, the method provided by the embodiments of this application includes: obtaining call information; determining a number of features based on the call information according to feature engineering and initializing the feature threshold of each feature; the feature engineering is used to represent the feature indicators for identifying fraud; updating the feature threshold according to the output quantity threshold, and constructing an identification model according to the updated feature threshold; the output quantity threshold is the number of warning numbers expected to be obtained based on the identification model; evaluating the target call information according to the identification model to determine the target warning number. The embodiments of this application determine the features related to information identification through feature engineering and construct an identification model according to the output quantity threshold for information identification, which is beneficial to improving the efficiency and accuracy of information identification.
[0123] Secondly, refer to the attached Figure 6 Describe an information identification system according to an embodiment of the present invention.
[0124] Figure 6 It is a schematic structural diagram of an information identification system according to an embodiment of the present invention. The system specifically includes:
[0125] A first module 610, configured to obtain call information;
[0126] A second module 620, configured to determine a number of features based on the call information according to feature engineering and initialize the feature threshold of each feature; the feature engineering is used to represent the feature indicators for identifying fraud;
[0127] A third module 630, configured to update the feature threshold according to the output quantity threshold and construct an identification model according to the updated feature threshold; the output quantity threshold is the number of warning numbers expected to be obtained based on the identification model;
[0128] The fourth module 640 is configured to evaluate the target call information according to the recognition model and determine the target warning number.
[0129] Optionally, in an embodiment of the present invention, the system further includes a fifth module configured to:
[0130] Determine the output quantity threshold according to the anti-fraud ranking of the operator;
[0131] Or determine the output quantity threshold according to the characteristics of holidays.
[0132] It can be seen that the content in the above method embodiments is applicable to the system embodiments. The functions specifically implemented by the system embodiments are the same as those in the above method embodiments, and the beneficial effects achieved are also the same as those in the above method embodiments.
[0133] Referring to Figure 7 , an embodiment of the present invention provides an information recognition device, including:
[0134] At least one processor 710;
[0135] At least one memory 720 for storing at least one program;
[0136] When the at least one program is executed by the at least one processor 710, the at least one processor 710 is caused to implement the information recognition method.
[0137] Similarly, the content in the above method embodiments is applicable to the device embodiments. The functions specifically implemented by the device embodiments are the same as those in the above method embodiments, and the beneficial effects achieved are also the same as those in the above method embodiments.
[0138] An embodiment of the present invention further provides a computer-readable storage medium, in which a program executable by a processor is stored, and the program executable by the processor is used to execute the above information recognition method when executed by the processor.
[0139] Similarly, the content in the above method embodiments is applicable to the storage medium embodiments. The functions specifically implemented by the storage medium embodiments are the same as those in the above method embodiments, and the beneficial effects achieved are also the same as those in the above method embodiments.
[0140] In some alternative embodiments, the functions / operations recited in the block diagrams may not occur in the order noted in the operational illustrations. For example, depending upon the functionality / operation involved, two blocks shown in succession may actually be executed substantially concurrently or the blocks may sometimes be executed in reverse order. Further, the embodiments presented and described in the flowcharts of the present invention are provided by way of example in order to provide a more thorough understanding of the technology. The disclosed methods are not limited to the operations and logic flows presented herein. Alternative embodiments are envisioned in which the order of various operations is altered and in which sub-operations described as part of a larger operation are performed independently.
[0141] In addition, although the present invention has been described in the context of functional modules, it should be understood that, unless otherwise stated to the contrary, one or more of the functions and / or features may be integrated in a single physical device and / or software module or one or more of the functions and / or features may be implemented in separate physical devices or software modules. It should also be understood that a detailed discussion of the actual implementation of each module is not necessary for an understanding of the present invention. Rather, given the attributes, functions, and internal relationships of the various functional modules in the devices disclosed herein, the actual implementation of the modules will be understood within the ordinary skill of an engineer. Thus, those of ordinary skill in the art will be able to implement the present invention as set forth in the claims without undue experimentation. It should also be understood that the particular concepts disclosed are illustrative only and are not intended to limit the scope of the present invention, which scope is determined by the full scope of the appended claims and their equivalents.
[0142] If the described functionality is implemented in the form of a software functional unit and sold or used as a stand-alone product, it may be stored in a computer-readable storage medium. Based on such understanding, the technical solution of the present invention, in essence or the part that contributes to the prior art or part of the technical solution, may be embodied in the form of a software product stored in a storage medium, including several programs for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a removable hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disc.
[0143] The logic and / or steps represented in the flowchart or otherwise described herein can, for example, be considered as a definitional sequence of executable instructions for implementing logical functions, which can be embodied in any computer-readable medium for use by or in connection with a program execution system, apparatus, or device, such as a computer-based system, a system including a processor, or other systems that can fetch and execute the instructions. As used in this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport the program for use by or in connection with the program execution system, apparatus, or device.
[0144] More specific examples of computer-readable media (a non-exhaustive list) include the following: an electrical connection having one or more wires (electronic device), a portable computer diskette (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber device, and a portable compact disc read-only memory (CDROM). Additionally, the computer-readable media can even be paper or other suitable media on which the program can be printed, as the program can be obtained electronically, for example, by optically scanning the paper or other media, followed by editing, interpretation, or other suitable processing as necessary, and then stored in a computer memory.
[0145] It should be understood that various parts of the present invention can be implemented in hardware, software, firmware, or a combination thereof. In the above-described embodiments, multiple steps or methods can be implemented in software or firmware stored in a memory and executed by a suitable program execution system. For example, if implemented in hardware, as in another embodiment, any one or a combination of the following techniques well known in the art can be used: discrete logic circuits having logic gates for implementing logical functions on data signals, application specific integrated circuits having appropriate combinational logic gates, programmable gate arrays (PGA), field programmable gate arrays (FPGA), and the like.
[0146] In the foregoing description of the present specification, the descriptions with reference to the terms "one embodiment / example", "another embodiment / example", or "certain embodiments / examples", etc., mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in any one or more embodiments or examples in a suitable manner.
[0147] Although embodiments of the present invention have been shown and described, those of ordinary skill in the art will understand that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the claims and their equivalents.
[0148] The above has specifically described the preferred embodiments of the present invention, but the present invention is not limited to the described embodiments. Those skilled in the art can also make various equivalent deformations or substitutions without violating the spirit of the present invention, and these equivalent deformations or substitutions are all included within the scope defined by the claims of the present invention.
Claims
1. An information recognition method, characterized in that, It includes the following steps: Obtain call information; According to feature engineering, determine several features based on the call information, and initialize the feature threshold for each of the features; The feature engineering is used to characterize the feature indicators for identifying fraud; According to the output quantity threshold, update the feature threshold, and build an identification model based on the updated feature threshold; the output quantity threshold is the number of warning numbers expected to be obtained based on the identification model; According to the identification model, evaluate the target call information to determine the target warning number.
2. The information recognition method according to claim 1, wherein The initialization of the feature threshold for each of the features includes: Traverse the work order numbers in the call information to determine the threshold range for each of the features; According to the threshold range, initialize the feature threshold for each of the features.
3. The information recognition method according to claim 1, wherein The update of the feature threshold according to the output quantity threshold includes: If the output quantity of the current identification model meets the output quantity threshold, determine that the updated feature threshold is the current feature threshold; Or, if the output quantity of the current identification model does not meet the output quantity threshold, and the output quantity no longer decreases with the update of the feature threshold, determine that the updated feature threshold is the current feature threshold; Or, if the output quantity of the current identification model does not meet the output quantity threshold, update the feature threshold.
4. The information recognition method according to claim 3, wherein The update of the feature threshold includes: According to the identification model, determine the feature output quantity corresponding to each feature threshold to obtain an output quantity set; According to the output quantity set, determine the first feature output quantity with the smallest value and the first feature corresponding to the first feature output quantity; Update the first feature threshold corresponding to the first feature.
5. The information recognition method according to claim 3, characterized in that The update of the feature threshold includes: The number of warning numbers obtained through the identification model based on the feature threshold before update is the first number; The number of warning numbers obtained through the identification model based on the updated feature threshold is the second number; the second number is less than or equal to the first number.
6. The information recognition method according to claim 1, characterized in that The output quantity threshold is determined through the following steps: Determine the output quantity threshold according to the anti-fraud ranking of the operator; Or, determine the output quantity threshold according to the characteristics of holidays.
7. The information recognition method according to claim 1, wherein The call information is determined through the following steps: Determine sample data; Determine the call information according to the dispatching time in the sample data.
8. An information recognition system, characterized in that, It includes: The first module is used to obtain call information; The second module is used to determine several features based on the call information according to feature engineering, and initialize the feature threshold for each of the features; the feature engineering is used to characterize the feature indicators for identifying fraud; The third module is used to update the feature threshold according to the output quantity threshold, and build an identification model based on the updated feature threshold; the output quantity threshold is the number of warning numbers expected to be obtained based on the identification model; The fourth module is used to evaluate the target call information according to the identification model to determine the target warning number.
9. An information recognition device, characterized in that, It includes: At least one processor; At least one memory for storing at least one program; When the at least one program is executed by the at least one processor, the at least one processor implements the information recognition method according to any one of claims 1 to 7.
10. A computer-readable storage medium storing a program executable by a processor, characterized in that, The program executable by the processor is used to implement the information recognition method according to any one of claims 1 to 7 when executed by the processor.