Industry classification and anomaly identification method and device, electronic equipment and storage medium
By continuously learning and using machine learning algorithms, combined with the experience of human experts, abnormal fluctuations in the outbound call numbers of micro and small enterprises can be identified, solving the problem of inaccurate identification in existing technologies and enabling timely early warning of malicious harassment and fraudulent calls.
Patent Information
- Application Number
- CN202211647545.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-21
- Publication Date
- 2026-02-10
- Estimated Expiration
- 2042-12-21
AI Technical Summary
Existing technologies are insufficient to effectively identify outbound call numbers of micro and small enterprises that deviate from the common and normal behavior of the industry they filled in when applying for the business, resulting in malicious harassment or illegal fraud calls. Existing models have problems with high false identification rate or low coverage in terms of screening conditions.
By using a continuous learning method based on call behavior, active enterprise numbers are screened, industry classifications and anomaly identification are performed, machine learning algorithms are used to continuously track and correct the normal behavioral characteristics of enterprises/numbers, and human expert experience is combined to monitor abnormal fluctuations in real time and issue early warnings.
It enables timely identification of enterprises/numbers, improves the timeliness and feasibility of risk prediction, reduces the number of unauthorized communications that slip through the net, and enhances the accuracy of anomaly identification.
Smart Images

Figure CN116233307B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of communications, and more specifically, to a method, apparatus, electronic device, and storage medium for industry classification and anomaly identification based on continuous learning of call behavior. Background Technology
[0002] In the process of providing outbound calling services to small and medium-sized enterprises (SMEs) and merchants, telecommunications operators have encountered situations where SMEs' outbound calling numbers are resold or stolen for illegal purposes. These calls may be used for harassment, fraud, or other malicious activities, and their data may be recorded in the operator's communication data. Such behavior often deviates from the standard practices reported by the applicant for the service, the expected practices for the services they provided, or shows significant fluctuations in their past calling behavior.
[0003] To address the potential issue of fraudulent and harassing calls during outbound calling services provided by telecom operators to small and medium-sized enterprises (SMEs), current mainstream solutions primarily rely on technical measures such as black-sample analysis of fraudulent / harassing calls to model and analyze such behavior. This approach extracts communication data records at a short-term time granularity (days / hours) for modeling, analysis, and prediction. Its drawback lies in the fact that short-term analysis models must accurately identify calls through strict screening criteria. Loose screening criteria lead to a high false positive rate, while strict criteria result in low coverage, making it difficult to prevent calls from numbers with high concealment or inconspicuous harassing or fraudulent behavior from slipping through the net.
[0004] With the development of technology, call behavior data analysis has been applied. For example, Chinese patent CN109274834B discloses a method for identifying express delivery numbers based on call behavior; Chinese patent CN112101046A discloses a method, device, and system for conversation analysis based on call behavior. When the outbound call numbers of micro and small enterprises deviate from the common and normal behaviors reported at the time of their business application, there is an urgent need to develop a method based on continuous learning of call behavior to achieve industry classification and abnormal fluctuation identification. This method can promptly identify abnormal fluctuations in the behavior of enterprises / numbers that differ from their previous normal behavior, thereby proactively detecting enterprises and numbers that may be exploited for unauthorized calls, tampering, or other illegal communications, and improving the timeliness and feasibility of risk prediction. Summary of the Invention
[0005] The technical problem this invention aims to solve is how to promptly identify and detect abnormal fluctuations in outbound call numbers that deviate from the common and normal behavior of the industry to which small and micro enterprises initially reported when applying for the business, resulting in calls being made for malicious harassment or illegal fraud.
[0006] To address the aforementioned technical problems, according to one aspect of the present invention, a method for industry classification and anomaly identification is provided. This method is based on call behavior and achieves industry classification and anomaly identification through continuous learning. The method includes the following steps: S1, sample screening: Selecting an industry to be identified as having anomalies, selecting multiple companies within the industry, and screening the registered phone numbers of these companies. The screened companies and their corresponding registered phone numbers must not have any reported records, and the numbers must be active for more than N months. Information on the industry and call behavior of the companies to which the numbers belong is collected. S2, continuous learning algorithm calculation: Performing normalized learning of communication behavior for the specified industry and companies, selecting white samples of industry / company objects from typical industries; applying machine learning algorithms to extract communication information records of the sample objects from the most recent 1 to N months, continuously tracking and training the communication characteristics of the industry / company samples, including the distribution of daily active call dates, active time periods, etc. Outbound / inbound call behavior characteristics, silent period distribution, etc.; combined with the experience of industry business experts, summarize the threshold range of significant characteristic indicators of normal corporate behavior, and output the normal behavioral habits of enterprises / industries / numbers; S3, learning result correction, select the number objects to be reviewed and corrected by random sampling, adopt multi-point correction, obtain the industry information and enterprise information of the number object to judge and identify its industry, and comprehensively correct whether the current industry and enterprise of the number are consistent with the continuous learning results. If there is a discrepancy, it is corrected based on the judgment results; based on the continuous learning and correction results, the confirmed credible numbers, enterprise objects and their normal behavioral characteristics information are stored in the database; S4, abnormal fluctuation detection, including industry abnormal behavior detection and enterprise abnormal behavior detection, continuously track and calculate the communication characteristics of the specified enterprise / enterprise number, and compare the deviation of the monitored object from the threshold of significant characteristic indicators of normal behavior of its enterprise or industry at regular intervals. Abnormal deviations are detected and identified in a timely manner for management and early warning.
[0007] According to an embodiment of the present invention, in step S1, the call behavior information may include 1 to N months of extracted call records, extracted access area records, and extracted SMS sending and receiving records of the registered phone number.
[0008] According to an embodiment of the present invention, in step S2, the machine learning algorithm adopts the idea of continuous learning, which can be to find a hyperplane to circle the positive samples in the sample, use this hyperplane to make decision prediction, and the samples inside the circle are the predicted target objects.
[0009] Furthermore, the continuous learning algorithm aims to find a hyperplane to enclose the positive samples in the sample set. This is achieved by setting the parameters of the generated hypersphere as its center o and corresponding hypersphere radius r > 0, and the hypersphere volume... (r) is minimized, and the center o is a linear combination of support vectors. Similar to the traditional SVM method, it can be required that the distance from all training data points x to the center is strictly less than r, where x = (x1, ..., x2) / (x3). m = (Call behavior characteristic factor set, industry / enterprise business attribute behavior characteristic factor set).
[0010] However, at the same time, a slack variable with a penalty coefficient of C is constructed. The optimization problem is as follows:
[0011]
[0012] After using Lagrange duality to solve, we can determine whether a new data point y is within the class. If the distance from y to the center is less than or equal to the radius... If it is outside the hypersphere, then it is not the target point.
[0013] According to an embodiment of the present invention, in step S3, multi-point correction may include the following steps: S31, obtaining industry information and enterprise information of the number object through manual telephone follow-up; S32, business experts analyze and identify the industry to which the number belongs by combining the number call record information and business attribute information; S33, the enterprise business manager confirms whether the current number behavior is consistent with the enterprise development service through self-check and review.
[0014] According to an embodiment of the present invention, step S4 may include the following steps:
[0015] S41. Continuously track and calculate the communication characteristics of a specified enterprise / enterprise number, including the latest call records, latest accessed regions records, and latest SMS sending and receiving records, and perform the latest behavior calculation;
[0016] S42. Behavioral comparison anomaly detection: compare the deviation of the monitored object from the threshold of significant characteristics of the normal behavior of its enterprise or industry.
[0017] S43. Abnormal reporting: timely detection and reporting of abnormal deviations, used for management and early warning.
[0018] According to a second aspect of the present invention, an apparatus for industry classification and anomaly identification is provided, comprising:
[0019] The sample screening module selects a specific industry to identify potential anomalies. Within this industry, it selects multiple companies and filters their registered phone numbers. The selected companies and their corresponding registered phone numbers must not have any reported activity, and the numbers must have been active for at least N months. The module collects industry information and call behavior data from the companies owning the numbers. The continuous learning algorithm module performs routine communication behavior training for specific industries and companies, selecting white samples from typical industries / companies. It applies machine learning algorithms to extract communication records from the sample objects over the past 1 to N months, continuously tracking and training the communication characteristics of the industry / company samples, including the distribution of daily active call dates, active time periods, outbound / inbound call behavior characteristics, and quiet time periods. Finally, it combines the experience of industry business experts to summarize significant typical corporate behaviors. The system includes a feature index threshold range, which outputs the normal behavioral habits of enterprises / industries / numbers; a learning result correction module, which selects numbers to be reviewed and corrected through random sampling, uses multi-point correction, obtains industry and enterprise information of the number to determine its industry affiliation, and comprehensively corrects whether the current number's industry and enterprise are consistent with the continuous learning results, correcting any inconsistencies based on the assessment results; and stores confirmed trustworthy numbers, enterprises, and their normal behavioral characteristics in a database based on the continuous learning and correction results; and an abnormal fluctuation detection module, which has the functions of detecting abnormal behavior in the industry and enterprises, continuously tracks and calculates the communication characteristics of specified enterprises / enterprise numbers, periodically compares the deviation of the monitored object from the threshold of significant characteristics of normal behavior of its enterprise or industry, and promptly detects and identifies abnormal deviations for management and early warning.
[0020] According to an embodiment of the present invention, the machine learning algorithm of the continuous learning algorithm module can adopt the idea of continuous learning algorithm to find a hyperplane to circle the positive samples in the sample, use this hyperplane to make decision prediction, and the samples inside the circle are the predicted target objects. By setting the parameters of the generated hypersphere as center o and the corresponding hypersphere radius r>0, the hypersphere volume is... (r) is minimized, and the center o is a linear combination of support vectors. Similar to the traditional SVM method, it can be required that the distance from all training data points x to the center is strictly less than r, where x = (x1, ..., x2) / (x3). m = (Call behavior characteristic factor set, industry / enterprise business attribute behavior characteristic factor set).
[0021] However, at the same time, a slack variable with a penalty coefficient of C is constructed. The optimization problem is as follows:
[0022]
[0023] After using Lagrange duality to solve, we can determine whether a new data point y is within the class. If the distance from y to the center is less than or equal to the radius... If it is outside the hypersphere, then it is not the target point.
[0024] According to a third aspect of the present invention, an electronic device is provided, comprising: a memory, a processor, and an industry classification and anomaly identification program stored in the memory and executable on the processor, wherein the industry classification and anomaly identification program, when executed by the processor, implements the steps of the above-described industry classification and anomaly identification method.
[0025] According to a fourth aspect of the present invention, a computer storage medium is provided, wherein the computer storage medium stores an industry classification and anomaly identification program, which, when executed by a processor, implements the steps of the above-described industry classification and anomaly identification method.
[0026] Compared with the prior art, the technical solution provided by the embodiments of the present invention can achieve at least the following beneficial effects:
[0027] This invention utilizes deep learning algorithms to proactively and continuously learn the routine calling behavior habits of small and medium-sized enterprises (SMEs), phone numbers, and their respective industries within the outbound calling services of telecom operators. This knowledge is then converted into personalized habitual behavior information for each enterprise and its phone number. Furthermore, the invention continuously tracks, monitors, and compares the latest communication actions to promptly identify abnormal fluctuations in the behavior of enterprises / phone numbers that deviate from their usual routines. This allows for the early detection of enterprises and phone numbers that may be exploited for unauthorized communication activities such as unauthorized calls or tampering.
[0028] This invention breaks with conventional methods in two ways. First, it changes short-term analysis to long-term information data calculation and analysis. Second, it changes the approach of using black numbers as samples for analysis and training to a normal, continuous learning approach based on industry and enterprise practices. This results in the continuous accumulation of industry objects, enterprise objects, number objects, and their reliable normal behavior knowledge, which is used to monitor abnormal fluctuations that deviate from normal behavior in a timely manner, identify problem objects promptly, and improve the timeliness and feasibility of risk prediction. Attached Figure Description
[0029] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings of the embodiments will be briefly described below. Obviously, the drawings described below only relate to some embodiments of the present invention and are not intended to limit the present invention.
[0030] Figure 1 This is a flowchart illustrating a method for industry classification and anomaly identification according to an embodiment of the present invention;
[0031] Figure 2 This is a schematic diagram illustrating a continuous learning algorithm according to an embodiment of the present invention;
[0032] Figure 3 This is a schematic diagram illustrating the application of the industry classification and anomaly identification method according to an embodiment of the present invention to an abnormal outbound call behavior monitoring and prevention system. Detailed Implementation
[0033] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the described embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0034] Unless otherwise defined, the technical or scientific terms used herein shall have the ordinary meaning as understood by one of ordinary skill in the art to which this invention pertains. The terms “first,” “second,” and similar terms used in the specification and claims of this patent application do not indicate any order, quantity, or importance, but are merely used to distinguish different components. Similarly, the terms “an” or “a” and similar terms do not indicate a limitation of quantity, but rather indicate the presence of at least one.
[0035] Figure 1 This is a flowchart illustrating a method for industry classification and anomaly identification according to an embodiment of the present invention.
[0036] The method for industry classification and anomaly detection is based on call behavior and achieves industry classification and anomaly detection through continuous learning. For example... Figure 1 As shown, the methods for industry classification and anomaly identification include the following steps:
[0037] S1. Sample screening: Select an industry to be identified for anomalies, select multiple companies within the industry, and screen the registered phone numbers of the companies. Among them, the screened companies and their corresponding registered phone numbers have no record of being reported, the numbers have been activated and have been continuously active for more than N months, and collect the industry information and call behavior information filled in by the companies to which the numbers belong.
[0038] S2. Continuous learning algorithm calculation: For specified industries and enterprises, normalized learning of communication behavior is carried out. White samples of industry / enterprise objects are selected from typical industries. Machine learning algorithms are applied to extract communication information records of sample objects from the most recent 1 to N months. The communication characteristics of industry / enterprise samples are continuously tracked and trained, including the distribution of daily active call dates, active time periods, outbound / inbound call behavior characteristics, and silent time period distribution. Combined with the experience of human experts in industry business direction, the threshold range of significant characteristic indicators of normal enterprise behavior is summarized, and the normal behavioral habits of enterprises / industries / numbers are output.
[0039] S3. Learning Result Correction: Select the number object to be reviewed and corrected by random sampling, and use multi-point correction to obtain the industry information and enterprise information of the number object to judge and identify its industry. The result is comprehensively corrected to see if the current industry and enterprise of the number are consistent with the continuous learning results. If there is any inconsistency, it is corrected based on the judgment result. Based on the continuous learning and correction results, the confirmed trustworthy numbers, enterprise objects and their normal behavior characteristics information are entered into the database.
[0040] S4. Abnormal fluctuation detection, including industry abnormal behavior detection and enterprise abnormal behavior detection, continuously tracks and calculates the communication characteristics of a specified enterprise / enterprise number, and periodically compares the deviation of the monitored object from the threshold of significant characteristics of the normal behavior of its enterprise or industry. Abnormal deviations are detected and identified in a timely manner for management and early warning.
[0041] This invention utilizes deep learning algorithms to proactively and continuously learn the routine calling behavior habits of small and medium-sized enterprises (SMEs), phone numbers, and their respective industries within the outbound calling services of telecom operators. This knowledge is then converted into personalized habitual behavior information for each enterprise and its phone number. Furthermore, the invention continuously tracks, monitors, and compares the latest communication actions to promptly identify abnormal fluctuations in the behavior of enterprises / phone numbers that deviate from their usual routines. This allows for the early detection of enterprises and phone numbers that may be exploited for unauthorized communication activities such as unauthorized calls or tampering.
[0042] According to one or more embodiments of the present invention, in step S1, the call behavior information includes retrieval of call records, retrieval of access area records, and retrieval of SMS sending and receiving records for 1 to N months of the registered phone number.
[0043] According to one or more embodiments of the present invention, in step S2, the machine learning algorithm adopts the idea of continuous learning to find a hyperplane to circle the positive samples in the sample, and uses this hyperplane to make decision prediction. The samples within the circle are the predicted target objects.
[0044] Figure 2 This is a schematic diagram illustrating a continuous learning algorithm according to an embodiment of the present invention.
[0045] like Figure 2 As shown, let the parameters of the generated hypersphere be the center o and the corresponding hypersphere radius r>0, and the hypersphere volume be... (r) is minimized, and the center o is a linear combination of support vectors. Similar to the traditional SVM method, it can be required that the distance from all training data points x to the center is strictly less than r.
[0046] Where x = (x1, ..., x) m = (Call behavior characteristic factor set, industry / enterprise business attribute behavior characteristic factor set).
[0047] However, at the same time, a slack variable with a penalty coefficient of C is constructed. The optimization problem is as follows:
[0048]
[0049] After using Lagrange duality to solve, we can determine whether a new data point y is within the class. If the distance from y to the center is less than or equal to the radius... If it is outside the hypersphere, then it is not the target point.
[0050] According to one or more embodiments of the present invention, in step S3, multi-point correction includes the following steps:
[0051] S31. Obtain industry and company information of the number holder through manual telephone follow-up;
[0052] S32. Business experts combine call record information and business attribute information to determine and identify the industry to which the number belongs.
[0053] S33. The business supervisor of the enterprise shall conduct a self-check and review to confirm whether the current number behavior is consistent with the enterprise development service.
[0054] According to one or more embodiments of the present invention, step S4 includes the following steps:
[0055] S41. Continuously track and calculate the communication characteristics of a specified enterprise / enterprise number, including the latest call records, latest accessed regions records, and latest SMS sending and receiving records, and perform the latest behavior calculation;
[0056] S42. Behavioral comparison anomaly detection: compare the deviation of the monitored object from the threshold of significant characteristics of the normal behavior of its enterprise or industry.
[0057] S43. Abnormal reporting: timely detection and reporting of abnormal deviations, used for management and early warning.
[0058] According to a second aspect of the present invention, an apparatus for industry classification and anomaly identification is provided, comprising: a sample screening module, a continuous learning algorithm module, a learning result correction module, and an anomaly fluctuation detection module.
[0059] The sample screening module is used to select a specific industry to identify whether there are any anomalies. Within the industry, multiple companies are selected, and the registered phone numbers of these companies are screened. Among these, the screened companies and their corresponding registered phone numbers have no record of being reported, and the numbers have been activated and continuously active for more than N months. The module also collects industry information and call behavior information filled in by the companies to which the numbers belong.
[0060] The continuous learning algorithm module is used to learn the normalized communication behavior of specified industries and enterprises. It selects white samples of industry / enterprise objects in typical industries; applies machine learning algorithms to extract the communication information records of sample objects from the most recent 1 to N months, and continuously tracks and trains the communication characteristics of industry / enterprise samples, including the distribution of daily active call dates, active time periods, outbound / inbound call behavior characteristics, and silent time period distribution; and combines the experience of human experts in industry business to summarize the threshold range of significant characteristic indicators of normal enterprise behavior, and outputs the normalized behavioral habits of enterprises / industries / numbers.
[0061] The learning result correction module is used to select numbers to be reviewed and corrected through random sampling, and to use multi-point correction to obtain industry information and enterprise information of the number to which it belongs, and to judge and identify its industry. The results are used to comprehensively correct whether the current industry and enterprise of the number are consistent with the continuous learning results. If there is any inconsistency, it is corrected based on the judgment results. Based on the continuous learning and correction results, the confirmed trustworthy numbers, enterprise objects and their normal behavioral characteristics information are stored in the database.
[0062] The abnormal fluctuation detection module has the functions of detecting abnormal behavior in the industry and abnormal behavior in enterprises. It is used to continuously track and calculate the communication characteristics of a specified enterprise / enterprise number, and periodically compare the deviation of the monitored object from the threshold of significant characteristics of normal behavior of its enterprise or industry. It can promptly detect and identify abnormal deviations for management and early warning.
[0063] According to one or more embodiments of the present invention, the machine learning algorithm of the continuous learning algorithm module adopts the idea of finding a hyperplane to circle the positive samples in the sample, using this hyperplane to make decision predictions, and the samples inside the circle are the predicted target objects.
[0064] By setting the generated hypersphere parameters as center o and corresponding hypersphere radius r>0, the hypersphere volume... (r) is minimized, and the center o is a linear combination of support vectors. Similar to the traditional SVM method, it can be required that the distance from all training data points x to the center is strictly less than r.
[0065] Where x = (x1, ..., x) m = (Call behavior characteristic factor set, industry / enterprise business attribute behavior characteristic factor set).
[0066] However, at the same time, a slack variable with a penalty coefficient of C is constructed. The optimization problem is as follows:
[0067]
[0068] After using Lagrange duality to solve, we can determine whether a new data point y is within the class. If the distance from y to the center is less than or equal to the radius... If it is outside the hypersphere, then it is not the target point.
[0069] This invention breaks with conventional methods in two ways. First, it changes short-term analysis to long-term information data calculation and analysis. Second, it changes the approach of using black numbers as samples for analysis and training to a normal, continuous learning approach based on industry and enterprise practices. This results in the continuous accumulation of industry objects, enterprise objects, number objects, and their reliable normal behavior knowledge, which is used to monitor abnormal fluctuations that deviate from normal behavior in a timely manner, identify problem objects promptly, and improve the timeliness and feasibility of risk prediction.
[0070] Figure 3 This is a schematic diagram illustrating the application of the industry classification and anomaly identification method according to an embodiment of the present invention to an abnormal outbound call behavior monitoring and prevention system.
[0071] like Figure 3 As shown, this patented method is applied to an abnormal outbound call behavior monitoring and prevention system. This system relies on the method of this invention to implement a normal learning module as a key function, while also realizing functions such as data information collection, multi-dimensional feature calculation, and a normal behavior feature knowledge base. This provides telecommunications operators with the ability to identify and monitor abnormal outbound calls. The system governance objectives are as follows:
[0072] Continuously accumulate industry, enterprise, and enterprise number whitelists, as well as industry, enterprise, and number feature databases; monitor enterprise entities that deviate from normal industry behavior fluctuations; monitor number entities that deviate from normal enterprise behavior.
[0073] According to another aspect of the present invention, an apparatus for industry classification and anomaly identification is provided, comprising: a memory, a processor, and an industry classification and anomaly identification program stored in the memory and executable on the processor, wherein the industry classification and anomaly identification program, when executed by the processor, implements the steps of the above-described industry classification and anomaly identification method.
[0074] The present invention also provides a computer storage medium.
[0075] The computer storage medium stores an industry classification and anomaly identification program, which, when executed by the processor, implements the steps of the aforementioned industry classification and anomaly identification method.
[0076] The method implemented when the industry classification and anomaly identification program running on the processor is executed can be referred to in various embodiments of the industry classification and anomaly identification method of the present invention, and will not be repeated here.
[0077] The present invention also provides a computer program product.
[0078] The computer program product of the present invention includes an industry classification and anomaly identification program, which, when executed by a processor, implements the steps of the industry classification and anomaly identification method as described above.
[0079] The method implemented when the industry classification and anomaly identification program running on the processor is executed can be referred to in various embodiments of the industry classification and anomaly identification method of the present invention, and will not be repeated here.
[0080] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) as described above, and includes several instructions to cause a terminal device (which may be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in the various embodiments of the present invention.
[0081] The above description is merely an exemplary embodiment of the present invention and is not intended to limit the scope of protection of the present invention, which is determined by the appended claims.
Claims
1. A method for industry classification and anomaly identification, wherein the method is based on call behavior and achieves industry classification and anomaly identification through continuous learning. The industry classification and anomaly identification method includes the following steps: S1. Sample screening: Select an industry to be identified for anomalies, select multiple companies within that industry, and then filter the registered phone numbers of these companies. The selected companies and their corresponding registered phone numbers have no record of being reported, the numbers have been activated and have been continuously active for more than N months, and the industry information and call behavior information of the companies to which the numbers belong are collected. S2. Continuous learning algorithm calculation, for the specified industries and enterprises, to perform normalized learning of communication behavior, and to select white samples of industry / enterprise objects in typical industries; Apply machine learning algorithms to extract communication information records of sample objects from the most recent 1 to N months, continuously track and train industry / enterprise sample communication characteristics, including daily call activity date distribution, active time period distribution, outbound / inbound call behavior characteristics, and silent time period distribution; combine industry business direction human expert experience to summarize the threshold range of significant characteristic indicators of normal enterprise behavior, and output the normal behavior habits of enterprises / industries / numbers; S3. Learning Result Correction: Select the number object to be reviewed and corrected by random sampling, and use multi-point correction to obtain the industry information and enterprise information of the number object to judge and identify its industry. The result is comprehensively corrected to see if the current industry and enterprise of the number are consistent with the continuous learning results. If there is any inconsistency, it is corrected based on the judgment result. Based on the continuous learning and correction results, the confirmed trustworthy numbers, enterprise objects and their normal behavior characteristics information are entered into the database. S4. Abnormal Fluctuation Detection: This includes industry-wide and enterprise-level abnormal behavior detection. It continuously tracks and calculates the communication characteristics of specified enterprise / enterprise numbers, periodically compares the monitored object with the threshold values of significant characteristics of normal behavior of its parent company or industry, and promptly detects and identifies abnormal deviations for management and early warning purposes. In step S2, the machine learning algorithm employs a continuous learning approach, which involves finding a hyperplane to enclose positive samples in the sample set. This hyperplane is then used for decision prediction, and the samples within the enclosed hyperplane are the predicted target objects. The continuous learning algorithm involves finding a hyperplane to enclose positive samples. This is achieved by defining the parameters of the generated hypersphere as its center o and corresponding radius r > 0, and its volume. (r) is minimized, and the center o is a linear combination of support vectors. Similar to the traditional SVM method, it can be required that the distance from all training data points x to the center is strictly less than r. Where x = (x1, ..., x) m = (Call behavior characteristic factor set, industry / enterprise business attribute behavior characteristic factor set). However, at the same time, a slack variable with a penalty coefficient of C is constructed. The optimization problem is as follows: After using Lagrange duality to solve, we can determine whether a new data point y is within the class. If the distance from y to the center is less than or equal to the radius... If it is outside the hypersphere, then it is not the target point.
2. The method as described in claim 1, wherein in step S1, the call behavior information includes retrieval of call records, retrieval of access area records, and retrieval of SMS sending and receiving records for 1 to N months of the registered phone number.
3. The method as described in claim 1, wherein step S3, the multi-point correction includes the following steps: S31. Obtain industry and company information of the number holder through manual telephone follow-up; S32. Business experts combine call record information and business attribute information to determine and identify the industry to which the number belongs. S33. The enterprise's business supervisor can conduct a self-check and review to confirm whether the current number's behavior is consistent with the enterprise's development services.
4. The method as described in claim 1, wherein step S4 comprises the following steps: S41. Continuously track and calculate the communication characteristics of a specified enterprise / enterprise number, including the latest call records, latest accessed regions records, and latest SMS sending and receiving records, and perform the latest behavior calculation; S42. Behavioral comparison anomaly detection: compare the deviation of the monitored object from the threshold of significant characteristics of the normal behavior of its enterprise or industry. S43. Abnormal reporting: timely detection and reporting of abnormal deviations, used for management and early warning.
5. An apparatus for industry classification and anomaly identification, comprising: The sample screening module is used to select a specific industry to be identified as having any anomalies. Within the industry, multiple companies are selected, and the registered phone numbers of these companies are screened. Among these, the screened companies and their corresponding registered phone numbers have no record of being reported, and the numbers have been activated and continuously active for more than N months. The module also collects industry information and call behavior information filled in by the companies to which the numbers belong. The continuous learning algorithm module is used to learn the normalized communication behavior of specified industries and enterprises. It selects white samples of industry / enterprise objects in typical industries; applies machine learning algorithms to extract the communication information records of sample objects from the most recent 1 to N months, and continuously tracks and trains the communication characteristics of industry / enterprise samples, including the distribution of daily call activity dates, active time periods, outbound / inbound call behavior characteristics, and silent time period distribution; and combines the experience of human experts in industry business to summarize the threshold range of significant characteristic indicators of normal enterprise behavior, and output the normalized behavior habits of enterprises / industries / numbers. The learning result correction module is used to select numbers to be reviewed and corrected through random sampling, and to perform multi-point correction. It obtains industry information and enterprise information of the number to which it belongs to, and analyzes and identifies the industry. The results are used to comprehensively correct whether the current industry and enterprise of the number are consistent with the continuous learning results. If there is any inconsistency, it is corrected based on the analysis results. Based on the continuous learning and correction results, the confirmed trustworthy numbers, enterprise objects and their normal behavioral characteristics information are stored in the database. The abnormal fluctuation detection module has the functions of detecting abnormal behavior in the industry and by enterprise. It is used to continuously track and calculate the communication characteristics of a specified enterprise / enterprise number, and periodically compare the deviation of the monitored object from the threshold of significant characteristics of normal behavior of its enterprise or industry. It can promptly detect and identify abnormal deviations for management and early warning. The machine learning algorithm in the continuous learning algorithm module adopts the idea of finding a hyperplane to enclose positive samples in the sample set, and using this hyperplane to make decision predictions. Samples within the enclosed hyperplane are the predicted target objects. By setting the generated hypersphere parameters as center o and corresponding hypersphere radius r>0, the hypersphere volume... (r) is minimized, and the center o is a linear combination of support vectors. Similar to the traditional SVM method, it can be required that the distance from all training data points x to the center is strictly less than r. Where x = (x1, ..., x) m = (Call behavior characteristic factor set, industry / enterprise business attribute behavior characteristic factor set). However, at the same time, a slack variable with a penalty coefficient of C is constructed. The optimization problem is as follows: After using Lagrange duality to solve, we can determine whether a new data point y is within the class. If the distance from y to the center is less than or equal to the radius... If it is outside the hypersphere, then it is not the target point.
6. An electronic device, comprising: The memory, the processor, and the industry classification and anomaly identification program stored on the memory and executable on the processor, wherein the industry classification and anomaly identification program, when executed by the processor, implements the steps of the industry classification and anomaly identification method as described in any one of claims 1 to 4.
7. A computer storage medium, wherein, The computer storage medium stores an industry classification and anomaly identification program, which, when executed by a processor, implements the steps of the industry classification and anomaly identification method as described in any one of claims 1 to 4.
Citation Information
Patent Citations
A method for identifying courier numbers based on call behavior
CN109274834B
Session analysis method, device and system based on call behavior
CN112101046A
Malicious registered enterprise behavior identification method and system based on isolated forest algorithm
CN112270553A
Abnormal enterprise identification method, device and equipment, medium and product
CN114997975A