Method and apparatus for data auditing

By obtaining intention information and performing feature calculations, and using the data audit prediction model to identify abnormal user services, the problem of inability to identify abnormal ordering in the existing technology is solved, reducing operator losses, and improving system efficiency and interactivity.

CN113706184BActive Publication Date: 2025-08-05HUAWEI TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202010442005.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-05-22
Publication Date
2025-08-05
Estimated Expiration
2040-05-22

AI Technical Summary

Technical Problem

The existing business audit methods cannot identify abnormal ordering situations in advance, resulting in user churn and operator revenue loss, and cannot predict before adverse consequences occur.

Method used

By obtaining intention information, performing characteristic calculations, analyzing user services using data audit prediction models, identifying abnormal user services, and reducing operator losses.

Benefits of technology

It realizes the identification of abnormal user services before adverse consequences, reduces potential customer complaints and revenue loss, and improves the digital efficiency of the CRM system and the interactiveness of the audit system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113706184B_ABST
    Figure CN113706184B_ABST
Patent Text Reader

Abstract

This application provides a method and device for data auditing, which can analyze the user services of an operator, identify abnormal user services before adverse consequences occur, and reduce the possible losses to the operator. The method includes: obtaining intention information, which is used to query abnormal user services; and obtaining the data to be predicted of at least one user service corresponding to the intention information; a processing unit, which is used to perform characteristic calculation on the data to be predicted to obtain the first characteristic data of the at least one user service; and inputting the first characteristic data into a data auditing prediction model to determine whether the at least one user service is abnormal.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of communications, and more particularly, to a method and apparatus for data auditing. Background Art

[0002] Under the digital age, telecom operator systems are basically implemented through large-scale distributed clusters. Applications are generally constructed by different modules, and these modules may be developed and maintained by different teams. In the customer relationship management (CRM) hierarchical architecture system in the telecom operator field, taking the business enable system (BES) as an example, it has significant characteristics such as a large number of users, a large volume of services, complex business rules, and high requirements for system stability.

[0003] One current way of business auditing is as follows: For the discovered abnormal subscription situations, monitor by writing structured query language (SQL) scripts. Each SQL script can correspond to an abnormal scenario. These abnormal scenarios are usually summarized by business experts after analyzing subscription records after user churn or user complaints. Therefore, the monitoring mode of the above auditing method lags behind and cannot detect abnormal subscription situations that have not yet caused any consequences in advance, that is, it cannot identify abnormal scenarios that have not occurred in the past in advance, which may lead to user churn and revenue loss of the operator. Summary of the Invention

[0004] This application provides a method and apparatus for data auditing, which can analyze the user services of the operator and identify abnormal user services before adverse consequences occur, reducing the possible losses to the operator.

[0005] In a first aspect, a method for data auditing is provided, including: obtaining intent information for querying abnormal user services; obtaining prediction data of at least one user service corresponding to the intent information; performing feature calculation on the prediction data to obtain first feature data of the at least one user service; and inputting the first feature data into a data auditing prediction model to determine whether the at least one user service is abnormal.

[0006] In the embodiments of the present application, the server may obtain the to-be-predicted data of at least one user service corresponding to the intent information, perform feature calculation on the to-be-predicted data to obtain the first feature data of at least one user service, and then input the first feature data of the at least one user service into the data audit prediction model to determine whether the at least one user service is abnormal. This auditing method can analyze the user services of the operator, identify abnormal user services before adverse consequences occur, that is, it can monitor and prevent abnormal user services before they are discovered, without writing SQL scripts for monitoring, and can effectively and timely predict abnormal user services, avoiding phenomena such as overcharging and zero charging caused by abnormal user services, thereby reducing the possible losses of the operator, such as potential customer complaints, customer churn, and revenue loss of the operator, which is beneficial to improving the digital efficiency of the CRM system. In addition, the data auditing method in the embodiments of the present application can analyze the possible auditing intent according to the text input by the operation and maintenance personnel, improving the interactivity of the auditing system and the efficiency of the middle platform.

[0007] It should be understood that the above to-be-predicted data may be for one user service or multiple user services, and the embodiments of the present application do not limit this. Correspondingly, the first feature data obtained by performing feature calculation on the to-be-predicted data may correspond to one user service or multiple user services.

[0008] Combined with the first aspect, in some implementation manners of the first aspect, the to-be-predicted data includes: the first subscription information of the main products of the telecom operator, the first subscription information of the package services, and the first subscription information of the package services.

[0009] In combination with the first aspect, in some implementations of the first aspect, the first feature data corresponds to at least one user. The first feature data of the first user among the at least one user includes at least one of the following: the first sub-feature data of the first user, which is used to indicate whether the first user has subscribed to the main product, and the first sub-feature data of the first user is determined based on the first subscription information of the main product in the data to be predicted; the second sub-feature data of the first user, which is used to indicate whether the first user has subscribed to at least two main products, and the second sub-feature data of the first user is determined based on the first subscription information of the main product in the data to be predicted; the third sub-feature data of the first user, which is used to indicate whether there is an overlap in time for at least two main products subscribed by the first user, and the third sub-feature data of the first user is determined based on the first subscription information of the main product in the data to be predicted. The first subscription information of the main product in the data to be predicted includes the start time and expiration time of the main product subscribed by the first user; the fourth sub-feature data of the first user, which is used to indicate the number of mandatory packages subscribed by the first user, and the fourth sub-feature data of the first user is determined based on the first subscription information of the package service in the data to be predicted; the fifth sub-feature data of the first user, which is used to indicate the number of packages subscribed by the first user, and the fifth sub-feature data of the first user is determined based on the first subscription information of the package service in the data to be predicted.

[0010] It should be understood that the first feature data may include all or part of the above five sub-feature data. The more dimensions the first feature data has, the richer the representation is, and the more accurate the prediction will be.

[0011] In an embodiment of the present application, the server inputs the first feature data of the above at least one user into the data auditing prediction model, and a prediction result can be obtained. The prediction result indicates whether the user service of the at least one user is abnormal. Exemplarily, the prediction result can be represented by Y, where Y = 1 indicates abnormal and Y = 0 indicates normal.

[0012] In combination with the first aspect, in some implementations of the first aspect, the intention information includes a query intention and application programming interface (API) parameters. The API parameters include at least one of a query period and a query user. The data to be predicted is data corresponding to the query intention and the API parameters.

[0013] It should be understood that the data to be predicted for at least one user service corresponding to the above intention information is the data to be predicted for at least one user service corresponding to the above query intention and API parameters.

[0014] In an embodiment of the present application, the above query intention is to query abnormal subscriptions. In one possible implementation, if the API parameters include a query period, then the data to be predicted is the subscription data of at least one user service within the query period; in another possible implementation, the API parameters include a query user, then the data to be predicted is the subscription data of at least one user service corresponding to the query user; in yet another possible implementation, the API parameters include a query period and a query user, then the data to be predicted is the subscription data of at least one user service corresponding to the query user within the query period.

[0015] In one possible implementation, the operation and maintenance personnel can directly input a standard intention, that is, the above intention information is a standard intention, and the server can directly obtain the data to be predicted according to the intention information. In another possible implementation, the operation and maintenance personnel do not input a standard intention, and the server can perform intention recognition on the input text of the operation and maintenance personnel to obtain the standard intention.

[0016] Combined with the first aspect, in some implementations of the first aspect, the obtaining of the intention information includes: obtaining the input text of the operation and maintenance personnel; performing intention recognition on the input text to obtain the intention information, and the intention recognition is used to select a standard intention with the highest similarity to the intention of the input text from multiple standard intentions.

[0017] Optionally, after the operation and maintenance personnel input text on the client, the client or the server can perform all or part of the operations such as spelling check, text tokenization, extraction of word vectors, extraction of sentence vectors, detection of unknown intentions, classification of business intentions, and sorting of intention confidence levels to return the finally recognized standard intention. In addition, the client or the server can also perform associated intention recommendation through an associated intention map.

[0018] Exemplarily, the above intention recognition may include three steps: word embedding, sentence embedding, and intention classification. Among them, word embedding is used to extract word vectors, specifically: according to the context relationship of words in the corpus, obtain the word vectors in the input text, and use the word vectors of the central vocabulary to predict the word vectors of adjacent words. Sentence embedding is used to extract sentence vectors, specifically: sum all the word vectors in the sentences of the input text (or calculate the term frequency–inverse document frequency (TF-IDF)) to obtain the sentence vector of the input text. Intention classification means: according to the sentence vector, calculate the similarity between the input text and a certain standard intention (or called standard description), and the similarity model can use algorithms such as support vector machine (SVM), K-nearest neighbor (KNN) classification algorithm, or other algorithms for intention classification, which is not limited here.

[0019] Combined with the first aspect, in some implementation manners of the first aspect, before obtaining the intention information, the method further includes: performing feature calculation on the to-be-trained data to obtain second feature data, where the second feature data is used to represent user service characteristics; marking whether the user service corresponding to the second feature data is abnormal to obtain third feature data corresponding to the second feature data; and obtaining the data auditing prediction model according to the second feature data and the third feature data.

[0020] It should be understood that the data auditing prediction model is obtained by the server through pre-training on historical user service data. In the embodiments of the present application, the second feature data represents the data obtained by performing feature calculation on the to-be-trained data in the training stage, and the first feature data represents the data obtained by performing feature calculation on the to-be-predicted data in the prediction stage.

[0021] In the training stage, the server can mark the user service corresponding to the second feature data according to historical experience information (for example, abnormal scenarios feedback by users). For example, an abnormal mark is 1, and a normal mark is 0. In this embodiment, the marking result is recorded as the third feature data. The server can use the second feature data and the third feature data as inputs for training to obtain the data auditing prediction model. The above training can be supervised training, for example, decision tree training, fusion classification training based on algorithms such as SVM, logistic regression (LR), random forest (RF), etc., which is not limited in the embodiments of the present application.

[0022] In combination with the first aspect, in certain implementations of the first aspect, the data to be trained includes: the second subscription information of the main products of the telecommunications operator, the second subscription information of the package services, and the second subscription information of the package plans.

[0023] It should be understood that the data to be predicted and the data to be trained are two different batches of data, but the data types of the data to be predicted and the data to be trained are the same, and both include the second subscription information of the main products of the telecommunications operator, the second subscription information of the package services, and the second subscription information of the package plans.

[0024] In combination with the first aspect, in certain implementations of the first aspect, the second feature data corresponds to at least one user, and the second feature data of the second user among the at least one user includes at least one of the following: the first sub-feature data of the second user, which is used to indicate whether the second user has subscribed to the main product, and the first sub-feature data of the second user is determined based on the second subscription information of the main product in the data to be trained; the second sub-feature data of the second user, which is used to indicate whether the second user has subscribed to at least two main products, and the second sub-feature data of the second user is determined based on the second subscription information of the main product in the data to be trained; the third sub-feature data of the second user, which is used to indicate whether there is an overlap in time for at least two main products subscribed by the second user, and the third sub-feature data of the second user is determined based on the second subscription information of the main product in the data to be trained, and the second subscription information of the main product in the data to be trained includes the start time and expiration time of the main product subscribed by the second user; the fourth sub-feature data of the second user, which is used to indicate the number of mandatory packages subscribed by the second user, and the fourth sub-feature data of the second user is determined based on the second subscription information of the package services in the data to be trained; the fifth sub-feature data of the second user, which is used to indicate the number of package plans subscribed by the second user, and the fifth sub-feature data of the second user is determined based on the second subscription information of the package plans in the data to be trained.

[0025] The second feature data is similar to the first feature data. The difference is that the second feature data is extracted by performing feature calculation on the data to be trained. It should be understood that the second feature data may include all or part of the above five sub-feature data. The more dimensions the second feature data has, the richer the representation, and the more accurate the prediction will be.

[0026] In combination with the first aspect, in certain implementations of the first aspect, the method further includes: obtaining the annotation information of the operation and maintenance personnel for the prediction result, where the annotation information is used to indicate whether the prediction result is accurate; and retraining the data auditing prediction model based on the annotation information.

[0027] In the embodiments of the present application, the operation and maintenance personnel can label whether the prediction result is correct according to experience, and the server can retrain based on the labeling result of the operation and maintenance personnel, update the experience information in a timely manner, so as to obtain a data auditing prediction model with higher accuracy.

[0028] In a second aspect, another method for data auditing is provided, including: performing feature calculation on the data to be trained to obtain second feature data, where the second feature data is used to represent user service characteristics; marking whether the user service corresponding to the second feature data is abnormal to obtain third feature data corresponding to the second feature data; and obtaining a data auditing prediction model based on the second feature data and the third feature data.

[0029] In combination with the second aspect, in some implementation manners of the second aspect, the data to be trained includes: second subscription information of the main products of a telecommunications operator, second subscription information of package services, and second subscription information of package services.

[0030] In combination with the second aspect, in some implementation manners of the second aspect, the second feature data corresponds to at least one user, and the second feature data of the second user among the at least one user includes at least one of the following: first sub-feature data of the second user, used to indicate whether the second user has subscribed to the main product, and the first sub-feature data of the second user is determined based on the second subscription information of the main product in the data to be trained; second sub-feature data of the second user, used to indicate whether the second user has subscribed to at least two main products, and the second sub-feature data of the second user is determined based on the second subscription information of the main product in the data to be trained; third sub-feature data of the second user, used to indicate whether there is an overlap in time between at least two main products subscribed by the second user, and the third sub-feature data of the second user is determined based on the second subscription information of the main product in the data to be trained, and the second subscription information of the main product in the data to be trained includes the start time and expiration time of the main product subscribed by the second user; fourth sub-feature data of the second user, used to indicate the number of mandatory packages subscribed by the second user, and the fourth sub-feature data of the second user is determined based on the second subscription information of the package service in the data to be trained; fifth sub-feature data of the second user, used to indicate the number of packages subscribed by the second user, and the fifth sub-feature data of the second user is determined based on the second subscription information of the package service in the data to be trained.

[0031] In combination with the second aspect, in some implementation manners of the second aspect, the method further includes: obtaining the annotation information of the operation and maintenance personnel on the prediction result, where the annotation information is used to indicate whether the prediction result is accurate; and retraining the data auditing prediction model based on the annotation information.

[0032] In a third aspect, a data auditing device is provided for performing the methods in any possible implementation manner of the above aspects. Specifically, the device includes units for performing the methods in any possible implementation manner of the above aspects.

[0033] In a fourth aspect, another data auditing device is provided, including a processor coupled to a memory and capable of executing instructions in the memory to implement the methods in any possible implementation manner of the above aspects. Optionally, the device further includes a memory. Optionally, the device further includes a communication interface, and the processor is coupled to the communication interface.

[0034] In one implementation manner, the data auditing device is a client. When the data auditing device is a client, the communication interface can be a transceiver or an input / output interface.

[0035] In another implementation manner, the data auditing device is a chip configured in a client. When the data auditing device is a chip configured in a client, the communication interface can be an input / output interface.

[0036] In one implementation manner, the data auditing device is a server. When the data auditing device is a server, the communication interface can be a transceiver or an input / output interface.

[0037] In another implementation manner, the data auditing device is a chip configured in a server. When the data auditing device is a chip configured in a server, the communication interface can be an input / output interface.

[0038] In a fifth aspect, a processor is provided, including: an input circuit, an output circuit, and a processing circuit. The processing circuit is used to receive a signal through the input circuit and transmit the signal through the output circuit, so that the processor executes the methods in any possible implementation manner of the above aspects.

[0039] In a specific implementation process, the above processor can be a chip, the input circuit can be an input pin, the output circuit can be an output pin, and the processing circuit can be transistors, gate circuits, flip-flops, and various logic circuits, etc. The input signal received by the input circuit can be received and input by, for example, but not limited to, a receiver, and the signal output by the output circuit can be output to, for example, but not limited to, a transmitter and transmitted by the transmitter, and the input circuit and the output circuit can be the same circuit, which is used as the input circuit and the output circuit at different times. The embodiments of the present application do not limit the specific implementation manners of the processor and various circuits.

[0040] In a sixth aspect, a processing device is provided, including a processor and a memory. The processor is configured to read instructions stored in the memory, and can receive signals through a receiver and transmit signals through a transmitter to execute the methods in any possible implementation manner in the above aspects.

[0041] Optionally, there is one or more processors and one or more memories.

[0042] Optionally, the memory can be integrated with the processor, or the memory is separately provided from the processor.

[0043] In a specific implementation process, the memory can be a non-transitory memory, such as a read only memory (ROM). It can be integrated with the processor on the same chip, or can be separately provided on different chips. The embodiments of the present application do not limit the type of the memory and the setting manner of the memory and the processor.

[0044] It should be understood that relevant data interaction processes, such as sending indication information, can be a process of outputting indication information from the processor, and receiving capability information can be a process of the processor receiving input capability information. Specifically, the processed output data can be output to the transmitter, and the input data received by the processor can come from the receiver. Among them, the transmitter and the receiver can be collectively referred to as a transceiver.

[0045] The processing device in the above sixth aspect can be a chip. The processor can be implemented by hardware or by software. When implemented by hardware, the processor can be a logic circuit, an integrated circuit, etc.; when implemented by software, the processor can be a general-purpose processor, which is implemented by reading software code stored in the memory. The memory can be integrated in the processor or can exist independently outside the processor.

[0046] In a seventh aspect, a computer program product is provided. The computer program product includes: a computer program (which can also be referred to as code or instruction). When the computer program is run, it causes the computer to execute the methods in any possible implementation manner in the above aspects.

[0047] In an eighth aspect, a computer-readable storage medium is provided. The computer-readable storage medium stores a computer program (which can also be referred to as code or instruction). When it runs on a computer, it causes the computer to execute the methods in any possible implementation manner in the above aspects.

[0048] In a ninth aspect, a communication system is provided, including the device for data auditing described above. BRIEF DESCRIPTION OF THE DRAWINGS

[0049] Figure 1 It is a schematic diagram of the system architecture of an embodiment of the present application.

[0050] Figure 2 It is a schematic flowchart of the data auditing method of an embodiment of the present application.

[0051] Figure 3 It is a schematic flowchart of another data auditing method of an embodiment of the present application.

[0052] Figure 4 It is a schematic diagram of another data auditing method of an embodiment of the present application.

[0053] Figure 5 It is a schematic block diagram of the data auditing device of an embodiment of the present application.

[0054] Figure 6 It is a schematic block diagram of another data auditing device of an embodiment of the present application. Detailed implementation manners

[0055] Next, the technical solutions in the present application will be described in conjunction with the accompanying drawings.

[0056] Telecom operator systems in the digital age are basically implemented through large-scale distributed clusters. Applications are generally built by different modules, and these modules may be developed and maintained by different teams. In the customer relationship management (CRM) hierarchical architecture system in the telecom operator field, taking the business enable system (BES) as an example, it has significant characteristics such as a large number of users, a large volume of business, complex business rules, and high requirements for system stability.

[0057] The current industry status is as follows: 1) There are black production teams in the industry that use various means to find system vulnerabilities and profit by generating orders that do not conform to business rule constraints; 2) Auditing discovers various abnormal data that do not meet the requirements of business rules, but are actually successfully accepted and stored in the database; 3) It is impossible to confirm the foreground operation scenario, there is no relevant error log retained, and there is no effective positioning method; 4) The problems continue despite repeated prohibitions, leading to escalated complaints from users, resulting in user and revenue losses, and a poor perception from the customer's senior management.

[0058] Based on the above problems, the operation and maintenance personnel can use verification methods for business auditing. There are two types of verification: front-end verification and back-end verification. Among them, the verification scope of front-end verification is limited, and in cases such as network jitter, it may cause the front-end verification to be bypassed, resulting in orders that break the front-end verification rules. The back-end verification will verify all product change operations, which affects the order creation time, especially in the case of main product changes involving a large number of product change operations. At the same time, verification requires cumulative calculation of the product data ordered by the user this time and the user data already archived in the database to determine whether there is an abnormal package for the user, and the logic is relatively complex.

[0059] Currently, a common method for business auditing is as follows: for the discovered abnormal subscription situations, the operation and maintenance personnel can monitor by writing Structured Query Language (SQL) scripts. Each SQL script can correspond to an abnormal scenario. These abnormal scenarios are usually summarized by the operation and maintenance personnel (or called business experts) after analyzing the subscription records after user churn or user complaints. Therefore, the monitoring mode of the above auditing method lags behind and cannot detect abnormal subscription situations that have not yet caused any consequences in advance, that is, it cannot identify abnormal scenarios that have not occurred in the past in advance, which may lead to user churn and revenue loss for the operator.

[0060] In view of this, the present application proposes a method and device for data auditing, which can analyze the user services of the operator and identify abnormal user services before adverse consequences occur, reducing the possible losses for the operator.

[0061] The technical solution of the embodiment of the present application can be applied to operation data management systems with different industry backgrounds. Figure 1 FIG. 100 is a schematic diagram of the system architecture of the embodiment of the present application. The system architecture 100 may include a client 110, a first server 120, and a second server 130.

[0062] Among them, the client 110 can also be called the user side, corresponding to the server, and can provide local services for users. The server can provide computing or application services for the client. The first server 120 and the second server 130 in the system architecture 100 of the present application can include a file server, a database server, an application server, a website (WEB) server, etc., and the embodiment of the present application does not limit this.

[0063] Figure 1 Exemplarily, one client and two servers are shown. Optionally, the system 100 may further include other numbers of clients and servers, and the embodiment of the present application does not limit this.

[0064] In an embodiment of the present application, the client 110 is used to obtain the information input by the operation and maintenance personnel and send it to the first server 120 or the second server 130. The first server 120 is used for prediction, and the second server 130 is used for offline training.

[0065] In a possible implementation manner, the client 110, the first server 120, and the second server 130 are three devices respectively.

[0066] In another possible implementation manner, the client 110 and the first server 120 are the same device. For example, the client 110 may have the functions of the first server 120.

[0067] In yet another possible implementation manner, the client 110, the first server 120, and the second server 130 are the same device. For example, the client 110 may have the functions of the first server 120 and the second server 130.

[0068] The above system architecture 100 is only an exemplary illustration for easy understanding, and the present application does not limit the actual physical forms of the above client and server.

[0069] In an embodiment of the present application, multiple clients can independently interact with the server (such as the above first server 120 and the above second server 130) for information. Exemplarily, the first server 120 may transmit information with multiple clients at the same time to determine the intention information sent by each client among the multiple clients, and perform data auditing according to the intention information. Since the process of the first server 120 performing business auditing for each client is similar, for easy understanding and illustration, the embodiment of the present application takes the process of the first server 120 performing business auditing based on the intention information obtained from one client as an example for description.

[0070] In an embodiment of the present application, a client or a server includes a hardware layer, an operating system layer running on the hardware layer, and an application layer running on the operating system layer. The hardware layer includes hardware such as a central processing unit (CPU), a memory management unit (MMU), and a memory (also referred to as main memory). The operating system can be any one or more computer operating systems that implement service processing through processes. For example, Linux operating system, Unix operating system, Android operating system, iOS operating system, or Windows operating system, etc. The application layer includes applications such as a browser, an address book, a word processing software, an instant messaging software, etc. Moreover, the embodiment of the present application does not particularly limit the specific structure of the execution subject of the method provided by the embodiment of the present application. As long as it can communicate according to the method provided by the embodiment of the present application by running a program that records the code of the method provided by the embodiment of the present application. For example, the execution subject of the method provided by the embodiment of the present application can be a client or a server, or a functional module in the client or the server that can call and execute the program.

[0071] In addition, various aspects or features of the present application can be implemented as a method, a device, or an article of manufacture using standard programming and / or engineering techniques. The term "article of manufacture" used in the present application covers a computer program accessible from any computer-readable device, carrier, or medium. For example, computer-readable media can include, but are not limited to: magnetic storage devices (such as hard disks, floppy disks, or magnetic tapes, etc.), optical discs (such as compact discs (CDs), digital versatile discs (DVDs), etc.), smart cards, and flash memory devices (such as erasable programmable read-only memories (EPROMs), cards, sticks, or key drives, etc.). In addition, various storage media described herein can represent one or more devices and / or other machine-readable media for storing information. The term "machine-readable media" can include, but is not limited to, wireless channels and various other media that can store, contain, and / or carry instructions and / or data.

[0072] For ease of understanding, the main products, packages, and packages involved in the embodiments of the present application will be introduced first below.

[0073] The main product refers to the main brand ordered by the user from the operator. For example, if the operator is China Mobile, its corresponding main products can include GoTone, Easyown, and M-Zone.

[0074] A package refers to a collection of goods under a specific main product. Packages are divided into two categories: mandatory packages and optional packages. Among them, a mandatory package is a service that users must subscribe to, and there are no restrictions on optional packages for users. Users can subscribe or not subscribe. For example, if the main product subscribed by a user is M-Zone, under M-Zone, there may be a data package (such as the 4G data self-selection package) or a voice package.

[0075] A plan refers to a specific product under a specific package. For example, a data package can include a 5G data plan, a 10G data plan, a 20G data plan, etc.; a voice package can include a 100-minute voice plan, a 200-minute voice plan, a 300-minute voice plan, etc.

[0076] In summary, the main product, package, and plan have a hierarchical relationship. The main product is at the top layer, and the plan is at the bottom layer. Users can first subscribe to the main product, then subscribe to packages under that main product, and then subscribe to plans from the subscribed packages.

[0077] In addition, before introducing the method provided in the embodiments of the present application, the following points are explained first.

[0078] First, in the embodiments of the present application, "pre-obtain" can be achieved by pre-saving corresponding codes, tables or other means that can be used to indicate relevant information in a device (such as including a client and a server). The present application does not limit its specific implementation method.

[0079] Second, in the embodiments shown below, each term and English abbreviation, such as the main product, intention information, etc., are all exemplary examples given for the convenience of description and should not constitute any limitation to the present application. The present application does not exclude the possibility of defining other terms that can achieve the same or similar functions in existing or future protocols.

[0080] Third, in the embodiments shown below, the first, second, and various numerical numbers are only for the convenience of description and are not used to limit the scope of the embodiments of the present application. For example, to distinguish different users, distinguish different feature data, etc.

[0081] Fourth, "at least one" means one or more, "at least two" and "multiple" mean two or more. "And / or" describes the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone, where A and B can be singular or plural. The character " / " generally indicates that the associated objects before and after are in an "or" relationship. "At least one of the following" or its similar expressions refer to any combination of these items, including any combination of single item or plural items. For example, at least one of a, b, and c can represent: a, or b, or c, or a and b, or a and c, or b and c, or a, b, and c, where a, b, and c can be single or multiple.

[0082] The method for data auditing provided by the present application will be described in detail below with reference to the accompanying drawings. Figure 2 FIG. shows a schematic flowchart of the data auditing method 200 provided by an embodiment of the present application. This method can be applied to Figure 1 the communication system shown. The method 200 can be executed by the client 110, the first server 120, and the second server 130 in the above system 100, but the embodiments of the present application do not limit this. The method 200 specifically includes the following steps:

[0083] S210, the first server 120 obtains intention information from the client 110, and this intention information is used to query abnormal user services.

[0084] The above intention information indicates that the operation and maintenance personnel (or management personnel) need to query abnormal user services, and this intention information can specifically be the input text of the operation and maintenance personnel. Exemplarily, when performing data auditing, the operation and maintenance personnel can input intention information on the client 110, and the first server 120 can obtain this intention information, so as to obtain the data to be predicted corresponding to this intention information, perform feature calculation on the data to be predicted, and then perform prediction based on the calculated first feature data and the pre-obtained data auditing prediction model.

[0085] As an optional embodiment, the intention information includes a query intention and application program interface (API) parameters. The API parameters include at least one of a query period and a query user, and the data to be predicted is data corresponding to the query intention and the API parameters.

[0086] The above query intention can represent the query content of the operation and maintenance personnel. In the embodiments of the present application, the query intention is to query abnormal subscriptions. The above API parameters may include a query period, or include a query user, or include a query period and a query user. Exemplarily, the query period may be within the most recent week, or within the most recent year, or a period input by the operation and maintenance personnel themselves (such as from February 1, 2019 to May 15, 2019). The query user may be a mobile phone number to be queried, or an ID number, etc. The embodiments of the present application do not limit this. The number of query periods and / or query users may be one or multiple, depending on the operation and maintenance personnel, and the embodiments of the present application do not limit this either.

[0087] In a possible implementation manner, the operation and maintenance personnel can directly input a standard intention on the client 110, that is, the above intention information is a standard intention, and the first server 120 can determine the query intention of the operation and maintenance personnel. In another possible implementation manner, the operation and maintenance personnel do not input a standard intention on the client 110. The client 110 can perform intention recognition on the input text of the operation and maintenance personnel to obtain a standard intention; or the client 110 sends the input text of the operation and maintenance personnel to the first server 120, and the first server 120 performs intention recognition on the input text.

[0088] As an optional embodiment, the obtaining of the intention information includes: obtaining the input text of the operation and maintenance personnel; performing intention recognition on the input text to obtain the intention information, and the intention recognition is used to select a standard intention with the highest similarity to the intention of the input text from multiple standard intentions.

[0089] Exemplarily, if the input text of the operation and maintenance personnel on the client 110 is "Query abnormal subscriptions within the last week" or "Check how many abnormal subscriptions there are within the last week", then the query intent of this query is "Query abnormal subscriptions", and the API parameter is the query period: "the last week"; Exemplarily, if the input text of the operation and maintenance personnel is "Query the abnormal subscriptions of 13152647565" or "Why does 13152647565 have two main products", then the query intent of this query is "Query abnormal subscriptions", and the API parameter is the query user: "13152647565". Among them, "Query abnormal subscriptions within the last week" and "Query the abnormal subscriptions of 13152647565" can be called the above standard intents or standard queries, and "Check how many abnormal subscriptions there are within the last week" and "Why does 13152647565 have two main products" can be called non-standard intents or extended queries. If the input text of the operation and maintenance personnel on the client 110 is the latter, the client 110 or the first server 120 can perform intent recognition on this input text, so as to select a standard intent with the highest similarity to the intent of the input text as the intent information of this query.

[0090] Optionally, after the operation and maintenance personnel input text on the client 110, the client 110 or the first server 120 can perform all or part of the operations such as spelling check, text tokenization, extracting word vectors, extracting sentence vectors, unknown intent detection, business intent classification, and intent confidence ranking, so as to return the finally recognized standard intent. In addition, the client 110 or the first server 120 can also perform associated intent recommendation through an associated intent graph.

[0091] Exemplarily, the above intent recognition can include three steps: word embedding, sentence embedding, and intent classification. Among them, word embedding is used to extract word vectors, specifically: according to the context relationship of words in the corpus, obtain the word vectors in the input text, and use the word vectors of the central vocabulary to predict the word vectors of adjacent vocabulary. Sentence embedding is used to extract sentence vectors, specifically: sum up all the word vectors in the sentences of the input text (or calculate the term frequency–inverse document frequency (TF-IDF)), to obtain the sentence vector of the input text. Intent classification means: according to the sentence vector, calculate the similarity between the input text and a certain standard intent (or called standard description), and the similarity model can use algorithms such as support vector machine (SVM), K-nearest neighbor (KNN) classification algorithm, or other algorithms for intent classification, which is not limited here.

[0092] S220, the first server 120 obtains the to-be-predicted data of at least one user service corresponding to the intent information.

[0093] It should be understood that the to-be-predicted data above can be for one user service or multiple user services, and the embodiments of the present application do not limit this. Correspondingly, the first feature data obtained by performing feature calculation on the to-be-predicted data can correspond to one user service or multiple user services.

[0094] In the case where the above intent information includes a query intent and API parameters, it should be understood that the to-be-predicted data of at least one user service corresponding to the above intent information is the to-be-predicted data of at least one user service corresponding to the query intent and API parameters.

[0095] In a possible implementation manner, if the API parameters include a query period, then the to-be-predicted data is the subscription data of at least one user service within the query period; in another possible implementation manner, the API parameters include a query user, then the to-be-predicted data is the subscription data of at least one user service corresponding to the query user; in yet another possible implementation manner, the API parameters include a query period and a query user, then the to-be-predicted data is the subscription data of at least one user service corresponding to the query user within the query period.

[0096] S230, the first server 120 performs feature calculation on the to-be-predicted data to obtain the first feature data of the at least one user service.

[0097] In the embodiments of the present application, the to-be-predicted data may include: the first subscription information of the main products of the telecommunications operator, the first subscription information of the package services, and the first subscription information of the package services. Optionally, the above subscription information may be stored in the database in the form of a subscription table. Exemplarily, the first subscription information of the above main products, the first subscription information of the package services, and the first subscription information of the package services may be obtained in advance through historical subscription records, package configuration tables, product configuration tables, etc.

[0098] In the process of feature calculation, the embodiments of the present application perform feature calculation on the to-be-predicted data based on the user granularity to extract the first feature data of each user among the at least one user corresponding to the to-be-predicted data.

[0099] As an optional embodiment, the first feature data corresponds to at least one user, and the first feature data of the first user among the at least one user includes at least one of the following:

[0100] The first sub-feature data of the first user is used to indicate whether the first user has subscribed to the main product. The first sub-feature data of the first user is determined based on the first subscription information of the main product in the data to be predicted;

[0101] The second sub-feature data of the first user is used to indicate whether the first user has subscribed to at least two main products. The second sub-feature data of the first user is determined based on the first subscription information of the main product in the data to be predicted;

[0102] The third sub-feature data of the first user is used to indicate whether there is an overlap in time among at least two main products subscribed by the first user. The third sub-feature data of the first user is determined based on the first subscription information of the main product in the data to be predicted. The first subscription information of the main product in the data to be predicted includes the start time and expiration time of the main product subscribed by the first user;

[0103] The fourth sub-feature data of the first user is used to indicate the number of mandatory packages subscribed by the first user. The fourth sub-feature data of the first user is determined based on the first subscription information of the package service in the data to be predicted;

[0104] The fifth sub-feature data of the first user is used to indicate the number of packages subscribed by the first user. The fifth sub-feature data of the first user is determined based on the first subscription information of the package service in the data to be predicted.

[0105] Exemplarily, the first server 120 can group the data to be predicted according to the identifier (such as ID) of the user, and then calculate the first feature data of each user. Taking the first user as an example, if the first user has subscribed to the main product, the above first sub-feature data can be recorded as X1 = 1, otherwise X1 = 0; if the first user has subscribed to at least two main products, the above second sub-feature data can be recorded as X2 = 1, otherwise X2 = 0; if X2 = 1, and there is an overlap in time among at least two main products subscribed by the first user, the above third sub-feature data can be recorded as X3 = 1, otherwise X3 = 0; the above fourth sub-feature data is determined according to the number of mandatory packages subscribed by the first user. If the number of mandatory packages subscribed by the first user is 3, the fourth sub-feature data can be recorded as X4 = 3; the above fifth sub-feature data is determined according to the number of packages subscribed by the first user. If the number of packages subscribed by the first user is 4, the fifth sub-feature data can be recorded as X5 = 4.

[0106] It should be understood that the first feature data may include all or part of the above five sub-feature data. The more dimensions the first feature data has, the richer the representation is, and the more accurate the prediction will be.

[0107] In an embodiment of the present application, the first server 120 inputs the first feature data of the at least one user into the data audit prediction model to obtain a prediction result, which indicates whether the user business of the at least one user is abnormal. For example, the prediction result can be represented by Y, where Y=1 indicates abnormal and Y=0 indicates normal.

[0108] S240: The first server 120 inputs the first feature data into a data audit prediction model to determine whether the at least one user service is abnormal.

[0109] In an embodiment of the present application, the server can obtain the data to be predicted of at least one user service corresponding to the intent information, perform characterization calculations on the data to be predicted, obtain the first characteristic data of at least one user service, and then input the first characteristic data of at least one user service into the data audit prediction model to determine whether the at least one user service is abnormal. This audit method can analyze the operator's user services and identify abnormal user services before any adverse consequences occur, that is, it can monitor and prevent abnormal user services before they are discovered, without the need to write SQL scripts for monitoring. It can effectively and timely predict abnormal user services and avoid over-deductions and zero deductions caused by abnormal user services, thereby reducing the operator's potential losses, such as potential customer complaints, customer churn, and operator revenue loss, etc., which is conducive to improving the digital efficiency of the CRM system.

[0110] In addition, the data audit method of the embodiment of the present application can analyze the possible audit intentions of the operation and maintenance personnel based on the text entered, thereby improving the interactivity of the audit system and the efficiency of the middle platform.

[0111] The data audit prediction model is obtained by pre-training the historical user business data by the second server 130. Therefore, as an optional embodiment, before the first server 120 uses the data audit prediction model to make a prediction based on the first feature data, that is, before S240, the embodiment of the present application may also include method 300, and method 300 is introduced in detail below.

[0112] Figure 3 A schematic flow chart of a data audit method 300 provided in an embodiment of the present application is shown. The method can be applied to Figure 1 In the communication system shown, the second server 130 may execute the method, but this embodiment of the present application does not limit this. The method 300 specifically includes the following steps:

[0113] S310: The second server 130 performs a characterization calculation on the training data to obtain second characteristic data, where the second characteristic data is used to represent user service characteristics.

[0114] The data auditing and prediction model is obtained by the second server 130 through pre-training on historical user service data. In the embodiments of the present application, the second feature data represents the data obtained by performing feature calculation on the data to be trained during the training phase.

[0115] As an optional embodiment, the data to be trained includes: the second subscription information of the main products of the telecommunications operator, the second subscription information of the package services, and the second subscription information of the package plans.

[0116] It should be understood that the data to be predicted and the data to be trained are two different batches of data, but the data types of the data to be predicted and the data to be trained are the same, and both include the second subscription information of the main products of the telecommunications operator, the second subscription information of the package services, and the second subscription information of the package plans.

[0117] As an optional embodiment, the second feature data corresponds to at least one user. The second feature data of the second user among the at least one user includes at least one of the following: the first sub-feature data of the second user, which is used to indicate whether the second user has subscribed to the main product, and the first sub-feature data of the second user is determined based on the second subscription information of the main product in the data to be trained; the second sub-feature data of the second user, which is used to indicate whether the second user has subscribed to at least two main products, and the second sub-feature data of the second user is determined based on the second subscription information of the main product in the data to be trained; the third sub-feature data of the second user, which is used to indicate whether there is an overlap in time for at least two main products subscribed by the second user, and the third sub-feature data of the second user is determined based on the second subscription information of the main product in the data to be trained, and the second subscription information of the main product in the data to be trained includes the start time and expiration time of the main product subscribed by the second user; the fourth sub-feature data of the second user, which is used to indicate the number of mandatory packages subscribed by the second user, and the fourth sub-feature data of the second user is determined based on the second subscription information of the package services in the data to be trained; the fifth sub-feature data of the second user, which is used to indicate the number of package plans subscribed by the second user, and the fifth sub-feature data of the second user is determined based on the second subscription information of the package plans in the data to be trained.

[0118] The second feature data is similar to the first feature data. The difference is that the second feature data is extracted by performing feature calculation on the data to be trained, which will not be elaborated here. It should be understood that the second feature data may include all or part of the above five sub-feature data. The more dimensions the second feature data has, the richer the representation will be, and the more accurate the prediction will be.

[0119] S320, the second server 130 marks whether the user service corresponding to the second feature data is abnormal to obtain the third feature data corresponding to the second feature data.

[0120] In the training phase, the second server 130 can mark the user service corresponding to the second feature data according to historical experience information (for example, abnormal scenarios feedback by users). For example, an abnormal mark is 1 and a normal mark is 0. In this embodiment, the marking result is recorded as the third feature data.

[0121] S330, the second server 130 obtains a data audit prediction model based on the second feature data and the third feature data.

[0122] The second server 130 can use the second feature data and the third feature data as inputs for training to obtain a data audit prediction model. The above training can be supervised training, for example, decision tree training, fusion classification training based on algorithms such as SVM, logistic regression (LR), and random forest (RF). The embodiments of the present application do not limit this.

[0123] As an optional embodiment, the method further includes: the second server 130 obtains annotation information from the operation and maintenance personnel of the client 110 for the prediction result, and the annotation information is used to indicate whether the prediction result is accurate; the second server 130 retrains the data audit prediction model based on the annotation information.

[0124] In the embodiments of the present application, the operation and maintenance personnel can mark whether the prediction result is correct according to experience. The second server 130 can retrain according to the marking result of the operation and maintenance personnel, update the experience information in a timely manner, and thus obtain a data audit prediction model with higher accuracy. Exemplarily, for a user with X1 = 1, X2 = 1, X3 = 1, X4 = 2, X5 = 1, the training result is abnormal, Y = 1. After inspection by the operation and maintenance personnel, it is found that the user service corresponding to this user is not abnormal. Then, the operation and maintenance personnel can mark the corresponding Y as 0. After the operation and maintenance personnel complete the marking, the second server 130 retrains, that is, re - finds the relationship between X and Y, and calculates a new data audit prediction model.

[0125] The data audit method in the embodiments of the present application can retrain according to the feedback of the operation and maintenance personnel after obtaining the prediction result to consolidate the latest experience model and effectively improve the model accuracy.

[0126] Figure 4 Shows a schematic diagram of another data audit method 400 provided by the embodiments of the present application. Figure 4Among them, the audit platform is used to obtain the input text of the operation and maintenance personnel, which is equivalent to the client 110 in the system architecture 100. The intent recognition server is used to recognize the intent of the input text, and the computing server is used to obtain the data to be predicted and perform online prediction; the intent recognition server and the computing server can be equivalent to the first server 120 in the system architecture 100. The training server is used to perform offline training of the prediction model, which is equivalent to the second server 130 in the system architecture 100. The storage server is used to store user business data.

[0127] In S401, the operation and maintenance personnel input text to the audit platform, which is the intent information, indicating to query abnormal user business.

[0128] In S402, the audit platform obtains the input text and sends the input text to the intent recognition server.

[0129] In S403, the intent recognition server recognizes the intent of the input text, obtains the standard intent information, and returns the standard intent information to the audit platform.

[0130] In S404, the audit platform sends the standard intent information to the computing server.

[0131] For the relevant details of S401 - S404, reference can be made to S210 in the above method 200, which will not be elaborated here.

[0132] In S405, the computing server determines the prediction range based on the standard intent information and sends the prediction range to the storage server.

[0133] In S406, the storage server obtains the prediction range sent by the computing server, uses the data within the prediction range as the data to be predicted, and sends the data to be predicted to the computing server.

[0134] For the relevant details of S405 - S406, reference can be made to S220 in the above method 200, which will not be elaborated here.

[0135] In S407, the computing server obtains the data to be predicted, performs feature calculation on the data to be predicted, obtains the first feature data of at least one user business, and uses the pre-trained data audit prediction model to perform prediction to determine whether the at least one user business is abnormal. For the relevant details of S407, reference can be made to S230 and S240 in the above method 200, which will not be elaborated here.

[0136] In S408, the computing server returns the prediction result to the audit platform so that the audit platform can display the prediction result.

[0137] Optionally, after S408, the above method 400 further includes:

[0138] In S409, the operation and maintenance personnel annotate the displayed prediction results according to experience.

[0139] In S410, the audit platform sends the annotation results of the operation and maintenance personnel to the computing server.

[0140] In S411, the computing server sends it to the training server again, and the training server can use the annotation results to retrain the data audit prediction model, thereby improving the prediction accuracy of the data audit prediction model.

[0141] Optionally, before S407, the above method 400 further includes:

[0142] In S412, the training server obtains a data audit prediction model based on the data to be trained and sends the data audit prediction model to the computing server. For the relevant details of S412, reference can be made to S310 - S330 in the above method 300, which will not be elaborated here.

[0143] In summary, for the relevant details of method 400, reference can be made to the above method 200, which will not be elaborated here.

[0144] The data audit method of the embodiment of the present application can analyze the user services of the operator, identify abnormal user services before adverse consequences occur, that is, it can monitor and prevent abnormal user services before they are discovered, without writing SQL scripts for monitoring, and can effectively and timely predict abnormal user services, avoiding phenomena such as overcharging and zero - charging caused by abnormal user services, thereby reducing the possible losses of the operator, such as potential customer complaints, customer churn, and revenue loss of the operator, etc., which is beneficial to improving the digital efficiency of the CRM system. In addition, the data audit method of the embodiment of the present application can analyze the possible audit intentions according to the text input by the operation and maintenance personnel, improving the interactivity of the audit system and the efficiency of the middle platform.

[0145] It should be understood that the magnitudes of the sequence numbers of the above - mentioned processes do not mean the order of execution. The order of execution of each process should be determined according to its function and internal logic, and should not constitute any limitation to the implementation process of the embodiment of the present application.

[0146] As described above in combination with Figures 1 to 4 ..., the data audit method according to the embodiment of the present application is described in detail. Next, in combination with Figures 5 to 6 ..., the data audit device according to the embodiment of the present application will be described in detail.

[0147] Figure 5 FIG. 500 shows a data audit device 500 provided by an embodiment of the present application. The device 500 includes: an acquisition unit 510 and a processing unit 520.

[0148] Among them, the obtaining unit 510 is used to obtain intent information for querying abnormal user services; and obtain prediction data to be predicted for at least one user service corresponding to the intent information; the processing unit 520 is used to perform feature calculation on the prediction data to be predicted to obtain first feature data of the at least one user service; and input the first feature data into a data auditing prediction model to determine whether the at least one user service is abnormal.

[0149] Optionally, the prediction data to be predicted includes: the first subscription information of the main products of a telecommunications operator, the first subscription information of package services, and the first subscription information of package plans.

[0150] Optionally, the first feature data corresponds to at least one user, and the first feature data of the first user among the at least one user includes at least one of the following: the first sub-feature data of the first user, used to indicate whether the first user has subscribed to the main product, and the first sub-feature data of the first user is determined based on the first subscription information of the main product in the prediction data to be predicted; the second sub-feature data of the first user, used to indicate whether the first user has subscribed to at least two main products, and the second sub-feature data of the first user is determined based on the first subscription information of the main product in the prediction data to be predicted; the third sub-feature data of the first user, used to indicate whether there is an overlap in time for at least two main products subscribed by the first user, and the third sub-feature data of the first user is determined based on the first subscription information of the main product in the prediction data to be predicted, and the first subscription information of the main product in the prediction data to be predicted includes the start time and expiration time of the main product subscribed by the first user; the fourth sub-feature data of the first user, used to indicate the number of mandatory packages subscribed by the first user, and the fourth sub-feature data of the first user is determined based on the first subscription information of the package services in the prediction data to be predicted; the fifth sub-feature data of the first user, used to indicate the number of package plans subscribed by the first user, and the fifth sub-feature data of the first user is determined based on the first subscription information of the package plans in the prediction data to be predicted.

[0151] Optionally, the intent information includes a query intent and application programming interface (API) parameters, the API parameters include at least one of a query period and query users, and the prediction data to be predicted is data corresponding to the query intent and the API parameters.

[0152] Optionally, the obtaining unit 510 is further used to: obtain the input text of an operation and maintenance personnel; the processing unit 520 is further used to:

[0153] Perform intent recognition on the input text to obtain the intent information. This intent recognition is used to select, from multiple standard intents, the one standard intent with the highest similarity to the intent of the input text.

[0154] Optionally, the processing unit 520 is further configured to: perform feature calculation on the training data to obtain second feature data, which is used to represent user service characteristics; mark whether the user service corresponding to the second feature data is abnormal to obtain third feature data corresponding to the second feature data; and obtain the data auditing prediction model based on the second feature data and the third feature data.

[0155] Optionally, the training data includes: second subscription information of the main products of a telecommunications operator, second subscription information of package services, and second subscription information of package plans.

[0156] Optionally, the second feature data corresponds to at least one user. The second feature data of the second user among the at least one user includes at least one of the following: the first sub-feature data of the second user, which is used to indicate whether the second user has subscribed to the main product, and the first sub-feature data of the second user is determined based on the second subscription information of the main product in the training data; the second sub-feature data of the second user, which is used to indicate whether the second user has subscribed to at least two main products, and the second sub-feature data of the second user is determined based on the second subscription information of the main product in the training data; the third sub-feature data of the second user, which is used to indicate whether there is an overlap in time for at least two main products subscribed by the second user, and the third sub-feature data of the second user is determined based on the second subscription information of the main product in the training data, and the second subscription information of the main product in the training data includes the start time and expiration time of the main product subscribed by the second user; the fourth sub-feature data of the second user, which is used to indicate the number of mandatory packages subscribed by the second user, and the fourth sub-feature data of the second user is determined based on the second subscription information of the package service in the training data; the fifth sub-feature data of the second user, which is used to indicate the number of package plans subscribed by the second user, and the fifth sub-feature data of the second user is determined based on the second subscription information of the package plan service in the training data.

[0157] Optionally, the obtaining unit 510 is further configured to: obtain the annotation information of the operation and maintenance personnel on the prediction result, and the annotation information is used to indicate whether the prediction result is accurate; the processing unit 520 is further configured to: re-train the data auditing prediction model based on the annotation information.

[0158] It should be understood that the device 500 herein is embodied in the form of functional units. The term "unit" herein may refer to an application specific integrated circuit (ASIC), an electronic circuit, a processor (such as a shared processor, a proprietary processor or a group of processors, etc.) for executing one or more software or firmware programs, a memory, a combined logic circuit and / or other suitable components supporting the described functions. In an alternative example, those skilled in the art can understand that the device 500 may specifically be the client or the server in the above embodiments, or alternatively, the functions of the client and the server in the above embodiments may be integrated in the device 500, and the device 500 may be used to execute each process and / or step corresponding to the client or the server in the above method embodiments. To avoid repetition, details are not described herein again.

[0159] The above device 500 has the function of implementing the corresponding steps executed by the client or the server in the above method; the above function may be implemented by hardware or by hardware executing corresponding software. The hardware or software includes one or more modules corresponding to the above functions. For example, the above obtaining unit 510 may be a communication interface, such as a transceiver interface.

[0160] In the embodiments of the present application, Figure 5 the device 500 in may also be a chip or a chip system, for example: a system on chip (SoC). Correspondingly, the obtaining unit 510 may be the transceiver circuit of the chip, which is not limited herein.

[0161] Figure 6 Fig. shows another device 600 for data auditing provided by the embodiments of the present application. The device 600 includes a processor 610, a transceiver 620 and a memory 630. Among them, the processor 610, the transceiver 620 and the memory 630 communicate with each other through an internal connection path. The memory 630 is used to store instructions, and the processor 610 is used to execute the instructions stored in the memory 630 to control the transceiver 620 to send signals and / or receive signals.

[0162] Among them, the transceiver 620 is used to obtain intention information, and the intention information is used to query abnormal user services; and, obtain the to-be-predicted data of at least one user service corresponding to the intention information; the processor 610 is used to perform feature calculation on the to-be-predicted data to obtain the first feature data of the at least one user service; and, input the first feature data into a data auditing prediction model to determine whether the at least one user service is abnormal.

[0163] It should be understood that the device 600 may specifically be the client or the server in the above embodiments. Alternatively, the functions of the client and the server in the above embodiments may be integrated in the device 600, and the device 600 may be used to execute the respective steps and / or processes corresponding to the client or the server in the above method embodiments. Optionally, the memory 640 may include a read-only memory and a random access memory, and provide instructions and data to the processor. A part of the memory may further include a non-volatile random access memory. For example, the memory may also store information about the device type. The processor 620 may be used to execute the instructions stored in the memory, and when the processor executes the instructions, the processor may execute the respective steps and / or processes corresponding to the client or the server in the above method embodiments.

[0164] It should be understood that in the embodiments of the present application, the processor may be a central processing unit (CPU), and the processor may also be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor, or the processor may also be any conventional processor, etc.

[0165] In the implementation process, the respective steps of the above method may be completed by the integrated logic circuit in the hardware of the processor or the instructions in the form of software. The steps of the method disclosed in combination with the embodiments of the present application may be directly embodied as being executed and completed by the hardware processor, or executed and completed by the combination of the hardware and software modules in the processor. The software module may be located in a mature storage medium in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, or an electrically erasable programmable memory, a register, etc. This storage medium is located in the memory, and the processor executes the instructions in the memory and combines its hardware to complete the steps of the above method. To avoid repetition, it will not be described in detail here.

[0166] Those of ordinary skill in the art can realize that the units and algorithm steps of the examples described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or by the combination of computer software and electronic hardware. Whether these functions are executed in hardware or software depends on the specific application and design constraints of the technical solution. The professional technician can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.

[0167] Those skilled in the art can clearly understand that for the convenience and conciseness of description, the specific working processes of the systems, devices, and units described above may refer to the corresponding processes in the foregoing method embodiments, and will not be described herein again.

[0168] In several embodiments provided in this application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of the devices or units can be in electrical, mechanical, or other forms.

[0169] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0170] In addition, in each embodiment of this application, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit.

[0171] If the functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art or part of this technical solution can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of this application. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs that can store program codes.

[0172] As described above, it is only the specific implementation manner of this application, but the protection scope of this application is not limited thereto. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed in this application, and all should be covered by the protection scope of this application. Therefore, the protection scope of this application should be subject to the protection scope of the claims.

Claims

1. A data audit method, characterized in that: include: Obtaining intent information, which is used to query abnormal user services; Acquire data to be predicted for at least one user service corresponding to the intent information; Performing characterization calculation on the data to be predicted to obtain first characteristic data of the at least one user service; The first feature data is input into a data audit prediction model to determine whether the at least one user service is abnormal, and the data audit prediction model is pre-trained based on historical user service data.

2. The method according to claim 1, characterized in that The data to be predicted includes: subscription information of the telecom operator's main products, subscription information of package services, and subscription information of set services.

3. The method according to claim 2, characterized in that The first feature data corresponds to at least one user, and the first feature data of a first user among the at least one user includes at least one of the following: The first sub-feature data of the first user is used to indicate whether the first user has ordered the main product, and the first sub-feature data of the first user is determined based on ordering information of the main product in the data to be predicted; The second sub-feature data of the first user is used to indicate whether the first user has ordered at least two main products, and the second sub-feature data of the first user is determined based on ordering information of the main products in the data to be predicted; The third sub-feature data of the first user is used to indicate whether there is temporal overlap between at least two main products ordered by the first user, the third sub-feature data of the first user being determined based on ordering information of the main products in the data to be predicted, the ordering information of the main products in the data to be predicted including a start time and an expiration time of the main products ordered by the first user; The fourth sub-characteristic data of the first user is used to indicate the quantity of required packages subscribed by the first user, the fourth sub-characteristic data of the first user being determined based on subscription information of the package service in the data to be predicted; The fifth sub-characteristic data of the first user is used to indicate the number of packages subscribed by the first user, and the fifth sub-characteristic data of the first user is determined based on the subscription information of the package services in the data to be predicted.

4. The method according to any one of claims 1 to 3, characterized in that The intent information includes query intent and application program interface (API) parameters, the API parameters include at least one of a query cycle and a query user, and the data to be predicted is data corresponding to the query intent and the API parameters.

5. The method according to any one of claims 1 to 3, characterized in that The acquisition intention information includes: Get the input text from the operation and maintenance personnel; Intent recognition is performed on the input text to obtain the intent information, wherein the intent recognition is used to select a standard intent having the highest similarity with the intent of the input text from a plurality of standard intents.

6. The method according to any one of claims 1 to 3, characterized in that Before acquiring the intention information, the method further includes: Performing a characterization calculation on the training data to obtain second characteristic data, where the second characteristic data is used to represent user service characteristics; Mark whether the user service corresponding to the second feature data is abnormal, so as to obtain third feature data corresponding to the second feature data; The data audit prediction model is obtained based on the second feature data and the corresponding third feature data.

7. The method according to any one of claims 1 to 3, characterized in that The method further comprises: Obtaining annotation information of the prediction result by the operation and maintenance personnel, where the annotation information is used to indicate whether the prediction result is accurate; Based on the labeled information, the data audit prediction model is retrained.

8. A data auditing device, characterized in that: include: an acquisition unit, configured to acquire intent information, wherein the intent information is used to query abnormal user services; and obtaining data to be predicted for at least one user service corresponding to the intention information; a processing unit, configured to perform a characterization calculation on the data to be predicted to obtain first characteristic data of the at least one user service; Furthermore, the first feature data is input into a data audit prediction model to determine whether the at least one user service is abnormal, and the data audit prediction model is pre-trained based on historical user service data.

9. The device according to claim 8, characterized in that The data to be predicted includes: subscription information of the telecom operator's main products, subscription information of package services, and subscription information of set services.

10. The device according to claim 9, characterized in that The first feature data corresponds to at least one user, and the first feature data of a first user among the at least one user includes at least one of the following: The first sub-feature data of the first user is used to indicate whether the first user has ordered the main product, and the first sub-feature data of the first user is determined based on ordering information of the main product in the data to be predicted; The second sub-feature data of the first user is used to indicate whether the first user has ordered at least two main products, and the second sub-feature data of the first user is determined based on ordering information of the main products in the data to be predicted; The third sub-feature data of the first user is used to indicate whether there is temporal overlap between at least two main products ordered by the first user, the third sub-feature data of the first user being determined based on ordering information of the main products in the data to be predicted, the ordering information of the main products in the data to be predicted including a start time and an expiration time of the main products ordered by the first user; The fourth sub-characteristic data of the first user is used to indicate the quantity of required packages subscribed by the first user, the fourth sub-characteristic data of the first user being determined based on subscription information of the package service in the data to be predicted; The fifth sub-characteristic data of the first user is used to indicate the number of packages subscribed by the first user, and the fifth sub-characteristic data of the first user is determined based on the subscription information of the package services in the data to be predicted.

11. The device according to any one of claims 8 to 10, characterized in that The intent information includes query intent and application program interface (API) parameters, the API parameters include at least one of a query cycle and a query user, and the data to be predicted is data corresponding to the query intent and the API parameters.

12. The device according to any one of claims 8 to 10, characterized in that The acquisition unit is further configured to: Get the input text from the operation and maintenance personnel; The processing unit is further configured to: Intent recognition is performed on the input text to obtain the intent information, wherein the intent recognition is used to select a standard intent having the highest similarity with the intent of the input text from a plurality of standard intents.

13. The device according to any one of claims 8 to 10, characterized in that The processing unit is further configured to: Performing a characterization calculation on the training data to obtain second characteristic data, where the second characteristic data is used to represent user service characteristics; Mark whether the user service corresponding to the second feature data is abnormal, so as to obtain third feature data corresponding to the second feature data; The data audit prediction model is obtained based on the second feature data and the corresponding third feature data.

14. The device according to any one of claims 8 to 10, characterized in that The acquisition unit is further configured to: Obtaining annotation information of the prediction result by the operation and maintenance personnel, where the annotation information is used to indicate whether the prediction result is accurate; The processing unit is further configured to: Based on the labeled information, the data audit prediction model is retrained.

15. A data auditing device, characterized in that: include: A processor is coupled to a memory, wherein the memory is used to store a program, and when the program is executed by the processor, the device executes the method according to any one of claims 1 to 7.

16. A computer-readable storage medium, characterized in that Used to store a computer program, the computer program comprising instructions for implementing the method according to any one of claims 1 to 7.

17. A chip, characterized in that: include: A processor and an interface, configured to call and run a computer program stored in a memory to execute the method according to any one of claims 1 to 7.