Business data classification method and apparatus, and storage medium

By automatically extracting and analyzing the characteristic fields of user business data through a business data classification device, the problem of low efficiency and accuracy in statistical analysis of user business types has been solved, and efficient data classification and visualization have been achieved.

CN117033730BActive Publication Date: 2025-12-16CHINA UNITED NETWORK COMM GRP CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310997032.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-08-08
Publication Date
2025-12-16
Estimated Expiration
2043-08-08

AI Technical Summary

Technical Problem

The efficiency and accuracy of existing user service type statistical analysis are low, mainly due to the rapid changes in user services and the significant delays in manual download and analysis methods.

Method used

The system automatically extracts feature fields from user business data using a business data classification device. It uses Python's re module and regular expression matching rules to extract user business type feature fields and identifier fields, and determines the user's business type based on these fields.

Benefits of technology

It improves the efficiency and accuracy of statistical analysis of user business types, reduces manual intervention, and achieves automated data classification and visualization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117033730B_ABST
    Figure CN117033730B_ABST
Patent Text Reader

Abstract

The present disclosure provides a service data classification method and device and a storage medium, relates to the technical field of communication, and solves the technical problems of low efficiency and low accuracy of user service type statistical analysis in related technologies. The method comprises the following steps: obtaining user service data; the user service data comprises a user service type feature field and a user identifier field; extracting a target first attribute field, a target second attribute field, a target third attribute field and a target user identifier field in the user service data; and determining the service type of a target user based on the target first attribute field, the target second attribute field, the target third attribute field and the target user identifier field. The present disclosure is used in the scene of service data classification statistics.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present disclosure relates to the technical field of communication, and particularly relates to a service data classification method and device and storage medium. BACKGROUND

[0002] At present, in the scenario of user service type statistics, user data needs to be manually logged in a user data backup server from multiple backup paths for downloading, and data analysis needs to be manually performed after the user data is downloaded.

[0003] Due to the rapid change of user services, the manual extraction and analysis method has a large delay, which leads to low efficiency and accuracy of user service type statistics and analysis. How to improve the efficiency and accuracy of user service type statistics and analysis has become a technical problem to be solved. SUMMARY

[0004] The present disclosure provides a service data classification method and device and storage medium, and solves the technical problem of low efficiency and accuracy of user service type statistics and analysis in the related art.

[0005] To achieve the above object, the present disclosure adopts the following technical solutions:

[0006] In a first aspect, a service data classification method is provided, which includes: obtaining user service data; the user service data includes a service type feature field of a user and a user identifier field; the service type feature field includes at least one of the following: a first attribute field, a second attribute field, and a third attribute field; the user identifier field is used to identify a user; the first attribute field is used to represent a calling service attribute of a user service; the second attribute field is used to represent a called service attribute of a user service; the third attribute field is used to represent a calling service attribute and a called service attribute of a voice over long term evolution (VoLTE) communication service of a user; extracting a target first attribute field, a target second attribute field, a target third attribute field, and a target user identifier field in the user service data; the target user identifier field is any one of the user identifier fields; the target first attribute field, the target second attribute field, and the target third attribute field are attribute fields of service data of a target user; and determining a service type of the target user based on the target first attribute field, the target second attribute field, the target third attribute field, and the target user identifier field.

[0007] With reference to the first aspect, in a possible implementation manner, the method specifically comprises: presetting a first extraction rule of the first attribute field, a second extraction rule of the second attribute field, a third extraction rule of the third attribute field, and a fourth extraction rule of the user identifier field; the first extraction rule, the second extraction rule, the third extraction rule, and the fourth extraction rule are all preset program scripts, and are used for extracting the attribute fields in the user service data; and the first attribute field, the second attribute field, the third attribute field, and the user identifier field in the user service data are extracted based on the first extraction rule, the second extraction rule, the third extraction rule, and the fourth extraction rule.

[0008] With reference to the first aspect, in a possible implementation manner, the method specifically comprises: determining a target user identified by the user identifier field based on the user identifier field; in a case where the service category characteristic field comprises one attribute field of the first attribute field, the second attribute field, and the third attribute field, determining that a service attribute corresponding to the one attribute field is a service category of the target user; in a case where the service category characteristic field comprises any two attribute fields of the first attribute field, the second attribute field, and the third attribute field, determining that a sum of two service attributes corresponding to the any two attribute fields is a service category of the target user; and in a case where the service category characteristic field comprises the first attribute field, the second attribute field, and the third attribute field, determining that a sum of a first service attribute corresponding to the first attribute field, a second service attribute corresponding to the second attribute field, and a third service attribute corresponding to the third attribute field is a service category of the target user.

[0009] With reference to the first aspect, in a possible implementation manner, the method specifically comprises: accessing the first server; the first server is used for storing compressed files of user service data; obtaining a directory of all compressed files in the first server; determining the compressed files of the user service data based on matching based on a first regular expression; and obtaining the user service data based on matching based on a second regular expression to determine a user service data text file path in the compressed files of the user service data.

[0010] With reference to the first aspect, in a possible implementation manner, the method further comprises: determining first service data classification statistical data based on the service data of each user in the first server; the first service data classification statistical data comprises a service category of each user; encapsulating the first service data classification statistical data as second service data classification statistical data; the second service data classification statistical data is the first service data classification statistical data encapsulated in a JSON format; and rendering the second service data classification statistical data as a visual chart.

[0011] In a second aspect, a service data classification apparatus is provided, which comprises a communication unit and a processing unit. The communication unit is configured to acquire user service data, wherein the user service data comprises a service type feature field of a user and a user identifier field; the service type feature field comprises at least one of a first attribute field, a second attribute field, and a third attribute field; and the user identifier field is used to identify the user. The processing unit is configured to extract a target first attribute field, a target second attribute field, a target third attribute field, and a target user identifier field from the user service data; the target user identifier field is any one of the user identifier fields; and the target first attribute field, the target second attribute field, and the target third attribute field are attribute fields of service data of a target user. The processing unit is further configured to determine a service type of the target user based on the target first attribute field, the target second attribute field, the target third attribute field, and the target user identifier field.

[0012] In a possible implementation of the second aspect, the processing unit is specifically configured to preset a first extraction rule of the first attribute field, a second extraction rule of the second attribute field, a third extraction rule of the third attribute field, and a fourth extraction rule of the user identifier field; the first extraction rule, the second extraction rule, the third extraction rule, and the fourth extraction rule are all preset program scripts used to extract the attribute fields in the user service data; and the first attribute field, the second attribute field, the third attribute field, and the user identifier field in the user service data are extracted based on the first extraction rule, the second extraction rule, the third extraction rule, and the fourth extraction rule.

[0013] In a possible implementation of the second aspect, the processing unit is specifically configured to determine a target user identified by the user identifier field based on the user identifier field; in a case where the service type feature field comprises one attribute field of the first attribute field, the second attribute field, and the third attribute field, determine that a service attribute corresponding to the one attribute field is the service type of the target user; in a case where the service type feature field comprises any two attribute fields of the first attribute field, the second attribute field, and the third attribute field, determine that a sum of two service attributes corresponding to the any two attribute fields is the service type of the target user; and in a case where the service type feature field comprises the first attribute field, the second attribute field, and the third attribute field, determine that a sum of a first service attribute corresponding to the first attribute field, a second service attribute corresponding to the second attribute field, and a third service attribute corresponding to the third attribute field is the service type of the target user.

[0014] In a possible implementation manner of the second aspect, the processing unit is specifically configured to: access the first server; the first server is configured to store compressed files of user service data; instruct the communication unit to acquire a directory of all compressed files in the first server; determine the compressed files of user service data based on matching of the first regular expression; determine paths of user service data text files in the compressed files of user service data based on matching of the second regular expression, and acquire the user service data.

[0015] In a possible implementation manner of the second aspect, the processing unit is further configured to: determine first service data classification statistical data based on service data of each user in the first server; the first service data classification statistical data includes service categories of each user; encapsulate the first service data classification statistical data into second service data classification statistical data; the second service data classification statistical data is the first service data classification statistical data encapsulated in a JSON format; and render the second service data classification statistical data into a visual chart.

[0016] A third aspect provides a service data classification apparatus, including a processor and a memory; the memory is configured to store computer execution instructions; when the service data classification apparatus is running, the processor executes the computer execution instructions stored in the memory, so that the service data classification apparatus executes the service data classification method as described in the first aspect and any possible implementation manner thereof.

[0017] A fourth aspect provides a computer readable storage medium, the computer readable storage medium stores instructions; when the instructions in the computer readable storage medium are executed by a processor of a service data classification apparatus, the service data classification apparatus executes the service data classification method as described in the first aspect and any possible implementation manner thereof.

[0018] A fifth aspect provides a chip, the chip includes a processor and a communication interface; the communication interface is coupled with the processor; the processor is configured to run a computer program or instructions, so as to realize the service data classification method as described in the first aspect and any possible implementation manner thereof.

[0019] In the present disclosure, the name of the service data classification apparatus does not constitute a limitation on the device or functional module itself, and in actual implementation, these devices or functional modules can appear in other names. As long as the functions of each device or functional module are similar to the present disclosure, it belongs to the scope of the present disclosure and its equivalent technology.

[0020] The technical scheme provided by the disclosure has at least the following beneficial effects: the business data classification device in the disclosure first acquires user business data through a preset program script; the user business data includes a user business type characteristic field and a user identifier field; the business type characteristic field includes at least one of a first attribute field, a second attribute field, and a third attribute field; the user identifier field is used to identify a user, that is, the business data classification device automatically extracts user business data stored in an FTP server through a preset program script; then the first attribute field, the second attribute field, the third attribute field, and the user identifier field in the user business data are extracted through a re module of Python; based on the first attribute field, the second attribute field, the third attribute field, and the user identifier field, the business type of a target user is determined; the target user is a user identified by the user identifier field; the first attribute field is used to represent the OCSI attribute of the user business; the second attribute field is used to represent the TCSI attribute of the user business; and the third attribute field is used to identify the SharediFCSetID attribute of the user business. The business data classification device automatically extracts user business data stored in an FTP server through a preset program script; according to a preset extraction rule, the re module of Python is used to extract the characteristic fields of the business data from the huge user business data, and the business types are classified based on the characteristic fields of the business data; manual login of a user data backup server, download of user data from multiple backup paths, and manual data analysis are not required; and therefore the efficiency and accuracy of user business type statistical analysis are improved. BRIEF DESCRIPTION OF DRAWINGS

[0021] In order to more clearly illustrate the technical scheme in the embodiments of the disclosure or the prior art, the drawings needed to be used in the embodiments or the prior art description will be briefly introduced.

[0022] Figure 1 A business data classification system structure schematic diagram provided by the embodiments of the disclosure;

[0023] Figure 2 A hardware structure schematic diagram of a business data classification device provided by the embodiments of the disclosure;

[0024] Figure 3 A flowchart of a business data classification method provided by the embodiments of the disclosure;

[0025] Figure 4 A flowchart of another business data classification method provided by the embodiments of the disclosure;

[0026] Figure 5 A flowchart of another business data classification method provided by the embodiments of the disclosure;

[0027] Figure 6 A flowchart of another service data classification method provided by an embodiment of the present disclosure is shown in FIG. 6.

[0028] Figure 7 A flowchart of another service data classification method provided by an embodiment of the present disclosure is shown in FIG. 6.

[0029] Figure 8 A flowchart of another service data classification method provided by an embodiment of the present disclosure is shown in FIG. 6.

[0030] Figure 9 A structural diagram of a service data classification apparatus provided by an embodiment of the present disclosure is shown in FIG. 7. DETAILED DESCRIPTION

[0031] A service data classification method, apparatus and storage medium provided by an embodiment of the present disclosure are described in detail below with reference to the accompanying drawings.

[0032] The term “and / or” in this document merely describes an association relationship of associated objects, and can represent three relationships, for example, A and / or B can represent three cases of existence of A alone, existence of A and B simultaneously, and existence of B alone.

[0033] The terms “first” and “second” and the like in the description of the present disclosure and the accompanying drawings are used to distinguish different objects or different processing of the same object, rather than to describe a specific order of the objects.

[0034] In addition, the terms “include” and “have” and any variations thereof mentioned in the description of the present disclosure are intended to cover the inclusions without exclusivity. For example, a process, method, system, product or device including a series of steps or units is not limited to the listed steps or units, but can optionally include other steps or units not listed or can optionally include other steps or units inherent to the process, method, product or device. It should be noted that in the embodiments of the present disclosure, the words “exemplary” or “for example” are used to mean serving as an example, instance or illustration. Any embodiment or design scheme described as “exemplary” or “for example” in the embodiments of the present disclosure should not be interpreted as more preferred or more advantageous than other embodiments or design schemes. Rather, the words “exemplary” or “for example” are used in the sense of presenting a related concept in a specific manner.

[0035] In the description of the present disclosure, the meaning of “a plurality of” is two or more, unless otherwise specified.

[0036] Currently, in the scenario of user service type statistics in a communication company, user service data needs to be analyzed. Although relevant user data is recorded when a user opens an account, due to continuous changes of user services, the accuracy and timeliness of front-end data are low, and therefore, user data analysis needs to extract user service data backed up by a home subscriber server (HSS) and a unified data management (UDM) network element.

[0037] The HSS and the UDM are user unified data management units, mainly responsible for management of user identification and user service subscription data. According to user service subscription information, basic service conditions of a user and user quantity of different service types can be determined. User data in the HSS and the UDM is backed up in the form of a compressed file every day, on a file transfer protocol (FTP) server. When user service classification statistics are performed, a user needs to manually log in to multiple backup path servers to download compressed files of user service data, decompress the files, and then manually analyze the user service data. However, because there are many types of services, the data analysis process is relatively complex, and user services change rapidly, manual extraction and analysis have a large delay, which results in low efficiency and accuracy of user service type statistical analysis.

[0038] To solve the above technical problems, the present disclosure provides a service data classification method and device and a storage medium, which are used to solve the technical problem of low efficiency and accuracy of user service type statistical analysis. The method comprises: obtaining user service data; the user service data comprises a user service type feature field and a user identifier field; the service type feature field comprises at least one of a first attribute field, a second attribute field, and a third attribute field; the user identifier field is used to identify a user; the first attribute field is used to represent a calling service attribute of a user service; the second attribute field is used to represent a called service attribute of a user service; the third attribute field is used to represent a calling service attribute and a called service attribute of a long term evolution voice bearer (volte) communication service of a user; extracting a target first attribute field, a target second attribute field, a target third attribute field, and a target user identifier field in the user service data; the target user identifier field is any one of the user identifier fields; the target first attribute field, the target second attribute field, and the target third attribute field are attribute fields of service data of a target user; and determining a service type of the target user based on the target first attribute field, the target second attribute field, the target third attribute field, and the target user identifier field.

[0039] In a possible implementation, the business data classification method can be applied to a business data classification system 100. Hereinafter, the business data classification method is described in combination with Figure 1 A business data classification system 100 provided by an embodiment of the present application is described in detail. As shown in Figure 1 Figure 1 A business data classification system 100 provided by an embodiment of the present application includes an FTP database 101 and a business data classification device 102. The FTP database 101 is configured to store compressed files of backup user business data. The business data classification device 102 is configured to obtain the compressed files of the user business data stored in the FTP database 101, decompress the compressed files, and obtain the user business data. The user business data includes a user business type feature field and a user identifier field. The business type feature field includes at least one of a first attribute field, a second attribute field, and a third attribute field. The user identifier field is configured to identify the user. The first attribute field, the second attribute field, the third attribute field, and the user identifier field in the user business data are extracted by using a re module of Python. The business type of the user is determined based on the first attribute field, the second attribute field, the third attribute field, and the user identifier field.

[0040] In a possible implementation, a hardware structure of the business data classification device 102 in the business data classification system 100 includes Figure 2 The hardware structure of the business data classification device 102 is described below by taking a business data classification device 200 shown in Figure 2 As shown in Figure 2 The business data classification device 200 includes at least one processor 201, a communication line 202, and at least one communication interface 204, and can further include a memory 203. The processor 201, the memory 203, and the communication interface 204 can be connected through the communication line 202.

[0041] The processor 201 can be a central processing unit (CPU), an application specific integrated circuit (ASIC), or one or more integrated circuits configured to implement one or more embodiments of the present application, for example, one or more digital signal processors (DSPs), or one or more field programmable gate arrays (FPGAs).

[0042] ​The communication line 202 can include a path for transmitting information between the above-mentioned components.

[0043] The communication interface 204, used for communicating with other devices or communication networks, can use any transceiver-like device, such as an Ethernet, a radio access network (RAN), a wireless local area networks (WLAN), etc.

[0044] The memory 203 can be a read-only memory (ROM) or other type of static storage device that can store static information and instructions, a random access memory (RAM) or other type of dynamic storage device that can store information and instructions, an electrically erasable programmable read-only memory (EEPROM), a compact disc read-only memory (CD-ROM) or other optical disk storage, a magnetic disk storage or other magnetic storage devices, or any other medium capable of storing desired program code in the form of instructions or data structures and that can be accessed by a computer, but is not limited thereto.

[0045] In one possible design, the memory 203 can exist independently of the processor 201, i.e., the memory 203 can be an external memory of the processor 201, and the memory 203 can be connected to the processor 201 through the communication line 202, used for storing execution instructions or application program codes, and controlled by the processor 201 to execute, to implement the service data classification method provided by the embodiments of the present disclosure. In another possible design, the memory 203 can also be integrated with the processor 201, i.e., the memory 203 can be an internal memory of the processor 201, for example, the memory 203 can be a cache, used for temporarily storing some data and instruction information, etc.

[0046] As one possible implementation, the processor 201 can include one or more CPUs, for example, the CPU0 and the CPU1 in the CPU 100. As another possible implementation, the service data classification apparatus 200 can include multiple processors, for example, the processor 201 and the processor 207 in the processor 200. As still another possible implementation, the service data classification apparatus 200 can further include the output device 205 and the input device 206. Figure 2 Figure 2 As one possible implementation, the processor 201 can include one or more CPUs, for example, the CPU0 and the CPU1 in the CPU 100. As another possible implementation, the service data classification apparatus 200 can include multiple processors, for example, the processor 201 and the processor 207 in the processor 200. As still another possible implementation, the service data classification apparatus 200 can further include the output device 205 and the input device 206. ​

[0047] The following provides a detailed description of the business data classification method provided in the embodiments of this disclosure.

[0048] like Figure 3 As shown, Figure 3 The business data classification method provided in this disclosure embodiment can be applied to, for example... Figure 2 The business data classification device shown includes the following methods S301-S303, which will be described in detail below.

[0049] S301, The business data classification device acquires user business data.

[0050] The user service data includes: a service type characteristic field and a user identifier field. The service type characteristic field includes at least one of the following: a first attribute field, a second attribute field, and a third attribute field. The user identifier field is used to identify the user. The first attribute field characterizes the calling service attribute of the user's service; the second attribute field characterizes the called service attribute of the user's service; and the third attribute field characterizes both the calling and called service attributes of the user's LTE voice carrier VoLTE communication service.

[0051] In one possible implementation, the user identification field can be either an International Mobile Subscriber Identity (IMSI) field or a Mobile Station International Subscriber Directory Number (MSISDN) field.

[0052] In one possible implementation, the first attribute field is the OCSI attribute field, the second attribute field is the TCSI attribute field, and the third attribute field is the SharediFCSetID attribute field.

[0053] It is understandable that the business type feature field may contain multiple attribute fields from the OCSI attribute field, TCSI attribute field, and SharediFCSetID attribute field, or it may contain only one attribute field.

[0054] S302, The business data classification device extracts the first attribute field, the second attribute field, the third attribute field, and the user identifier field from the user business data.

[0055] In an example, the service data classification apparatus can analyze the user service data by using the re module of Python, and extract the first attribute field, the second attribute field, the third attribute field, and the user identification field from the user service data by using a regular expression with a preset matching rule.

[0056] In an example, the service data classification apparatus can analyze the user service data by using the re module of Python, and extract the OCSI attribute field, the TCSI attribute field, the SharediFCSetID attribute field, and the IMSI field and the MSISDN field from the user service data by using a regular expression with a preset matching rule.

[0057] S303, the service data classification apparatus determines the service type of the target user based on the target first attribute field, the target second attribute field, the target third attribute field, and the target user identification field.

[0058] In an example, the service data classification apparatus determines the target user based on the user identification field, determines the target first service attribute corresponding to the target first attribute field by using the target first attribute field in the target user service data, determines the target second service attribute corresponding to the target second attribute field by using the target second attribute field, determines the target third service attribute corresponding to the target third attribute field by using the target third attribute field, and determines the service type of the target user as the target first service attribute-target second service attribute-target third service attribute.

[0059] In an example, the service data classification apparatus determines the target first service attribute as the one-number multi-terminal service based on the target first attribute field OCSI = 1, determines the target second service attribute as the VOLTE called anchor based on the target second attribute field TCSI = 1, determines the target third service attribute as the in-province 5G new media based on the target third attribute field SharediFCSetID = 3510, and determines the service type of the target user as one-number multi-terminal-VOLTE called anchor-in-province 5G new media.

[0060] The technical solution provided by the above embodiments can bring at least the following beneficial effects: The business data classification device first obtains user business data through a preset program script; the user business data includes: a user business type feature field and a user identifier field; the business type feature field includes at least one of the following: a first attribute field, a second attribute field, and a third attribute field; the user identifier field is used to identify the user, that is, the business data classification device automatically extracts the user business data stored in the FTP server through the preset program script; then, the first attribute field, the second attribute field, the third attribute field, and the user identifier field in the user business data are extracted through the re module of Python; based on the first attribute field, the second attribute field, the third attribute field, and the user identifier field, the business type of the target user is determined; the target user is the user identified by the user identifier field; the first attribute field is used to characterize the OCSI attribute of the user business; the second attribute field is used to characterize the TCSI attribute of the user business; the third attribute field is used to identify the SharediFCSetID attribute of the user business. The business data classification device automatically extracts user business data stored on an FTP server using a pre-set program script; based on pre-set extraction rules, it extracts feature fields from the massive user business data using Python's re module, and classifies business types based on these feature fields; it eliminates the need for manual login to the user data backup server to download user data from multiple backup paths, and manual data analysis; thus improving the efficiency and accuracy of user business type statistical analysis.

[0061] In one possible implementation, combining Figure 3 ,like Figure 4 As shown in the figure, the process of the business data classification device extracting the target first attribute field, target second attribute field, target third attribute field, and target user identifier field from the user business data in the above S302 can be specifically implemented through the following S401-S402, which will be explained in detail below.

[0062] S401, the business data classification device presets a first extraction rule for the first attribute field, a second extraction rule for the second attribute field, a third extraction rule for the third attribute field, and a fourth extraction rule for the user identifier field.

[0063] Among them, the first extraction rule, the second extraction rule, the third extraction rule, and the fourth extraction rule are all preset program scripts used to extract attribute fields from user business data.

[0064] In one possible implementation, the business data classification device uses Python to preset the first extraction rule for the first attribute field "OCSI=(\d+)"; the second extraction rule for the second attribute field "TCSI=(\d+)"; the third extraction rule for the third attribute field "SharediFCSetID=(\d+)"; and the fourth extraction rule for the user identifier field "IMSI=(\d+)".

[0065] S402, the business data classification device extracts the first attribute field, the second attribute field, the third attribute field, and the user identifier field from the user business data based on the first extraction rule, the second extraction rule, the third extraction rule, and the fourth extraction rule.

[0066] In one possible implementation, the business data classification device extracts the OCSI attribute field from the user business data using a first extraction rule "OCSI=(\d+)" for the first attribute field preset by Python; extracts the TCSI attribute field from the user business data using a second extraction rule "TCSI=(\d+)" for the second attribute field; extracts the SharediFCSetID attribute field from the user business data using a third extraction rule "SharediFCSetID=(\d+)" for the third attribute field; and extracts the IMSI field from the user business data using a fourth extraction rule "IMSI=(\d+)" for the user identifier field.

[0067] The technical solution provided by the above embodiments can bring at least the following beneficial effects: the business data classification device presets a first extraction rule for the first attribute field, a second extraction rule for the second attribute field, a third extraction rule for the third attribute field, and a fourth extraction rule for the user identifier field, and extracts the first attribute field, the second attribute field, and the third attribute field from the user business data according to the preset first extraction rule, second extraction rule, third extraction rule, and fourth extraction rule; so that the business data classification device can determine the user's business type based on the first attribute field, the second attribute field, and the third attribute field.

[0068] In one possible implementation, combining Figure 4 ,like Figure 5 As shown, the process by which the business data classification device determines the business type of a target user based on the first attribute field, the second attribute field, the third attribute field, and the user identifier field in S303 above can be specifically implemented through the following S501-S504, which will be explained in detail below.

[0069] S501, The business data classification device determines the target user identified by the user identification field based on the user identification field.

[0070] In a possible implementation, the user identifier field is an IMSI field or an MSISDN field, and the international mobile subscriber identity of the user can be determined through the IMSI field or the MSISDN field, so as to determine the target user.

[0071] S502, the service data classification apparatus determines, in a case where the service type feature field includes one of the first attribute field, the second attribute field, and the third attribute field, that a service attribute corresponding to the one attribute field is the service type of the target user.

[0072] In a possible implementation, the service data classification apparatus determines that the user service data of the target user only includes one of an OCSI attribute field, a TCSI attribute field, and a SharediFCSetID attribute field, and at this time, a service attribute corresponding to the attribute field is the service type of the target user.

[0073] For example, a table of service attributes corresponding to the OCSI attribute field is shown in Table 1 below, a table of service attributes corresponding to the TCSI attribute field is shown in Table 2 below, and a table of service attributes corresponding to the SharediFCSetID attribute field is shown in Table 3 below.

[0074] Table 1 Table of service attributes corresponding to the OCSI attribute field

[0075]

[0076] Table 2 Table of service attributes corresponding to the TCSI attribute field

[0077]

[0078]

[0079] Table 3 Table of service attributes corresponding to the SharediFCSetID attribute field

[0080]

[0081] For example, the service data classification apparatus determines that the user service data of the target user only includes the OCSI attribute field OCSI = 3, and at this time, the service data classification apparatus determines that the service type of the target user is the Shanghai Anti-fraud Platform.

[0082] S503, the service data classification apparatus determines, in a case where the service type feature field includes any two of the first attribute field, the second attribute field, and the third attribute field, that a sum of two service attributes corresponding to the any two attribute fields is the service type of the target user.

[0083] In a possible implementation, the service data classification apparatus determines that the user service data of the target user contains two attribute fields of the OCSI attribute field, the TCSI attribute field and the SharediFCSetID attribute field, and the service attributes corresponding to the two attribute fields are the service type of the target user.

[0084] For example, the service data classification apparatus determines that the user service data of the target user contains the OCSI attribute field OCSI=3 and the SharediFCSetID attribute field SharediFCSetID=200, and determines that the service type of the target user is Shanghai Anti-fraud Platform-SMS service.

[0085] S504, in the case where the service type feature field includes the first attribute field, the second attribute field and the third attribute field, the service data classification apparatus determines the sum of the first service attribute corresponding to the first attribute field, the second service attribute corresponding to the second attribute field and the third service attribute corresponding to the third attribute field as the service type of the target user.

[0086] In a possible implementation, the service data classification apparatus determines that the user service data of the target user contains three attribute fields of the OCSI attribute field, the TCSI attribute field and the SharediFCSetID attribute field, and the service attributes corresponding to the three attribute fields are the service type of the target user.

[0087] For example, the service data classification apparatus determines that the user service data of the target user contains the OCSI attribute field OCSI=502, the TCSI attribute field TCSI=501 and the SharediFCSetID attribute field SharediFCSetID=310, and determines that the service type of the target user is Womai Card-One Machine Multi-number-High Definition Ring.

[0088] The technical scheme provided by the above embodiments can bring at least the following beneficial effects: the service data classification apparatus determines a target user identified by the user identification field based on the user identification field; and in the case that the service type feature field includes any one of the first attribute field, the second attribute field and the third attribute field, determines that the service attribute corresponding to any one of the attribute fields is the service type of the target user; in the case that the service type feature field includes any two of the first attribute field, the second attribute field and the third attribute field, determines that the sum of the two service attributes corresponding to the any two attribute fields is the service type of the target user; and in the case that the service type feature field includes the first attribute field, the second attribute field and the third attribute field, determines that the sum of the first service attribute corresponding to the first attribute field, the second service attribute corresponding to the second attribute field and the third service attribute corresponding to the third attribute field is the service type of the target user.

[0089] In a possible implementation manner, the method is combined with Figure 5 As shown in the following S301, the service data classification apparatus obtains user service data, which can be implemented through the following S601-S604, which are described in detail as follows. Figure 6

[0090] S601, the service data classification apparatus accesses a first server.

[0091] The first server is configured to store compressed files of user service data.

[0092] In a possible implementation manner, the user service data in the HSS and the UDM network element is backed up in the FTP server in the form of compressed files every day, and the service data classification apparatus can connect to the FTP server and log in to the FTP server through the ftplib library in Python.

[0093] S602, the service data classification apparatus obtains a directory of all compressed files in the first server.

[0094] In a possible implementation manner, after logging in to the FTP server, the service data classification apparatus obtains all storage directories under the FTP server through the nlst() method.

[0095] S603, the service data classification apparatus performs matching based on a first regular expression to determine the compressed file of the user service data.

[0096] S604, the service data classification apparatus performs matching based on a second regular expression to determine a user service data text file path in the compressed file of the user service data, and obtains the user service data.

[0097] ​In one possible implementation, after the business data classification device obtains the compressed file of the user's business data, it can decompress the compressed file of the user's business data using decompression software to obtain the text document of the user's business data.

[0098] The technical solution provided by the above embodiments can bring at least the following beneficial effects: the business data classification device accesses the FTP server through Python, obtains the compressed file of user business data and decompresses it; that is, the business data classification device automatically extracts the compressed file of user business data and decompresses it into a text document form through a preset program script, so that the business data classification device can classify and statistically analyze the user's business data.

[0099] In one possible implementation, combining Figure 3 ,like Figure 7 As shown, after the business data classification device determines the business type of the target user based on the target first attribute field, the target second attribute field, the target third attribute field, and the target user identifier field in the above S303, it is also necessary to determine the business data classification statistics result. This process can be implemented through the following S701-S703, which will be explained in detail below.

[0100] S701, The business data classification device determines the first business data classification statistics based on the business data of each user in the first server.

[0101] The first business data category statistics include the business types for each user.

[0102] In one possible implementation, user service data in the HSS and UDM network elements is backed up daily to an FTP server in the form of compressed files. The service data classification device can connect to the FTP server using the ftplib library in Python and log in using the login() method. It can then obtain all storage directories under the FTP server using the nlst() method, thereby obtaining the compressed files of user service data. Finally, it can decompress the compressed files of user service data using decompression software to obtain a text document containing all user service data.

[0103] In one possible implementation, the business data classification device can perform the above steps S401-S402 for each user's business data to determine the business type of each user.

[0104] S702, The business data classification device encapsulates the first business data classification statistics into the second business data classification statistics.

[0105] The second business data category statistics are the first business data category statistics encapsulated in JSON format.

[0106] In a possible implementation, the business data classification apparatus can create a view function in the Django framework, obtain the first business data classification statistical data, and encapsulate the first business data classification statistical data into JSON format data.

[0107] S703, the business data classification apparatus renders the second business data classification statistical data into a visual chart.

[0108] In a possible implementation, after determining the business data classification statistical result, the business data classification apparatus can determine the home location of each user based on the IMSI and the MSISDN number segment, and aggregate the business data classification statistical result according to the home location.

[0109] In a possible implementation, the business data classification apparatus can render the second business data classification statistical data into a column chart or a chart by using the ECharts framework in an HTML template file, so as to intuitively display the business data classification result.

[0110] The technical solutions provided by the above embodiments can at least bring the following beneficial effects: the business data classification apparatus obtains the business data of each user in the FTP server, determines the business data classification statistical result based on the business data of each user, and intuitively displays the business data classification statistical result by building a visual chart through the Django framework.

[0111] In a possible implementation, a flowchart of the business data classification method is shown in FIG. 8. Figure 8 As shown in FIG. 8, the process of business data classification mainly includes four processes: backup data obtaining, feature field extraction, data classification analysis, and data visualization.

[0112] It can be seen that the technical solutions provided by the embodiments of the present disclosure are mainly introduced from the perspective of the method. In order to realize the above functions, it contains the corresponding hardware structure and / or software module for executing each function. Those skilled in the art should easily realize that, in combination with the modules and algorithm steps of the examples described in the embodiments disclosed in the present text, the embodiments of the present disclosure can be realized in the form of hardware or a combination of hardware and computer software. Whether a certain function is executed by hardware or computer software driven hardware depends on the specific application and design constraints of the technical solutions. Professional technicians can use different methods to realize the described functions for each specific application, but such implementation should not be considered beyond the scope of the present disclosure.

[0113] The embodiments of the present disclosure can divide the function modules of the service data classification apparatus according to the above method examples. For example, each function module can be divided according to each function, or two or more functions can be integrated in one processing module. The above integrated module can be realized in the form of hardware or in the form of a software function module. Optionally, the division of the modules in the embodiments of the present disclosure is illustrative, and is only a logical function division. In actual implementation, another division mode can be used.

[0114] In a possible implementation manner, as shown in Figure 9 Figure 9 FIG. 9 is a structural schematic diagram of a service data classification apparatus 900 provided by the present disclosure.

[0115] The service data classification apparatus 900 includes a communication unit 901 and a processing unit 902. The communication unit 901 is configured to acquire user service data. The user service data includes a service type feature field of a user and a user identifier field. The service type feature field includes at least one of a first attribute field, a second attribute field, and a third attribute field. The user identifier field is used to identify the user. The first attribute field is used to represent a calling service attribute of the user service. The second attribute field is used to represent a called service attribute of the user service. The third attribute field is used to represent a calling service attribute and a called service attribute of a voice over long term evolution (VoLTE) communication service of the user. The processing unit 902 is configured to extract a target first attribute field, a target second attribute field, a target third attribute field, and a target user identifier field in the user service data. The target user identifier field is any one of the user identifier fields. The target first attribute field, the target second attribute field, and the target third attribute field are attribute fields of service data of a target user. The processing unit 902 is configured to determine a service type of the target user based on the target first attribute field, the target second attribute field, the target third attribute field, and the target user identifier field.

[0116] In a possible implementation manner, the processing unit 902 is specifically configured to preset a first extraction rule of the first attribute field, a second extraction rule of the second attribute field, a third extraction rule of the third attribute field, and a fourth extraction rule of the user identifier field. The first extraction rule, the second extraction rule, the third extraction rule, and the fourth extraction rule are all preset program scripts, and are used to extract the attribute fields in the user service data. The first attribute field, the second attribute field, the third attribute field, and the user identifier field in the user service data are extracted based on the first extraction rule, the second extraction rule, the third extraction rule, and the fourth extraction rule.

[0117] ​In a possible implementation, the processing unit 902 is specifically configured to: determine a target user identified by the user identification field based on the user identification field; in a case where the service category feature field includes one of the first attribute field, the second attribute field and the third attribute field, determine that a service attribute corresponding to the one attribute field is a service category of the target user; in a case where the service category feature field includes any two of the first attribute field, the second attribute field and the third attribute field, determine that a sum of two service attributes corresponding to the any two attribute fields is the service category of the target user; and in a case where the service category feature field includes the first attribute field, the second attribute field and the third attribute field, determine that a sum of a first service attribute corresponding to the first attribute field, a second service attribute corresponding to the second attribute field and a third service attribute corresponding to the third attribute field is the service category of the target user.

[0118] In a possible implementation, the processing unit 902 is specifically configured to: access a first server; the first server is configured to store compressed files of user service data; instruct the communication unit 901 to acquire a directory of all compressed files in the first server; determine the compressed files of user service data based on matching of a first regular expression; and acquire user service data based on matching of a second regular expression to determine a path of a user service data text file in the compressed files of user service data.

[0119] In a possible implementation, the processing unit 902 is further configured to: determine first service data classification statistical data based on the service data of each user in the first server; the first service data classification statistical data includes a service category of each user; encapsulate the first service data classification statistical data as second service data classification statistical data; the second service data classification statistical data is the first service data classification statistical data encapsulated in a JSON format; and render the second service data classification statistical data as a visual chart.

[0120] The embodiments of the present disclosure further provide a service data classification apparatus, which comprises a processor and a memory; the memory is configured to store computer execution instructions; when the service data classification apparatus is running, the processor executes the computer execution instructions stored in the memory, so that the service data classification apparatus executes the service data classification method described in the embodiments of the present disclosure.

[0121] The embodiments of the present disclosure provide a computer program product comprising instructions which, when executed on a computer, cause the computer to carry out the service data classification method in the method embodiments.

[0122] Embodiments of the present disclosure provide a chip, comprising a processor and a communication interface, the communication interface and the processor are coupled, and the processor is configured to run computer programs or instructions to implement the service data classification method in the above method embodiments.

[0123] In some embodiments, a computer readable storage medium can be used. The computer readable storage medium can be a tangible medium that can retain the programs for later use or referred to as a computer readable storage medium. The computer readable storage medium, for example, can be, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or apparatus, or any combination of the above. More specific examples (a non-exhaustive list) of the computer readable storage medium include an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), a register, a hard disk, an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any other suitable combination of the above, or any other medium from which a computer can read. In one example, a storage medium is coupled to a processor such that the processor can read information from the storage medium and write information to the storage medium. Of course, the storage medium can be a component of the processor. The processor and the storage medium can be located in an application-specific integrated circuit (ASIC). In the embodiments of the present disclosure, the computer readable storage medium can be any tangible medium that contains or stores a program that can be used by or in connection with an instruction execution system, apparatus, or device.

[0124] Since the device, equipment, computer readable storage medium, computer program product in the embodiments of the present disclosure can be applied to the above method, the technical effects it can obtain can also be referred to the above method embodiments, and the present disclosure embodiments will not be repeated here.

[0125] The above shows only a specific embodiment of the present disclosure, but the protection scope of the present disclosure is not limited thereto, any change or replacement within the technical scope disclosed by the present disclosure should be covered in the protection scope of the present disclosure. Therefore, the protection scope of the present disclosure should be subject to the protection scope of the claims.

Claims

1. A method of classifying service data, characterized by, The method comprises: acquiring user service data; the user service data comprises a user service type feature field and a user identifier field; the service type feature field comprises at least one of a first attribute field, a second attribute field, and a third attribute field; the user identifier field is used for identifying a user; the first attribute field is used for representing a calling service attribute of a user service; the second attribute field is used for representing a called service attribute of a user service; and the third attribute field is used for representing a calling service attribute and a called service attribute of a long term evolution voice bearer (VoLTE) communication service of a user; extracting a target first attribute field, a target second attribute field, a target third attribute field, and a target user identifier field from the user service data; the target user identifier field is any one of the user identifier fields; and the target first attribute field, the target second attribute field, and the target third attribute field are attribute fields of service data of a target user; determining a service type of the target user based on the target first attribute field, the target second attribute field, the target third attribute field, and the target user identifier field; the determination of the service type of the target user based on the target first attribute field, the target second attribute field, the target third attribute field, and the target user identifier field comprises: determining a target user identified by the user identifier field; in a case where the service type feature field comprises one of the first attribute field, the second attribute field, and the third attribute field, determining a service attribute corresponding to the one attribute field as the service type of the target user; in a case where the service type feature field comprises any two of the first attribute field, the second attribute field, and the third attribute field, determining a sum of two service attributes corresponding to the any two attribute fields as the service type of the target user; in a case where the service type feature field comprises the first attribute field, the second attribute field, and the third attribute field, determining a sum of a first service attribute corresponding to the first attribute field, a second service attribute corresponding to the second attribute field, and a third service attribute corresponding to the third attribute field as the service type of the target user.

2. The method of claim 1, wherein, the extraction of the target first attribute field, the target second attribute field, the target third attribute field, and the target user identifier field from the user service data comprises: presetting a first extraction rule of the first attribute field, a second extraction rule of the second attribute field, a third extraction rule of the third attribute field, and a fourth extraction rule of the user identifier field; the first extraction rule, the second extraction rule, the third extraction rule, and the fourth extraction rule are all preset program scripts used for extracting attribute fields in the user service data. Extract the first attribute field, the second attribute field, the third attribute field, and the user identifier field in the user service data based on the first extraction rule, the second extraction rule, the third extraction rule, and the fourth extraction rule.

3. The method according to any of claims 1-2, characterized in that, The user service data is obtained, including: Accessing a first server; the first server is used to store compressed files of user service data; Obtaining the directory of all compressed files in the first server; Based on the first regular expression matching, the compressed file of the user service data is determined; Based on the second regular expression matching, the user service data text file path in the compressed file of the user service data is determined, and the user service data is obtained.

4. The method of claim 3, wherein, The method further comprises: Based on the service data of each user in the first server, the first service data classification statistical data is determined; the first service data classification statistical data includes the service type of each user; The first service data classification statistical data is encapsulated as second service data classification statistical data; the second service data classification statistical data is the first service data classification statistical data encapsulated in JSON format; The second service data classification statistical data is rendered into a visual chart.

5. A service data classification apparatus characterized by comprising: Including: A communication unit and a processing unit; The communication unit is used to obtain user service data; The user service data includes a user service type feature field and a user identifier field; the service type feature field includes at least one of the following: a first attribute field, a second attribute field, and a third attribute field; the user identifier field is used to identify a user; the first attribute field is used to represent the calling service attribute of user service; the second attribute field is used to represent the called service attribute of user service; the third attribute field is used to represent the calling service attribute and the called service attribute of the long term evolution voice bearer (volte) communication service of a user; The processing unit is used to extract the target first attribute field, the target second attribute field, the target third attribute field, and the target user identifier field in the user service data; the target user identifier field is any one of the user identifier fields; the target first attribute field, the target second attribute field, and the target third attribute field are the attribute fields of the service data of a target user; The processing unit is used to determine the service type of the target user based on the target first attribute field, the target second attribute field, the target third attribute field, and the target user identifier field; The processing unit is specifically used to: Determine the target user identified by the user identifier field based on the user identifier field; In the case that the service type feature field includes one attribute field of the first attribute field, the second attribute field, and the third attribute field, determine that the service attribute corresponding to the one attribute field is the service type of the target user; In a case where the service category characteristic field includes any two of the first attribute field, the second attribute field, and the third attribute field, a sum of two service attributes corresponding to the any two attribute fields respectively is determined as the service category of the target user. In a case where the service category characteristic field includes the first attribute field, the second attribute field, and the third attribute field, a sum of a first service attribute corresponding to the first attribute field, a second service attribute corresponding to the second attribute field, and a third service attribute corresponding to the third attribute field is determined as the service category of the target user.

6. The apparatus of claim 5, wherein, The processing unit is specifically configured to: preset a first extraction rule of the first attribute field, a second extraction rule of the second attribute field, a third extraction rule of the third attribute field, and a fourth extraction rule of the user identifier field; the first extraction rule, the second extraction rule, the third extraction rule, and the fourth extraction rule are all preset program scripts for extracting attribute fields in the user service data; extract the first attribute field, the second attribute field, the third attribute field, and the user identifier field in the user service data based on the first extraction rule, the second extraction rule, the third extraction rule, and the fourth extraction rule.

7. A service data classification apparatus characterized by comprising: comprise: a processor and a memory; wherein the memory is configured to store computer execution instructions, when the service data classification device is running, the processor executes the computer execution instructions stored in the memory, so that the service data classification device executes the service data classification method in any one of claims 1-4.

8. A computer-readable storage medium, characterized in that, The computer readable storage medium stores instructions, when the instructions in the computer readable storage medium are executed by the processor of the service data classification, the service data classification device executes the service data classification method in any one of claims 1-4.

Citation Information

Patent Citations

  • Automatic feature online processing method and device for log type data, machine readable medium and equipment

    CN112560353A

  • Service recognition method, device and system and calculation equipment

    CN113010510A