Business data processing method and device, electronic equipment and readable storage medium
By calculating massive business data, the connected graphs are generated and the maximum connected sub-graph is determined, the problem of inefficient manual statistics and management is solved, and fast and accurate data processing and statistics are achieved, and the overall business processing efficiency is improved.
Patent Information
- Application Number
- CN202311763616.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-20
- Publication Date
- 2025-06-20
AI Technical Summary
With the emergence of massive business data, it has become difficult to use manual methods to count and manage business data tables, which are inefficient and prone to errors.
By obtaining target service data, performing graph calculations to generate a connectivity graph, determining the maximum connected sub-graph of each node, and statistics are performed based on these sub-graphs, and statistical results are output.
It realizes the rapid and accurate processing and statistics of massive business data, and improves the overall efficiency of business processing.
Smart Images

Figure CN120179863A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of data processing. Specifically, it relates to graph computing technology in the technical field of data processing. More specifically, it relates to a method, apparatus, electronic device, and readable storage medium for processing business data. Background Art
[0002] Business data refers to various information generated and collected during daily business operations. It can be used to record historical information, analyze development trends, etc. Usually, business data can be stored in a business data table with preset fields, and the entries in the business data table increase as the business data continuously increases.
[0003] However, with the emergence of a large amount of business data, it has become very difficult to manually count and manage business data tables. The whole process is not only inefficient but also error-prone. Summary of the Invention
[0004] To solve the above technical problems, this application provides a method, apparatus, electronic device, and readable storage medium for processing business data to achieve the purpose of improving the overall efficiency of business processing.
[0005] To achieve the above technical objectives, the embodiments of this application provide the following technical solutions:
[0006] In a first aspect, an embodiment of this application provides a method for processing business data. The method for processing business data includes: obtaining target business data; where the target business data includes: multiple pieces of business data with the same parameter items; performing graph computing based on the target business data to generate a connected graph; where each node in the connected graph represents a piece of business data, and each edge represents the association relationship between the business data represented by the connected nodes; for each node in the connected graph, determining the largest connected subgraph containing the node; and based on the largest connected subgraphs of each node, counting the target business data and outputting a statistical result.
[0007] Optionally, the performing graph computing based on the target business data to generate a connected graph includes: generating a point-edge relationship table based on the target business data; where the point-edge relationship table includes: multiple pairs of nodes and their association information, and the association information of each pair of nodes includes the parameter items with similar parameter values in the business data represented by the two nodes, and the similarity value of the similar parameter values; generating the connected graph according to the point-edge relationship table, where the length of the edge in the connected graph is positively correlated with the similarity value of the similar parameter values of the pair of nodes connected by the edge.
[0008] Optionally, when the similar parameter values are composed of Chinese characters, the similarity values of the similar parameter values are obtained through an edit distance algorithm; when the similar parameter values are composed of numbers and / or letters, the similarity values of the similar parameter values are Boolean values.
[0009] Optionally, the identical parameter items include at least one of company name, credit code, file identifier, and settlement card number; wherein, the file identifier is the identifier of the file established for the customer, and the settlement card number is the account number used by the customer for expense settlement.
[0010] Optionally, after obtaining the target business data, the method further includes: when the identical parameter item includes the company name, removing the invalid data of the parameter values under the company name in each piece of the business data according to the invalid data cleaning strategy; wherein, the invalid data cleaning strategy includes at least one of removing geographical location words, removing company attribute words, and removing punctuation marks.
[0011] In a second aspect, an embodiment of the present application provides a processing device for business data. The processing device for business data includes: an acquisition module, configured to acquire target business data; wherein, the target business data includes: multiple pieces of business data with identical parameter items; a graph calculation module, configured to perform graph calculation based on the target business data to generate a connected graph; wherein, each node of the connected graph represents a piece of business data, and each edge represents the association relationship between the business data represented by the connected nodes; a sub-graph calculation module, configured to determine, for each node in the connected graph, the largest connected sub-graph containing the node; and an output module, configured to count the target business data based on the largest connected sub-graphs of the nodes and output a statistical result.
[0012] Optionally, the graph calculation module includes: a relationship table unit, configured to generate a point-edge relationship table based on the target business data; wherein, the point-edge relationship table includes: multiple pairs of nodes and their association information, and the association information of each pair of nodes includes the parameter items with similar parameter values in the business data represented by the two nodes, and the similarity value of the similar parameter values; a graph calculation unit, configured to generate the connected graph according to the point-edge relationship table, wherein the length of the edge in the connected graph is positively correlated with the similarity value of the similar parameter values of the pair of nodes connected by the edge.
[0013] In a third aspect, an embodiment of the present application provides an electronic device, including: a processor and a memory; wherein, the memory is connected to the processor, and the memory is used to store a computer program; the processor is configured to implement the method for processing business data as described in the first aspect by running the computer program stored in the memory.
[0014] Fourthly, an embodiment of the present application provides a computer-readable storage medium, on which a computer program is stored. When the computer program is run by a processor, it implements the method for processing service data as described in the first aspect above.
[0015] Fifthly, an embodiment of the present application provides a computer program product or a computer program. The computer program product includes a computer program, and the computer program is stored in a computer-readable storage medium; a processor of the computer device reads the computer program from the computer-readable storage medium, and when the processor executes the computer program, it implements the steps of the method for processing service data as described in the first aspect.
[0016] In the method for processing service data provided by the present application, after obtaining the target service data to be processed, graph calculation is performed on the target service data to generate a connected graph. Since the target service data includes multiple service data with the same parameter items. Therefore, by means of graph calculation, the association relationship between multiple service data can be represented in the form of a graph. Then, for each node in the connected graph, the largest connected subgraph containing the node is determined; the service data represented by each node in each largest connected subgraph is regarded as a group of service data. Thus, based on the largest connected subgraphs of each node, the statistics of the target service data are realized. Finally, the statistical result is output. In the case of dealing with a large amount of service data that needs to be processed / statistic, compared with the manual processing method, the present application can obtain the statistical result quickly and accurately, thereby improving the overall efficiency of service processing. Description of the Drawings
[0017] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained according to the provided drawings without creative efforts.
[0018] Figure 1 It is a schematic flowchart of a method for processing service data provided by an embodiment of the present application;
[0019] Figure 2 It is an actual application flowchart of a method for processing service data provided by an embodiment of the present application;
[0020] Figure 3 It is one of the schematic diagrams of the connected graph provided by an embodiment of the present application;
[0021] Figure 4 It is another schematic diagram of the connected graph provided by an embodiment of the present application;
[0022] Figure 5 Schematic diagram of the output page provided by an embodiment of the present application;
[0023] Figure 6 Block diagram of the structure of a processing device for service data provided by an embodiment of the present application;
[0024] Figure 7 Schematic diagram of the structure of an electronic device provided by an embodiment of the present application. Detailed implementation manners
[0025] Unless otherwise defined, the technical terms or scientific terms used in the embodiments of the present application should have the ordinary meanings understood by those of ordinary skill in the art to which the present application belongs. The "first", "second" and similar terms used in the embodiments of the present application do not denote any order, quantity or importance, but are only used to avoid confusion of components.
[0026] Unless otherwise required by the context, throughout the specification, "a plurality of" means "at least two", and "including" is interpreted as an open and inclusive meaning, that is, "including, but not limited to". In the description of the specification, the terms "an embodiment", "some embodiments", "exemplary embodiments", "examples", "specific examples" or "some examples", etc. are intended to indicate that specific features, structures, materials or characteristics related to the embodiment or example are included in at least one embodiment or example of the present application. The schematic representations of the above terms do not necessarily refer to the same embodiment or example.
[0027] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.
[0028] Exemplary method
[0029] The embodiments of the present application provide a method for processing service data, as Figure 1 shown, the method for processing service data includes:
[0030] Step S101: Obtain target service data; wherein, the target service data includes: multiple pieces of service data with the same parameter items.
[0031] In this step, the business system can be accessed to obtain the target business data in the business system. For example, if the business system uses a certain business data table to record all business data, the business data table can be obtained through the corresponding interface of the business system, so as to obtain the target business data, but it is not limited to this. It can be understood that the business system usually records business data according to preset parameter items. Therefore, each piece of business data has the same parameter items. Here, the parameter items can also be regarded as fields in the business data table. The specific data content of the business data is not limited in this embodiment.
[0032] Step S102: Perform graph calculation based on the target business data to generate a connected graph.
[0033] In this step, graph calculation can be understood as modeling data in the form of a graph to obtain results that are difficult to obtain from a flattened perspective in the past. The result is the graph, where the graph can be regarded as an abstract data structure used to represent the association relationship between objects, and is described using vertices and edges: vertices represent objects, and edges represent the relationship between objects. Here, the process of graph calculation will not be elaborated in detail. It can be understood that in this embodiment, the connected graph obtained through the foregoing graph calculation is equivalent to the graph obtained by performing graph calculation. The nodes of the connected graph are equivalent to the vertices of the foregoing graph, and the edges of the connected graph are equivalent to the edges of the foregoing graph. Therefore, each node of the connected graph represents a piece of business data, and each edge represents the association relationship between the business data represented by the connected nodes. In this way, the connected graph can not only represent each piece of business data, but also represent the association relationship between each piece of business data.
[0034] Step S103: For each node in the connected graph, determine the largest connected subgraph containing the node.
[0035] In this step, the largest connected subgraph algorithm can be used to calculate the largest connected subgraph of each node in the connected graph. Among them, the largest connected subgraph algorithm includes various algorithms for solving the largest connected subgraph, which will not be elaborated here. It can be understood that the largest connected subgraph of a certain node mentioned above is the largest connected subgraph containing the node. For any largest connected subgraph, all nodes within the largest connected subgraph have the same / similar parameter values under the same parameter items. For example, the same parameter items of multiple pieces of business data include the credit code of the enterprise. Then, the nodes with the same credit code in the connected graph are connected to each other. After determining the corresponding largest connected subgraph, the nodes included in the largest connected subgraph have the same credit code.
[0036] Step S104: Based on the largest connected subgraph of each node, count the target business data and output the statistical result.
[0037] In this step, statistics are separately performed for each maximum connected subgraph, and each node in each maximum connected subgraph is counted as a group of service data. Then, based on each group of service data, the statistical result is determined and output. The statistical result includes, but is not limited to, the number of entries in the service data.
[0038] In the embodiment of the present application, after obtaining the target service data to be processed, graph calculation is performed on the target service data to generate a connected graph. Since the target service data includes multiple service data with the same parameter items, the association relationship between the multiple service data can be represented in the form of a graph by means of graph calculation. Then, for each node in the connected graph, the maximum connected subgraph containing the node is determined; the service data represented by each node in each maximum connected subgraph is regarded as a group of service data. Thus, based on the maximum connected subgraphs of each node, the statistics of the target service data are realized. Finally, the statistical result is output. In the case of dealing with a large amount of service data to be processed / statistical, compared with the manual processing method, the present application can quickly and accurately obtain the statistical result, thereby improving the overall efficiency of service processing.
[0039] In some embodiments, performing graph calculation based on the target service data to generate a connected graph includes:
[0040] Based on the target service data, a point-edge relationship table is generated; the point-edge relationship table includes: multiple pairs of nodes and their association information, and the association information of each pair of nodes includes the parameter items with similar parameter values and the similarity value of the similar parameter values in the service data represented by the two nodes;
[0041] A connected graph is generated according to the point-edge relationship table, where the length of the edge in the connected graph is positively correlated with the similarity value of the similar parameter values of the pair of nodes connected by the edge.
[0042] It should be noted that, for the convenience of generating a connected graph, this embodiment first prepares the data content required for generating the connected graph, that is, nodes and edges (association relationships). Each piece of service data in the target service data is abstracted into a node. The association relationship between the service data represented by a pair (two) of nodes can be abstracted into an edge. Furthermore, a point-edge relationship table including multiple pairs of nodes and their association information is obtained. Then, according to the point-edge relationship table, a connected graph that can represent the content contained in the point-edge relationship table is generated. The connected graph can also be regarded as another form of representation of the point-edge relationship table.
[0043] It can be understood that among a pair of nodes, similar parameter values under the same parameter item can be regarded as the association relationship of this pair of nodes. For example, if the same parameter item is "company name", then if a pair of nodes have similar company names, the association relationship between the two can be having similar company names. In addition, to facilitate the determination of the maximum connected subgraph, the length of the edge can be set according to the similarity value of the similar parameter values. Specifically, the greater the similarity value of the similar parameter values, the longer the corresponding edge length. Conversely, the smaller the similarity value of the similar parameter values, the shorter the corresponding edge length. When determining the maximum connected subgraph, the connected subgraph with the longest total edge length among the connected subgraphs of the connected graph is determined as the maximum connected subgraph.
[0044] In the embodiments of the present application, first, the data content required to generate the connected graph is prepared, that is, nodes and edges (association relationships). Then, using these data contents, the connected graph can be quickly generated.
[0045] In some embodiments, when the similar parameter values are composed of Chinese, the similarity value of the similar parameter values is obtained through the edit distance algorithm;
[0046] When the similar parameter values are composed of numbers and / or letters, the similarity value of the similar parameter values is a Boolean type value.
[0047] It should be noted that when the similar parameter values are composed of Chinese, the similarity value between two Chinese parameter values is calculated by the edit distance (Levenshtein Distance) algorithm. Among them, the edit distance algorithm is an algorithm for comparing the similarity between two strings. For example, if the similar parameter values are XXX Life Services and XXX Smart Property Services respectively. Then, the edit distance dis of the similar parameter values = 3, and the similarity value = 1 - dis / mx(s, t) = 1 - 3 / 9 = 0.67. Where s and t are the similar parameter values.
[0048] The Boolean type value usually has only two results, true and false, which can be represented by 1 and 0 respectively. Therefore, when the similarity value of the similar parameter values is a Boolean type value, a similarity value of 0 can indicate that the similar parameter values are different; a similarity value of 1 can indicate that the similar parameter values are the same.
[0049] In the embodiments of the present application, different methods can be used to measure the similarity value for similar parameter values in different data forms.
[0050] In some embodiments, the same parameter item includes at least one of: company name, credit code, file identifier, and settlement card number;
[0051] Among them, the file identifier is the identifier of the file established for the customer, and the settlement card number is the account used by the customer for fee settlement.
[0052] It should be noted that the credit code is the enterprise credit code and is unique. The file identifier is the identifier of the file established for each customer to which each piece of business data belongs in the business system. Among them, the company name, credit code, file identifier, and settlement card number are all related to the business system corresponding to the target business data. For example, in a certain business system, usually for each cooperative customer, a file is established to record the file identifier, the credit code of the customer, the company name of the customer, and the account used by the customer during transactions. These data contents are stored in the business system as target business data. For these data contents, they can be processed by the method provided in the above embodiments to realize data statistics and management.
[0053] In the embodiments of the present application, for target business data including at least one of the company name, credit code, file identifier, and settlement card number, statistics and management can be realized.
[0054] In order to quickly determine the business data belonging to the same company / enterprise / group, in some embodiments of the present application, after obtaining the target business data, the method further includes:
[0055] When the same parameter item includes the company name, according to the invalid data cleaning strategy, remove the invalid data of the parameter value under the company name in each piece of business data;
[0056] Among them, the invalid data cleaning strategy includes at least one of removing geographical location words, removing company attribute words, and removing punctuation marks.
[0057] It should be noted that during the process of counting the target business data, it is necessary to count the business data belonging to the same company / enterprise / group together. Therefore, these business data should have the same parameter value under the company name parameter item. However, when entering business data, usually for the company name parameter item, the actual parameter value is entered, and even if they belong to the same company, the parameter values may be different. For example, the branch of Company A in place a and the branch of Company A in place b both belong to Company A. When entering business data for the two, the parameter values under the company name parameter item of the two pieces of business data are usually: the branch of Company A in place a and the branch of Company A in place b. In view of this, it is necessary to clean the parameter values under the company name and remove the invalid data that causes interference.
[0058] Specifically, for the removal of geographic location words in the invalid data cleaning strategy, after the invalid data is removed, the remaining content will no longer contain geographic location words. Among them, geographic location words include: the full names or abbreviations of multiple provinces, cities, and counties. For example, the company name before data cleaning is A Company A Branch or A Company B Branch, and the company name after data cleaning is A Company. Similarly, for the removal of company attribute words in the invalid data cleaning strategy, after the invalid data is removed, the remaining content will no longer contain company attribute words. Among them, company attribute words include: branch, sub-branch, bank, bank, business department, stationed in + city, company, group, joint-stock company, limited company, limited liability company. Similarly, for the removal of punctuation marks in the invalid data cleaning strategy, after the invalid data is removed, the remaining content will no longer contain punctuation marks. Among them, punctuation marks include multiple punctuation marks in Chinese and English, such as commas, semicolons, periods, brackets, spaces, etc., but are not limited to this.
[0059] In the embodiment of the present application, by performing data cleaning on the parameter values under the company name parameter item, the business data belonging to the same company / enterprise / group can be quickly determined.
[0060] In some embodiments, the plurality of business data further include a target parameter item whose parameter value is a monthly settlement card number. After outputting the statistical results, the method further includes:
[0061] For each node’s maximum connected subgraph, output the customer group identification;
[0062] Among them, the customer group identifier is the file identifier with the largest number of monthly settlement card numbers in the largest connected subgraph.
[0063] It should be noted that the monthly settlement card number is an account number opened for each customer in the business system for transactions. For example, the customers of a logistics company (cooperative enterprises or commercial users) often need to use a large number of express services, so paying the express fee immediately after using the service may increase management costs. Therefore, the logistics company will provide these customers with a "monthly settlement card number". The enterprise or customer can use this "monthly settlement card number" as a payment account, and then pay all express fees at one time on the agreed date of each month.
[0064] In the above situation, it is usually impossible to directly determine the relationships between customers. For example, each company under a certain group corresponds to different monthly settlement card numbers in the business system. However, it is impossible to directly determine from the business system which companies belong to the same organization. In addition, when companies belonging to the same group conduct transactions / settlements, the monthly settlement card numbers used may not be the ones corresponding to themselves. This situation makes the management of monthly settlement card numbers for the group more difficult. Therefore, it is necessary to use the method provided in the above embodiments to count business data, so as to facilitate management. In particular, it is convenient to manage the monthly settlement card numbers of each company under the group.
[0065] It can be understood that the customers corresponding to all nodes in the largest connected subgraph belong to the same company / enterprise / group. When determining the identifier (customer group identifier) that identifies the company / enterprise / group, the file identifier with the largest number of monthly settlement card numbers can be selected as the customer group identifier directly, but it is not limited to this. A unique customer group identifier can also be generated based on a preset rule. In this embodiment, different customer groups have different customer group identifiers, and the customer group identifier is a unique identifier.
[0066] In the embodiments of the present application, the customer group identifier can be output to facilitate subsequent statistics and management of business data at the group level.
[0067] Next, a specific example will be used to illustrate the method for processing business data provided in the present application. As Figure 2 shown. The method for processing business data includes:
[0068] Step S201: Obtain project code information from the first business system, master data information from the second business system, and monthly settlement customer information from the third business system respectively. Among them, the project code information includes the relevant information of customers who cannot open monthly settlement card numbers; the master data information includes the main data information of customers. The monthly settlement customer information includes the relevant information of customers who have opened monthly settlement card numbers. The information obtained from the three business systems is equivalent to the target business data in the above embodiments of the application. Here, the target business data is demonstrated with specific data content as an example. Among them, the project code information is shown in Table 1 (only part of the data is shown), the master data information is shown in Table 2 (only part of the data is shown), and the monthly settlement customer information is shown in Table 3 (only part of the data is shown).
[0069]
[0070] Table 1
[0071]
[0072] Table 2
[0073]
[0074] Table 3
[0075] Step S202: Generate a point-edge relationship table based on the project code information, master data information, and monthly settlement customer information. The generation process is the same as that of generating the point-edge relationship table in the above embodiments and will not be elaborated here. In this embodiment, when comparing for the same master file, the same settlement card number, and the same credit code, if the two parties being compared are exactly the same, the similarity is 1; otherwise, it is 0. When comparing for the same company, if the target ratio is less than 1, the similarity value = 1 - the target ratio; otherwise, the similarity value is 0. The target ratio is the ratio of the edit distance between the two parties being compared to the maximum length of the two parties being compared. Optionally, additional processing can also be performed according to specific business requirements. For example, for the settlement method of the monthly settlement card number being unified settlement, multiple pre-set monthly settlement affiliated companies can be excluded. Then, determine the company with the largest number of monthly settlement card numbers under the master file number to which the monthly settlement card number belongs as the company name under this master file number, and use it as the company name of the company to which the monthly settlement card number belongs.
[0076] For the point-edge relationship table, it can be as shown in Table 4:
[0077] First node Second node Relationship Similarity value c-1 c-2 Same company 1 c-2 c-3 Same master file 1 c-3 c-4 Same master file 1 c-4 c-5 Same settlement card number 1 c-5 m-1 Same company 1 m-2 c-1 Same credit code 1 m-3 c-4 Same credit code 1 p-1 c-1 Same company 0.57 p-1 c-2 Same company 0.57 p-1 c-4 Same company 0.57 p-2 c-3 Same company 1 p-3 c-2 Same credit code 1 p-3 c-1 Same company 0.57 p-3 c-2 Same company 0.57 p-3 c-4 Same company 0.57 p-3 c-1 Same company 0.57
[0078] Table 4
[0079] Step S203: Generate a connected graph based on the point-edge relationship table. This process is the same as that of generating the connected graph in the above application embodiments and will not be elaborated here. Among them, the generated connected graph can be as Figure 3 shown, where each node is represented by a monthly settlement number, a file number, and a project number. Of course, Figure 3 only shows the connected graph composed of partial business data, which can represent each node and its associated relationships. The relationship of the next business data in the group dimension can be as Figure 4 shown.
[0080] Step S204: Generate the corresponding maximum connected subgraph based on each node. This process is the same as that of generating the maximum connected subgraph in the above application embodiments and will not be elaborated here.
[0081] Step S205: Generate a group identifier. This process is the same as that of generating the customer group identifier in the above application embodiments and will not be elaborated here.
[0082] It should be noted that after generating the group identifier, the statistical results of the target business data based on each maximum connected subgraph will be output / displayed on the same page as the group identifier. The output page is as Figure 5As shown, but not limited to this. Optionally, after adding new business data, if the new business data belongs to the group corresponding to the existing group identifier, there is no need to perform the above steps S201 - S205 for the new business data. If the new business data does not belong to the group corresponding to the existing group identifier, the corresponding statistics and output can be achieved for the new business data according to the above steps.
[0083] In the embodiments of the present application, through the output results, the cooperation situation of each customer under the group in the company's entire network can be understood at the group level; the quotes of each company under the group can be pulled through for early warning; the customer attribution can be clarified at the group level to resolve disputes over competing for customers in multiple regions.
[0084] Exemplary device
[0085] Some embodiments of the present application also provide a processing device for business data, as Figure 6 shown. The processing device for business data includes: an acquisition module 61 for acquiring target business data; where the target business data includes: multiple pieces of business data with the same parameter items; a graph calculation module 62 for performing graph calculation based on the target business data to generate a connected graph; where each node of the connected graph represents a piece of business data, and each edge represents the association relationship between the business data represented by the connected nodes; a sub - graph calculation module 63 for determining, for each node in the connected graph, the largest connected sub - graph containing the node; an output module 64 for statistically analyzing the target business data based on the largest connected sub - graphs of each node and outputting the statistical results.
[0086] In some embodiments, the graph calculation module 62 includes: a relationship table unit for generating a point - edge relationship table based on the target business data; where the point - edge relationship table includes: multiple pairs of nodes and their association information, and the association information of each pair of nodes includes the parameter items with similar parameter values and the similarity value of the similar parameter values in the business data represented by the two nodes; a graph calculation unit for generating a connected graph according to the point - edge relationship table, where the length of the edge in the connected graph is positively correlated with the similarity value of the similar parameter values of the pair of nodes connected by the edge.
[0087] In some embodiments, when the similar parameter values are composed of Chinese characters, the similarity value of the similar parameter values is obtained through an edit distance algorithm; when the similar parameter values are composed of numbers and / or letters, the similarity value of the similar parameter values is a Boolean - type value.
[0088] In some embodiments, the same parameter items include at least one of: company name, credit code, file identifier, and settlement card number; where the file identifier is the identifier of the file established for the customer, and the settlement card number is the account number used by the customer for fee settlement.
[0089] In some embodiments, the apparatus further includes: a data cleaning module, configured to, when the same parameter item includes a company name, remove invalid data of the parameter value under the company name in each piece of business data according to an invalid data cleaning policy; wherein, the invalid data cleaning policy includes at least one of: removing geographical location words, removing company attribute words, and removing punctuation marks.
[0090] In some embodiments, the multiple pieces of business data further include a target parameter item whose parameter value is a monthly settlement card number: the apparatus further includes: a group identification module, configured to output a customer group identification for the largest connected subgraph of each node; wherein, the customer group identification is the file identification with the largest number of monthly settlement card numbers in the largest connected subgraph.
[0091] The processing apparatus for business data provided by the embodiments of the present application belongs to the same inventive concept as the method for processing business data provided by the above embodiments of the present application. For technical details not described in detail in this embodiment, reference may be made to the specific processing content of the method for processing business data provided by the above embodiments of the present application, which will not be elaborated herein.
[0092] Exemplary electronic device
[0093] Another embodiment of the present application further proposes an electronic device. Refer to Figure 7 As shown, an exemplary embodiment of the present application further provides an electronic device, including: a memory and a processor, the memory stores a computer program, and when the processor executes the computer program, it executes the steps in the method for processing business data according to various embodiments of the present application described in the above embodiments of the present application.
[0094] The internal structure of the electronic device may be as Figure 7 As shown, the electronic device includes a processor, a memory, a network interface, and an input device connected through a system bus. Among them, the processor of the electronic device is used to provide computing and control capabilities. The memory of the electronic device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the electronic device is used to communicate with an external terminal through a network connection. The computer program, when executed by the processor, performs the steps in the method for processing business data according to various embodiments of the present application described in the above embodiments of the present application.
[0095] The processor may include a main processor, and may also include a baseband chip, a modem, etc.
[0096] The memory stores a program for implementing the technical solution of the present invention, and may also store an operating system and other critical services. Specifically, the program may include program code, and the program code includes computer operation instructions. More specifically, the memory may include a read-only memory (ROM), other types of static storage devices that can store static information and instructions, a random access memory (RAM), other types of dynamic storage devices that can store information and instructions, a disk memory, a flash memory, and so on.
[0097] The processor may be a general-purpose processor, such as a general-purpose central processing unit (CPU), a microprocessor, etc., or an application-specific integrated circuit (ASIC), or one or more integrated circuits for controlling the execution of the program of the present invention. It may also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.
[0098] The input device may include a device for receiving user input data and information, such as a keyboard, a mouse, a camera, a scanner, a light pen, a voice input device, a touch screen, a pedometer or a gravity sensor, etc.
[0099] The output device may include a device for allowing outputting information to the user, such as a display screen, a printer, a speaker, etc.
[0100] The communication interface may include a device using any transceiver type to communicate with other devices or communication networks, such as Ethernet, a radio access network (RAN), a wireless local area network (WLAN), etc.
[0101] The processor executes the program stored in the memory and calls other devices, which can be used to implement each step of any one of the service data processing methods provided in the above embodiments of the present application.
[0102] The electronic device may further include a display component and a voice component. The display component may be a liquid crystal display screen or an electronic ink display screen. The input device of the electronic device may be a touch layer covered on the display component, or a button, a trackball or a touchpad provided on the housing of the electronic device, or an external keyboard, a touchpad or a mouse, etc.
[0103] Those skilled in the art can understand, Figure 7The structure shown is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the electronic device to which the solution of this application is applied. The specific electronic device may include more or fewer components than those shown in the figure, or combine some components, or have a different component layout.
[0104] Exemplary computer program product and storage medium
[0105] In addition to the above methods and devices, the method for processing service data provided by the embodiments of this application may also be a computer program product, which includes computer program instructions. When the computer program instructions are run by a processor, the processor is caused to execute the steps in the method for processing service data according to various embodiments of this application described in the "Exemplary Method" section above of this application.
[0106] The computer program product can be written in any combination of one or more programming languages to write program code for performing the operations of the embodiments of this application. The programming languages include object-oriented programming languages such as Java, C++, etc., and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computing device, partially on the user's device, executed as an independent software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server.
[0107] In addition, the embodiments of this application also provide a computer-readable storage medium, on which a computer program is stored. The computer program is executed by a processor to perform the steps in the method for processing service data according to various embodiments of this application described in the "Exemplary Method" section above of this application.
[0108] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the embodiments provided in the present application can include non-volatile and / or volatile memories. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM), etc.
[0109] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in the present application.
[0110] The above-described embodiments only represent several implementation manners of the present application. The description is relatively specific and detailed, but it should not be construed as a limitation on the scope of the solutions provided by the embodiments of the present application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the patent of the present application should be subject to the appended claims.
Claims
1. A method for processing service data, characterized in that, The method for processing the service data includes: Obtaining target service data; wherein, the target service data includes: multiple pieces of service data with the same parameter items; Performing graph calculation based on the target service data to generate a connected graph; wherein, each node of the connected graph represents a piece of service data, and each edge represents the association relationship between the service data represented by the connected nodes; For each node in the connected graph, determining the maximum connected subgraph containing the node; Based on the maximum connected subgraphs of each node, statistically analyzing the target service data and outputting a statistical result.
2. The method according to claim 1, characterized in that, The performing graph calculation based on the target service data to generate a connected graph includes: Generating a point-edge relationship table based on the target service data; wherein, the point-edge relationship table includes: multiple pairs of nodes and their association information, and the association information of each pair of nodes includes the parameter items with similar parameter values in the service data represented by the two nodes, and the similarity value of the similar parameter values; Generating the connected graph according to the point-edge relationship table, wherein the length of the edge in the connected graph is positively correlated with the similarity value of the similar parameter values of the pair of nodes connected by the edge.
3. The method according to claim 2, characterized in that, When the similar parameter values are composed of Chinese characters, the similarity value of the similar parameter values is obtained through an edit distance algorithm; When the similar parameter values are composed of numbers and / or letters, the similarity value of the similar parameter values is a Boolean-type value.
4. The method according to claim 1, characterized in that, The same parameter items include at least one of: company name, credit code, file identifier, and settlement card number; Wherein, the file identifier is the identifier of the file established for the customer, and the settlement card number is the account number used by the customer for fee settlement.
5. The method according to claim 4, characterized in that, After obtaining the target service data, the method further includes: When the same parameter item includes the company name, removing the invalid data of the parameter value under the company name in each piece of the service data according to the invalid data cleaning strategy; Wherein, the invalid data cleaning strategy includes at least one of: removing geographical location words, removing company attribute words, and removing punctuation marks.
6. The method according to claim 4, characterized in that, The multiple pieces of service data further include a target parameter item with a monthly settlement card number as the parameter value: After outputting the statistical result, the method further includes: For the maximum connected subgraph of each node, outputting a customer group identifier; Wherein, the customer group identifier is the file identifier with the largest number of monthly settlement card numbers in the maximum connected subgraph.
7. A device for processing service data, characterized in that, The service data processing device includes: An obtaining module, configured to obtain target service data; wherein, the target service data includes: multiple pieces of service data with the same parameter items; A graph calculation module, configured to perform graph calculation based on the target service data to generate a connected graph; wherein, each node of the connected graph represents a piece of service data, and each edge represents the association relationship between the service data represented by the connected nodes; A subgraph calculation module, configured to, for each node in the connected graph, determine the maximum connected subgraph containing the node; An output module, configured to statistically analyze the target service data based on the maximum connected subgraphs of each node and output a statistical result.
8. The device according to claim 7, characterized in that, The graph calculation module includes: A relationship table unit for generating a point-edge relationship table based on the target service data; wherein, the point-edge relationship table includes: multiple pairs of nodes and their associated information, and the associated information of each pair of nodes includes the parameter items with similar parameter values in the service data represented by the two nodes, and the similarity value of the similar parameter values; A graph calculation unit for generating the connected graph according to the point-edge relationship table, wherein the length of the edge in the connected graph is positively correlated with the similarity value of the similar parameter values of a pair of nodes connected by the edge.
9. An electronic device, characterized in that, Comprising: A processor and a memory; Wherein, the memory is connected to the processor, and the memory is used to store a computer program; The processor is used to implement the processing method of the service data as described in any one of claims 1 to 6 by running the computer program stored in the memory.
10. A computer-readable storage medium, characterized in that,A computer program is stored on the computer-readable storage medium, and when the computer program is run by the processor, the processing method of the service data as described in any one of claims 1 to 6 is implemented.