Data analysis method and system for multi-source business data
By determining the key business data and data mining description vectors of multi-source business data, combining data analysis neural networks for vector aggregation, and screening out reliable source business data, the problem of low reliability of user management operations is solved and the reliability of data analysis is improved.
Patent Information
- Application Number
- CN202311282251.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-10-07
- Publication Date
- 2025-09-09
- Estimated Expiration
- 2043-10-07
AI Technical Summary
In the existing technology, the reliability of user management and control operations for multi-source business data is not high, and the reliability of data analysis is insufficient.
By determining the key business data of each source business data, mining the first and second data mining description vectors, combining data analysis neural networks for vector aggregation, and screening out reliable source business data for user management.
The reliability of user management and control operations is improved, the reliability of data analysis is enhanced, and the problem of low reliability in existing technologies is improved.
Smart Images

Figure CN117194525B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a data analysis method and system for multi-source business data. Background Art
[0002] Performing user management and control operations on business users is a common application, for example, user risk management and control for loan business users, and user security management and control for query business users. The basis for performing user management and control operations, i.e., the corresponding business data, generally exists from multiple data sources, resulting in a large amount of business data. However, in existing technologies, user management and control operations are directly performed using this business data, resulting in low reliability of user management and control, and therefore low reliability of data analysis. Summary of the Invention
[0003] In view of this, an object of the present invention is to provide a data analysis method and system for multi-source business data, so as to improve the reliability of data analysis to a certain extent.
[0004] To achieve the above objectives, the embodiments of the present invention adopt the following technical solutions:
[0005] A data analysis method for multi-source business data, comprising:
[0006] Determining key business data corresponding to each of at least two source business data, the key business data including first-category local business data and second-category local business data, a first screening ratio corresponding to the first-category local business data being greater than a second screening ratio corresponding to the second-category local business data, the first-category local business data being screened out from the source business data based on the first screening ratio, and the second-category local business data being screened out from the source business data based on the second screening ratio, the at least two source business data being formed by collecting user information of target business users through corresponding at least two data source platforms, each of the source business data including at least one of image data, audio data, and text data;
[0007] mining, based on the first type of local business data corresponding to each of the at least two source business data, a first data mining description vector corresponding to each of the at least two source business data, wherein the first data mining description vector is formed by performing a deep mining operation based on the description vector to be processed, and the description vector to be processed is formed by performing a feature mining operation on the first type of local business data;
[0008] mining a second data mining description vector corresponding to each of the at least two source business data based on the key business data corresponding to each of the at least two source business data and the description vector to be processed corresponding to each of the at least two source business data;
[0009] Analyzing the filtered source business data from the at least two source business data based on the first data mining description vector corresponding to each source business data of the at least two source business data and the second data mining description vector corresponding to each source business data of the at least two source business data;
[0010] Based on the filtered source business data, user management and control operations are performed on the target business user.
[0011] In some preferred embodiments, in the above-mentioned data analysis method for multi-source business data, the step of mining a first data mining description vector corresponding to each of the at least two source business data based on the first type of local business data corresponding to each of the at least two source business data includes:
[0012] Performing a feature mining operation on a first type of local business data corresponding to the source business data to be processed to form a first number of first description vectors to be processed of the source business data to be processed, where the source business data to be processed is any one of the at least two source business data;
[0013] Analyzing the first importance parameters of the first number of first description vectors to be processed based on the first type of local service data corresponding to the source service data to be processed;
[0014] A first data mining description vector corresponding to the source business data to be processed is determined based on the first number of first description vectors to be processed and the first importance parameters of the first number of first description vectors to be processed.
[0015] In some preferred embodiments, in the above-mentioned data analysis method for multi-source business data, the step of mining a second data mining description vector corresponding to each of the at least two source business data based on the key business data corresponding to each of the at least two source business data and the description vector to be processed corresponding to each of the at least two source business data includes:
[0016] Performing a feature mining operation on the key business data corresponding to the source business data to be processed to form a second number of second description vectors to be processed of the source business data to be processed;
[0017] Determining, based on the key business data corresponding to the source business data to be processed, second importance parameters of the second number of second description vectors to be processed and second importance parameters of the first number of first description vectors to be processed;
[0018] Based on the second importance parameters of the second number of second description vectors to be processed, the second importance parameters of the first number of first description vectors to be processed, the second number of second description vectors to be processed and the first number of first description vectors to be processed, the second data mining description vector corresponding to the source business data to be processed is determined.
[0019] In some preferred embodiments, in the above-mentioned data analysis method for multi-source business data, the step of determining the second importance parameters of the second number of second description vectors to be processed and the second importance parameters of the first number of first description vectors to be processed based on the key business data corresponding to the source business data to be processed includes:
[0020] Analyzing the second importance parameters of the second number of second description vectors to be processed and the second importance parameters of the first number of first description vectors to be processed based on the key business data corresponding to the source business data to be processed and the data source application parameters corresponding to the source business data to be processed, wherein the data source application parameters are used to reflect the data application status of the data source of the corresponding source business data;
[0021] The step of analyzing the second importance parameters of the second number of second description vectors to be processed and the second importance parameters of the first number of first description vectors to be processed based on the key business data corresponding to the source business data to be processed and the data source application parameters corresponding to the source business data to be processed includes:
[0022] Performing a combination operation on the key business data corresponding to the source business data to be processed and the data source application parameters corresponding to the source business data to be processed to form combined data to be processed corresponding to the source business data to be processed;
[0023] The second importance parameters of the second number of second description vectors to be processed and the second importance parameters of the first number of first description vectors to be processed are analyzed based on the combination data to be processed corresponding to the source business data to be processed.
[0024] In some preferred embodiments, in the above-mentioned data analysis method for multi-source business data, the step of analyzing the filtered source business data from the at least two source business data based on the first data mining description vector corresponding to each source business data of the at least two source business data and the second data mining description vector corresponding to each source business data of the at least two source business data includes:
[0025] performing a vector aggregation operation on a first data mining description vector corresponding to each source business data of the at least two source business data and a second data mining description vector corresponding to each source business data of the at least two source business data to form an aggregated data mining description vector corresponding to each source business data of the at least two source business data;
[0026] Analyzing, based on the aggregated data mining description vector corresponding to each of the at least two source business data, a reliability analysis result corresponding to each of the at least two source business data, the reliability analysis result being a predictive characterization parameter of the reliability of performing a user management and control operation on the target business user based on the source business data;
[0027] According to the reliability analysis result corresponding to each source business data of the at least two source business data, the filtered source business data of the at least two source business data are analyzed.
[0028] In some preferred embodiments, in the above-mentioned data analysis method for multi-source business data, the step of performing a vector aggregation operation on the first data mining description vector corresponding to each of the at least two source business data and the second data mining description vector corresponding to each of the at least two source business data to form an aggregated data mining description vector corresponding to each of the at least two source business data includes:
[0029] Determining, based on key business data corresponding to source business data to be analyzed, a third importance parameter of a second data mining description vector corresponding to the source business data to be analyzed, wherein the source business data to be analyzed is any one of the at least two source business data;
[0030] Based on the third importance parameter of the second data mining description vector corresponding to the source business data to be analyzed, the first data mining description vector corresponding to the source business data to be analyzed and the second data mining description vector corresponding to the source business data to be analyzed are aggregated to form an aggregated data mining description vector corresponding to the source business data to be analyzed.
[0031] In some preferred embodiments, in the above-mentioned data analysis method for multi-source business data, the step of determining the third importance parameter of the second data mining description vector corresponding to the source business data to be analyzed based on the key business data corresponding to the source business data to be analyzed includes:
[0032] Analyzing a third importance parameter of the second data mining description vector corresponding to the source business data to be analyzed based on the key business data corresponding to the source business data to be analyzed and the data source application parameter corresponding to the source business data to be analyzed, wherein the data source application parameter is used to reflect the data application status of the data source of the corresponding source business data;
[0033] The step of analyzing the third importance parameter of the second data mining description vector corresponding to the source business data to be analyzed based on the key business data corresponding to the source business data to be analyzed and the data source application parameters corresponding to the source business data to be analyzed includes:
[0034] Performing a combination operation on the key business data corresponding to the source business data to be analyzed and the data source application parameters of the source business data to be analyzed to form combined data to be analyzed corresponding to the source business data to be analyzed;
[0035] According to the combined data to be analyzed corresponding to the source business data to be analyzed, a third importance parameter of the second data mining description vector corresponding to the source business data to be analyzed is analyzed.
[0036] In some preferred embodiments, in the above-mentioned data analysis method for multi-source business data, the step of aggregating the first data mining description vector corresponding to the source business data to be analyzed and the second data mining description vector corresponding to the source business data to be analyzed based on the third importance parameter of the second data mining description vector corresponding to the source business data to be analyzed to form an aggregated data mining description vector corresponding to the source business data to be analyzed includes:
[0037] performing a multiplication operation on the second data mining description vector of the source business data to be analyzed according to the third importance parameter of the second data mining description vector corresponding to the source business data to be analyzed, so as to form an adjusted data mining description vector corresponding to the source business data to be analyzed;
[0038] A superposition operation is performed on the adjusted data mining description vector corresponding to the source business data to be analyzed and the first data mining description vector corresponding to the source business data to be analyzed to form an aggregated data mining description vector corresponding to the source business data to be analyzed.
[0039] In some preferred embodiments, in the above-mentioned data analysis method for multi-source business data, the step of mining a first data mining description vector corresponding to each of the at least two source business data based on the first type of local business data corresponding to each of the at least two source business data includes:
[0040] Using a front-end feature mining unit in a data analysis neural network, performing a feature mining operation on the first type of local business data corresponding to each of the at least two source business data to form a first data mining description vector corresponding to each of the at least two source business data;
[0041] The step of mining a second data mining description vector corresponding to each of the at least two source business data based on the key business data corresponding to each of the at least two source business data and the description vector to be processed corresponding to each of the at least two source business data includes:
[0042] Using a back-end feature mining unit in the data analysis neural network, a fusion mining operation is performed on the key business data corresponding to each of the two source business data and the description vector to be processed corresponding to each of the at least two source business data, so as to output a second data mining description vector corresponding to each of the at least two source business data;
[0043] The step of performing a vector aggregation operation on the first data mining description vector corresponding to each of the at least two source business data and the second data mining description vector corresponding to each of the at least two source business data to form an aggregated data mining description vector corresponding to each of the at least two source business data includes:
[0044] Utilizing a vector aggregation unit in the data analysis neural network, performing an aggregation operation on a first data mining description vector corresponding to each source business data of the at least two source business data and a second data mining description vector corresponding to each source business data of the at least two source business data to form an aggregated data mining description vector corresponding to each source business data of the at least two source business data;
[0045] The step of analyzing the reliability analysis result corresponding to each of the at least two source business data based on the aggregated data mining description vector corresponding to each of the at least two source business data includes:
[0046] Utilizing the data possibility analysis unit in the data analysis neural network, the aggregated data mining description vector corresponding to each of the at least two source business data is analyzed and output to obtain a reliability analysis result corresponding to each of the at least two source business data.
[0047] An embodiment of the present invention also provides a data analysis system for multi-source business data, including a processor and a memory, wherein the memory is used to store a computer program, and the processor is used to execute the computer program to implement the above-mentioned data analysis method for multi-source business data.
[0048] The data analysis method and system for multi-source business data provided by the embodiment of the present invention can first determine the key business data corresponding to each source business data; mine the first data mining description vector corresponding to each source business data based on the first type of local business data corresponding to each source business data; mine the second data mining description vector corresponding to each source business data based on the key business data and the description vector to be processed corresponding to each source business data; analyze the filtered source business data in at least two source business data based on the first data mining description vector corresponding to each source business data and the second data mining description vector corresponding to each source business data; and perform user management operations on target business users based on the filtered source business data. Based on the above content, since the filtered source business data will be filtered out before the user management operation is performed, the basis for the user management operation is more reliable. In addition, since the key business data of the first type of local business data and the second type of local business data will be referenced during the screening process of the filtered source business data, the reliability of the screening is also higher. Therefore, the reliability of data analysis (the reliability of user management) is improved to a certain extent, thereby improving the problem of low reliability in the existing technology.
[0049] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, preferred embodiments are given below and described in detail with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0050] Figure 1 This is a structural block diagram of a data analysis system for multi-source business data provided by an embodiment of the present invention.
[0051] Figure 2 A flowchart of the steps of a data analysis method for multi-source business data provided in an embodiment of the present invention.
[0052] Figure 3 A schematic diagram of various modules included in a data analysis device for multi-source business data provided by an embodiment of the present invention. Implementation Method
[0053] To make the objectives, technical solutions, and advantages of the embodiments of the present invention more clear, the technical solutions of the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Generally, the components of the embodiments of the present invention described and shown in the drawings herein can be arranged and designed in various different configurations.
[0054] Therefore, the following detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the invention as claimed, but rather merely represents selected embodiments of the present invention. All other embodiments derived by persons of ordinary skill in the art based on the embodiments of the present invention without creative effort shall fall within the scope of protection of the present invention.
[0055] like Figure 1 As shown, an embodiment of the present invention provides a data analysis system for multi-source business data. The data analysis system may include a memory and a processor.
[0056] Specifically, the memory and processor are electrically connected, directly or indirectly, to enable data transmission or interaction. For example, the electrical connection may be achieved via one or more communication buses or signal lines. The memory may store at least one software function module (computer program) in the form of software or firmware. The processor may be configured to execute the executable computer program stored in the memory, thereby implementing the data analysis method for multi-source business data provided in an embodiment of the present invention.
[0057] It should be understood that, in one possible embodiment, the memory may be, but is not limited to, random access memory (RAM), read only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), etc.
[0058] It should be understood that, in one possible implementation, the processor may be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), a system on chip (SoC), etc.; it may also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components.
[0059] It should be understood that, in a possible implementation, the data analysis system for multi-source business data may be a server with data processing capabilities.
[0060] Combine Figure 2 Embodiments of the present invention further provide a data analysis method for multi-source business data, which can be applied to the aforementioned data analysis system for multi-source business data. The method steps defined in the process related to the data analysis method for multi-source business data can be implemented by the data analysis system for multi-source business data.
[0061] The following will Figure 2 The specific process shown is explained in detail.
[0062] Step S110 : determining the key business data corresponding to each of at least two source business data.
[0063] In an embodiment of the present invention, the data analysis system for multi-source business data can determine key business data corresponding to each of at least two source business data. The key business data includes a first type of local business data and a second type of local business data, wherein a first screening ratio corresponding to the first type of local business data is greater than a second screening ratio corresponding to the second type of local business data, the first type of local business data is screened from the source business data based on the first screening ratio, and the second type of local business data is screened from the source business data based on the second screening ratio. The at least two source business data are formed by collecting user information from target business users through at least two corresponding data source platforms, and each source business data includes at least one of image data, audio data, and text data.
[0064] Step S120 : mining a first data mining description vector corresponding to each source business data in the at least two source business data according to the first type of local business data corresponding to each source business data in the at least two source business data.
[0065] In an embodiment of the present invention, the data analysis system for multi-source business data can mine a first data mining description vector corresponding to each of the at least two source business data based on the first type of local business data corresponding to each of the at least two source business data. The first data mining description vector is formed by performing a deep mining operation based on the description vector to be processed, which is formed by performing a feature mining operation on the first type of local business data.
[0066] Step S130, based on the key business data corresponding to each of the at least two source business data and the description vector to be processed corresponding to each of the at least two source business data, mine the second data mining description vector corresponding to each of the at least two source business data.
[0067] In an embodiment of the present invention, the data analysis system for multi-source business data can mine a second data mining description vector corresponding to each of the at least two source business data based on the key business data corresponding to each of the at least two source business data and the description vector to be processed corresponding to each of the at least two source business data. Thus, when mining the second data mining description vector for the at least two source business data, in addition to using the key business data corresponding to the at least two source business data, the description vector to be processed corresponding to the at least two source business data is also utilized, thereby enabling the learning of multi-level information in the source business data.
[0068] Step S140: Analyze the filtered source business data from the at least two source business data based on the first data mining description vector corresponding to each source business data in the at least two source business data and the second data mining description vector corresponding to each source business data in the at least two source business data.
[0069] In an embodiment of the present invention, the data analysis system for multi-source business data can analyze the filtered source business data from the at least two source business data based on the first data mining description vector corresponding to each of the at least two source business data and the second data mining description vector corresponding to each of the at least two source business data.
[0070] Step S150: performing user management and control operations on the target business user based on the filtered source business data.
[0071] In an embodiment of the present invention, the data analysis system for multi-source business data can perform user management operations on the target business users based on the filtered source business data. For example, user risk analysis can be performed on the filtered source business data based on a user risk analysis neural network formed through network optimization to obtain corresponding user risk characterization data. Then, user management operations can be performed on the target business users based on the user risk characterization data, such as marking the target business users as risky users or normal users.
[0072] Based on the aforementioned content (i.e., the aforementioned steps S110-S150), since the source business data will be filtered out before the user control operation is performed, the basis for the user control operation is more reliable. In addition, since the key business data of the first category of local business data and the second category of local business data will be referenced during the screening process of the source business data, the reliability of the screening is also higher. Therefore, the reliability of data analysis (the reliability of user control) is improved to a certain extent, thereby improving the problem of low reliability existing in the existing technology.
[0073] It should be understood that, in a possible implementation, step S110 described above, i.e., the step of determining the key business data corresponding to each of the at least two source business data, may further include the following specific implementation steps:
[0074] For any one of the at least two source business data, performing a feature space mapping operation on the source business data to form a corresponding business data mapping vector;
[0075] Based on the first filter matrix, performing a sliding window filtering operation on the service data mapping vector to form a corresponding first sliding window filtering vector;
[0076] performing a sliding window filtering operation on the service data mapping vector based on a second filter matrix to form a corresponding second sliding window filter vector, wherein a size of the first filter matrix is smaller than a size of the second filter matrix (a first screening ratio corresponding to the first type of local service data is larger than a second screening ratio corresponding to the second type of local service data), such that the number of vector parameters included in the first sliding window filter vector is larger than the number of vector parameters included in the second sliding window filter vector;
[0077] The first sliding window filter vector is used as the first type of local business data included in the key business data, and the second sliding window filter vector is used as the second type of local business data included in the key business data to form the key business data corresponding to the source business data.
[0078] It should be understood that, in a possible implementation, step S120 described above, i.e., the step of mining a first data mining description vector corresponding to each of the at least two source business data based on the first type of local business data corresponding to each of the at least two source business data, may further include the following specific implementation steps:
[0079] Performing a feature mining operation on first-category local business data corresponding to the source business data to be processed to form a first number of first description vectors to be processed of the source business data to be processed, where the source business data to be processed is any one of the at least two source business data. For example, each of the at least two source business data can be used sequentially or in parallel as the source business data to be processed to form corresponding processing. For example, for the first number of first description vectors to be processed, a first number of feature mining units can be used to perform feature mining operations on the first-category local business data, such as performing a convolution operation to obtain them.
[0080] Analyze the first importance parameters of the first number of first description vectors to be processed based on the first type of local business data corresponding to the source business data to be processed. Exemplarily, the first number of feature mining units may correspond to a first number of importance analysis units. Thus, the first number of importance analysis units may be used to analyze and output the first type of local business data respectively to obtain the first importance parameters of the first number of first description vectors to be processed. The importance analysis unit may include a function such as softmax.
[0081] Based on the first number of first description vectors to be processed and the first importance parameters of the first number of description vectors to be processed, the first data mining description vector corresponding to the source business data to be processed is determined. For example, based on the first importance parameters of the first number of first description vectors to be processed, the first number of first description vectors to be processed can be subjected to a superposition operation (i.e., the first importance parameters are used as weighting coefficients for weighted superposition).
[0082] It should be understood that, in a possible implementation scheme, step S130 described above, i.e., the step of mining a second data mining description vector corresponding to each of the at least two source business data based on the key business data corresponding to each of the at least two source business data and the description vector to be processed corresponding to each of the at least two source business data, may further include the following specific implementation steps:
[0083] Performing a feature mining operation on the key business data corresponding to the source business data to be processed to form a second number of second description vectors to be processed of the source business data to be processed. Exemplarily, a second number of feature mining units may be used to perform feature mining operations on the key business data to form a second number of second description vectors to be processed.
[0084] Determining, based on the key business data corresponding to the source business data to be processed, second importance parameters of the second number of second description vectors to be processed and second importance parameters of the first number of first description vectors to be processed;
[0085] Based on the second importance parameters of the second number of second description vectors to be processed, the second importance parameters of the first number of first description vectors to be processed, the second number of second description vectors to be processed and the first number of first description vectors to be processed, the second data mining description vector corresponding to the source business data to be processed is determined. For example, based on the second importance parameters of the second number of second description vectors to be processed and the second importance parameters of the first number of first description vectors to be processed, a weighted superposition operation can be performed on the second number of second description vectors to be processed and the first number of first description vectors to be processed.
[0086] It should be understood that, in a possible implementation scheme, the step of determining the second importance parameters of the second number of second description vectors to be processed and the second importance parameters of the first number of first description vectors to be processed based on the key business data corresponding to the source business data to be processed may further include the following specific implementation steps:
[0087] Based on the key business data corresponding to the source business data to be processed and the data source application parameters corresponding to the source business data to be processed, the second importance parameters of the second number of second description vectors to be processed and the second importance parameters of the first number of first description vectors to be processed are analyzed. The data source application parameters are used to reflect the data application status of the data source of the corresponding source business data, such as the positive correlation between the number of times the data source has been used as a screening source business data in history.
[0088] It should be understood that, in a possible implementation scheme, the step of analyzing the second importance parameters of the second number of second description vectors to be processed and the second importance parameters of the first number of first description vectors to be processed based on the key business data corresponding to the source business data to be processed and the data source application parameters corresponding to the source business data to be processed may further include the following specific implementation steps:
[0089] Performing a combination operation on the key business data corresponding to the source business data to be processed and the data source application parameters corresponding to the source business data to be processed to form the combined data to be processed corresponding to the source business data to be processed; when the key business data belongs to a vector, the data source application parameters can be encoded into a vector, and then the vector can be aggregated, such as a cascade combination operation, to form a corresponding aggregated vector as the combined data to be processed; when the key business data does not belong to a vector, the key business data and the data source application parameters can be combined, and then the obtained data combination is encoded to form the combined data to be processed corresponding to the source business data to be processed;
[0090] Based on the processed combined data corresponding to the processed source business data, the second importance parameters of the second number of second description vectors to be processed and the second importance parameters of the first number of first description vectors to be processed are analyzed. For example, based on the configured multiple importance analysis units, the processed combined data can be analyzed and output respectively to obtain the second importance parameters of the second number of second description vectors to be processed and the second importance parameters of the first number of first description vectors to be processed. In an embodiment of the present invention, the specific parameters of each importance analysis unit can be implemented in the corresponding network optimization process.
[0091] It should be understood that, in a possible implementation scheme, step S140 in the above description, i.e., the step of analyzing the filtered source business data from the at least two source business data based on the first data mining description vector corresponding to each source business data in the at least two source business data and the second data mining description vector corresponding to each source business data in the at least two source business data, may further include a specific implementation step described below:
[0092] performing a vector aggregation operation on a first data mining description vector corresponding to each source business data of the at least two source business data and a second data mining description vector corresponding to each source business data of the at least two source business data to form an aggregated data mining description vector corresponding to each source business data of the at least two source business data, that is, performing a vector aggregation operation on the corresponding first data mining description vector and the second data mining description vector;
[0093] Analyzing, based on the aggregated data mining description vector corresponding to each of the at least two source business data, a reliability analysis result corresponding to each of the at least two source business data, the reliability analysis result being a predictive characterization parameter of the reliability of performing a user management and control operation on the target business user based on the source business data. For example, the aggregated data mining description vector may be analyzed based on a processing unit (which may include a softmax function, etc.) included in a corresponding neural network to obtain a corresponding reliability analysis result.
[0094] Based on the reliability analysis results corresponding to each of the at least two source business data, the filtered source business data in the at least two source business data are analyzed. For example, one or more source business data with the largest prediction characterization parameters represented by the corresponding reliability analysis results can be used as the filtered source business data.
[0095] It should be understood that, in a possible implementation scheme, the step of performing a vector aggregation operation on the first data mining description vector corresponding to each of the at least two source business data and the second data mining description vector corresponding to each of the at least two source business data to form an aggregated data mining description vector corresponding to each of the at least two source business data may further include a specific implementation step described below:
[0096] determining, based on key business data corresponding to source business data to be analyzed, a third importance parameter of a second data mining description vector corresponding to the source business data to be analyzed, wherein the source business data to be analyzed is any one of the at least two source business data, and each of the at least two source business data can be used sequentially or in parallel as the source business data to be analyzed for subsequent corresponding processing;
[0097] Based on the third importance parameter of the second data mining description vector corresponding to the source business data to be analyzed, the first data mining description vector corresponding to the source business data to be analyzed and the second data mining description vector corresponding to the source business data to be analyzed are aggregated to form an aggregated data mining description vector corresponding to the source business data to be analyzed. In this way, the aggregated data mining description vector can carry information of both the first data mining description vector and the second data mining description vector, and has stronger expressive ability.
[0098] It should be understood that, in a possible implementation scheme, the step of determining, based on the key business data corresponding to the source business data to be analyzed, the third importance parameter of the second data mining description vector corresponding to the source business data to be analyzed may further include the following specific implementation steps:
[0099] Based on the key business data corresponding to the source business data to be analyzed and the data source application parameters corresponding to the source business data to be analyzed, the third importance parameter of the second data mining description vector corresponding to the source business data to be analyzed is analyzed. The data source application parameter is used to reflect the data application status of the data source of the corresponding source business data, as described above.
[0100] It should be understood that, in a possible implementation scheme, the step of analyzing the third importance parameter of the second data mining description vector corresponding to the source business data to be analyzed based on the key business data corresponding to the source business data to be analyzed and the data source application parameters corresponding to the source business data to be analyzed may further include the following specific implementation steps:
[0101] Performing a combination operation on the key business data corresponding to the source business data to be analyzed and the data source application parameters of the source business data to be analyzed to form combined data to be analyzed corresponding to the source business data to be analyzed, as described above;
[0102] According to the combined data to be analyzed corresponding to the source business data to be analyzed, the third importance parameter of the second data mining description vector corresponding to the source business data to be analyzed is analyzed. The specific processing process can refer to the relevant description in the above text.
[0103] It should be understood that, in a possible implementation scheme, the step of aggregating the first data mining description vector corresponding to the source business data to be analyzed and the second data mining description vector corresponding to the source business data to be analyzed based on the third importance parameter of the second data mining description vector corresponding to the source business data to be analyzed to form an aggregated data mining description vector corresponding to the source business data to be analyzed may further include a specific implementation step described below:
[0104] performing a multiplication operation on the second data mining description vector of the source business data to be analyzed according to the third importance parameter of the second data mining description vector corresponding to the source business data to be analyzed, so as to form an adjusted data mining description vector corresponding to the source business data to be analyzed;
[0105] A superposition operation is performed on the adjusted data mining description vector corresponding to the source business data to be analyzed and the first data mining description vector corresponding to the source business data to be analyzed to form an aggregated data mining description vector corresponding to the source business data to be analyzed.
[0106] It should be understood that, in a possible implementation scheme, the step of mining out the first data mining description vector corresponding to each of the at least two source business data based on the first type of local business data corresponding to each of the at least two source business data includes: utilizing the front-end feature mining unit in the data analysis neural network to perform feature mining operations on the first type of local business data corresponding to each of the at least two source business data to form a first data mining description vector corresponding to each of the at least two source business data. The specific processing process can be as described above.
[0107] It should be understood that, in a possible implementation scheme, the step of mining out the second data mining description vector corresponding to each of the at least two source business data based on the key business data corresponding to each of the at least two source business data and the description vector to be processed corresponding to each of the at least two source business data includes: utilizing the back-end feature mining unit in the data analysis neural network to perform a fusion mining operation on the key business data corresponding to each of the two source business data and the description vector to be processed corresponding to each of the at least two source business data, so as to output the second data mining description vector corresponding to each of the at least two source business data. The specific processing process can be as described above.
[0108] It should be understood that, in a possible implementation scheme, the step of performing a vector aggregation operation on the first data mining description vector corresponding to each of the at least two source business data and the second data mining description vector corresponding to each of the at least two source business data to form an aggregated data mining description vector corresponding to each of the at least two source business data includes: utilizing the vector aggregation unit in the data analysis neural network to perform an aggregation operation on the first data mining description vector corresponding to each of the at least two source business data and the second data mining description vector corresponding to each of the at least two source business data to form an aggregated data mining description vector corresponding to each of the at least two source business data. The specific processing process can be as described above.
[0109] It should be understood that, in a possible implementation scheme, the step of analyzing the reliability analysis results corresponding to each of the at least two source business data based on the aggregated data mining description vector corresponding to each of the at least two source business data includes: using the data possibility analysis unit in the data analysis neural network to analyze and output the aggregated data mining description vector corresponding to each of the at least two source business data to obtain the reliability analysis results corresponding to each of the at least two source business data. The specific processing process can be as described above.
[0110] It should be understood that, in a possible implementation, before the step of determining the key business data corresponding to each of the at least two source business data, that is, before step S110, the data analysis method for multi-source business data may further include the following specific implementation steps:
[0111] Determine the key business data corresponding to the exemplary source business data, such as the relevant content above;
[0112] Using the front-end feature mining unit, perform a feature mining operation on the first type of local business data corresponding to the exemplary source business data to form a first data mining description vector corresponding to the exemplary source business data, as described above;
[0113] Using the back-end feature mining unit, a feature mining operation is performed on the key business data corresponding to the exemplary source business data and the description vector to be processed corresponding to the exemplary source business data to form a second data mining description vector corresponding to the exemplary source business data, as described above.
[0114] Utilizing the vector aggregation unit, performing an aggregation operation on the first data mining description vector corresponding to the exemplary source business data and the second data mining description vector corresponding to the exemplary source business data to form an aggregated data mining description vector corresponding to the exemplary source business data, as described above;
[0115] Utilizing the data possibility analysis unit in the data analysis neural network, the aggregated data mining description vector corresponding to the exemplary source business data is analyzed and output to obtain a reliability analysis result corresponding to the exemplary source business data, as described above.
[0116] Analyze the corresponding network optimization error index based on the reliability analysis result corresponding to the exemplary source business data, the actual reliability parameter corresponding to the exemplary source business data, and the optimization importance parameter of the exemplary source business data. There is a negative correlation between the optimization importance parameter and the data source application parameter of the exemplary source business data. The actual reliability parameter is an actual representation parameter of the reliability of performing user management and control operations on the corresponding exemplary business user based on the exemplary source business data.
[0117] Based on the network optimization error index, the data analysis neural network is subjected to a network optimization operation to form an optimized data analysis neural network. That is to say, the network parameters of the data analysis neural network can be optimized and adjusted in the direction of reducing the network optimization error index to form an optimized data analysis neural network.
[0118] It should be understood that, in one possible implementation, the step of analyzing the corresponding network optimization error index based on the reliability analysis result corresponding to the exemplary source service data, the actual reliability parameter corresponding to the exemplary source service data, and the optimization importance parameter of the exemplary source service data may further include the following specific implementation steps:
[0119] Extracting a pre-configured first parameter and a second parameter, illustratively, the first parameter may be equal to 1, and the second parameter may be equal to 0;
[0120] Calculating a difference between the first parameter and a reliability analysis result corresponding to the exemplary source business data to obtain a first difference, and calculating a difference between the first parameter and an actual reliability parameter corresponding to the exemplary source business data to obtain a second difference;
[0121] Performing a logarithm operation on a reliability analysis result corresponding to the exemplary source service data to obtain a corresponding first logarithm operation result, and performing a de-logarithm operation on the first difference to obtain a corresponding second logarithm operation result;
[0122] Using the actual reliability parameter corresponding to the exemplary source service data as a weighting coefficient of the first logarithmic operation result, and using the second difference as the second logarithmic operation result, to perform a weighted summation calculation to obtain a local optimization error index corresponding to the exemplary source service data;
[0123] In the case of multiple exemplary source business data, the optimization importance parameters corresponding to each of the exemplary source business data are extracted, and based on the optimization importance parameters corresponding to each of the exemplary source business data, a weighted summation calculation is performed on the local optimization error index corresponding to each of the exemplary source business data, and based on the result of the weighted summation calculation and the second parameter, the corresponding network optimization error index is calculated, for example, the sum of the network optimization error index and the result of the weighted summation calculation is equal to the second parameter.
[0124] It should be understood that, in a possible implementation, the step of calculating the optimized importance parameter corresponding to the exemplary source service data may include:
[0125] Analyze data application of the data source of the exemplary source business data to obtain the number of times the data source has been used as screening source business data in history;
[0126] The number of times and the first parameter are summed to obtain a target sum, and a negatively correlated optimization importance parameter is calculated based on the target sum. For example, a power operation (such as 0.5 power, etc.) is performed on the target sum, and then the ratio between the first parameter and the power operation result is calculated to obtain the corresponding optimization importance parameter.
[0127] Combine Figure 3 The embodiment of the present invention further provides a data analysis device for multi-source business data, which can be applied to the above-mentioned data analysis system for multi-source business data. The data analysis device for multi-source business data may include:
[0128] a key data determination module, configured to determine key business data corresponding to each of at least two source business data, the key business data including a first category of local business data and a second category of local business data, the first category of local business data corresponding to a first screening ratio being greater than a second screening ratio corresponding to the second category of local business data, the first category of local business data being screened from the source business data based on the first screening ratio, and the second category of local business data being screened from the source business data based on the second screening ratio, the at least two source business data being formed by collecting user information of target business users via corresponding at least two data source platforms, each of the source business data including at least one of image data, audio data, and text data;
[0129] a first data mining module, configured to mine, based on first-category local business data corresponding to each of the at least two source business data, a first data mining description vector corresponding to each of the at least two source business data, wherein the first data mining description vector is formed by performing a deep mining operation on the description vector to be processed, and the description vector to be processed is formed by performing a feature mining operation on the first-category local business data;
[0130] a second data mining module, configured to mine a second data mining description vector corresponding to each of the at least two source business data based on the key business data corresponding to each of the at least two source business data and the description vector to be processed corresponding to each of the at least two source business data;
[0131] a business data screening module, configured to analyze the screened source business data from the at least two source business data based on a first data mining description vector corresponding to each source business data of the at least two source business data and a second data mining description vector corresponding to each source business data of the at least two source business data;
[0132] The user management and control module is used to perform user management and control operations on the target business users based on the filtered source business data.
[0133] In summary, the data analysis method and system for multi-source business data provided by the present invention can first determine the key business data corresponding to each source business data; based on the first type of local business data corresponding to each source business data, mine the first data mining description vector corresponding to each source business data; based on the key business data and the description vector to be processed corresponding to each source business data, mine the second data mining description vector corresponding to each source business data; based on the first data mining description vector corresponding to each source business data and the second data mining description vector corresponding to each source business data, analyze the filtered source business data in at least two source business data; based on the filtered source business data, perform user management operations on the target business user. Based on the above content, since the filtered source business data will be filtered out before the user management operation is performed, the basis for the user management operation is more reliable. In addition, since the key business data of the first type of local business data and the second type of local business data will be referenced during the screening process of the filtered source business data, the reliability of the screening is also higher. Therefore, the reliability of data analysis (the reliability of user management) is improved to a certain extent, thereby improving the problem of low reliability in the existing technology.
[0134] The foregoing description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Those skilled in the art will readily appreciate that various modifications and variations of the present invention are possible. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present invention are intended to be within the scope of protection of the present invention.
Claims
1. A data analysis method for multi-source business data, characterized in that: include: Determining key business data corresponding to each of at least two source business data, the key business data including first-category local business data and second-category local business data, a first screening ratio corresponding to the first-category local business data being greater than a second screening ratio corresponding to the second-category local business data, the first-category local business data being screened out from the source business data based on the first screening ratio, and the second-category local business data being screened out from the source business data based on the second screening ratio, the at least two source business data being formed by collecting user information of target business users through corresponding at least two data source platforms, each of the source business data including at least one of image data, audio data, and text data; mining, based on the first type of local business data corresponding to each of the at least two source business data, a first data mining description vector corresponding to each of the at least two source business data, wherein the first data mining description vector is formed by performing a deep mining operation based on the description vector to be processed, and the description vector to be processed is formed by performing a feature mining operation on the first type of local business data; mining a second data mining description vector corresponding to each of the at least two source business data based on the key business data corresponding to each of the at least two source business data and the description vector to be processed corresponding to each of the at least two source business data; Analyzing the filtered source business data from the at least two source business data based on a first data mining description vector corresponding to each source business data in the at least two source business data and a second data mining description vector corresponding to each source business data in the at least two source business data, the steps comprising: Performing a vector aggregation operation on a first data mining description vector corresponding to each source business data of the at least two source business data and a second data mining description vector corresponding to each source business data of the at least two source business data to form an aggregated data mining description vector corresponding to each source business data of the at least two source business data, the steps comprising: Determining, based on key business data corresponding to source business data to be analyzed, a third importance parameter of a second data mining description vector corresponding to the source business data to be analyzed, wherein the source business data to be analyzed is any one of the at least two source business data; The steps of aggregating the first data mining description vector corresponding to the source business data to be analyzed and the second data mining description vector corresponding to the source business data to be analyzed based on the third importance parameter of the second data mining description vector corresponding to the source business data to be analyzed to form an aggregated data mining description vector corresponding to the source business data to be analyzed include: performing a multiplication operation on the second data mining description vector of the source business data to be analyzed according to the third importance parameter of the second data mining description vector corresponding to the source business data to be analyzed, so as to form an adjusted data mining description vector corresponding to the source business data to be analyzed; Performing a superposition operation on the adjusted data mining description vector corresponding to the source business data to be analyzed and the first data mining description vector corresponding to the source business data to be analyzed to form an aggregated data mining description vector corresponding to the source business data to be analyzed; Analyzing, based on the aggregated data mining description vector corresponding to each of the at least two source business data, a reliability analysis result corresponding to each of the at least two source business data, the reliability analysis result being a predictive characterization parameter of the reliability of performing a user management and control operation on the target business user based on the source business data; Analyzing the filtered source business data from the at least two source business data according to the reliability analysis result corresponding to each source business data of the at least two source business data; Based on the filtered source business data, user management and control operations are performed on the target business user.
2. The data analysis method for multi-source business data according to claim 1, characterized in that: The step of mining a first data mining description vector corresponding to each of the at least two source business data based on the first type of local business data corresponding to each of the at least two source business data includes: Performing a feature mining operation on a first type of local business data corresponding to the source business data to be processed to form a first number of first description vectors to be processed of the source business data to be processed, where the source business data to be processed is any one of the at least two source business data; Analyzing the first importance parameters of the first number of first description vectors to be processed based on the first type of local service data corresponding to the source service data to be processed; A first data mining description vector corresponding to the source business data to be processed is determined based on the first number of first description vectors to be processed and the first importance parameters of the first number of first description vectors to be processed.
3. The data analysis method for multi-source business data according to claim 2, characterized in that: The step of mining a second data mining description vector corresponding to each of the at least two source business data based on the key business data corresponding to each of the at least two source business data and the description vector to be processed corresponding to each of the at least two source business data includes: Performing a feature mining operation on the key business data corresponding to the source business data to be processed to form a second number of second description vectors to be processed of the source business data to be processed; Determining, based on the key business data corresponding to the source business data to be processed, second importance parameters of the second number of second description vectors to be processed and second importance parameters of the first number of first description vectors to be processed; Based on the second importance parameters of the second number of second description vectors to be processed, the second importance parameters of the first number of first description vectors to be processed, the second number of second description vectors to be processed and the first number of first description vectors to be processed, the second data mining description vector corresponding to the source business data to be processed is determined.
4. The data analysis method for multi-source business data according to claim 3, characterized in that: The step of determining the second importance parameters of the second number of second description vectors to be processed and the second importance parameters of the first number of first description vectors to be processed based on the key business data corresponding to the source business data to be processed includes: Analyzing the second importance parameters of the second number of second description vectors to be processed and the second importance parameters of the first number of first description vectors to be processed based on the key business data corresponding to the source business data to be processed and the data source application parameters corresponding to the source business data to be processed, wherein the data source application parameters are used to reflect the data application status of the data source of the corresponding source business data; The step of analyzing the second importance parameters of the second number of second description vectors to be processed and the second importance parameters of the first number of first description vectors to be processed based on the key business data corresponding to the source business data to be processed and the data source application parameters corresponding to the source business data to be processed includes: Performing a combination operation on the key business data corresponding to the source business data to be processed and the data source application parameters corresponding to the source business data to be processed to form combined data to be processed corresponding to the source business data to be processed; The second importance parameters of the second number of second description vectors to be processed and the second importance parameters of the first number of first description vectors to be processed are analyzed based on the combination data to be processed corresponding to the source business data to be processed.
5. The data analysis method for multi-source business data according to claim 1, characterized in that: The step of determining the third importance parameter of the second data mining description vector corresponding to the source business data to be analyzed based on the key business data corresponding to the source business data to be analyzed includes: Analyzing a third importance parameter of the second data mining description vector corresponding to the source business data to be analyzed based on the key business data corresponding to the source business data to be analyzed and the data source application parameter corresponding to the source business data to be analyzed, wherein the data source application parameter is used to reflect the data application status of the data source of the corresponding source business data; The step of analyzing the third importance parameter of the second data mining description vector corresponding to the source business data to be analyzed based on the key business data corresponding to the source business data to be analyzed and the data source application parameters corresponding to the source business data to be analyzed includes: Performing a combination operation on the key business data corresponding to the source business data to be analyzed and the data source application parameters of the source business data to be analyzed to form combined data to be analyzed corresponding to the source business data to be analyzed; According to the combined data to be analyzed corresponding to the source business data to be analyzed, a third importance parameter of the second data mining description vector corresponding to the source business data to be analyzed is analyzed.
6. The data analysis method for multi-source business data according to claim 1, characterized in that: The step of mining a first data mining description vector corresponding to each of the at least two source business data based on the first type of local business data corresponding to each of the at least two source business data includes: Using a front-end feature mining unit in a data analysis neural network, performing a feature mining operation on the first type of local business data corresponding to each of the at least two source business data to form a first data mining description vector corresponding to each of the at least two source business data; The step of mining a second data mining description vector corresponding to each of the at least two source business data based on the key business data corresponding to each of the at least two source business data and the description vector to be processed corresponding to each of the at least two source business data includes: Using a back-end feature mining unit in the data analysis neural network, a fusion mining operation is performed on the key business data corresponding to each of the two source business data and the description vector to be processed corresponding to each of the at least two source business data, so as to output a second data mining description vector corresponding to each of the at least two source business data; The step of performing a vector aggregation operation on the first data mining description vector corresponding to each of the at least two source business data and the second data mining description vector corresponding to each of the at least two source business data to form an aggregated data mining description vector corresponding to each of the at least two source business data includes: Using a vector aggregation unit in the data analysis neural network, aggregating a first data mining description vector corresponding to each source business data of the at least two source business data and a second data mining description vector corresponding to each source business data of the at least two source business data to form an aggregated data mining description vector corresponding to each source business data of the at least two source business data; The step of analyzing the reliability analysis result corresponding to each of the at least two source business data based on the aggregated data mining description vector corresponding to each of the at least two source business data includes: Utilizing the data possibility analysis unit in the data analysis neural network, the aggregated data mining description vector corresponding to each of the at least two source business data is analyzed and output to obtain a reliability analysis result corresponding to each of the at least two source business data.
7. A data analysis system for multi-source business data, characterized in that: The method comprises a processor and a memory, wherein the memory is used to store a computer program, and the processor is used to execute the computer program to implement the method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Data processing method and device, electronic equipment and storage medium
CN115393849A
Business big data mining method and system and cloud platform
CN115640336A