A multi-dimensional customer data analysis and precise customer search method

By cleaning and extracting features from multi-dimensional customer data, combined with adjustments to classification algorithms and machine learning models, the inaccuracy problem caused by data noise is resolved, and the accuracy of finding potential customers and predictions is improved.

CN119312149BActive Publication Date: 2025-09-12BEIJING CROSSTEK TECH INC
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202411858441.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-17
Publication Date
2025-09-12
Estimated Expiration
2044-12-17

AI Technical Summary

Technical Problem

In the existing technology, due to the presence of data noise during the data collection process, the noise is mistakenly identified as data features, resulting in inaccurate data, which in turn reduces the accuracy of finding potential customers.

Method used

By cleaning, integrating and extracting features from multi-dimensional customer data, using classification algorithms for classification, and adjusting according to the amount of noise data, the update frequency of the machine learning model, the collection interval of customer data and the amount of data transmitted in a single batch, the accuracy of data cleaning and the prediction accuracy of the model can be improved.

Benefits of technology

By increasing the update frequency of the machine learning model, reducing the customer data collection interval and the amount of data transmitted in a single batch, the accuracy of noise recognition is improved, more data details are obtained, transmission delays are reduced, and the accuracy of finding potential customers is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119312149B_ABST
    Figure CN119312149B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of data processing technology, and in particular to a multi-dimensional customer data analysis and precision customer search method, comprising: transmitting customer data collected from multiple data sources to a location to be processed, and sequentially performing cleaning, integration, and feature extraction operations on the customer data to output preprocessed data; analyzing customer behavior based on the classification to output analysis results, training a machine learning model based on the analysis results, updating the machine learning model based on newly collected customer data, and using the machine learning model to search for potential customers; determining the accuracy of searching for potential customers based on the proportion of noise data in the customer data; adjusting the update frequency of the machine learning model if the accuracy does not meet the requirements; and adjusting the collection interval of customer data if the prediction accuracy does not meet the requirements. The present invention improves the accuracy of searching for potential customers.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data processing technology, and in particular to a multi-dimensional customer data analysis and precise customer search method. Background Art

[0002] In existing technology, multi-dimensional customer data refers to a collection of customer information collected from multiple perspectives. These perspectives include, but are not limited to, basic customer information, transaction behavior, consumption preferences, social interactions, and service feedback. Each aspect contains numerous specific data points. By integrating these different dimensions of data, a comprehensive and three-dimensional customer profile can be constructed.

[0003] Chinese Patent Publication No. CN112667911A discloses a method for finding potential customers using social software big data, comprising the following steps: S1, creating a screening model 1 and a screening model 2 based on historical customer information as data, S2, setting screening conditions for the screening model 1, the screening conditions including educational background, field of education, current company employed, years of service, position held, and whether any 3D or CAD software is used, and the screening model 1 automatically searches for new customers that meet the screening conditions on the Internet based on the screening conditions, S3, inputting the information of the new customers found by the screening model 1 into the screening model 2, and adding the industry and company size as the screening conditions of the screening model 2 for further screening, S4, sending a customized message to contact the new customer if the conditions are met. It can be seen from this that the method for finding potential customers using social software big data has the problem that due to the presence of data noise during the data collection process, the noise is mistakenly identified as a data feature, and the data is not cleaned up properly, resulting in inaccurate data and a decrease in the accuracy of finding potential customers. Summary of the Invention

[0004] To this end, the present invention provides a multi-dimensional customer data analysis and precise customer search method to overcome the problem in the prior art that due to the presence of data noise in the data collection process, the noise is mistakenly identified as a data feature, and the data is not cleaned up thoroughly, resulting in inaccurate data and thus a decrease in the accuracy of finding potential customers.

[0005] To achieve the above-mentioned objectives, the present invention provides a multi-dimensional customer data analysis and precise customer search method, comprising: transmitting customer data collected from multiple data sources to a location to be processed, and sequentially cleaning, integrating, and extracting features from the customer data to output preprocessed data, and classifying the preprocessed data using a classification algorithm; analyzing customer behavior according to the classification to output analysis results, and training a machine learning model based on the analysis results, and updating the machine learning model based on newly collected customer data, and using the machine learning model to search for potential customers; respectively obtaining the amount of noise data in the customer data and the total amount of customer data; determining the accuracy of searching for potential customers based on the proportion of noise data in the customer data; if the accuracy does not meet the requirements, adjusting the update frequency of the machine learning model, or determining the prediction accuracy of the machine learning model based on the missing rate of features within several prediction cycles; if the prediction accuracy does not meet the requirements, adjusting the collection interval of the customer data, or adjusting the amount of data transmitted in a single batch of customer data based on the average delay time of customer data processing.

[0006] Furthermore, determining the accuracy of finding potential customers includes:

[0007] Compare the proportion of noise data in the customer data with the preset first proportion;

[0008] If the proportion of noise data in the customer data is greater than the preset first proportion, it is determined that the accuracy of finding potential customers does not meet the requirements.

[0009] Further, determining the prediction accuracy of the machine learning model includes:

[0010] Comparing the noise data ratio in the customer data with the preset first ratio and the preset second ratio respectively;

[0011] If the proportion of noise data in the customer data is greater than the preset first proportion and less than or equal to the preset second proportion, it is preliminarily determined that the prediction accuracy of the machine learning model does not meet the requirements, and whether the prediction accuracy of the learning model meets the requirements is determined based on the missing rate of features in several prediction cycles.

[0012] Furthermore, adjusting the update frequency of the machine learning model includes:

[0013] Comparing the noise data ratio in the customer data with the preset second ratio;

[0014] If the proportion of noise data in the customer data is greater than the preset second proportion, the update frequency of the machine learning model is increased.

[0015] Furthermore, the increase in the update frequency of the machine learning model is determined by the difference between the proportion of noise data in the customer data and a preset second proportion.

[0016] Furthermore, adjusting the collection interval of the customer data includes:

[0017] Comparing the missing rates of the features in the plurality of prediction periods with a preset first missing rate and a preset second missing rate respectively;

[0018] If the missing rate of features within the plurality of prediction cycles is greater than the preset first missing rate, it is determined that the prediction accuracy of the machine learning model does not meet the requirements;

[0019] If the missing rate of the features within the plurality of prediction periods is greater than a preset first missing rate and less than or equal to a preset second missing rate, reducing the collection interval of the customer data;

[0020] If the missing rate of the features within the several prediction cycles is greater than the preset second missing rate, it is preliminarily determined that the real-time processing of the customer data does not meet the requirements, and whether the real-time processing of the customer data meets the requirements is determined based on the average delay time of the customer data processing.

[0021] Furthermore, the reduction range of the customer data collection interval is determined by the difference between the feature missing rate in a number of prediction periods and a preset first missing rate.

[0022] Furthermore, adjusting the data volume of a single batch of the customer data transmission includes:

[0023] Comparing the average delay time of the customer data processing with a preset delay time;

[0024] If the average delay time of the customer data processing is greater than the preset delay time, it is determined that the real-time processing of the customer data does not meet the requirements, and the data volume of a single batch transmission of the customer data is reduced.

[0025] Furthermore, the average delay duration of the client data processing is a ratio of the total delay duration of the client data processing in a plurality of processing cycles to the number of processing cycles.

[0026] Furthermore, the reduction extent of the data volume of a single batch of client data transmission is determined by the difference between an average delay time of client data processing and a preset delay time.

[0027] Compared with the prior art, the beneficial effect of the present invention is that the method of the present invention adjusts the update frequency of the machine learning model according to the proportion of noise data in the customer data. Since data noise is mixed in the process of data collection, the noise is mistakenly identified as a data feature, resulting in the data not being cleaned up properly, resulting in inaccurate data. By increasing the update frequency of the machine learning model, the model can adapt to these changes more quickly, relearn the effective features and new noise patterns in the data, so as to improve the accuracy of noise recognition, thereby improving the effect of data cleaning. The collection interval of customer data is adjusted according to the missing rate of features in several prediction cycles. Since when extracting features from customer data of multiple data types, it is necessary to merge data of different modalities into a complete feature, which may result in loss of features. Some key information makes the final feature representation unable to fully reflect the complex characteristics of the customer, which makes the trained model unable to predict accurately. By reducing the collection interval of customer data, more data details can be obtained to make up for the missing parts, so that the model can learn more complete customer behavior patterns, thereby improving prediction accuracy. The single batch transmission data volume of customer data is adjusted according to the average delay time of customer data processing. Since data may need to be collected from multiple collection points, when transmitting these data, it may cause data transmission delays, resulting in increased delays when processing data. By reducing the single batch transmission data volume of customer data, it is possible to more flexibly find transmission channels in crowded networks, reduce waiting time, complete transmission faster, reduce delays, and improve the accuracy of finding potential customers.

[0028] Furthermore, the method of the present invention adjusts the update frequency of the machine learning model by setting a preset first ratio and a preset second ratio. Since data noise is mixed in the process of data collection, the noise is mistakenly identified as a data feature, resulting in the data not being cleaned up properly, resulting in inaccurate data. By increasing the update frequency of the machine learning model, the model can adapt to these changes more quickly, relearn the effective features and new noise patterns in the data, so as to improve the accuracy of noise recognition, thereby improving the effect of data cleaning and further improving the accuracy of finding potential customers.

[0029] Furthermore, the method of the present invention adjusts the collection interval of customer data by setting a preset first missing rate and a preset second missing rate. Since when extracting features from customer data of multiple data types, it is necessary to fuse data of different modalities into a complete feature, which may result in the loss of some key information, making the final feature representation unable to fully reflect the complex characteristics of the customer, thereby causing the trained model to be unable to predict accurately. By reducing the collection interval of customer data, more data details can be obtained to make up for the missing parts, so that the model can learn a more complete customer behavior pattern, thereby improving the prediction accuracy and further improving the accuracy of finding potential customers.

[0030] Furthermore, the method of the present invention adjusts the amount of data transmitted in a single batch of customer data by setting a preset delay time. Since data may need to be collected from multiple collection points, data transmission may be delayed when these data are transmitted, resulting in increased delays when processing the data. By reducing the amount of data transmitted in a single batch of customer data, it is possible to more flexibly find transmission channels in a crowded network, reduce waiting time, complete transmission faster, reduce delays, and further improve the accuracy of finding potential customers. BRIEF DESCRIPTION OF THE DRAWINGS

[0031] Figure 1 This is an overall flow chart of the multi-dimensional customer data analysis and precise customer search method according to an embodiment of the present invention;

[0032] Figure 2 This is a logic flow chart of a multi-dimensional customer data analysis and precise customer search method according to an embodiment of the present invention;

[0033] Figure 3 This is a specific flow chart of the process of adjusting the update frequency of the machine learning model in the multi-dimensional customer data analysis and precise customer search method according to an embodiment of the present invention;

[0034] Figure 4 This is a specific flow chart of the process of adjusting the customer data collection interval in the multi-dimensional customer data analysis and precise customer search method according to an embodiment of the present invention. DETAILED DESCRIPTION

[0035] In order to make the objects and advantages of the present invention more clearly understood, the present invention is further described below in conjunction with embodiments; it should be understood that the specific embodiments described herein are merely used to explain the present invention and are not intended to limit the present invention.

[0036] The preferred embodiments of the present invention are described below with reference to the accompanying drawings. It should be understood by those skilled in the art that these embodiments are only used to explain the technical principles of the present invention and are not intended to limit the scope of protection of the present invention.

[0037] It should be noted that, in the description of the present invention, terms such as "up", "down", "left", "right", "inside", and "outside" indicating directions or positional relationships are based on the directions or positional relationships shown in the accompanying drawings. This is only for the convenience of description and does not indicate or imply that the device or element must have a specific orientation, be constructed and operated in a specific orientation. Therefore, it cannot be understood as a limitation on the present invention.

[0038] Furthermore, it should be noted that, in the description of the present invention, unless otherwise expressly specified or limited, the terms "mounted," "connected," and "connected" should be understood in a broad sense. For example, they may refer to fixed connections, detachable connections, or integral connections; mechanical connections or electrical connections; direct connections or indirect connections through an intermediate medium; and internal communication between two components. Those skilled in the art will understand the specific meanings of the above terms in the present invention based on specific circumstances.

[0039] See also Figure 1 、 Figure 2 、 Figure 3 as well as Figure 4 As shown, they are respectively an overall flow chart, a logic flow chart, a specific flow chart of the process of adjusting the update frequency of the machine learning model, and a specific flow chart of the process of adjusting the collection interval of customer data of the multi-dimensional customer data analysis and precise customer search method according to an embodiment of the present invention. A multi-dimensional customer data analysis and precise customer search method according to the present invention includes:

[0040] Step S1: transferring the collected customer data from multiple data sources to a location to be processed, and sequentially performing cleaning, integration, and feature extraction operations on the customer data to output pre-processed data, and classifying the pre-processed data using a classification algorithm;

[0041] Step S2: Analyze the customer's behavior according to the classification to output analysis results, train a machine learning model based on the analysis results, update the machine learning model based on newly collected customer data, and use the machine learning model to search for potential customers;

[0042] Step S3, respectively obtaining the amount of noise data in the customer data and the total amount of customer data;

[0043] Step S4, determining the accuracy of finding potential customers based on the proportion of noise data in the customer data;

[0044] Step S5: If the accuracy does not meet the requirements, the update frequency of the machine learning model is adjusted, or the prediction accuracy of the machine learning model is determined based on the missing rate of features within several prediction cycles;

[0045] Step S6: If the prediction accuracy does not meet the requirement, the collection interval of the customer data is adjusted, or the data volume of a single batch of customer data transmission is adjusted based on the average delay time of the customer data processing.

[0046] Specifically, multiple data sources include e-commerce platforms, social media platforms, and suppliers.

[0047] Specifically, customer data includes the names of products purchased by the customer.

[0048] Specifically, the characteristics include the frequency with which customers inquire about products, the number of times customers purchase products, and the time interval between customer purchases.

[0049] Specifically, the preprocessed data included malformed purchase times.

[0050] Specifically, classification algorithms include decision tree, support vector machine algorithm, and naive Bayes algorithm.

[0051] Specifically, the analysis results include the purchase frequency of customers purchasing the product, the usage frequency of customers using the product, and the purchase level of customers purchasing the product.

[0052] Specifically, machine learning models include linear regression models, logistic regression models, and random forest models.

[0053] Specifically, the missing rate of a feature is the ratio of the number of data points with a specific data feature missing to the total number of data points in a multidimensional customer dataset.

[0054] In implementation, the method of the present invention adjusts the update frequency of the machine learning model according to the proportion of noise data in the customer data. Since data noise is mixed in the process of data collection, the noise is mistakenly identified as a data feature, resulting in the data not being cleaned up properly, resulting in inaccurate data. By increasing the update frequency of the machine learning model, the model can adapt to these changes more quickly, relearn the effective features and new noise patterns in the data, so as to improve the accuracy of noise recognition, thereby improving the effect of data cleaning. The collection interval of customer data is adjusted according to the missing rate of features in several prediction cycles. Since when extracting features from customer data of multiple data types, it is necessary to merge data of different modalities into a complete feature, which may result in the loss of some key information. The final feature representation cannot fully reflect the complex characteristics of the customer, which makes the trained model unable to predict accurately. By reducing the collection interval of customer data, more data details can be obtained to make up for the missing parts, so that the model can learn more complete customer behavior patterns, thereby improving prediction accuracy. The single batch transmission data volume of customer data is adjusted according to the average delay time of customer data processing. Since data may need to be collected from multiple collection points, data transmission may be delayed when transmitting these data, resulting in increased delay in data processing. By reducing the single batch transmission data volume of customer data, it is possible to more flexibly find transmission channels in crowded networks, reduce waiting time, complete transmission faster, reduce delays, and improve the accuracy of finding potential customers.

[0055] Specifically, determining the accuracy of finding potential customers includes:

[0056] Obtain the amount of noise data in the customer data and the total amount of customer data respectively, and calculate the proportion of the noise data in the customer data;

[0057] Comparing the noise data ratio in the customer data with a preset first ratio;

[0058] If the proportion of noise data in the customer data is greater than the preset first proportion, it is determined that the accuracy of finding potential customers does not meet the requirements.

[0059] Specifically, determining the prediction accuracy of the machine learning model includes:

[0060] Comparing the noise data ratio in the customer data with the preset first ratio and the preset second ratio respectively;

[0061] If the proportion of noise data in the customer data is greater than the preset first proportion and less than or equal to the preset second proportion, it is preliminarily determined that the prediction accuracy of the machine learning model does not meet the requirements, and whether the prediction accuracy of the learning model meets the requirements is determined based on the missing rate of features in several prediction cycles.

[0062] It can be understood that the three intervals corresponding to the preset first proportion and the preset second proportion correspond to three situations respectively:

[0063] The first interval is when the proportion of noise data in the customer data is less than or equal to the preset first proportion, corresponding to the situation where the accuracy of finding potential customers meets the requirements;

[0064] The second interval is when the proportion of noise data in the customer data is greater than the preset first proportion and less than or equal to the preset second proportion. This corresponds to the situation where some key information may be lost when extracting features from customer data of multiple data types, making the final feature representation unable to fully reflect the complex characteristics of the customer, and thus causing the trained model to be unable to make accurate predictions.

[0065] The third interval is when the proportion of noise data in the customer data is greater than the preset second proportion. This corresponds to the situation where data noise is mixed in the data collection process, resulting in the noise being mistakenly identified as a data feature, resulting in the data not being cleaned up properly, and causing inaccurate data.

[0066] In practice, the preset first ratio is generally selected in the range of [0.12, 0.14], and the preset second ratio is generally selected in the range of [0.15, 0.17].

[0067] Preferably, the preferred embodiment of the preset first ratio is 0.13, and the preferred embodiment of the preset second ratio is 0.16.

[0068] In implementation, the method of the present invention determines the accuracy of finding potential customers by setting a preset first ratio and a preset second ratio, thereby reducing the impact of the decreased stability of finding potential customers due to inaccurate determination of the accuracy of finding potential customers, and further improving the accuracy of finding potential customers.

[0069] Specifically, adjusting the update frequency of the machine learning model includes:

[0070] Comparing the noise data ratio in the customer data with the preset second ratio;

[0071] If the proportion of noise data in the customer data is greater than the preset second proportion, the update frequency of the machine learning model is increased.

[0072] Specifically, the increase in the update frequency of the machine learning model is determined by the difference between the proportion of noise data in the customer data and a preset second proportion.

[0073] Specifically, when the difference between the proportion of noise data in customer data and the preset second proportion is within 0.02, the update frequency of the machine learning model is increased to 1 times the original; when the difference between the proportion of noise data in customer data and the preset second proportion exceeds 0.02, the update frequency of the machine learning model increases by 2 times / day for every 0.01 that exceeds. For example, the difference between the proportion of noise data in customer data and the preset second proportion is 0.04, the current update frequency of the machine learning model is 1 time / day, and the increased update frequency of the machine learning model is 1×1+2×2=5 times / day.

[0074] In implementation, the method of the present invention adjusts the update frequency of the machine learning model by setting a preset first ratio and a preset second ratio. Since data noise is mixed in the process of data collection, the noise is mistakenly identified as a data feature, resulting in the data not being cleaned up properly, resulting in inaccurate data. By increasing the update frequency of the machine learning model, the model can adapt to these changes more quickly, relearn the effective features and new noise patterns in the data, so as to improve the accuracy of noise recognition, thereby improving the effect of data cleaning and further improving the accuracy of finding potential customers.

[0075] Specifically, adjusting the collection interval of the customer data includes:

[0076] Obtain the number of missing features and the total number of features, and calculate the missing rate of features within several prediction cycles;

[0077] Comparing the missing rates of the features in the plurality of prediction periods with a preset first missing rate and a preset second missing rate respectively;

[0078] If the missing rate of features within the plurality of prediction cycles is greater than the preset first missing rate, it is determined that the prediction accuracy of the machine learning model does not meet the requirements;

[0079] If the missing rate of the features within the plurality of prediction periods is greater than a preset first missing rate and less than or equal to a preset second missing rate, reducing the collection interval of the customer data;

[0080] If the missing rate of the features within the several prediction cycles is greater than the preset second missing rate, it is preliminarily determined that the real-time processing of the customer data does not meet the requirements, and whether the real-time processing of the customer data meets the requirements is determined based on the average delay time of the customer data processing.

[0081] It can be understood that the three intervals corresponding to the preset first missing rate and the preset second missing rate correspond to three situations respectively:

[0082] The first interval is when the feature missing rate within a certain number of prediction cycles is less than or equal to the preset first missing rate, corresponding to the situation where the prediction accuracy of the machine learning model meets the requirements;

[0083] The second interval refers to the situation where the feature missing rate within a certain number of prediction cycles is greater than the preset first missing rate and less than or equal to the preset second missing rate. This corresponds to the situation where some key information may be lost when extracting features from customer data of various data types, making the final feature representation unable to fully reflect the complex characteristics of the customer, thereby causing the trained model to be unable to make accurate predictions.

[0084] The third interval is when the feature missing rate within several prediction cycles is greater than the preset second missing rate. Correspondingly, data collection may need to be performed from multiple collection points. When transmitting this data, data transmission may be delayed, resulting in increased delays in data processing.

[0085] In practice, the preset first missing rate is generally selected from the range of [0.16, 0.18], and the preset second missing rate is generally selected from the range of [0.19, 0.21].

[0086] Preferably, the preferred embodiment of the preset first missing rate is 0.17, and the generally selected range of the preset second missing rate is 0.2.

[0087] In implementation, the method of the present invention determines the prediction accuracy of the machine learning model by setting a preset first missing rate and a preset second missing rate, thereby reducing the impact of the decreased accuracy in finding potential customers due to inaccurate determination of the prediction accuracy of the machine learning model, and further improving the accuracy in finding potential customers.

[0088] Specifically, the reduction range of the collection interval of the customer data is determined by the difference between the missing rate of the feature in several prediction periods and a preset first missing rate.

[0089] Specifically, when the difference between the missing rate of features in several prediction cycles and the preset first missing rate is within 0.03, the customer data collection interval is reduced to 0.9 times the original time; when the difference between the missing rate of features in several prediction cycles and the preset first missing rate exceeds 0.02, the customer data collection interval is reduced by 1.5 minutes for every 0.02. For example, if the difference between the missing rate of features in several prediction cycles and the preset first missing rate is 0.07, the current customer data collection interval is 10 minutes, and the reduced customer data collection interval is 10×0.9-1.5×2=6 minutes.

[0090] In implementation, the method of the present invention adjusts the collection interval of customer data by setting a preset first missing rate and a preset second missing rate. Since when extracting features from customer data of multiple data types, it is necessary to fuse data of different modalities into a complete feature, which may result in the loss of some key information, making the final feature representation unable to fully reflect the complex characteristics of the customer, thereby causing the trained model to be unable to predict accurately. By reducing the collection interval of customer data, more data details can be obtained to make up for the missing parts, so that the model can learn a more complete customer behavior pattern, thereby improving the prediction accuracy and further improving the accuracy of finding potential customers.

[0091] Specifically, adjusting the data volume of a single batch of customer data transmission includes:

[0092] Comparing the average delay time of the customer data processing with a preset delay time;

[0093] If the average delay time of the customer data processing is greater than the preset delay time, it is determined that the real-time processing of the customer data does not meet the requirements, and the data volume of a single batch transmission of the customer data is reduced.

[0094] It is understandable that the two intervals corresponding to the preset delay time correspond to two situations:

[0095] The first interval is when the average delay in processing customer data is less than or equal to the preset delay, which corresponds to the situation where the real-time processing of customer data meets the requirements;

[0096] The second interval is when the average delay in processing customer data is longer than the preset delay. This means that data may need to be collected from multiple collection points. When transmitting this data, data transmission may be delayed, resulting in increased delay in data processing.

[0097] In practice, the preset delay time is generally selected in the range of [450ms, 550ms].

[0098] Preferably, the preferred embodiment of the preset delay time length is 500ms.

[0099] In implementation, the method of the present invention determines the real-time processing of customer data by setting a preset delay time, thereby reducing the impact of the reduced accuracy of finding potential customers due to inaccurate determination of the real-time processing of customer data, and further improving the accuracy of finding potential customers.

[0100] Specifically, the average delay duration of the client data processing is the ratio of the total delay duration of the client data processing in a number of processing cycles to the number of processing cycles.

[0101] Specifically, the reduction extent of the data volume of a single batch of client data transmission is determined by the difference between the average delay time of client data processing and the preset delay time.

[0102] Specifically, when the difference between the average delay time of customer data processing and the preset delay time is within 100ms, the data volume of a single batch of customer data transmission is reduced to 0.92 times the original; when the difference between the average delay time of customer data processing and the preset delay time exceeds 100ms, the data volume of a single batch of customer data transmission is reduced by 200KB for every 50ms exceeding it. For example, if the difference between the average delay time of customer data processing and the preset delay time is 200ms, the current data volume of a single batch of customer data transmission is 3000KB, the reduced data volume of a single batch of customer data transmission is 3000KB.

[0103] During implementation, the method of the present invention adjusts the amount of data transmitted in a single batch of customer data by setting a preset delay time. Since data may need to be collected from multiple collection points, data transmission may be delayed when these data are transmitted, resulting in increased delays when processing the data. By reducing the amount of data transmitted in a single batch of customer data, it is possible to more flexibly find transmission channels in a crowded network, reduce waiting time, complete transmission faster, reduce delays, and further improve the accuracy of finding potential customers.

[0104] Thus far, the technical solutions of the present invention have been described in conjunction with the preferred embodiments shown in the accompanying drawings. However, it will be readily understood by those skilled in the art that the scope of protection of the present invention is obviously not limited to these specific embodiments. Without departing from the principles of the present invention, those skilled in the art may make equivalent changes or substitutions to the relevant technical features, and the technical solutions after such changes or substitutions will fall within the scope of protection of the present invention.

Claims

1. A multi-dimensional customer data analysis and precise customer search method, characterized in that: include: Transmitting the customer data collected from multiple data sources to a location to be processed, and sequentially performing cleaning, integration, and feature extraction operations on the customer data to output pre-processed data, and classifying the pre-processed data using a classification algorithm; Analyze customer behavior according to the classification to output analysis results, train a machine learning model based on the analysis results, update the machine learning model based on newly collected customer data, and use the machine learning model to find potential customers; respectively obtaining the amount of noise data in the customer data and the total amount of customer data; If the proportion of noise data in the customer data is less than or equal to the preset first proportion, it is determined that the accuracy of finding potential customers meets the requirements; If the proportion of noise data in the customer data is greater than the preset first proportion and less than or equal to the preset second proportion, it is preliminarily determined that the prediction accuracy of the machine learning model does not meet the requirements, and whether the prediction accuracy of the machine learning model meets the requirements is determined based on the missing rate of features in several prediction cycles; wherein the missing rate of a feature is the ratio of the number of data points with a specific data feature missing to the total number of data points in the multi-dimensional customer data set; If the proportion of noise data in the customer data is greater than a preset second proportion, increasing the update frequency of the machine learning model; The preset second proportion is greater than the preset first proportion; Determine the prediction accuracy of the machine learning model based on the missing rate of features over several prediction periods, including: If the missing rate of features within a number of prediction cycles is less than or equal to a preset first missing rate, it is determined that the prediction accuracy of the machine learning model meets the requirements; If the missing rate of the feature within a number of prediction periods is greater than a preset first missing rate and less than or equal to a preset second missing rate, then reducing the collection interval of the customer data; If the missing rate of the features within a number of prediction periods is greater than a preset second missing rate, it is preliminarily determined that the real-time processing of the customer data does not meet the requirements, and whether the real-time processing of the customer data meets the requirements is determined based on the average delay time of the customer data processing; wherein the preset second missing rate is greater than the preset first missing rate; Determine whether the real-time processing of customer data meets the requirements based on the average delay in processing customer data, including: If the average delay time of customer processing is longer than the preset delay time, the amount of data transmitted in a single batch of customer data will be reduced; If the average delay time of customer processing is less than or equal to the preset delay time, it is determined that the real-time processing of customer data meets the requirements; The multiple data sources include e-commerce platforms, social media platforms, and suppliers; The customer data includes the name of the product purchased by the customer; The characteristics include the frequency of customer inquiries about the product, the number of times a customer purchases the product, and the time interval between customer purchases; The analysis results include the purchase frequency of the product by the customer, the usage frequency of the product by the customer, and the purchase level of the product by the customer.

2. The multi-dimensional customer data analysis and precise customer search method according to claim 1, characterized in that: The increase in the update frequency of the machine learning model is determined by the difference between the proportion of noise data in the customer data and a preset second proportion.

3. The multi-dimensional customer data analysis and precise customer search method according to claim 2, characterized in that: The reduction range of the collection interval of the customer data is determined by the difference between the missing rate of the feature in several prediction periods and a preset first missing rate.

4. The multi-dimensional customer data analysis and precise customer search method according to claim 3, characterized in that: The average delay duration of the client data processing is a ratio of the total delay duration of the client data processing in a plurality of processing cycles to the number of processing cycles.

5. The multi-dimensional customer data analysis and precise customer search method according to claim 4, characterized in that: The reduction range of the single batch transmission data volume of the client data is determined by the difference between the average delay time of the client data processing and the preset delay time.

Citation Information

Patent Citations

  • Method for searching potential customers by utilizing social software big data

    CN112667911A

  • Potential customer mining and active registration guiding method and system based on big data analysis

    CN118569906A

  • Application software recommendation method and system based on big data

    CN118708820A

  • Federal learning-based privacy protection type large-scale model training and deployment method

    CN118734360A