Data source grading-based marketing lead scoring method and system

WO2026166005A1PCT designated stage Publication Date: 2026-08-13BEIJING JIZHI TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-04-17
Publication Date
2026-08-13

Smart Images

  • Figure CN2025089524_13082026_PF_FP_ABST
    Figure CN2025089524_13082026_PF_FP_ABST
Patent Text Reader

Abstract

The present invention provides a data source grading-based marketing lead scoring method and system. The data source grading-based marketing lead scoring method comprises: extracting data stream feature data information and a data source classification scheme of data sources corresponding to obtained marketing leads; using the data stream feature data information and the data source classification scheme to classify the data sources corresponding to the obtained marketing leads, to obtain different categories of data sources; using target grading coefficients corresponding to the different categories of data sources, in combination with a preset initial data collection frequency, to set data collection frequencies, and obtaining marketing lead data; determining data quality and security of the marketing lead data to obtain marketing lead data satisfying a preset security criterion; and calling a deep learning model for scoring, and using the deep learning model for scoring to obtain a score corresponding to the marketing lead data that satisfies the preset security criterion. The system comprises modules corresponding to the steps of the method.
Need to check novelty before this filing date? Find Prior Art

Description

A marketing lead scoring method and system based on data source hierarchy Technical Field

[0001] This invention proposes a marketing lead scoring method and system based on data source hierarchy, belonging to the field of marketing real-time scoring and control technology. Background Technology

[0002] With the rapid development of internet technology, the marketing field has undergone unprecedented changes. To stand out in a highly competitive market, companies are seeking efficient and precise marketing strategies. In this process, marketing leads, as the bridge connecting potential customers and actual sales, are undeniably crucial. The quality of marketing leads directly impacts the conversion rate, customer retention rate, and ultimately, performance of marketing campaigns. Traditional marketing lead evaluation methods often rely on manual screening or simple rule filtering, which is not only inefficient but also ill-suited to handling massive amounts of complex marketing lead data. With the rise of big data technology, companies have begun to explore using big data analytics to improve the efficiency and accuracy of marketing lead evaluation. However, simple big data analytics still faces challenges such as insufficient data real-time performance and poor model adaptability, especially in the face of rapidly changing market environments. In recent years, the rapid development of deep learning technology has provided new solutions to these problems. Deep learning models can automatically learn complex features in data, efficiently process large-scale data, and possess a certain degree of self-optimization capability. However, when applying deep learning to marketing lead evaluation, how to effectively handle data from different sources and of different qualities, and how to ensure the real-time performance and accuracy of the model, have become urgent technical challenges. Summary of the Invention

[0003] This invention provides a marketing lead scoring method and system based on data source hierarchy to solve the technical problems existing in the prior art. The technical solution adopted is as follows:

[0004] A marketing lead scoring method based on data source hierarchy, the marketing lead scoring method based on data source hierarchy includes:

[0005] Extract data stream feature information and data source classification scheme from the data sources corresponding to the marketing leads; wherein, the data source classification scheme includes a first classification scheme and a second classification scheme;

[0006] The data stream feature data information and data source classification scheme are used to classify the data sources corresponding to the marketing leads to obtain different categories of data sources;

[0007] By combining the target grading coefficients corresponding to different categories of data sources with the preset initial data collection frequency, the data collection frequency corresponding to the different categories of data sources is set; and marketing lead data is collected according to the set data collection frequency to obtain marketing lead data;

[0008] The marketing lead data is assessed for data quality and security to obtain marketing lead data that meets preset security standards.

[0009] The deep learning model used for scoring is retrieved, and the scoring score corresponding to the marketing lead data that meets the preset security standards is obtained using the deep learning model used for scoring.

[0010] Furthermore, the data stream feature information and data source classification scheme are used to classify the data sources corresponding to the marketing leads, obtaining different categories of data sources, including:

[0011] Extract the data flow characteristics of each data source, wherein the data flow characteristic data information includes the data volume, data integrity, data generation or update frequency and data redundancy ratio corresponding to the data flow per unit time.

[0012] Extract the amount of data corresponding to the data stream per unit time for each data source;

[0013] Data flow parameters are obtained by utilizing the amount of data flowing through each data source per unit time.

[0014] The data traffic parameters are obtained using the following formula:

[0015] Where P represents the data flow parameter; n represents the number of time intervals in which the data source runs, and the time interval is 1 second; D i D represents the amount of data corresponding to the data stream in the i-th unit of time; n D represents the total amount of data corresponding to a data stream spanning n units of time; fmax D represents the maximum value of the change in data volume corresponding to a data stream over n units of time; bi D represents the standard deviation of the data volume corresponding to the data stream at the i-th unit of time; max This represents the maximum amount of data corresponding to a data stream spanning n units of time.

[0016] The data traffic parameters are compared with preset data traffic parameter thresholds:

[0017] When the data flow parameters exceed the preset data flow parameter threshold, the data source is classified using the first classification scheme.

[0018] When the data flow parameters do not exceed the preset data flow parameter threshold, the data source is classified using the second classification scheme.

[0019] Furthermore, the first rating scheme is as follows:

[0020] When the data flow parameters exceed the preset data flow parameter threshold, the data integrity, data generation or update frequency and data redundancy ratio contained in the data flow feature data information are retrieved.

[0021] Retrieve data flow parameters;

[0022] The first data classification coefficient is obtained by using the data traffic parameters, data integrity, data generation or update frequency, and data redundancy ratio.

[0023] The first data grading coefficient is obtained by the following formula:

[0024] Among them, S 01 The first data classification coefficient is represented by ; P represents the data flow parameter; n represents the number of time intervals in which the data source runs, and the time interval is 1 second; f i R represents the data generation or update frequency corresponding to the i-th unit of time; i D represents the percentage of data redundancy corresponding to the i-th unit of time; i W represents the amount of data in the data stream corresponding to the i-th unit of time; i Indicates the data integrity corresponding to the i-th unit of time; f b W represents the standard deviation of the data generation or update frequency corresponding to n units of time; bi This represents the standard deviation of data integrity corresponding to the i-th unit of time;

[0025] Compare the first data grading coefficient with a preset first grading coefficient threshold;

[0026] When the first data classification coefficient exceeds the preset first classification coefficient threshold, the data source corresponding to the first data classification coefficient exceeding the preset first classification coefficient threshold is regarded as the first-level data source.

[0027] When the first data classification coefficient does not exceed the preset first classification coefficient threshold, the data source corresponding to the first data classification coefficient not exceeding the preset first classification coefficient threshold is regarded as the secondary data source.

[0028] Furthermore, the second rating scheme is as follows:

[0029] When the data flow parameters do not exceed the preset data flow parameter threshold, retrieve the data integrity, data generation or update frequency and data redundancy ratio contained in the data flow feature data information;

[0030] Retrieve data flow parameters;

[0031] Retrieve the rate of change of the number of data types in the data source per unit of time;

[0032] The second data classification coefficient is obtained by combining the data flow parameters with the rate of change of data type quantity per unit time, data integrity, data generation or update frequency, and data redundancy ratio.

[0033] The second data grading coefficient is obtained using the following formula:

[0034] Among them, S 02 P represents the data classification coefficient; P represents the data flow parameter; P y This represents a preset data flow parameter threshold; n represents the number of time intervals in which the data source runs, and the time interval is 1 second; f i R represents the data generation or update frequency corresponding to the i-th unit of time; i D represents the percentage of data redundancy corresponding to the i-th unit of time; i W represents the amount of data in the data stream corresponding to the i-th unit of time; i Indicates the data integrity corresponding to the i-th unit of time; f b W represents the standard deviation of the data generation or update frequency corresponding to n units of time; bi N represents the standard deviation of data integrity corresponding to the i-th unit of time; i N represents the rate of change in the number of data types corresponding to the i-th unit of time; b The standard deviation of the rate of change of the number of data types over n units of time;

[0035] Compare the second data grading coefficient with a preset second grading coefficient threshold;

[0036] When the second data classification coefficient exceeds the preset second classification coefficient threshold, the data source corresponding to the second data classification coefficient exceeding the preset second classification coefficient threshold is regarded as the first-level data source.

[0037] When the second data grading coefficient does not exceed the preset second grading coefficient threshold, the data source corresponding to the second data grading coefficient not exceeding the preset second grading coefficient threshold is regarded as the secondary data source.

[0038] Furthermore, by combining the target grading coefficients corresponding to different categories of data sources with a preset initial data collection frequency, the data collection frequency corresponding to the different categories of data sources is set, including:

[0039] Retrieve the first data classification coefficient or the second data classification coefficient corresponding to the primary data source as the first target classification coefficient;

[0040] Retrieve the preset initial data collection frequency from the database;

[0041] The data acquisition frequency corresponding to the first-level data source is obtained by combining the first target classification coefficient with the preset initial data acquisition frequency;

[0042] The data acquisition frequency corresponding to the primary data source is obtained by the following formula:

[0043] Among them, F 01 Indicates the data collection frequency corresponding to the primary data source; F c Indicates the preset initial data acquisition frequency; S m01 S represents the first objective classification coefficient; y01 This represents the threshold value of the grading coefficient corresponding to the first target grading coefficient.

[0044] Furthermore, by combining the target grading coefficients corresponding to different categories of data sources with a preset initial data collection frequency, the method further includes setting the data collection frequency corresponding to the different categories of data sources.

[0045] Retrieve the first data classification coefficient or the second data classification coefficient corresponding to the secondary data source as the second target classification coefficient;

[0046] Retrieve the preset initial data collection frequency from the database;

[0047] The data acquisition frequency corresponding to the secondary data source is obtained by combining the second target grading coefficient with the preset initial data acquisition frequency;

[0048] The data acquisition frequency corresponding to the secondary data source is obtained by the following formula:

[0049] Among them, F 02 Indicates the data collection frequency corresponding to the secondary data source; F c Indicates the preset initial data acquisition frequency; S m02 S represents the second objective classification coefficient; y02 This represents the threshold value of the grading coefficient corresponding to the second objective grading coefficient.

[0050] Furthermore, the security of the marketing lead data is determined based on its data quality, and marketing lead data that meets security requirements is obtained, including:

[0051] Extract the data characteristic parameters of each collected marketing lead data, wherein the data characteristic parameters include the encryption rate of the marketing lead data collected per unit time, the encryption resource utilization rate, and the data backup frequency;

[0052] The security evaluation coefficient is obtained by using the encryption rate, encryption resource utilization rate, and data backup frequency of the marketing lead data collected per unit time.

[0053] The safety evaluation coefficient is obtained using the following formula:

[0054] Where K represents the safety evaluation coefficient; m represents the number of times marketing lead data was collected; J i U represents the encryption rate corresponding to the marketing lead data collected in the i-th instance; i B represents the encrypted resource utilization rate corresponding to the marketing lead data collected in the i-th instance; i B represents the data backup frequency corresponding to the marketing lead data collected in the i-th instance; c This indicates the preset data backup frequency reference value;

[0055] The safety evaluation coefficient is compared with a preset safety coefficient threshold.

[0056] When the security evaluation coefficient exceeds the preset security coefficient threshold, the collected marketing lead data is determined to meet the security requirements.

[0057] Further, the deep learning model used for scoring is retrieved, and the scoring score corresponding to the marketing lead data that meets the preset security standards is obtained using the deep learning model used for scoring, including:

[0058] Retrieve the trained and optimized deep learning model from the database;

[0059] Extract the optimization completion time of the most recent optimization of the deep learning model;

[0060] The optimization completion time is obtained based on the optimization completion time and the current model call time.

[0061] The optimization completion time is compared with a preset time threshold.

[0062] When the optimization completion time exceeds the preset time threshold, the deep learning model is re-optimized to obtain the optimized deep learning model.

[0063] The marketing lead data that meets the security requirements is input into the optimized deep learning model;

[0064] The deep learning model is used to score marketing lead data that meets security requirements, and the corresponding scores for the marketing lead data are obtained.

[0065] Furthermore, the structure of the deep learning model is as follows:

[0066] The input layer is used to classify the security-compliant marketing lead data, obtain structured and unstructured data, and preprocess the structured and unstructured data to obtain preprocessed data information. The structured data includes basic customer information (such as age, gender, and geographic location), historical purchase records, and interaction records. The unstructured data includes text data (such as customer feedback and social media comments) and image data (such as product images). The preprocessing includes standardization and normalization, scaling of numerical features, and word segmentation and vectorization of text data (such as Word2Vec, TF-IDF, or BERT embedding).

[0067] The feature extraction layer is used to extract and construct features corresponding to the preprocessed data information so that the model can learn more effectively.

[0068] The model inference layer is used by the deep neural network (DNN) model to generate scoring results based on the features corresponding to the data information.

[0069] The output layer is used to output the scoring results.

[0070] A marketing lead scoring system based on data source hierarchy, the marketing lead scoring system based on data source hierarchy includes:

[0071] The data retrieval module is used to extract data flow feature information and data source classification schemes from the data sources corresponding to marketing leads; wherein, the data source classification schemes include a first classification scheme and a second classification scheme;

[0072] The data source classification module is used to classify the data sources corresponding to the marketing leads by using the data flow feature data information and the data source classification scheme, and to obtain data sources of different categories.

[0073] The data collection frequency setting module is used to set the data collection frequency corresponding to the different categories of data sources by combining the target grading coefficients corresponding to the different categories of data sources with the preset initial data collection frequency; and to collect marketing lead data according to the set data collection frequency to obtain marketing lead data.

[0074] The security assessment module is used to assess the data quality and security of the marketing lead data and obtain marketing lead data that meets preset security standards.

[0075] The scoring control module is used to retrieve the deep learning model used for scoring and to obtain the scoring score corresponding to the marketing lead data that meets the preset security standards using the deep learning model used for scoring.

[0076] Beneficial effects of this invention:

[0077] This invention proposes a marketing lead scoring method and system based on data source hierarchy. Through automated and intelligent scoring, businesses can quickly filter out high-value marketing leads, reducing the manual screening costs of large numbers of low-quality leads and thus improving marketing efficiency. Deep learning models can automatically learn complex features in the data and accurately score marketing leads based on these features. This enables businesses to more accurately identify potential customers and develop more targeted marketing strategies. During data processing, strict quality control and security checks ensure that the data input to the model is reliable and valid. This helps avoid erroneous scoring results due to input data quality issues, thereby enhancing data security. This method can process and analyze marketing leads in the data stream in real time, providing strong support for real-time decision-making by businesses. This enables businesses to quickly respond to market changes, seize business opportunities, and improve market competitiveness. Attached Figure Description

[0078] Figure 1 is a flowchart of the method described in this invention;

[0079] Figure 2 is a system block diagram of the system described in this invention. Detailed Implementation

[0080] The preferred embodiments of the present invention will be described below with reference to the accompanying drawings. It should be understood that the preferred embodiments described herein are for illustration and explanation only and are not intended to limit the present invention.

[0081] This invention proposes a marketing lead scoring method based on data source hierarchy, as shown in Figure 1. The marketing lead scoring method based on data source hierarchy includes:

[0082] S1. Extract the data flow feature data information and data source classification scheme of the data source corresponding to the marketing leads; wherein, the data source classification scheme includes a first classification scheme and a second classification scheme;

[0083] S2. Classify the data sources corresponding to the marketing leads using the data stream feature data information and data source classification scheme to obtain data sources of different categories;

[0084] S3. Using the target grading coefficients corresponding to different types of data sources and the preset initial data collection frequency, set the data collection frequency corresponding to the different types of data sources; and collect marketing lead data according to the set data collection frequency to obtain marketing lead data.

[0085] S4. Perform a data quality and security assessment on the marketing lead data to obtain marketing lead data that meets preset security standards;

[0086] S5. Retrieve the deep learning model used for scoring, and use the deep learning model used for scoring to obtain the scoring score corresponding to the marketing lead data that meets the preset security standards.

[0087] The working principle of the above technical solution is as follows: This step mainly focuses on the data flow characteristics of the data source, namely the source, flow mode, and update frequency of the data. Combining the first and second classification schemes, the data source is meticulously divided to obtain data sources of different levels. The purpose of this classification is to provide targeted guidance for subsequent data collection, ensuring that the collected data is both comprehensive and representative.

[0088] Different data collection frequencies are set according to different levels of data sources. This ensures that the collected marketing lead data is both timely and meets the actual needs of the enterprise. Before inputting marketing lead data into the deep learning model, its data quality needs to be evaluated. Data quality evaluation may include checks on data completeness, accuracy, consistency, etc. Only data that meets security requirements, that is, data whose quality reaches a certain standard, will be used in subsequent scoring stages. The purpose of this step is to ensure that the data input into the model is reliable and valid, thereby avoiding incorrect scoring results generated by the model due to input data quality issues.

[0089] Marketing lead data that has passed security checks is fed into a trained and optimized deep learning model. The deep learning model automatically learns complex features from the data and scores the marketing leads based on these features. The scoring reflects the potential value of the marketing lead, i.e., its likelihood of converting into an actual sale. Businesses can then rank and filter marketing leads based on the scoring results, thereby developing more effective marketing strategies.

[0090] The effects of the above technical solution are as follows: Through automated and intelligent scoring methods, enterprises can quickly filter out high-value marketing leads, reducing the manual screening costs of a large number of low-quality leads, thereby improving marketing efficiency. Deep learning models can automatically learn complex features in data and accurately score marketing leads based on these features. This enables enterprises to more accurately identify potential customers and develop more targeted marketing strategies. During data processing, strict quality control and security checks ensure that the data input to the model is reliable and valid. This helps avoid erroneous scoring results due to input data quality issues, thus enhancing data security. This method can process and analyze marketing leads in the data stream in real time, providing strong support for real-time decision-making by enterprises. This enables enterprises to quickly respond to market changes, seize business opportunities, and improve market competitiveness.

[0091] In summary, this technical solution achieves real-time and accurate evaluation of marketing leads through steps such as data source classification, data collection, data security assessment, and deep learning model scoring, providing strong support for enterprises' marketing decisions.

[0092] In one embodiment of the present invention, the data source corresponding to the marketing lead is classified using the data stream feature data information and the data source classification scheme to obtain different categories of data sources, including:

[0093] S201. Extract the data flow characteristics of each data source, wherein the data flow characteristic data information includes the data volume, data integrity, data generation or update frequency and data redundancy ratio corresponding to the data flow per unit time.

[0094] S202. Extract the amount of data corresponding to the data stream per unit time for each data source;

[0095] S203. Obtain data flow parameters by utilizing the data volume corresponding to the data flow per unit time of each data source;

[0096] The data traffic parameters are obtained using the following formula:

[0097] Where P represents the data flow parameter; n represents the number of time intervals in which the data source runs, and the time interval is 1 second; D i D represents the amount of data corresponding to the data stream in the i-th unit of time; n D represents the total amount of data corresponding to a data stream spanning n units of time; fmax D represents the maximum value of the change in data volume corresponding to a data stream over n units of time; bi D represents the standard deviation of the data volume corresponding to the data stream at the i-th unit of time; maxThis represents the maximum amount of data corresponding to a data stream spanning n units of time.

[0098] S204. Compare the data traffic parameters with a preset data traffic parameter threshold:

[0099] S205. When the data flow parameters exceed the preset data flow parameter threshold, the data source is classified using the first classification scheme.

[0100] S206. When the data flow parameters do not exceed the preset data flow parameter threshold, the data source is classified using the second classification scheme.

[0101] The working principle of the above technical solution is as follows: This step first focuses on the data flow characteristics of the data source, including the data volume corresponding to the data flow per unit time, data integrity, data generation or update frequency, and data redundancy ratio. These characteristics can comprehensively reflect the status and quality of the data source, providing an important basis for subsequent classification. The data volume corresponding to the data flow within a unit time (e.g., 1 second) of each data source is extracted. Based on these data volumes, a data flow parameter P is obtained through complex calculation formulas (involving data volume, maximum data volume variation, data volume standard deviation, maximum data volume, etc.). This parameter is a comprehensive indicator that can quantify the data flow characteristics of the data source, facilitating subsequent comparison and classification. The calculated data flow parameter P is compared with a preset data flow parameter threshold. The purpose of this step is to divide the data sources into two categories: one is data sources with high data flow parameters, and the other is data sources with low data flow parameters. For data sources whose data flow parameters exceed the threshold, the first classification scheme is used for classification. For data sources whose data flow parameters do not exceed the threshold, the second classification scheme is used for classification. These two classification strategies may be based on different standards and methods to adapt to the characteristics and needs of different data sources.

[0102] The above technical solution achieves the following results: By deeply analyzing the data flow characteristics of the data source and combining them with complex data flow parameter calculation formulas, this solution enables precise classification of data sources. This helps enterprises better understand and manage data sources, improving data utilization efficiency and accuracy. The solution employs automated data processing and classification workflows, reducing the impact of manual intervention and subjective judgment. This improves the efficiency and accuracy of classification processing, reducing the enterprise's labor and time costs. The solution provides two classification strategies (a first classification scheme and a second classification scheme) to adapt to the characteristics and needs of different data sources. This increases the flexibility and applicability of the solution, enabling its widespread application in various scenarios and fields. By precisely classifying data sources, enterprises can formulate different data collection, storage, and processing strategies based on different levels of data sources. This helps optimize resource allocation and improve the efficiency and effectiveness of data processing.

[0103] On the other hand, by extracting characteristics such as data volume, data integrity, data generation or update frequency, and data redundancy ratio corresponding to the data stream per unit time, and combining them with complex data flow parameter calculation formulas, the characteristics of the data source can be quantified more accurately. This helps eliminate the subjectivity and uncertainty in manual classification, improving the accuracy and objectivity of classification. Based on the comparison results of data flow parameters with preset thresholds, a first or second classification scheme can be flexibly selected to adapt to the characteristics and needs of different data sources. This dynamically adjusted classification strategy helps optimize resource allocation and improve the efficiency and accuracy of classification processing. The entire classification process adopts an automated workflow, including data extraction, calculation, comparison, and classification steps, reducing manual intervention and repetitive work. This helps shorten processing time, improve processing efficiency, and reduce the enterprise's labor and time costs. This technical solution can extract and analyze the data flow characteristic data information of the data source in real time and make classification decisions quickly. This helps enterprises respond promptly to market changes and data source fluctuations, improving the real-time performance and flexibility of data processing. Through precise classification, enterprises can formulate different data collection, storage, and processing strategies according to different levels of data sources. This helps optimize resource allocation, avoid resource waste and redundancy, and improve resource utilization efficiency. Automated processing and precise data classification help reduce the labor and time costs for enterprises in data processing and classification. Simultaneously, optimized resource allocation can further reduce storage and maintenance costs. This technical solution is adaptable to data sources of different types, sizes, and domains, exhibiting strong adaptability and versatility. This allows enterprises to flexibly apply the solution in various scenarios, improving system stability and reliability. During data processing and classification, the solution fully considers the characteristics and needs of the data source and takes corresponding measures to address potential problems and anomalies. This helps enhance the system's fault tolerance and robustness, improving the accuracy and reliability of data processing.

[0104] In summary, the technical benefits of the above-mentioned solution in terms of performance indicators are mainly reflected in improved classification accuracy, increased processing efficiency, resource optimization and cost reduction, and enhanced system stability and reliability. These technical benefits help enterprises better manage and utilize data sources, improve data processing efficiency and accuracy, reduce labor and time costs, and provide strong support for business development and decision-making. Furthermore, by deeply analyzing the data flow characteristics of the data source and combining complex data flow parameter calculation formulas with two classification strategies, this solution achieves precise classification and automated processing of the data source. This helps improve data utilization efficiency and accuracy, optimize resource allocation, and reduce enterprise labor and time costs.

[0105] In one embodiment of the present invention, the first classification scheme is as follows:

[0106] Step 1a: When the data flow parameters exceed the preset data flow parameter threshold, retrieve the data integrity, data generation or update frequency and data redundancy ratio contained in the data flow feature data information;

[0107] Step 2a: Retrieve data flow parameters;

[0108] Step 3a: Obtain the first data classification coefficient using the data traffic parameters, data integrity, data generation or update frequency, and data redundancy ratio;

[0109] The first data grading coefficient is obtained by the following formula:

[0110] Among them, S 01 The first data classification coefficient is represented by ; P represents the data flow parameter; n represents the number of time intervals in which the data source runs, and the time interval is 1 second; f i R represents the data generation or update frequency corresponding to the i-th unit of time; i D represents the percentage of data redundancy corresponding to the i-th unit of time; i W represents the amount of data in the data stream corresponding to the i-th unit of time; i Indicates the data integrity corresponding to the i-th unit of time; f b W represents the standard deviation of the data generation or update frequency corresponding to n units of time; bi This represents the standard deviation of data integrity corresponding to the i-th unit of time;

[0111] Step 4a: Compare the first data grading coefficient with the preset first grading coefficient threshold;

[0112] Step 5a: When the first data classification coefficient exceeds the preset first classification coefficient threshold, the data source corresponding to the first data classification coefficient exceeding the preset first classification coefficient threshold is taken as the first-level data source.

[0113] Step 6a: When the first data classification coefficient does not exceed the preset first classification coefficient threshold, the data source corresponding to the first data classification coefficient not exceeding the preset first classification coefficient threshold is taken as the secondary data source.

[0114] The working principle of the above technical solution is as follows: When the data flow parameters exceed a preset threshold, key parameters such as data integrity, data generation or update frequency, and data redundancy ratio contained in the data flow characteristic data information are first retrieved. Simultaneously, previously calculated data flow parameters are also retrieved for subsequent calculation of the first data classification coefficient. Using the above formula, combined with the data flow parameters, data integrity, data generation or update frequency, data redundancy ratio, and their changes over different units of time (such as standard deviation), the first data classification coefficient S01 is calculated. This coefficient is a comprehensive indicator that fully reflects the performance of the data source across multiple dimensions. The calculated first data classification coefficient S01 is compared with the preset first classification coefficient threshold. The purpose of this step is to classify the data sources into two categories: one is data sources with higher classification coefficients (primary data sources), and the other is data sources with lower classification coefficients (secondary data sources). Data sources with classification coefficients exceeding the threshold are marked as primary data sources. Data sources with classification coefficients below the threshold are marked as secondary data sources. This classification method helps enterprises better understand and manage data sources, providing guidance for subsequent data collection, storage, and processing.

[0115] The above technical solution achieves the following effects: By comprehensively considering multiple dimensions (such as data flow, data integrity, data generation or update frequency, and data redundancy ratio), it enables precise classification of data sources. This helps enterprises better understand the quality and characteristics of data sources, providing strong support for subsequent data processing. The entire classification process is automated, reducing the impact of manual intervention and subjective judgment. This improves the efficiency and accuracy of classification, and reduces the enterprise's labor and time costs. By accurately classifying data sources, enterprises can formulate different data collection, storage, and processing strategies based on different levels of data sources. This helps optimize resource allocation and improve the efficiency and effectiveness of data processing. By comprehensively considering multiple dimensions to evaluate the quality of data sources, this technical solution helps to screen out high-quality data sources, improving the accuracy and reliability of data processing. This technical solution can adapt to data sources of different types, sizes, and fields, exhibiting strong adaptability and flexibility. This helps enterprises flexibly apply this technical solution in different scenarios, improving system stability and reliability.

[0116] On the other hand, this technical solution achieves a comprehensive evaluation of data sources by considering multiple dimensions such as data flow parameters, data integrity, data generation or update frequency, and data redundancy ratio. This multi-dimensional evaluation method helps to more accurately reflect the actual quality and characteristics of data sources, thereby improving the accuracy of classification. This technical solution can dynamically adjust the classification strategy based on actual data flow characteristics to adapt to the characteristics and needs of different data sources. This dynamic adjustment mechanism helps maintain the accuracy and effectiveness of classification, avoiding classification deviations caused by changes in data sources. The entire classification process adopts an automated workflow, including data extraction, calculation, comparison, and classification steps, reducing manual intervention and repetitive work. This helps to shorten processing time, improve processing efficiency, and reduce the enterprise's labor and time costs. This technical solution utilizes efficient calculation formulas and algorithms to quickly calculate the first data classification coefficient. This helps to improve the efficiency of classification processing while ensuring classification accuracy. By accurately classifying data sources, enterprises can formulate different data collection, storage, and processing strategies based on different levels of data sources. This helps to optimize resource allocation, avoid resource waste and redundancy, and improve resource utilization efficiency. Automated processing and precise data classification help reduce the labor and time costs for enterprises in data processing and classification. Simultaneously, optimized resource allocation can further reduce storage and maintenance costs. This technical solution is adaptable to data sources of different types, sizes, and domains, exhibiting strong adaptability and versatility. This allows enterprises to flexibly apply the solution in various scenarios, improving system stability and reliability. During data processing and classification, the solution fully considers the characteristics and needs of the data source and takes corresponding measures to address potential problems and anomalies. This helps enhance the system's fault tolerance and robustness, improving the accuracy and reliability of data processing.

[0117] In summary, the technical benefits of the above-mentioned solution in terms of performance indicators are mainly reflected in improved classification accuracy, increased processing efficiency, optimized resource allocation, and enhanced system stability. These benefits help enterprises better manage and utilize data sources, improve data processing efficiency and accuracy, reduce labor and time costs, and provide strong support for business development and decision-making. Furthermore, the first classification scheme of this technical solution comprehensively considers multiple dimensions to evaluate the quality of data sources and achieves accurate classification of data sources. This helps enterprises better understand and manage data sources, optimize resource allocation, and improve the efficiency and accuracy of data processing.

[0118] In one embodiment of the present invention, the second classification scheme is as follows:

[0119] Step 1b: When the data flow parameters do not exceed the preset data flow parameter threshold, retrieve the data integrity, data generation or update frequency, and data redundancy ratio contained in the data flow feature data information;

[0120] Step 2b: Retrieve data flow parameters;

[0121] Step 3b: Retrieve the rate of change of the number of data types in the data source per unit of time;

[0122] Step 4b: Using the data flow parameters combined with the rate of change of data type quantity per unit time, data integrity, data generation or update frequency, and data redundancy ratio, obtain the second data classification coefficient;

[0123] The second data grading coefficient is obtained using the following formula:

[0124] Among them, S 02 P represents the data classification coefficient; P represents the data flow parameter; P y This represents a preset data flow parameter threshold; n represents the number of time intervals in which the data source runs, and the time interval is 1 second; f i R represents the data generation or update frequency corresponding to the i-th unit of time; i D represents the percentage of data redundancy corresponding to the i-th unit of time; i W represents the amount of data in the data stream corresponding to the i-th unit of time; i Indicates the data integrity corresponding to the i-th unit of time; f b W represents the standard deviation of the data generation or update frequency corresponding to n units of time; bi N represents the standard deviation of data integrity corresponding to the i-th unit of time; i N represents the rate of change in the number of data types corresponding to the i-th unit of time; b The standard deviation of the rate of change of the number of data types over n units of time;

[0125] Step 5b: Compare the second data grading coefficient with the preset second grading coefficient threshold;

[0126] Step 6b: When the second data classification coefficient exceeds the preset second classification coefficient threshold, the data source corresponding to the second data classification coefficient exceeding the preset second classification coefficient threshold is taken as the first-level data source.

[0127] Step 7b: When the second data classification coefficient does not exceed the preset second classification coefficient threshold, the data source corresponding to the second data classification coefficient not exceeding the preset second classification coefficient threshold is taken as the secondary data source.

[0128] The working principle of the above technical solution is as follows: the strategy is triggered when the data flow parameters do not exceed the preset data flow parameter threshold. Subsequently, key parameters such as data integrity, data generation or update frequency, and data redundancy ratio contained in the data flow characteristic data information are retrieved. Simultaneously, the data flow parameters and the rate of change of the number of data types per unit time of the data source are also retrieved. Using a complex formula, combining data flow parameters, the rate of change of the number of data types, and multiple dimensions such as data integrity, data generation or update frequency, and data redundancy ratio, a second data classification coefficient S02 is calculated. This coefficient is a comprehensive indicator that can fully reflect the performance of the data source in multiple dimensions and takes into account the impact of the rate of change of the number of data types on data quality. The calculated second data classification coefficient S02 is compared with the preset second classification coefficient threshold. Based on the comparison results, the data sources are divided into two categories: one is data sources with higher classification coefficients (primary data sources), and the other is data sources with lower classification coefficients (secondary data sources). Data sources with classification coefficients exceeding the threshold are marked as primary data sources, indicating that their data quality is high and deserves special attention and priority processing. For data sources whose grading coefficient does not exceed the threshold, they are marked as secondary data sources, indicating that their data quality is relatively low, but they may still have some utilization value.

[0129] The above technical solution achieves the following effects: By comprehensively considering multiple dimensions such as data flow parameters, the rate of change in the number of data types, data integrity, data generation or update frequency, and data redundancy ratio, this solution enables a comprehensive assessment of data source quality. This helps enterprises more accurately understand the actual quality and characteristics of data sources, providing strong support for subsequent data collection, storage, and processing. Introducing the rate of change in the number of data types as one of the evaluation dimensions makes the grading strategy more detailed and precise. This helps enterprises better identify high-quality data sources and prioritize their processing and utilization. Through precise grading of data sources, enterprises can formulate different data collection, storage, and processing strategies based on different levels of data sources. This helps optimize resource allocation, avoid resource waste and redundancy, and improve resource utilization efficiency. This technical solution is adaptable to data sources of different types, sizes, and fields, exhibiting strong adaptability and flexibility. This helps enterprises flexibly apply this technical solution in different scenarios, improving system stability and reliability. Precise data source grading helps enterprises better conduct data governance and decision analysis. Enterprises can formulate more reasonable data management and utilization strategies based on the grading results, improving data quality and data value.

[0130] On the other hand, the technical solution introduces the rate of change in the number of data types as a new evaluation dimension, which, together with data flow parameters, data integrity, data generation or update frequency, and data redundancy ratio, constitutes a comprehensive evaluation system. This multi-dimensional comprehensive evaluation method can more accurately reflect the performance and characteristics of data sources in different aspects, thereby improving the accuracy of classification. The classification strategy in the technical solution can dynamically adjust the evaluation indicators and weights according to the actual performance of the data source to adapt to data sources of different types, sizes, and fields. This dynamic adaptability helps maintain the accuracy and effectiveness of classification, accurately assessing its quality and value even when the data source changes. The calculation formulas and algorithms in the technical solution have been optimized to quickly calculate the second data classification coefficient in a short time. This helps reduce the time consumption in the evaluation process, improves processing efficiency, and thus supports enterprises to make decisions and respond faster. The entire classification process adopts an automated workflow, including data extraction, calculation, comparison, and classification steps, reducing manual intervention and repetitive work. The automated processing workflow not only improves processing efficiency but also reduces the risk of human error, ensuring the accuracy and consistency of classification. Based on the classification results, enterprises can formulate different data collection, storage, and processing strategies for data sources of different levels. For high-quality data sources (primary data sources), priority can be given to collection and processing to ensure data accuracy and timeliness; for lower-quality data sources (secondary data sources), appropriate processing measures or further analysis and screening can be taken based on the actual situation. Through precise hierarchical and optimized configuration strategies, enterprises can more effectively utilize limited resources and avoid waste and redundancy. This helps reduce the enterprise's data collection, storage, and processing costs, and improves overall economic efficiency.

[0131] The tiered strategy in the technical solution fully considers the performance and characteristics of data sources in different aspects, and can tolerate a certain degree of data fluctuation and anomalies. This helps enhance the system's fault tolerance and robustness, ensuring stable operation even when data sources change or anomalies occur. The technical solution is designed with good scalability, allowing for flexible adjustments and optimizations as the enterprise's business develops and needs change. This helps maintain the system's advanced nature and competitiveness, providing strong support for the enterprise's future business development.

[0132] In summary, the technical benefits of the above-mentioned solutions in terms of performance indicators are mainly reflected in significantly improved classification accuracy, optimized processing efficiency, further optimized resource utilization, and enhanced system stability and reliability. These technical effects help enterprises better manage and utilize data sources, improve the efficiency and accuracy of data processing, reduce enterprise costs, and provide strong support for business development and decision-making. Furthermore, the second classification scheme in the above-mentioned technical solutions comprehensively considers multiple dimensions to evaluate the quality of data sources and achieves accurate classification of data sources. This helps enterprises better understand and manage data sources, optimize resource allocation, and improve the efficiency and accuracy of data processing.

[0133] One embodiment of the present invention utilizes target grading coefficients corresponding to different categories of data sources in combination with a preset initial data collection frequency to set data collection frequencies corresponding to the different categories of data sources, including:

[0134] S301a. Retrieve the first data classification coefficient or the second data classification coefficient corresponding to the primary data source as the first target classification coefficient;

[0135] S302a. Retrieve the preset initial data acquisition frequency from the database;

[0136] S303a. Obtain the data acquisition frequency corresponding to the first-level data source by combining the first target classification coefficient with the preset initial data acquisition frequency;

[0137] The data acquisition frequency corresponding to the primary data source is obtained by the following formula:

[0138] Among them, F 01 Indicates the data collection frequency corresponding to the primary data source; F c Indicates the preset initial data acquisition frequency; S m01 S represents the first objective classification coefficient; y01 This represents the threshold value of the grading coefficient corresponding to the first target grading coefficient.

[0139] The working principle of the above technical solution is as follows: When it is necessary to set the data collection frequency for a primary data source, the first data grading coefficient or the second data grading coefficient corresponding to that data source is first retrieved (in this solution, it is represented by the first target grading coefficient Sm01). This grading coefficient is the result of a comprehensive evaluation of multiple dimensions such as data source quality, importance, and change frequency, and can reflect the actual value and needs of the data source. The preset initial data collection frequency Fc is retrieved from the database. This initial frequency is a baseline value preset based on factors such as system requirements, resource constraints, and data processing capabilities. Using the first target grading coefficient and the preset initial data collection frequency, the data collection frequency corresponding to the primary data source is calculated using a specific formula.

[0140] The effects of the above technical solution are as follows: By setting different data collection frequencies for different levels of data sources, it ensures that important and high-quality data sources are collected more frequently and promptly, thereby improving the efficiency and accuracy of data collection. Simultaneously, for data sources of lower quality or with infrequent changes, the collection frequency can be appropriately reduced to minimize unnecessary resource waste and data processing burden. For Tier 1 data sources, due to their high data quality and importance, setting a higher data collection frequency ensures that the system can obtain the latest data in a timely manner, enabling rapid response and decision-making. This helps improve the overall performance and response speed of the system, providing enterprises with more timely and accurate data support. By comprehensively considering multiple dimensions such as the quality, importance, and frequency of change of the data source when setting the data collection frequency, it ensures that the collected data is more reliable and accurate. This helps enterprises better understand business conditions and market dynamics, providing stronger support for decision-making. By optimizing the data collection frequency, enterprises can utilize resources more rationally, avoiding unnecessary waste and redundancy. This helps reduce the enterprise's data collection, storage, and processing costs, improving the enterprise's overall economic efficiency.

[0141] On the other hand, the technical solution can dynamically adjust the data collection frequency based on the primary target grading coefficient (i.e., data grading coefficient) of the primary data source. When the quality, importance, or frequency of change of the data source changes, the collection frequency can be adjusted accordingly to adapt to actual needs. This dynamic adjustment mechanism helps improve the efficiency and accuracy of data collection, ensuring that the system can acquire important and valuable data in a timely manner. For data sources with lower quality or infrequent changes, the technical solution can reduce their data collection frequency, thereby reducing unnecessary redundant collection. This helps alleviate the system's processing burden, improve resource utilization efficiency, and reduce the enterprise's data collection costs. By setting a higher data collection frequency for primary data sources, the technical solution can ensure that the system can acquire the latest data in real time. This helps improve the system's real-time performance, enabling enterprises to respond quickly to market changes and business needs. The acquisition of real-time data provides stronger support for enterprise decision-making. Enterprises can analyze and predict based on the latest data, thereby making more accurate and timely decisions. By comprehensively considering multiple dimensions such as the quality, importance, and frequency of change of the data source when setting the data collection frequency, the technical solution helps improve the quality and reliability of the collected data. This helps enterprises better understand business conditions and market dynamics, providing more accurate data support for decision-making. A high data acquisition frequency ensures data integrity and continuity, reducing the risk of data loss or omission. This helps enterprises build more complete and accurate datasets, providing strong support for subsequent data analysis and mining. The technical solution optimizes resource allocation by dynamically adjusting the data acquisition frequency. The system can rationally allocate resources according to actual needs, avoiding waste and redundancy. This helps reduce operating costs and improve overall economic efficiency. The technical solution is designed with good scalability, allowing for flexible adjustment and optimization as the enterprise's business develops and needs change. This helps maintain the system's advanced nature and competitiveness, providing strong support for the enterprise's future business development.

[0142] In summary, the technical benefits of the above-mentioned solutions in terms of performance indicators are mainly reflected in the optimization of data acquisition efficiency, the improvement of system response speed, the enhancement of data reliability and accuracy, and the improvement of resource utilization. These effects help enterprises better manage and utilize data resources, thereby enhancing their competitiveness and market position. Furthermore, by setting different data acquisition frequencies for different levels of data sources, the above-mentioned solutions achieve multiple technical benefits, including optimized data acquisition efficiency, improved system response speed, enhanced data reliability and accuracy, and reduced enterprise costs. These effects also help enterprises better manage and utilize data resources, thereby enhancing their competitiveness and market position.

[0143] In one embodiment of the present invention, the data acquisition frequency corresponding to the different categories of data sources is set by combining the target classification coefficients corresponding to the different categories of data sources with a preset initial data acquisition frequency, and further includes:

[0144] S301b: Retrieve the first data classification coefficient or the second data classification coefficient corresponding to the secondary data source as the second target classification coefficient;

[0145] S302b: Retrieve the preset initial data acquisition frequency from the database;

[0146] S303b: Obtain the data acquisition frequency corresponding to the secondary data source by combining the second target classification coefficient with the preset initial data acquisition frequency;

[0147] The data acquisition frequency corresponding to the secondary data source is obtained by the following formula:

[0148] Among them, F 02 Indicates the data collection frequency corresponding to the secondary data source; F c Indicates the preset initial data acquisition frequency; S m02 S represents the second objective classification coefficient; y02 This represents the threshold value of the grading coefficient corresponding to the second objective grading coefficient.

[0149] The working principle of the above technical solution is as follows: First, the system needs to retrieve the first or second data grading coefficient corresponding to the secondary data source as the second target grading coefficient. This step is to evaluate dimensions such as the importance, frequency of change, or data quality of the secondary data source in order to set an appropriate data collection frequency for it. Next, the system retrieves a preset initial data collection frequency from the database. This initial frequency is a baseline value used for subsequent adjustments based on the specific circumstances of the data source. Using the second target grading coefficient and the preset initial data collection frequency, the system calculates the data collection frequency corresponding to the secondary data source using a specific formula. This formula considers the impact of the grading coefficient on the data collection frequency, as well as the role of the grading coefficient threshold. The grading coefficient threshold is a key parameter used to determine the specific degree of influence of the grading coefficient on the data collection frequency.

[0150] The above technical solution achieves the following effects: By setting different data collection frequencies for different levels of data sources, the system can collect and process data more efficiently. For secondary data sources, setting the frequency based on their importance, frequency of change, or data quality helps reduce unnecessary redundant collection and improves the targeting and efficiency of data collection. A reasonable data collection frequency setting can reduce the system's processing burden and improve overall system performance. For secondary data sources, if their data changes infrequently or is relatively low in importance, reducing the collection frequency can reduce system resource consumption and improve system stability and response speed. Setting the data collection frequency by comprehensively considering factors such as the importance and frequency of change of the data source helps improve the reliability and accuracy of the collected data. For secondary data sources, if their data quality is high or changes frequently, increasing the collection frequency can ensure the timeliness and completeness of the data, thereby enhancing the reliability and accuracy of the data. A reasonable data collection frequency setting helps reduce the enterprise's data collection, storage, and processing costs. For secondary data sources, if their data changes infrequently or is low in importance, reducing the collection frequency can reduce the need for data storage and processing, thereby reducing the enterprise's operating costs. This technical solution has good flexibility and scalability. As a business grows and its needs change, the data collection frequency settings of different data sources can be easily adjusted to adapt to new business scenarios and requirements.

[0151] On the other hand, this technical solution significantly improves data acquisition efficiency by setting different data acquisition frequencies for different levels of data sources (such as secondary data sources). For data sources with frequent changes or high importance, increasing the acquisition frequency ensures the timeliness and accuracy of the data; for data sources with infrequent changes or lower importance, decreasing the acquisition frequency reduces redundant data collection and improves resource utilization. By rationally setting the data acquisition frequency, this technical solution can reduce resource consumption during data acquisition, storage, and processing. This not only helps reduce the company's operating costs but also improves the overall energy efficiency of the system. A reasonable data acquisition frequency setting can shorten the time for the system to obtain the latest data, thereby improving the system's response speed. This is particularly important for business scenarios that require real-time data processing and analysis. By increasing the data acquisition frequency, this technical solution ensures that the system can obtain the latest data in a timely manner, thereby enhancing the system's real-time performance. This is of great significance for business scenarios that require rapid response, such as monitoring and early warning. This technical solution helps improve the quality of the collected data by comprehensively considering factors such as the importance and frequency of change of the data source. This not only reduces data errors and omissions but also improves the accuracy and completeness of the data. Setting a reasonable data collection frequency ensures data timeliness and continuity, thereby enhancing data reliability. This is especially important for business scenarios that require long-term data tracking and analysis. This technical solution allows for flexible adjustment of the data collection frequency based on the actual situation of the data source and business needs. It can adapt to different levels of data sources and can be dynamically adjusted as business develops and needs change. The solution is designed with good scalability and maintainability. As business grows and new data sources emerge, new data sources can be easily added and corresponding collection frequencies set without requiring large-scale modifications or reconstruction of the system.

[0152] In summary, the technical benefits of the above-mentioned solution in terms of performance indicators are mainly reflected in improved data acquisition efficiency and resource utilization, enhanced system response speed and real-time performance, improved data quality and reliability, and increased flexibility and scalability. These effects help enterprises better manage and utilize data resources, thereby enhancing their competitiveness and market position. Furthermore, by setting different data acquisition frequencies for different levels of data sources, this solution achieves multiple technical benefits, including optimized data acquisition efficiency, improved system performance, enhanced data reliability and accuracy, and reduced enterprise costs. These effects also help enterprises better manage and utilize data resources, thereby enhancing their competitiveness and market position.

[0153] One embodiment of the present invention involves performing a data quality and security assessment on the marketing lead data to obtain marketing lead data that meets preset security standards, including:

[0154] S401. Extract the data characteristic parameters of each collected marketing lead data, wherein the data characteristic parameters include the encryption rate of the marketing lead data collected per unit time, the encryption resource utilization rate, and the data backup frequency.

[0155] S402. Obtain a security evaluation coefficient by using the encryption rate, encryption resource utilization rate and data backup frequency of the marketing lead data collected per unit time.

[0156] The safety evaluation coefficient is obtained using the following formula:

[0157] Where K represents the safety evaluation coefficient; m represents the number of times marketing lead data was collected; J i U represents the encryption rate corresponding to the marketing lead data collected in the i-th instance; i B represents the encrypted resource utilization rate corresponding to the marketing lead data collected in the i-th instance; i B represents the data backup frequency corresponding to the marketing lead data collected in the i-th instance; c This indicates the preset data backup frequency reference value;

[0158] S403. Compare the safety evaluation coefficient with the preset safety coefficient threshold.

[0159] S404. When the security evaluation coefficient exceeds the preset security coefficient threshold, the collected marketing lead data is determined to meet the security requirements.

[0160] The working principle of the above technical solution is as follows: Data characteristic parameters are extracted from each collected marketing lead data set. These parameters include the encryption rate, encryption resource utilization rate, and data backup frequency of the marketing lead data collected per unit time. These parameters reflect data confidentiality, resource utilization efficiency, and data recovery capability, and are important indicators for assessing data security. Using the extracted data characteristic parameters, a security evaluation coefficient (K) is calculated using a specific formula. This formula comprehensively considers the impact of encryption rate, encryption resource utilization rate, and data backup frequency on data security and provides a quantitative evaluation result. The calculated security evaluation coefficient K is compared with a preset security coefficient threshold. If the K value exceeds the threshold, the collected marketing lead data is deemed to meet security requirements; otherwise, it does not meet the requirements, and further security measures or re-collection of data may be necessary.

[0161] The above technical solution achieves the following results: By extracting data characteristic parameters and calculating security evaluation coefficients, this solution can quantitatively assess the security of marketing lead data. This provides enterprises with an objective and accurate evaluation standard, helping them to better manage and protect data. By comparing the security evaluation coefficient with a preset security threshold, this solution can promptly identify potential data security risks. This helps enterprises take timely measures, such as strengthening data encryption, improving resource utilization, or increasing data backup frequency, thereby enhancing data security capabilities. Furthermore, this solution can optimize data collection strategies based on feedback from the security evaluation coefficients. For example, for data collection batches with lower security, the collection method can be adjusted, the collection frequency increased, or more advanced data encryption technologies adopted to improve data quality and security. By ensuring the security and accuracy of marketing lead data, this solution helps enterprises formulate more effective marketing strategies and decisions. This not only improves customer satisfaction and loyalty but also enhances the enterprise's market competitiveness, bringing greater business value.

[0162] Meanwhile, this technical solution achieves precise quantitative assessment of data security by introducing data characteristic parameters (such as encryption rate, encryption resource utilization, and data backup frequency) and security evaluation coefficients. Compared to traditional subjective judgment or qualitative assessment methods, this quantitative assessment method is more accurate and objective, helping enterprises to understand the security status of their data more clearly. By automatically extracting data characteristic parameters and calculating security evaluation coefficients, this technical solution significantly improves the efficiency of data security assessment. Enterprises can obtain data security assessment results more quickly, thereby taking corresponding security measures in a timely manner to reduce data security risks. Based on the feedback results of the security evaluation coefficients, enterprises can optimize their data acquisition strategies. For example, for data acquisition batches with lower security, they can adjust the acquisition method, increase the acquisition frequency, or adopt more advanced data encryption technologies to improve data quality and security. By focusing on parameters such as data backup frequency, this technical solution helps enterprises optimize their data storage strategies. A reasonable data backup frequency not only ensures data integrity and availability but also reduces unnecessary storage resource consumption and improves storage efficiency. Because this technical solution can promptly identify potential data security risks and take corresponding security measures to prevent them, it helps reduce system downtime or failure time caused by data security issues. This improves the system's processing speed and stability to some extent. By quantitatively assessing data security, businesses can more quickly identify and resolve data security incidents. This helps enhance a company's emergency response capabilities, ensuring that swift action can be taken to address data security incidents and minimize losses. Ensuring the security and accuracy of marketing lead data is crucial for a company's market competitiveness. This technology solution, by improving data security protection capabilities, helps companies develop more effective marketing strategies and decisions, thereby enhancing their market competitiveness. Data security is an important component of a company's reputation. By implementing this technology solution, companies can demonstrate their commitment to and emphasis on data security, thereby strengthening customer trust and loyalty.

[0163] In summary, the technical benefits of this solution in terms of performance indicators are mainly reflected in the quantitative assessment of data security, optimization of data collection and storage, enhanced system response and processing capabilities, and improved corporate competitiveness and reputation. These effects collectively constitute the significant performance advantages of this technical solution. Furthermore, by quantitatively assessing data security, improving data security protection capabilities, optimizing data collection strategies, and enhancing corporate market competitiveness, this technical solution provides enterprises with a comprehensive and effective data security solution.

[0164] In one embodiment of the present invention, a deep learning model for scoring is invoked, and the scoring score corresponding to the marketing lead data that meets the preset security standard is obtained using the deep learning model for scoring, including:

[0165] S501. Retrieve the trained and optimized deep learning model from the database;

[0166] S502. Extract the optimization completion time of the most recent optimization of the deep learning model;

[0167] S503. Obtain the optimization completion time length based on the optimization completion time and the current model call time;

[0168] S504. Compare the optimization completion time with a preset time length threshold.

[0169] S505. When the optimization completion time exceeds a preset time length threshold, the deep learning model is re-optimized to obtain the optimized deep learning model.

[0170] S506. Input the marketing lead data that meets the security requirements into the optimized deep learning model;

[0171] S507. The marketing lead data that meets the security requirements is scored using the deep learning model to obtain the corresponding score of the marketing lead data.

[0172] The structure of the deep learning model is as follows:

[0173] The input layer is used to classify the security-compliant marketing lead data, obtain structured and unstructured data, and preprocess the structured and unstructured data to obtain preprocessed data information. The structured data includes basic customer information (such as age, gender, and geographic location), historical purchase records, and interaction records. The unstructured data includes text data (such as customer feedback and social media comments) and image data (such as product images). The preprocessing includes standardization and normalization, scaling of numerical features, and word segmentation and vectorization of text data (such as Word2Vec, TF-IDF, or BERT embedding).

[0174] The feature extraction layer is used to extract and construct features corresponding to the preprocessed data information so that the model can learn more effectively.

[0175] The model inference layer is used by the deep neural network (DNN) model to generate scoring results based on the features corresponding to the data information.

[0176] The output layer is used to output the scoring results.

[0177] The working principle of the above technical solution is as follows: A trained and optimized deep learning model is retrieved from the database. The optimization completion time of the most recent optimization of the deep learning model is extracted. The optimization completion time is calculated based on the optimization completion time and the current model call time. The optimization completion time is compared with a preset time length threshold. If the optimization completion time exceeds the preset time length threshold, the deep learning model is re-optimized to obtain an updated model; otherwise, the currently optimized model is used. Marketing lead data that meets security requirements is input into the optimized deep learning model. Inside the model, the input layer first classifies the input marketing lead data, obtaining structured and unstructured data. Structured data (such as basic customer information, historical purchase records, and interaction records) and unstructured data (such as text data and image data) are preprocessed, including standardization, normalization, and text data segmentation and vectorization, to obtain preprocessed data information. The feature extraction layer extracts and constructs corresponding features based on the preprocessed data information. The model inference layer uses a deep neural network (DNN) model to generate scoring results based on the extracted features. The output layer outputs the scoring results for subsequent analysis or decision-making.

[0178] The above technical solution achieves the following effects: By comparing the optimized completion time with a preset time threshold, it ensures the deep learning model maintains a certain level of timeliness during use. This helps avoid performance degradation or inaccurate scoring due to model obsolescence. The input layer classifies and preprocesses the input marketing lead data, improving the efficiency and quality of data processing. Through standardization, normalization, and text data segmentation and vectorization, the data becomes more suitable for processing and analysis by the deep learning model. The feature extraction layer accurately extracts and constructs corresponding features from the preprocessed data. These features reflect the essential attributes and key information of the data, helping the model learn and reason more effectively. The model inference layer uses a deep neural network (DNN) model to generate scoring results based on the extracted features, ensuring the objectivity and accuracy of the scoring process. This helps reduce the influence of human factors on the scoring results, improving the fairness and credibility of the scoring. The deep learning model in this technical solution is scalable and flexible. As business develops and needs change, the model's structure and parameters can be easily adjusted to adapt to new data features and scoring requirements.

[0179] In summary, this technical solution provides strong support for the scoring and analysis of marketing lead data by ensuring the timeliness of the model, improving data processing efficiency, accurately extracting features, generating objective scoring results, and providing scalability and flexibility.

[0180] An embodiment of the present invention proposes a marketing lead scoring system based on data source hierarchy, as shown in Figure 2. The marketing lead scoring system based on data source hierarchy includes:

[0181] The data retrieval module is used to extract data flow feature data information and data source classification schemes corresponding to the marketing leads; wherein, the data source classification schemes include a first classification scheme and a second classification scheme;

[0182] The data source classification module is used to classify the data sources corresponding to the marketing leads by using the data flow feature data information and the data source classification scheme, and to obtain data sources of different categories.

[0183] The data collection frequency setting module is used to set the data collection frequency corresponding to the different categories of data sources by combining the target grading coefficients corresponding to the different categories of data sources with the preset initial data collection frequency; and to collect marketing lead data according to the set data collection frequency to obtain marketing lead data.

[0184] The security assessment module is used to assess the data quality and security of the marketing lead data and obtain marketing lead data that meets preset security standards.

[0185] The scoring control module is used to retrieve the deep learning model used for scoring and to obtain the scoring score corresponding to the marketing lead data that meets the preset security standards using the deep learning model used for scoring.

[0186] The working principle of the above technical solution is as follows: This step mainly focuses on the data flow characteristics of the data source, namely the source, flow mode, and update frequency of the data. Combining the first and second classification schemes, the data source is meticulously divided to obtain data sources of different levels. The purpose of this classification is to provide targeted guidance for subsequent data collection, ensuring that the collected data is both comprehensive and representative.

[0187] Different data collection frequencies are set according to different levels of data sources. This ensures that the collected marketing lead data is both timely and meets the actual needs of the enterprise. Before inputting marketing lead data into the deep learning model, its data quality needs to be evaluated. Data quality evaluation may include checks on data completeness, accuracy, consistency, etc. Only data that meets security requirements, that is, data whose quality reaches a certain standard, will be used in subsequent scoring stages. The purpose of this step is to ensure that the data input into the model is reliable and valid, thereby avoiding incorrect scoring results generated by the model due to input data quality issues.

[0188] Marketing lead data that has passed security checks is fed into a trained and optimized deep learning model. The deep learning model automatically learns complex features from the data and scores the marketing leads based on these features. The scoring reflects the potential value of the marketing lead, i.e., its likelihood of converting into an actual sale. Businesses can then rank and filter marketing leads based on the scoring results, thereby developing more effective marketing strategies.

[0189] The effects of the above technical solution are as follows: Through automated and intelligent scoring methods, enterprises can quickly filter out high-value marketing leads, reducing the manual screening costs of a large number of low-quality leads, thereby improving marketing efficiency. Deep learning models can automatically learn complex features in data and accurately score marketing leads based on these features. This enables enterprises to more accurately identify potential customers and develop more targeted marketing strategies. During data processing, strict quality control and security checks ensure that the data input to the model is reliable and valid. This helps avoid erroneous scoring results due to input data quality issues, thus enhancing data security. This method can process and analyze marketing leads in the data stream in real time, providing strong support for real-time decision-making by enterprises. This enables enterprises to quickly respond to market changes, seize business opportunities, and improve market competitiveness.

[0190] In summary, this technical solution achieves real-time and accurate evaluation of marketing leads through steps such as data source classification, data collection, data security assessment, and deep learning model scoring, providing strong support for enterprises' marketing decisions.

[0191] Obviously, those skilled in the art can make various modifications and variations to this invention without departing from its spirit and scope. Therefore, if these modifications and variations fall within the scope of the claims of this invention and their equivalents, this invention also intends to include these modifications and variations.

Claims

A marketing lead scoring method based on data source hierarchy, characterized in that, The marketing lead scoring method based on data source hierarchy includes: Extract data stream feature information and data source classification schemes corresponding to the data sources used to acquire marketing leads; wherein, the data source classification schemes include a first classification scheme and a second classification scheme; The data stream feature data information and data source classification scheme are used to classify the data sources corresponding to the marketing leads to obtain different categories of data sources; By combining the target grading coefficients corresponding to different categories of data sources with the preset initial data collection frequency, the data collection frequency corresponding to the different categories of data sources is set; and marketing lead data is collected according to the set data collection frequency to obtain marketing lead data; The marketing lead data is assessed for data quality and security to obtain marketing lead data that meets preset security standards. The deep learning model used for scoring is retrieved, and the scoring score corresponding to the marketing lead data that meets the preset security standards is obtained using the deep learning model used for scoring. The marketing lead scoring method based on data source hierarchy as described in claim 1 is characterized in that, The data stream feature information and data source classification scheme are used to classify the data sources corresponding to the marketing leads, and different categories of data sources are obtained, including: Extract the data flow characteristics of each data source, wherein the data flow characteristic data information includes the data volume, data integrity, data generation or update frequency and data redundancy ratio corresponding to the data flow per unit time. Extract the amount of data corresponding to the data stream per unit time for each data source; Data flow parameters are obtained by utilizing the amount of data flowing through each data source per unit time. The data traffic parameters are obtained using the following formula: Where P represents the data flow parameter; n represents the number of time intervals in which the data source runs, and the time interval is 1 second; D i D represents the amount of data corresponding to the data stream in the i-th unit of time; n D represents the total amount of data corresponding to a data stream spanning n units of time; fmax D represents the maximum value of the change in data volume corresponding to a data stream over n units of time; bi D represents the standard deviation of the data volume corresponding to the data stream at the i-th unit of time; max This represents the maximum amount of data corresponding to a data stream spanning n units of time. The data traffic parameters are compared with preset data traffic parameter thresholds: When the data flow parameters exceed the preset data flow parameter threshold, the data source is classified using the first classification scheme. When the data flow parameters do not exceed the preset data flow parameter threshold, the data source is classified using the second classification scheme. The marketing lead scoring method based on data source hierarchy as described in claim 2 is characterized in that, The first rating scheme is as follows: When the data flow parameters exceed the preset data flow parameter threshold, the data integrity, data generation or update frequency and data redundancy ratio contained in the data flow feature data information are retrieved. Retrieve data flow parameters; The first data classification coefficient is obtained by using the data traffic parameters, data integrity, data generation or update frequency, and data redundancy ratio. The first data grading coefficient is obtained by the following formula: Among them, S 01 The first data classification coefficient is represented by ; P represents the data flow parameter; n represents the number of time intervals in which the data source runs, and the time interval is 1 second; f i R represents the data generation or update frequency corresponding to the i-th unit of time; i D represents the percentage of data redundancy corresponding to the i-th unit of time; i W represents the amount of data in the data stream corresponding to the i-th unit of time; i Indicates the data integrity corresponding to the i-th unit of time; f b W represents the standard deviation of the data generation or update frequency corresponding to n units of time; bi This represents the standard deviation of data integrity corresponding to the i-th unit of time; Compare the first data grading coefficient with a preset first grading coefficient threshold; When the first data classification coefficient exceeds the preset first classification coefficient threshold, the data source corresponding to the first data classification coefficient exceeding the preset first classification coefficient threshold is regarded as the first-level data source. When the first data classification coefficient does not exceed the preset first classification coefficient threshold, the data source corresponding to the first data classification coefficient not exceeding the preset first classification coefficient threshold is regarded as the secondary data source. The marketing lead scoring method based on data source hierarchy as described in claim 2 is characterized in that, The second rating scheme is as follows: When the data flow parameters do not exceed the preset data flow parameter threshold, retrieve the data integrity, data generation or update frequency and data redundancy ratio contained in the data flow feature data information; Retrieve data flow parameters; Retrieve the rate of change of the number of data types in the data source per unit of time; The second data classification coefficient is obtained by combining the data flow parameters with the rate of change of data type quantity per unit time, data integrity, data generation or update frequency, and data redundancy ratio. The second data grading coefficient is obtained using the following formula: Among them, S 02 P represents the data classification coefficient; P represents the data flow parameter; P y This represents a preset data flow parameter threshold; n represents the number of time intervals in which the data source runs, and the time interval is 1 second; f i R represents the data generation or update frequency corresponding to the i-th unit of time; i D represents the percentage of data redundancy corresponding to the i-th unit of time; i W represents the amount of data in the data stream corresponding to the i-th unit of time; i Indicates the data integrity corresponding to the i-th unit of time; f b W represents the standard deviation of the data generation or update frequency corresponding to n units of time; bi N represents the standard deviation of data integrity corresponding to the i-th unit of time; i N represents the rate of change in the number of data types corresponding to the i-th unit of time; b The standard deviation of the rate of change of the number of data types over n units of time; Compare the second data grading coefficient with a preset second grading coefficient threshold; When the second data classification coefficient exceeds the preset second classification coefficient threshold, the data source corresponding to the second data classification coefficient exceeding the preset second classification coefficient threshold is regarded as the first-level data source. When the second data grading coefficient does not exceed the preset second grading coefficient threshold, the data source corresponding to the second data grading coefficient not exceeding the preset second grading coefficient threshold is regarded as the secondary data source. The marketing lead scoring method based on data source hierarchy according to claim 3 is characterized in that, Using target grading coefficients corresponding to different categories of data sources, combined with a preset initial data collection frequency, the data collection frequency corresponding to the different categories of data sources is set, including: Retrieve the first data classification coefficient or the second data classification coefficient corresponding to the primary data source as the first target classification coefficient; Retrieve the preset initial data collection frequency from the database; The data acquisition frequency corresponding to the first-level data source is obtained by combining the first target classification coefficient with the preset initial data acquisition frequency; The data acquisition frequency corresponding to the primary data source is obtained by the following formula: Among them, F 01 Indicates the data collection frequency corresponding to the primary data source; F c Indicates the preset initial data acquisition frequency; S m01 S represents the first objective classification coefficient; y01 This represents the threshold value of the grading coefficient corresponding to the first target grading coefficient. The marketing lead scoring method based on data source hierarchy according to claim 3 is characterized in that, The method further includes setting the data collection frequency corresponding to the different categories of data sources by combining the target classification coefficients corresponding to different categories of data sources with a preset initial data collection frequency, and also includes: Retrieve the first data classification coefficient or the second data classification coefficient corresponding to the secondary data source as the second target classification coefficient; Retrieve the preset initial data collection frequency from the database; The data acquisition frequency corresponding to the secondary data source is obtained by combining the second target grading coefficient with the preset initial data acquisition frequency; The data acquisition frequency corresponding to the secondary data source is obtained by the following formula: Among them, F 02 Indicates the data collection frequency corresponding to the secondary data source; F c Indicates the preset initial data acquisition frequency; S m02 S represents the second objective classification coefficient; y02 This represents the threshold value of the grading coefficient corresponding to the second objective grading coefficient. The marketing lead scoring method based on data source hierarchy as described in claim 1 is characterized in that, The marketing lead data is subjected to data quality and security assessment to obtain marketing lead data that meets preset security standards, including: Extract the data characteristic parameters of each collected marketing lead data, wherein the data characteristic parameters include the encryption rate of the marketing lead data collected per unit time, the encryption resource utilization rate, and the data backup frequency; The security evaluation coefficient is obtained by using the encryption rate, encryption resource utilization rate, and data backup frequency of the marketing lead data collected per unit time. The safety evaluation coefficient is obtained using the following formula: Where K represents the safety evaluation coefficient; m represents the number of times marketing lead data was collected; J i U represents the encryption rate corresponding to the marketing lead data collected in the i-th instance; i B represents the encrypted resource utilization rate corresponding to the marketing lead data collected in the i-th instance; i B represents the data backup frequency corresponding to the marketing lead data collected in the i-th instance; c This indicates the preset data backup frequency reference value; The safety evaluation coefficient is compared with a preset safety coefficient threshold. When the security evaluation coefficient exceeds the preset security coefficient threshold, the collected marketing lead data is determined to meet the security requirements. The marketing lead scoring method based on data source hierarchy as described in claim 1 is characterized in that, Retrieve the deep learning model used for scoring, and use the deep learning model for scoring to obtain the scoring score corresponding to the marketing lead data that meets the preset security standards, including: Retrieve the trained and optimized deep learning model from the database; Extract the optimization completion time of the most recent optimization of the deep learning model; The optimization completion time is obtained based on the optimization completion time and the current model call time. The optimization completion time is compared with a preset time threshold. When the optimization completion time exceeds the preset time threshold, the deep learning model is re-optimized to obtain the optimized deep learning model. Marketing lead data that meets security requirements is input into the optimized deep learning model; The deep learning model is used to score marketing lead data that meets security requirements, and the corresponding scores for the marketing lead data are obtained. The marketing lead scoring method based on data source hierarchy as described in claim 1 or 8 is characterized in that, The structure of the deep learning model is as follows: The input layer is used to classify the marketing lead data that meets security requirements, obtain structured and unstructured data, and preprocess the structured and unstructured data to obtain preprocessed data information; wherein, the structured data includes basic customer information, historical purchase records, and interaction records; the unstructured data includes text data and image data; The feature extraction layer is used to extract and construct features corresponding to the preprocessed data information based on the preprocessed data information; The model inference layer is used by the deep neural network (DNN) model to generate scoring results based on the features corresponding to the data information. The output layer is used to output the scoring results. A marketing lead scoring system based on data source hierarchy, characterized in that, The marketing lead scoring system based on data source hierarchy includes: The data retrieval module is used to extract data flow feature data information and data source classification schemes corresponding to the marketing leads; wherein, the data source classification schemes include a first classification scheme and a second classification scheme; The data source classification module is used to classify the data sources corresponding to the marketing leads by using the data flow feature data information and the data source classification scheme, and to obtain data sources of different categories. The data collection frequency setting module is used to set the data collection frequency corresponding to the different categories of data sources by combining the target grading coefficients corresponding to the different categories of data sources with the preset initial data collection frequency; and to collect marketing lead data according to the set data collection frequency to obtain marketing lead data. The security assessment module is used to assess the data quality and security of the marketing lead data and obtain marketing lead data that meets preset security standards. The scoring control module is used to retrieve the deep learning model used for scoring and to obtain the scoring score corresponding to the marketing lead data that meets the preset security standards using the deep learning model used for scoring.