An e-commerce system big data processing method and system

By prioritizing and extracting fields from interactive data from e-commerce platforms, storing it in a categorized storage system, and determining data access interfaces, the problem of high latency and low efficiency in data processing under the live shopping function of e-commerce platforms is solved, enabling efficient and accurate data access and real-time decision support.

CN122633767APending Publication Date: 2026-08-25SHENZHEN FUXIAOMI TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610782441.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-02
Publication Date
2026-08-25

AI Technical Summary

Technical Problem

After introducing live shopping functionality, existing e-commerce platforms struggle to handle the rapid and ever-changing data streams in highly interactive scenarios. This results in high data processing latency, low efficiency, and poor accuracy, impacting the accuracy of real-time dashboards, operational reports, and recommendation systems. Consequently, live streaming strategies cannot be adjusted in a timely manner, leading to a decline in user interaction experience.

Method used

By acquiring interactive data from e-commerce systems, prioritizing and extracting fields, storing the data in a categorized storage system, determining data access interfaces, and responding to the acquisition needs of downstream applications through matching processing, efficient and accurate data retrieval is achieved.

Benefits of technology

It significantly improves the real-time performance, accuracy, and resource utilization of data processing, avoids message backlog and data inconsistency, supports real-time business decision-making and user experience, and optimizes data storage structure and retrieval process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122633767A_ABST
    Figure CN122633767A_ABST
Patent Text Reader

Abstract

The application relates to the technical field of data processing, in particular to an e-commerce system big data processing method and system. The method comprises the following steps: acquiring e-commerce interaction data received by an e-commerce system in each time period; performing priority classification on the e-commerce interaction data in each time period to obtain each type of interaction data; after storing each type of interaction data into a preset classified storage system, determining a data calling interface of each type of interaction data stored in the preset classified storage system; in response to acquisition demand information of a downstream application, matching the acquisition demand information with a corresponding data calling interface to obtain matched data of the acquisition demand information, so as to complete calling processing of the e-commerce interaction data. The application aims to solve the problem that the existing data processing method is difficult to effectively process big data generated by an e-commerce platform, leading to high delay and low efficiency of data processing, and low accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data processing technology, specifically to a method and system for big data processing in an e-commerce system. Background Technology

[0002] After e-commerce platforms launched live shopping features, the number of users and the variety of products surged. With the introduction of highly interactive features such as live shopping, e-commerce platforms experienced a surge in data streams with rapidly changing content. Existing data processing methods proved inadequate in the data collection, real-time processing, and storage stages, making it difficult to ensure that the data was processed and used in a timely and accurate manner.

[0003] Meanwhile, the explosive growth and dynamic data streams generated by live-stream shopping present delays in existing data processing methods. This results in discontinuous and significantly lagging data streams received by the database, preventing real-time dashboards, operational reports, and recommendation systems that rely on this data from obtaining the latest and most accurate live-stream interaction information. Furthermore, data inconsistencies can arise due to duplicate or missing data processing. This negatively impacts the platform's user experience and operational efficiency. Operational teams are unable to monitor anomalies in real time, adjust live-stream strategies promptly, experience a decline in user interaction, and strategies relying on real-time interactive data become ineffective due to untimely and inaccurate data. Existing big data processing methods struggle to effectively adapt to the high-volume data streams generated by live-stream shopping, resulting in high processing latency and low efficiency and accuracy. Summary of the Invention

[0004] The purpose of this invention is to provide a big data processing method and system for e-commerce systems, which solves the problems that existing data processing methods are unable to effectively process the big data generated by e-commerce platforms, resulting in high data processing latency, low efficiency, and low data processing accuracy.

[0005] To achieve the above objectives, the present invention adopts the following technical solution: a big data processing method for an e-commerce system, comprising: The system receives e-commerce interaction data for each time period; this data includes order information, interaction information, and payment information. The e-commerce interaction data for each time period is prioritized and categorized to obtain each type of interaction data; After storing each type of interactive data in a preset classification storage system, determine the data retrieval interface for each type of interactive data stored in the preset classification storage system; In response to the acquisition request information from downstream applications, the acquisition request information is matched with the corresponding data call interface to obtain matching data for the acquisition request information, so as to complete the call processing of e-commerce interactive data.

[0006] Furthermore, the steps of prioritizing and classifying the e-commerce interaction data for each time period to obtain each type of interaction data include: Field extraction processing is performed on the e-commerce interaction data for each time period to obtain the data field information for each time period; The data field information for each time period is sorted according to field type to obtain sorted data field information; The sorted data fields are prioritized to obtain each type of interactive data.

[0007] Furthermore, the steps of extracting fields from the e-commerce interaction data for each time period to obtain the data field information for each time period include: Based on the e-commerce interaction data for each time period, determine the interaction information for each time period; The interaction information for each time period is processed by field extraction to obtain the original data field information for each time period; Information correction processing is performed on the original information of the data fields for each time period to obtain the data field information for each time period.

[0008] Further, the step of prioritizing and classifying the sorted data fields to obtain each type of interactive data includes: The sorted data fields are prioritized to obtain categorized data; The categorized data is then subjected to data type correction processing to obtain interactive data for each category.

[0009] Furthermore, after storing each type of interactive data in a preset categorized storage system, the step of determining the data access interface for each type of interactive data stored in the preset categorized storage system includes: After storing each type of interactive data in a preset classification storage system, determine the storage timestamp and storage type of each type of interactive data; The confidence level of each type of interactive data is determined based on the storage timestamp and storage type of each type of interactive data; Based on the confidence level of each type of interactive data, determine the data retrieval interface for each type of interactive data stored in the preset classification storage system.

[0010] Furthermore, the step of determining the data retrieval interface for each type of interactive data stored in the preset classification storage system based on the confidence level of each type of interactive data includes: The confidence level of each type of interaction data is compared with the preset confidence level to obtain a confidence level comparison value; Based on the confidence level comparison value and the preset calling interface model, the calling interface scheme is determined; Using the aforementioned API call scheme, the API call of the preset classification storage system is matched to obtain the data call interface for each type of interactive data.

[0011] Furthermore, the step of determining the API call scheme based on the confidence level comparison value and the preset API call model includes: Using the preset call interface model, numerical matching is performed on the confidence comparison value to obtain all call interface candidate schemes and the matching degree of each call interface candidate scheme; By utilizing the matching degree of each candidate scheme for the calling interface, the matching results of all candidate schemes for the calling interface are filtered to obtain the calling interface scheme.

[0012] Furthermore, in response to the acquisition request information from downstream applications, the acquisition request information is matched with the corresponding data call interface to obtain matching data for the acquisition request information, thereby completing the process of calling e-commerce interactive data. The steps include: In response to the demand information from downstream applications, a demand risk assessment is performed on the demand information to obtain a risk assessment value. Using the risk assessment value, the acquisition requirement information is matched with the corresponding data call interface to obtain matching data for the acquisition requirement information, thereby completing the call processing of e-commerce interactive data.

[0013] Furthermore, using the risk assessment value, the steps of matching the acquisition demand information with the corresponding data call interface to obtain matching data for the acquisition demand information, in order to complete the call processing of e-commerce interactive data, include: Based on the risk assessment value, all candidate interfaces for data invocation that match the information required for acquisition are determined; The obtained demand information and all the data call candidate interfaces are respectively subjected to simulated matching processing to obtain each simulated matching data; Perform data type matching processing on each of the simulated matching data to obtain the data type matching degree; Based on the matching degree of the data type and all data call candidate interfaces, after determining the data call interface corresponding to the information to be obtained, the matching data for obtaining the information to be obtained is obtained to complete the call processing of e-commerce interactive data.

[0014] This invention also provides a big data processing system for e-commerce systems, the system comprising: The data receiving module is used to acquire e-commerce interaction data received by the e-commerce system for each time period; the e-commerce interaction data includes order information, interactive communication information, and payment information. The data classification module is used to prioritize and classify the e-commerce interaction data for each time period to obtain each type of interaction data; The interface determination module is used to determine the data call interface for each type of interactive data stored in the preset classification storage system after storing each type of interactive data in the preset classification storage system. The data matching module is used to respond to the acquisition request information of downstream applications, match the acquisition request information with the corresponding data call interface, and obtain the matching data of the acquisition request information to complete the call processing of e-commerce interactive data.

[0015] Compared with existing technologies, the e-commerce system big data processing method and system of the present invention have the following advantages: This invention acquires e-commerce interaction data received by the e-commerce system for each time period and prioritizes and categorizes this data. This effectively addresses the rapid and variable data streams generated by e-commerce platforms, especially in highly interactive scenarios such as live-streaming shopping. When responding to data acquisition requests from downstream applications, it matches the request information with the corresponding data call interface, thereby efficiently and accurately retrieving e-commerce interaction data. This significantly improves the real-time performance, accuracy, and resource utilization of data processing, avoiding message backlog, processing delays, and data inconsistencies, effectively supporting real-time business decision-making and user experience. Attached Figure Description

[0016] To more clearly illustrate the specific embodiments of the present invention, the accompanying drawings used in the specific embodiments will be briefly described below. In all the drawings, the elements or parts are not necessarily drawn to scale.

[0017] Figure 1 This is a flowchart of a big data processing method for an e-commerce system according to the present invention.

[0018] Figure 2 This is a structural block diagram of an e-commerce system big data processing system according to the present invention.

[0019] In the diagram: 210, data receiving module; 220, data classification module; 230, interface determination module; 240, data matching module.

[0020] The implementation and advantages of the present invention will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation

[0021] The following drawings disclose several embodiments of the present invention. For clarity, many practical details will be described in the following description. However, it should be understood that these practical details are not intended to limit the invention. That is, in some embodiments of the invention, these practical details are not essential. Furthermore, for the sake of simplicity, some conventional structures and components will be shown in the drawings in a simple schematic manner.

[0022] It should be noted that all directional indications (such as up, down, left, right, front, back, etc.) in the embodiments of the present invention are only used to explain the relative positional relationship and movement of each component in a certain specific posture (as shown in the figure). If the specific posture changes, the directional indication will also change accordingly.

[0023] Furthermore, in this invention, the use of terms such as "first" and "second" is for descriptive purposes only and does not specifically refer to any order or sequence, nor is it intended to limit the invention. They are merely used to distinguish components or operations described using the same technical terms, and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Therefore, a feature defined with "first" or "second" may explicitly or implicitly include at least one of those features. Additionally, the technical solutions of various embodiments can be combined with each other, but only if they are feasible for those skilled in the art. If a combination of technical solutions is contradictory or impossible to implement, such a combination should be considered nonexistent and not within the scope of protection claimed by this invention.

[0024] To further understand the content, features, and effects of this invention, the following embodiments are provided, and detailed descriptions are given below in conjunction with the accompanying drawings: Please see Figure 1 This invention provides a big data processing method for an e-commerce system, comprising the following steps: S100. Acquire e-commerce interaction data received by the e-commerce system for each time period. This data includes order information, interactive communication information, and payment information. Specifically, e-commerce interaction data refers to all user interaction data received by the e-commerce system within a specific time period. It covers various user behaviors on the platform, such as order information generated from user purchases, interactive communication information like comments and inquiries made by users in live streams or product detail pages, and payment information from completed payments. This data forms the basis for e-commerce platform operation analysis and decision-making. This step achieves data acquisition through various methods. For example, a data acquisition agent can be configured to monitor the e-commerce system's internal message queues, such as Kafka or RabbitMQ, in real time. When new orders, interactive communication, or payment events occur, this data is immediately captured and transmitted to the data processing module. Alternatively, scheduled tasks can be used, such as every 5 or 10 minutes, to batch extract newly added or updated order information, interactive communication information, and payment information from the e-commerce system's business database. Furthermore, API interfaces can be used to allow the e-commerce system to proactively push data to a big data processing platform. The acquired data, whether real-time streaming data or batch data, contains key behavioral information of users on e-commerce platforms for subsequent processing.

[0025] S200. Prioritize and classify the e-commerce interaction data for each time period to obtain each type of interaction data. Priority classification refers to dividing the e-commerce interaction data into different types based on their importance or urgency according to preset rules or algorithms, thus categorizing the raw data into different interaction data types. For example, payment information may be given higher priority to ensure timely transaction processing. This step is based on preset static rules for classification. For example, all payment information can be set as the highest priority, order information as the second highest priority, and interactive communication information as a general priority. After receiving the data, it is classified into the corresponding priority category according to the specific identifier of the data content (such as data type fields). Dynamic classification can also be performed using machine learning models. By training the classification model and inputting the characteristics of the e-commerce interaction data (such as data source, data volume, and data timeliness requirements), the model outputs the priority category of the data. For example, for the instantaneous surge of comment data in a live shopping scenario, the model may identify it as high priority to ensure a real-time interactive experience.

[0026] S300. After storing each type of interactive data in a preset categorized storage system, determine the data access interface for each type of interactive data stored in the preset categorized storage system. The categorized storage system refers to a pre-configured system capable of storing data according to data type or priority. This system can be a distributed file system, relational database, non-relational database, or data warehouse, etc., and its purpose is to optimize the data storage structure for easier subsequent retrieval and processing. Specifically, the storage of data and the determination of the interface can be performed in the following ways: After the data is categorized, it can be written to different storage areas or database tables in the preset categorized storage system according to its priority and type. For example, high-priority data may be stored in a high-performance in-memory database or SSD storage, while low-priority data may be stored in a lower-cost HDFS or object storage. After data storage is completed, the corresponding data access interface is automatically generated or searched based on the data's storage location, storage format, and data type. For example, if payment information is stored in the "payment_transactions" table of a relational database, an SQL query interface is generated; if interactive communication information is stored in a NoSQL database, a RESTful API interface may be generated. Alternatively, you can predefine the storage locations and corresponding API templates for different data types and priorities. Once the data is stored, select the most suitable API from the templates based on the actual data storage situation and populate the necessary parameters to determine the final data API.

[0027] S400: In response to the acquisition request information from downstream applications, the acquisition request information is matched with the corresponding data call interface to obtain matching data for the acquisition request information, thereby completing the e-commerce interactive data call processing. Here, the data call interface refers to a method or protocol that provides external applications or services with access to specific types of data in the categorized storage system. Through the data call interface, downstream applications can efficiently and accurately obtain the required data according to their own needs. Downstream applications refer to various applications that rely on the big data processing results of the e-commerce system, such as real-time recommendation systems, operational reporting systems, risk control systems, and user behavior analysis systems. These applications obtain processed data through the data call interface to support their business functions. Specifically, when a downstream application sends acquisition request information, the acquisition request information includes metadata such as the type and time range of the required data. After receiving the request, all identified data call interfaces are traversed, and the request information is compared with the metadata of each interface. For example, if the request is to obtain payment information for the past hour, an interface that can provide payment information and supports time range queries is searched. After a successful match, the corresponding interface is called to obtain the required data from the categorized storage system and return it to the downstream application. Alternatively, a data service layer can be built, which maintains a registry of data call interfaces. When a downstream application issues a data retrieval request, the data service layer intelligently selects the most matching interface from the registry based on the request content and executes the data call. For example, for the data requirements of a real-time recommendation system, the service layer might prioritize interfaces that provide low-latency and high-throughput data streams.

[0028] In this embodiment, by acquiring e-commerce interaction data received by the e-commerce system for each time period, the comprehensiveness and real-time nature of the data source are ensured. Next, this data is prioritized and categorized, allowing data of different importance to be treated differently, laying the foundation for subsequent differentiated storage and processing. Subsequently, the categorized data is stored in a pre-defined categorized storage system, and corresponding data access interfaces are determined. This not only optimizes the data storage structure but also provides a standardized data access path for downstream applications. Finally, in response to the data acquisition needs of downstream applications, the system matches the demand information with the data access interfaces, achieving accurate and efficient data retrieval. This invention, through priority classification, effectively mitigates the impact of instantaneous high-concurrency writes on the message queue, ensuring that raw event data is promptly and completely ingested into the system. Simultaneously, the intelligent data access interface determination mechanism enables downstream applications to acquire the required data more flexibly and efficiently, avoiding business decision-making errors caused by data lag or inaccuracy. For example, a real-time recommendation system can adjust its recommendation strategy in a timely manner based on the latest and most accurate interaction data, improving user conversion rates. Furthermore, by optimizing data storage and retrieval, system resource consumption is effectively reduced, overall data processing efficiency is improved, and solid technical support is provided for the sustainable development of the e-commerce platform.

[0029] In some embodiments of this application described above, the step of prioritizing and classifying e-commerce interaction data for each time period to obtain each type of interaction data includes: The e-commerce interaction data for each time period is processed by field extraction to obtain the data field information for that time period. This step involves identifying and separating data units with specific meanings and structures from the raw unstructured or semi-structured e-commerce interaction data. For example, predefined parsing rules, regular expressions, or machine learning models can be used to extract fields such as product ID, purchase quantity, and transaction amount from order information; user ID, message content, and timestamp from interactive communication information; and payment method, payment status, and payment amount from payment information. The purpose is to transform the raw, heterogeneous e-commerce interaction data into structured and processable data field information, laying the foundation for subsequent data processing and classification.

[0030] The data fields for each time period are sorted by field type to obtain sorted data field information. This step involves organizing and arranging the extracted data fields according to their inherent attributes or preset classification criteria. For example, all numeric fields (such as transaction amount and purchase quantity) can be grouped into one category, text fields (such as message content and product description) into another, and timestamp fields (such as order creation time and payment time) into yet another. Fields can also be initially sorted according to their importance, sensitivity, or processing timeliness based on business needs; for example, fields involving user privacy or financial security can be given higher priority. The purpose is to perform preliminary structuring and standardization of the data fields, facilitating subsequent priority classification operations.

[0031] The sorted data fields are then prioritized to obtain different categories of interactive data. This step involves meticulously dividing the sorted data fields according to preset classification rules and priority standards, resulting in interactive data categories with varying priorities. For example, different data fields or combinations of fields can be assigned different priorities based on dimensions such as business importance (e.g., order status and payment results), data timeliness (e.g., real-time chat messages and latest browsing history), and data sensitivity (e.g., user identity information and bank card numbers). Higher-priority data may require faster processing speeds, stricter storage strategies, or more frequent updates. The aim is to ensure that critical data is processed and managed promptly and effectively, while optimizing system resource allocation.

[0032] This invention extracts fields from e-commerce interaction data for each time period, breaking down the original, complex data into manageable and structured data fields. By sorting these fields according to their type, the invention further standardizes the data organization, enabling the systematic identification and classification of different data types. Finally, by prioritizing the sorted data fields, different data can be assigned processing priorities based on preset business rules and importance standards. This step-by-step, refined processing approach allows e-commerce systems to more accurately identify and manage massive amounts of interaction data, avoiding the resource waste and inefficiency of indiscriminately processing all data.

[0033] In some embodiments of this application described above, the step of extracting fields from the e-commerce interaction data for each time period to obtain data field information for each time period includes: Based on the e-commerce interaction data for each time period, the interaction information for each time period is determined. This step involves filtering out information directly related to user interaction behavior from the raw e-commerce interaction data, for example, filtering out non-interactive data such as system logs and advertising pushes. The aim is to focus on core interaction data, providing a cleaner data source for subsequent field extraction. This interaction information can include user behavior data on the e-commerce platform, such as browsing, clicking, searching, adding to cart, favoriting, placing orders, making payments, leaving reviews, and making inquiries.

[0034] The interaction information for each time period undergoes field extraction processing to obtain the raw data field information for each time period. This step can utilize preset parsing rules, regular expressions, or machine learning models to identify and extract key data fields from the interaction information, such as order number, product ID, user ID, comment content, and payment amount. The extracted data at this stage may still contain some formatting errors or redundant information, hence it is referred to as the raw data field information.

[0035] Information correction processing is performed on the raw data fields for each time period to obtain the data field information for that time period. Specifically, information correction processing involves cleaning, standardizing, and deduplicating the extracted raw data field information. For example, date formats are standardized, text content is segmented and tagged with parts of speech, and numerical data undergoes range validation or missing value filling. The purpose is to eliminate data noise, standardize data formats, and ensure the accuracy and consistency of data field information, providing high-quality input for subsequent data sorting and classification.

[0036] Specifically, an e-commerce system receives a large amount of e-commerce interaction data within a certain time period, including user browsing history, shopping cart additions, order payments, customer service inquiries, and some internal system log information. First, based on the e-commerce interaction data for each time period, pre-defined rules or models are used to identify and determine interaction information directly related to user interaction. For example, system logs are filtered out, retaining only user browsing, shopping cart additions, payments, and inquiries. Next, the filtered interaction information undergoes field extraction processing. For example, order payment information extracts fields such as order number, product ID, payment amount, and payment time; customer service inquiries extract fields such as user ID, inquiry content, and inquiry time. At this stage, the raw data fields may contain issues such as inconsistent date formats (e.g., 2023-01-01 vs. 2023 / 01 / 01), missing product IDs, or incorrect user ID formats. Subsequently, the raw data fields undergo information correction processing. For example, all date formats are standardized to YYYY-MM-DD, missing product IDs are marked or supplemented through related data, and user ID formats are validated and corrected. In this way, the final data field information is cleaned and standardized, which can more accurately reflect the user's e-commerce interaction behavior and provide a high-quality data foundation for subsequent priority classification and storage.

[0037] This invention improves processing efficiency and accuracy by identifying core interactive information from raw e-commerce interaction data, avoiding the extraction of fields from large amounts of irrelevant or low-value data. Furthermore, after extracting the raw data field information, information correction processing effectively identifies and corrects errors, inconsistencies, or missing data, ensuring that the final data field information is high-quality and standardized. This phased and refined processing approach allows subsequent data sorting and priority classification to be based on more reliable data, thereby enhancing the robustness and effectiveness of the entire e-commerce system's big data processing method.

[0038] In some embodiments of this application described above, the step of prioritizing and classifying the sorted data field information to obtain each type of interactive data includes: The sorted data fields are then prioritized to obtain categorized data. This step involves dividing the sorted data fields into different categories based on preset priority rules or models. For example, data fields can be categorized into high-priority, medium-priority, and low-priority data based on dimensions such as field importance, sensitivity, update frequency, or business relevance. Note that the categorized data represents a preliminary classification result and may not have undergone rigorous data type validation.

[0039] The categorized data undergoes data type correction processing to obtain interactive data for each category. This step involves a detailed data type check and adjustment of the data in each category, based on the initial classification. For example, if a category is defined as storing numerical data but contains string data, conversion or labeling is required. Data type correction ensures that the data in each category conforms to its expected type specification, thereby improving data quality and usability. Data type correction includes, but is not limited to, data format standardization, outlier detection and handling, missing value imputation, and data type conversion (such as string to number conversion, date format standardization), etc. Its purpose is to eliminate data type inconsistencies or errors in the categorized data, making the final interactive data for each category more standardized and accurate.

[0040] Specifically, after prioritizing the sorted data fields, a category of high-priority order amount data is obtained. However, in the actual collected raw data, due to differences in different merchants or payment channels, some order amounts in this category may be stored as strings (e.g., 123.45 yuan) instead of standard floating-point numbers. Without processing, subsequent statistical analysis or financial settlement systems may encounter errors or require additional conversion steps when calling this data. This is where data type correction comes in. Each data item in the high-priority order amount data is checked. When the string "123.45 yuan" is detected, it is converted to the standard floating-point number 123.45 according to preset correction rules. Simultaneously, if missing amount data is found, it can be filled according to business rules (e.g., filled with 0 or predicted by an algorithm). Through data type correction, the final high-priority order amount data will all be standardized floating-point numbers, ensuring data consistency and availability, greatly simplifying the calling and processing flow of downstream applications, and improving the accuracy of data analysis.

[0041] This invention effectively avoids inconsistencies or non-standardization of data types by introducing data type correction processing after initial priority classification. Specifically, the sorted data fields are prioritized to obtain preliminary categorized data, ensuring that the data is divided according to importance or business needs at a macro level. Based on this, data type correction processing is applied to the categorized data to specifically identify and correct data type deviations within each category. For example, data that does not conform to the preset data type is converted, cleaned, or labeled, thereby ensuring that the data in each category also has a high degree of type consistency and accuracy at the micro level. Due to this two-stage processing method, the final interactive data for each category not only has a clear priority division but also possesses high-quality data type specifications, laying a solid foundation for subsequent data storage and retrieval.

[0042] In some embodiments of this application described above, after storing each type of interactive data in a preset categorized storage system, the step of determining the data access interface for each type of interactive data stored in the preset categorized storage system includes: After storing each type of interactive data in a preset categorized storage system, the storage timestamp and storage type of each type of interactive data are determined. The storage timestamp refers to the specific time record when each type of interactive data is stored in the preset categorized storage system, reflecting the timeliness or freshness of the data. The storage type can be understood as the specific classification attribute of each type of interactive data during storage, such as the data format (e.g., structured and unstructured), the importance level of the data (e.g., core data and auxiliary data), the sensitivity of the data (e.g., public data and private data), or the characteristics of its storage medium.

[0043] Based on the storage timestamp and storage type of each type of interactive data, the confidence level of each type of interactive data is determined. Confidence level is a quantitative indicator obtained by comprehensively evaluating the importance, reliability, or sensitivity of each type of interactive data, and its purpose is to provide a decision-making basis for subsequent data access interface selection. The determination of confidence level is based on the data's storage timestamp and storage type. For example, recently stored data marked as highly important or highly sensitive will typically have a higher confidence level.

[0044] Based on the confidence level of each type of interactive data, determine the data retrieval interface for each type of interactive data stored in the preset classification storage system.

[0045] Specifically, the e-commerce system receives two types of interactive data: real-time user payment order data and historical user browsing history data. When user payment order data is stored in the categorized storage system, its storage timestamp (e.g., current time) and storage type (e.g., high priority, structured, or sensitive data) are recorded. Based on this information, the confidence level of the payment order data is determined to be high. Based on this high confidence level, an API interface with high throughput and strict encryption mechanisms is assigned to it to ensure fast and secure access to the payment data. For historical user browsing history data, its storage timestamp may be older, and the storage type is non-sensitive and semi-structured data; therefore, its confidence level is determined to be lower. An API interface with lower access frequency restrictions and general permissions is assigned to it. In this way, data of different importance and timeliness can obtain calling interfaces that match their characteristics, thereby optimizing the efficiency and security of overall data processing and access.

[0046] This invention first records the storage timestamp and storage type after each type of interactive data is stored, thereby comprehensively understanding the basic attributes and timeliness of the data. Based on these attributes, the confidence level of each type of interactive data is further derived, reflecting the importance and sensitivity of the data. Because of the introduction of this dynamic evaluation metric, the determination of subsequent data access interfaces is no longer a static preset, but can be intelligently adjusted according to the actual characteristics of the data. For example, for high-confidence data, an interface with a higher security level or better access performance can be assigned; while for low-confidence data, an interface with stricter access permissions or lower priority may be assigned. This dynamic interface determination mechanism based on data confidence effectively avoids the problems of a single and inflexible interface determination method.

[0047] In some embodiments of this application described above, the step of determining the data retrieval interface for each type of interactive data stored in a preset categorized storage system based on the confidence level of each type of interactive data includes: The confidence level of each type of interactive data is compared with a preset confidence level to obtain a confidence level comparison value. The preset confidence level is a threshold or a series of thresholds pre-set based on business needs, data sensitivity, or system security policies. For example, for highly sensitive data, the preset confidence level might be set to a higher value to ensure that only highly reliable data can be accessed through specific API calls.

[0048] Based on the confidence comparison value and the preset call interface model, a call interface scheme is determined. The confidence comparison value reflects the degree of deviation or conformity between the current data confidence level and the preset standard. The preset call interface model can be a rule set, lookup table, or algorithm model that includes parameters such as various call interface types, access permissions, encryption protocols, and transmission rates. This model intelligently selects or generates the call interface scheme most suitable for the current data characteristics and security requirements based on the confidence comparison value. For example, when the confidence comparison value indicates that the data has extremely high reliability, the model may recommend an interface scheme that supports high concurrency, low latency, and advanced encryption functions.

[0049] Using the aforementioned API call scheme, API call matching is performed on a pre-defined categorized storage system to obtain the data call interface for each type of interactive data. API call matching refers to finding or configuring the most suitable actual data call interface within the pre-defined categorized storage system based on the determined API call scheme. This ensures that the determined interface not only meets the characteristics and security requirements of the data itself but is also compatible with the actual capabilities and configuration of the storage system.

[0050] Specifically, the e-commerce system receives payment information for a batch of high-value orders. After classification, the confidence level of this type of payment data is determined to be 0.95. This 0.95 confidence level is then compared to a preset confidence level. For example, the preset confidence level is set to 0.90, indicating that payment data above this value is considered highly trustworthy and sensitive. Since 0.95 is higher than 0.90, a positive confidence comparison value is obtained. Further, based on the confidence comparison value, a preset API call model is activated. This model may include a series of rules, such as: if the confidence comparison value indicates that the data is highly sensitive, then an API interface scheme based on the OAuth2.0 protocol, with end-to-end encryption and a transmission rate of no less than 100Mbps is recommended. Based on this rule, a highly secure and high-performance API call scheme is determined. Finally, using this API call scheme, the API call of the preset classified storage system storing the payment information is matched. If the storage system provides multiple API interfaces, it automatically filters and matches the most suitable API interface based on the protocol, encryption standard, and performance requirements defined in the solution, such as an interface named SecurePaymentAPI_v3.0. Consequently, downstream applications requesting this type of payment data will be directed to use the rigorously filtered and matched highly secure interface for data calls, thus ensuring the secure transmission and processing of sensitive payment information.

[0051] This invention effectively solves the problem of transforming abstract confidence levels into concrete, operable data retrieval interfaces adapted to storage systems by introducing mechanisms such as confidence level comparison, interface modeling, and interface matching. By comparing the confidence level of each type of interactive data with a pre-set confidence level, the reliability or importance of the data can be quantitatively assessed, providing an objective basis for subsequent interface selection. Secondly, the pre-set interface model can generate the optimal interface scheme based on this quantitative assessment result, avoiding the bias and inefficiency that may result from manual judgment. Finally, by matching this interface scheme to the categorized storage system, the compatibility and efficiency of the selected interface with the actual storage environment are ensured, making the data retrieval process more accurate, secure, and efficient.

[0052] In some embodiments of this application described above, the step of determining the API call scheme based on the confidence comparison value and the preset API call model includes: Using the pre-defined API call model, numerical matching is performed on the confidence comparison value to obtain all API call candidate schemes and the matching degree of each candidate scheme. Specifically, the pre-defined API call model is a knowledge base or algorithm set containing various API call rules, parameter configurations, and matching logic. This model aims to provide potential API call options based on different data characteristics and call requirements. The confidence comparison value refers to the numerical value obtained by comparing the confidence of each type of interactive data with a preset confidence level. This value reflects the degree of difference between the data quality or reliability and the preset standard. Numerical matching refers to using the pre-defined API call model, taking the confidence comparison value as input, and calculating and evaluating it through the algorithm logic within the model to generate possible API call schemes. For example, this numerical matching process can be implemented based on rule engines, machine learning algorithms, or fuzzy logic to ensure that all API call candidate schemes that meet certain conditions can be comprehensively identified. All API call candidate schemes refer to all potential and selectable API call configurations or strategies generated during the numerical matching process based on the confidence comparison value and the API call model. The matching degree of each candidate API call solution is a quantitative indicator that measures the degree of fit between each candidate solution and the current confidence level, as well as the potential call requirements. The matching degree can be calculated based on various factors, such as API compatibility, data transmission efficiency, security, cost, and alignment with business needs.

[0053] By utilizing the matching degree of each candidate API call scheme, a scheme filtering process is performed on the matching results of all candidate API call schemes to obtain the API call scheme. The scheme filtering process refers to the process of evaluating, ranking, and selecting these candidate schemes according to a preset filtering strategy or optimization objective after obtaining all candidate API call schemes and their matching degrees. For example, the filtering strategy may include selecting the scheme with the highest matching degree, selecting the scheme that meets specific performance indicators (such as low latency and high throughput), or selecting the scheme with the best cost-effectiveness. Through the scheme filtering process, an optimal or most suitable API call scheme can ultimately be determined from multiple candidate schemes. The API call scheme refers to the specific interface configuration and strategy that is ultimately selected for actual data invocation after filtering.

[0054] Specifically, the e-commerce system needs to access interactive data with a medium confidence level (e.g., 0.6). The pre-defined API call model receives this confidence level and performs numerical matching. For example, the model might identify three candidate API call schemes: Scheme A: Using a high-security, medium-speed API, with a matching score of 0.85 (security is the primary consideration due to high data sensitivity). Scheme B: Using a medium-security, high-speed direct database connection, with a matching score of 0.70 (suitable for data with high real-time requirements but slightly lower security requirements). Scheme C: Using a low-security, low-cost third-party data service interface, with a matching score of 0.60 (suitable for non-core data with tolerable latency). After obtaining the candidate schemes and their matching scores, a pre-defined filtering strategy is used to filter the schemes. If the filtering strategy prioritizes the scheme with the highest security and a matching score of at least 0.8, then Scheme A will be selected as the final API call scheme. If the filtering strategy prioritizes the scheme with the fastest transmission speed and a matching score of at least 0.7, then Scheme B might be selected. This allows for the flexible and intelligent determination of the optimal data call interface based on specific business needs and data characteristics, enabling more refined and efficient data management.

[0055] When this invention receives a confidence level comparison value, the preset call interface model no longer simply outputs a fixed interface scheme. Instead, it comprehensively identifies all potential candidate call interface schemes that meet the criteria by performing fine-grained numerical matching on the comparison value. The multi-scheme generation mechanism ensures that no possible optimization options are overlooked in complex and ever-changing data call scenarios. Furthermore, by calculating the matching degree of each candidate call interface scheme, the degree of fit between each scheme and the current data characteristics and call requirements can be quantified. This quantitative evaluation provides an objective basis for subsequent decision-making. Finally, by using the matching degree to screen all candidate schemes, the most suitable call interface scheme can be intelligently selected from multiple options based on preset optimization objectives (e.g., selecting the scheme with the highest matching degree, optimal performance, or lowest cost). Therefore, this invention ensures that the determined call interface scheme not only has higher accuracy but also exhibits stronger adaptability and robustness when facing different confidence level comparison values ​​and diverse call requirements.

[0056] In some embodiments of this application described above, the steps of matching the acquisition request information with the corresponding data call interface in response to the downstream application's acquisition request information to obtain matching data for the acquisition request information, thereby completing the e-commerce interactive data call processing, include: Upon receiving a data acquisition request from a downstream application, a risk assessment is performed on that request to obtain a risk assessment value. Specifically, this risk assessment involves comprehensively considering the security, compliance, and sensitivity of the data acquisition request from the downstream application to quantify its potential risk level. This risk assessment can be based on various factors, such as the type and sensitivity level of the requested data (e.g., user personal identification information and payment transaction records), the identity and permissions of the request initiator, the frequency and historical behavior patterns of the request, and the business scenario in which the request occurs. By analyzing and calculating these factors, a risk assessment value can be obtained, which can be a numerical score, a risk level (e.g., low, medium, high risk), or a set of risk tags. The purpose is to provide a risk basis for subsequent data access interface matching and ensure the security of data access.

[0057] Using the aforementioned risk assessment value, the required information is matched with the corresponding data call interface to obtain matched data for the required information, thereby completing the data call processing for e-commerce interaction. This step refers to incorporating the previously obtained risk assessment value into the selection of the data call interface. For example, when the risk assessment value is high, a stricter matching strategy can be adopted, such as matching only interfaces that provide de-identified or anonymized data, or requiring downstream applications to perform additional authentication or authorization confirmation. Conversely, when the risk assessment value is low, interfaces that provide more complete or real-time data can be matched. Thus, the risk assessment value, as an important parameter in the matching process, can guide the system to select the most suitable and secure data call interface.

[0058] Specifically, an e-commerce system receives a request from a downstream application to obtain historical order details of a specific user, including sensitive information such as the user's shipping address and payment method. First, in response to this request, a risk assessment is performed. This involves analyzing the identity of the requester (e.g., whether it is an internal high-privilege application or an external partner), its historical data access behavior, request frequency, and the sensitivity of the requested data. If the assessment indicates a high risk (e.g., the requester has low privileges but requests sensitive data or the request frequency is abnormal), a high risk assessment value is obtained. Subsequently, this risk assessment value is used to match the request with the corresponding data access interface. For example, if the risk assessment value is high, an interface providing anonymized order data might be prioritized. This interface hides the user's specific shipping address and payment method, providing only order product information and transaction amount. Alternatively, the requester might be required to undergo secondary authentication, or, in extremely high-risk cases, the data access request might be directly rejected, generating a corresponding security alert. Even if downstream applications send requests containing sensitive data, the system can intelligently adjust data access policies based on risk assessment results, ensuring that data is only provided when specific security conditions are met or through specific security interfaces, thereby effectively preventing potential data leakage or misuse risks.

[0059] This invention, when a downstream application issues a data access request, no longer blindly matches interfaces, but first performs a comprehensive risk assessment of the request. By evaluating factors such as the sensitivity of the request and the credibility of the initiator, a quantified risk assessment value can be obtained. This risk assessment value is then used to guide the matching process of data access interfaces. For example, for high-risk requests, interfaces with stricter access controls, data anonymization, or auditing functions can be prioritized for matching, and requests that do not comply with security policies can even be rejected. Because of the introduction of the risk assessment mechanism, the matching of data access interfaces is no longer merely a functional correspondence, but incorporates security and compliance, thereby fundamentally improving the security of data access.

[0060] In some embodiments of this application described above, the steps of matching the acquisition demand information with the corresponding data call interface using the risk assessment value to obtain matching data for the acquisition demand information, thereby completing the e-commerce interactive data call processing, include: Based on the risk assessment value, all candidate data call interfaces matching the acquisition requirement information are determined. This step refers to, after receiving the acquisition requirement information from the downstream application and completing the requirement risk assessment, using the risk assessment value, combined with preset interface security policies, data access permissions, and interface function descriptions, to initially filter out a set of interfaces that meet the security requirements and basic functional matching conditions from all available data call interfaces as potential candidate interfaces. For example, if the risk assessment value is high, interfaces with stricter security authentication or finer-grained access control will be given priority.

[0061] The obtained requirement information is then simulated and matched against all the candidate data call interfaces to obtain simulated matching data for each interface. This step involves a virtual data call attempt for each candidate interface. Specifically, this simulated matching process includes, but is not limited to: checking the availability of the interface, simulating data transmission paths, evaluating the response time of the interface, and preliminarily verifying whether the format of the data returned by the interface is compatible with the requirement information. Its purpose is to pre-evaluate the potential performance and compatibility of each candidate interface before the actual data call, thereby providing foundational data for subsequent accurate matching.

[0062] For each simulated matching data point, a data type matching process is performed to obtain the data type matching degree. This step involves in-depth analysis of the data structure, data format, and data field types that the interface represented by the simulated matching data can provide, and comparing it with the explicit or implicit data type requirements in the demand information. For example, if the demand information requires a numerical order amount, but an interface can only provide a string-type order amount, its matching degree will be correspondingly lower. The data type matching degree can be a percentage value or a rating scale, the purpose of which is to quantify the degree to which each candidate interface matches the demand information in terms of data content and format.

[0063] Based on the data type matching degree and all candidate data call interfaces, after determining the data call interface corresponding to the required information, matching data for the required information is obtained to complete the e-commerce interactive data call processing. Specifically, after obtaining the data type matching degree of all candidate interfaces, these matching degrees and other factors (e.g., interface performance indicators, cost, and initial risk assessment values) are comprehensively considered. Through preset decision logic or algorithms, the data call interface that best meets the required information is selected from the candidate interfaces. For example, the interface with the highest matching degree and an acceptable risk assessment value can be selected. Once the corresponding data call interface is determined, it can be used to perform actual data calls and obtain matching data that meets the required information.

[0064] Specifically, a downstream application needs to obtain a list of users who purchased a specific product and their total purchase amount within the past week. First, after receiving this request, the e-commerce system performs a risk assessment, obtaining a risk assessment value. For example, if the request involves user privacy data, the risk assessment value is high. Next, based on this risk assessment value, all available data access interfaces are initially screened for candidate interfaces. For example, user behavior analysis, order detail query, and aggregation statistics interfaces might be selected as candidate interfaces because they may all provide relevant data and meet the security policies for high-risk requests. Then, the request information is simulated and matched against these three candidate interfaces. For example, a simulated user behavior analysis interface might return user IDs and behavioral events, a simulated order detail query interface might return detailed order records, and a simulated aggregation statistics interface might directly return summarized user IDs and total purchase amount data. This yields each simulated matching data. Subsequently, data type matching is performed on each simulated matching data to obtain the data type matching degree. For example, the user behavior analysis interface requires further processing to obtain the total purchase amount, and its data type matching degree may be 70%; the order detail query interface requires summarizing and calculating multiple orders, and its data type matching degree may be 85%; while the aggregation statistics interface directly provides the user ID and total purchase amount, which highly matches the data type of the required information, and its data type matching degree may be 95%. Finally, based on the data type matching degree and all candidate data call interfaces, the aggregation statistics interface was determined to be the data call interface corresponding to obtaining the required information because it has the highest data type matching degree and meets the risk assessment requirements. Through this interface, the list of users who purchased specific products in the past week and their matching data of total purchase amount can be obtained efficiently and accurately, thereby completing the processing of e-commerce interaction data.

[0065] This invention effectively avoids the inaccuracies and inefficiencies that can result from relying solely on a single risk assessment value for matching by introducing a multi-stage, refined matching mechanism. First, candidate interfaces for data invocation are initially screened using risk assessment values ​​to ensure basic security and functional compatibility. Second, simulated matching of candidate interfaces allows for advance prediction of their actual performance and compatibility, avoiding blind invocation. Furthermore, by performing data type matching on the simulated matching data, the structural and formatal fit between the interface-provided data and the required data is quantified, making the matching process more accurate. Finally, by considering data type matching and other factors, the most suitable, secure, and efficient data invocation interface is intelligently selected from multiple candidate interfaces, significantly improving the accuracy and reliability of data invocation.

[0066] Based on any of the above embodiments, please refer to the following e-commerce system big data processing method. Figure 2 The present invention also provides an e-commerce system big data processing system, which includes a data receiving module 210, a data classification module 220, an interface determination module 230 and a data matching module 240.

[0067] The data receiving module 210 is used to acquire e-commerce interaction data received by the e-commerce system for each time period; among which, e-commerce interaction data includes order information, interactive communication information and payment information.

[0068] The data classification module 220 is used to prioritize and classify the e-commerce interaction data for each time period to obtain each type of interaction data.

[0069] The interface determination module 230 is used to determine the data call interface for each type of interactive data stored in the preset classification storage system after storing each type of interactive data in the preset classification storage system.

[0070] The data matching module 240 is used to respond to the acquisition demand information of the downstream application, match the acquisition demand information with the corresponding data call interface, and obtain the matching data of the acquisition demand information to complete the call processing of e-commerce interactive data.

[0071] In this embodiment, the data receiving module 210 comprehensively acquires e-commerce interactive data, the data classification module prioritizes the data 220, the interface determination module 230 intelligently manages data storage and access interfaces, and the data matching module 240 enables accurate and efficient data retrieval. This system provides a standardized and efficient data access path for downstream applications, significantly improving the real-time performance, accuracy, and system resource utilization of data processing.

[0072] The above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features therein. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention, and they should all be covered within the scope of the present invention specification.

Claims

1. A big data processing method for an e-commerce system, characterized in that, include: The system receives e-commerce interaction data for each time period; this data includes order information, interaction information, and payment information. The e-commerce interaction data for each time period is prioritized and categorized to obtain each type of interaction data; After storing each type of interactive data in a preset classification storage system, determine the data retrieval interface for each type of interactive data stored in the preset classification storage system; In response to the acquisition request information from downstream applications, the acquisition request information is matched with the corresponding data call interface to obtain matching data for the acquisition request information, so as to complete the call processing of e-commerce interactive data.

2. The big data processing method for an e-commerce system according to claim 1, characterized in that, The steps for prioritizing and classifying e-commerce interaction data for each time period to obtain each type of interaction data include: Field extraction processing is performed on the e-commerce interaction data for each time period to obtain the data field information for each time period; The data field information for each time period is sorted according to field type to obtain sorted data field information; The sorted data fields are prioritized to obtain each type of interactive data.

3. The big data processing method for an e-commerce system according to claim 2, characterized in that, The steps for extracting fields from the e-commerce interaction data for each time period to obtain the data field information for each time period include: Based on the e-commerce interaction data for each time period, determine the interaction information for each time period; The interaction information for each time period is processed by field extraction to obtain the original data field information for each time period; Information correction processing is performed on the original information of the data fields for each time period to obtain the data field information for each time period.

4. The big data processing method for an e-commerce system according to claim 2, characterized in that, The steps for prioritizing and classifying the sorted data fields to obtain each type of interactive data include: The sorted data fields are prioritized to obtain categorized data; The categorized data is then subjected to data type correction processing to obtain interactive data for each category.

5. The big data processing method for an e-commerce system according to claim 1, characterized in that, After storing each type of interactive data in a preset categorized storage system, the step of determining the data access interface for each type of interactive data stored in the preset categorized storage system includes: After storing each type of interactive data in a preset classification storage system, determine the storage timestamp and storage type of each type of interactive data; The confidence level of each type of interactive data is determined based on the storage timestamp and storage type of each type of interactive data; Based on the confidence level of each type of interactive data, determine the data retrieval interface for each type of interactive data stored in the preset classification storage system.

6. The big data processing method for an e-commerce system according to claim 5, characterized in that, The steps for determining the data retrieval interface for each type of interactive data stored in the preset classification storage system based on the confidence level of each type of interactive data include: The confidence level of each type of interaction data is compared with the preset confidence level to obtain a confidence level comparison value; Based on the confidence level comparison value and the preset calling interface model, the calling interface scheme is determined; Using the aforementioned API call scheme, the API call of the preset classification storage system is matched to obtain the data call interface for each type of interactive data.

7. The big data processing method for an e-commerce system according to claim 6, characterized in that, The steps for determining the API call scheme based on the confidence level comparison value and the preset API call model include: Using the preset call interface model, numerical matching is performed on the confidence comparison value to obtain all call interface candidate schemes and the matching degree of each call interface candidate scheme; By utilizing the matching degree of each candidate scheme for the calling interface, the matching results of all candidate schemes for the calling interface are filtered to obtain the calling interface scheme.

8. The big data processing method for an e-commerce system according to claim 1, characterized in that, In response to the acquisition request information from downstream applications, the acquisition request information is matched with the corresponding data call interface to obtain matching data for the acquisition request information, thereby completing the process of calling e-commerce interactive data. The steps include: In response to the demand information from downstream applications, a demand risk assessment is performed on the demand information to obtain a risk assessment value. Using the risk assessment value, the acquisition requirement information is matched with the corresponding data call interface to obtain matching data for the acquisition requirement information, thereby completing the call processing of e-commerce interactive data.

9. The big data processing method for an e-commerce system according to claim 8, characterized in that, Using the risk assessment value, the steps of matching the acquisition demand information with the corresponding data call interface to obtain matching data for the acquisition demand information, in order to complete the call processing of e-commerce interactive data, include: Based on the risk assessment value, all candidate interfaces for data invocation that match the information required for acquisition are determined; The obtained demand information and all the data call candidate interfaces are respectively subjected to simulated matching processing to obtain each simulated matching data; Perform data type matching processing on each of the simulated matching data to obtain the data type matching degree; Based on the matching degree of the data type and all data call candidate interfaces, after determining the data call interface corresponding to the information to be obtained, the matching data for obtaining the information to be obtained is obtained to complete the call processing of e-commerce interactive data.

10. A big data processing system for an e-commerce system, characterized in that, The system includes: The data receiving module is used to acquire e-commerce interaction data received by the e-commerce system for each time period; the e-commerce interaction data includes order information, interactive communication information, and payment information. The data classification module is used to prioritize and classify the e-commerce interaction data for each time period to obtain each type of interaction data; The interface determination module is used to determine the data call interface for each type of interactive data stored in the preset classification storage system after storing each type of interactive data in the preset classification storage system. The data matching module is used to respond to the acquisition request information of downstream applications, match the acquisition request information with the corresponding data call interface, and obtain the matching data of the acquisition request information to complete the call processing of e-commerce interactive data.