A method, device, equipment and medium for analyzing buried point data

By collecting and analyzing front-end embedded data from ride-hailing driver clients, constructing timestamp paths, and calculating multi-dimensional indicators, the problems of data fragmentation and insufficient timeliness were solved. This enabled accurate analysis of client operation characteristics and real-time fault location, improving client stability and the effectiveness of operational strategies.

CN122152623APending Publication Date: 2026-06-05BEIJING BAILONG MAYUN TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
BEIJING BAILONG MAYUN TECH CO LTD
Filing Date
2026-01-12
Publication Date
2026-06-05

AI Technical Summary

Technical Problem

Existing data processing and analysis solutions for ride-hailing driver clients suffer from data fragmentation, insufficient timeliness, and missing analysis dimensions. They are unable to accurately trace the root cause of problems, capture performance degradation trends in real time, and locate the causes of failures. As a result, the accuracy, comprehensiveness, and timeliness of client operation feature analysis results are greatly reduced, and they cannot effectively support client version iteration and operational strategy optimization.

Method used

Collect front-end event tracking data from ride-hailing driver clients, construct a complete user operation path with timestamps, identify erroneous event tracking data through multi-dimensional indicator calculation and differential feature analysis, establish a complete causal chain of operations, generate optimized operation strategies and perform version optimization, and achieve real-time data analysis and fault location.

Benefits of technology

It improves the accuracy and timeliness of client operation feature analysis, enabling rapid identification of fault causes, supporting real-time adjustments to client version optimization and operational strategies, and enhancing the stability of client operation and driver user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122152623A_ABST
    Figure CN122152623A_ABST
Patent Text Reader

Abstract

The present application relates to the field of data analysis, and particularly relates to a method and device for analyzing buried point data, equipment and medium. The present application firstly collects front-end buried point data in full, integrates scattered data such as page stay duration, resource loading and API performance, breaks the data fragmentation barrier, can establish a complete operation causal link, and accurately traces the problem source. Secondly, based on real-time buried point data direct analysis, the batch processing mode is abandoned, the performance degradation trend can be captured in real time, and the timeliness deficiency pain point is solved. Then, the analysis process constructs a user operation path with a time stamp, associates resource loading failure with corresponding page elements, fills in the analysis dimension blank, and accurately locates the fault inducement. Finally, the version optimization forms a closed loop, fundamentally improves the accuracy and timeliness of client operation analysis, and guarantees the stable operation of the client.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data analysis, and specifically to a method, apparatus, equipment, and medium for analyzing embedded data. Background Technology

[0002] In the technical practice of ride-hailing platforms for managing the operational status and optimizing the performance of driver-side clients, existing data processing and analysis solutions have significant technical shortcomings, making it difficult to meet the needs of precise and real-time client optimization. Specifically, firstly, there is the problem of data fragmentation. Core operational data such as page dwell time, resource loading status, and API performance are scattered across different data modules, making it impossible to establish a complete causal chain of "button click → API call → page rendering" through a unified association identifier. This makes it difficult to quickly trace the root cause after an anomaly occurs. Secondly, there is insufficient timeliness. Existing technologies mostly use batch data processing mechanisms, resulting in a minimum delay of one hour in the analysis of client operational data. This makes it impossible to capture the dynamic degradation trend of page performance in real time and to trigger early warning and intervention mechanisms in a timely manner.

[0003] Finally, there is a lack of analytical dimensions. The analysis process lacks a continuous view of user operation paths centered on timestamps, and there are technical gaps in the correlation analysis between resource loading failure events and corresponding page elements, making it impossible to accurately locate performance bottlenecks and causes of failures. These technical deficiencies significantly reduce the accuracy, comprehensiveness, and timeliness of client operation feature analysis results, failing to provide effective support for client version iteration and operational strategy optimization, and hindering the operational stability of ride-hailing driver clients and the improvement of driver user experience. Summary of the Invention

[0004] In view of this, embodiments of the present invention provide a method, apparatus, device and medium for analyzing embedded data, in order to solve the technical defects of existing online ride-hailing driver client operation data processing and analysis solutions, such as data fragmentation, insufficient timeliness and lack of analysis dimensions, which make it impossible to accurately trace the root cause of the problem, capture the trend of performance degradation in real time and locate the cause of the failure.

[0005] In a first aspect, embodiments of the present invention provide a method for analyzing embedded data, the method comprising: Collect front-end embedded data generated when ride-hailing drivers perform terminal operations on the corresponding client; The front-end embedded data was analyzed to obtain analysis results on the operation and usage characteristics of the driver client from different dimensions; Based on the analysis results, an optimized operation strategy for the corresponding client of the ride-hailing driver is generated; Based on the optimized operation strategy, the client version corresponding to the ride-hailing driver is optimized.

[0006] Furthermore, the analysis of the front-end embedded data yields analysis results of the driver client's operation and usage characteristics from different dimensions, including: Based on the aforementioned front-end data, a complete user operation path with timestamps is constructed, and a structured dataset is output. Perform multi-dimensional indicator calculations on the dataset according to a preset time window to obtain basic indicator data; Identify erroneous data points from the aforementioned front-end data points; The differences between the basic indicator data and the error tracking data are analyzed, and the analysis results of the client operation and usage characteristics are constructed based on the differences.

[0007] Furthermore, the analysis of the differential characteristics between the basic indicator data and the error tracking data, and the construction of analysis results of the driver client's operation and usage characteristics based on the differential characteristics, includes: The basic indicator data and error tracking data are associated with each other according to preset items to obtain an associated dataset. Event sequence matching is performed on the data in the associated dataset to obtain the key event data corresponding to each error event. Based on the key event data for each error event, analyze and calculate the numerical and performance differences of the abnormal deviation feature dimensions, and output differentiated features; Based on the differentiated features, the triggering logic and the correlation between the impact links of each sub-feature in the operating feature domain of the client are determined, and the analysis results of the driver client's operating and usage features are obtained.

[0008] Furthermore, the step of generating the optimized operation strategy for the ride-hailing driver's corresponding client based on the analysis results includes: Based on the analysis results, the problems of each sub-feature within the running feature domain are identified, and a problem location list is obtained; For different types of problems in the problem location list, the corresponding optimization operation rule library is matched to obtain the initial operation strategy corresponding to different types of problems; The initial operation strategies are prioritized based on the ride-hailing operation objectives, and the initial operation strategies with higher priority than the preset priority are selected as candidate operation strategies. The feasibility of the candidate operation strategies is verified, and the candidate operation strategies that pass the verification are then used as the optimization operation strategies.

[0009] Furthermore, based on the analysis results, the problem of determining each sub-feature within the operational feature domain is obtained, resulting in a problem location list, including: The analysis results of the driver client's operation and usage characteristics are broken down to obtain a dimensional feature subset including page usage characteristics, interaction behavior patterns, operating performance status, and abnormal fault characteristics; Based on the dimensional feature subset, a set of abnormal feature items exceeding the baseline threshold range is obtained; The set of abnormal features is validated to obtain a valid set of abnormal features; Based on the effective anomaly feature set, a structured list of driver client problem locations with clear problem types and associated business scenarios is obtained.

[0010] Furthermore, the optimization of the client version for the ride-hailing driver based on the optimized operation strategy includes: The optimized operation strategy was broken down and transformed to obtain a structured list of version optimization requirements; Based on the aforementioned version optimization requirement list, a solution is designed to obtain a version optimization document that includes the development scope, technology selection, and time nodes. The optimization document for this version was developed and tested to obtain a client-optimized version installation package that passed the tests. The optimized client-side installation package was used for canary rollout iterations to obtain a version optimization report.

[0011] Furthermore, the method also includes: Data evaluation was conducted based on runtime data from the optimized client version to obtain quantitative evaluation results; Based on the quantitative evaluation results, the strategy is iterated to obtain personalized optimization sub-strategies for specific driver segments.

[0012] Secondly, embodiments of the present invention provide an analysis device for embedded data, the device comprising: The data collection module is used to collect front-end embedded data generated when ride-hailing drivers perform terminal operations on the corresponding client. The analysis module is used to analyze the front-end embedded data to obtain analysis results of the driver client's operation and usage characteristics from different dimensions; The generation module is used to generate optimized operation strategies for the corresponding client of the ride-hailing driver based on the analysis results; The optimization module is used to optimize the version of the client corresponding to the ride-hailing driver based on the optimized operation strategy.

[0013] Thirdly, embodiments of the present invention provide a computer device, including: a memory and a processor, the memory and the processor being communicatively connected to each other, the memory storing computer instructions, and the processor executing the computer instructions to perform the method described in the first aspect or any corresponding embodiment thereof.

[0014] Fourthly, embodiments of the present invention provide a computer-readable storage medium storing computer instructions that cause a computer to perform the method described in the first aspect or any of its corresponding embodiments.

[0015] This application first collects all front-end event tracking data, integrating scattered data such as page dwell time, resource loading, and API performance. This breaks down data silos, establishing a complete causal chain for operations and accurately tracing the root cause of problems. Second, it directly analyzes real-time event tracking data, abandoning batch processing methods, enabling real-time capture of performance degradation trends and addressing the pain point of insufficient timeliness. Third, the analysis process constructs timestamped user operation paths, associating resource loading failures with corresponding page elements, filling gaps in the analysis and accurately pinpointing the causes of failures. Finally, the optimized implementation forms a closed loop, fundamentally improving the accuracy and timeliness of client-side operational analysis and ensuring stable client operation. Attached Figure Description

[0016] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the drawings used in the description of the specific embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0017] Figure 1 This is a flowchart illustrating a method for analyzing embedded data according to some embodiments of the present invention; Figure 2 This is a flowchart illustrating another method for analyzing embedded data according to some embodiments of the present invention; Figure 3 This is a flowchart illustrating another method for analyzing embedded data according to some embodiments of the present invention; Figure 4 This is a flowchart illustrating another method for analyzing embedded data according to some embodiments of the present invention; Figure 5 This is a structural block diagram of an analysis device for embedded data according to an embodiment of the present invention; Figure 6 This is a schematic diagram of the hardware structure of a computer device according to an embodiment of the present invention. Detailed Implementation

[0018] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0019] According to embodiments of the present invention, a method, apparatus, device, and medium for analyzing embedded data are provided. It should be noted that the steps shown in the flowcharts in the accompanying drawings can be executed in a computer, such as a set of computer-executable instructions. Furthermore, although a logical order is shown in the flowcharts, in some cases, the steps shown or described may be executed in a different order than that shown here.

[0020] This embodiment provides a method for analyzing embedded data. Figure 1 This is a flowchart of a method for analyzing embedded data according to an embodiment of the present invention, such as... Figure 1 As shown, the process includes the following steps: Step S101: Collect front-end embedded data generated when ride-hailing drivers perform terminal operations on the corresponding client.

[0021] In this embodiment, the full collection of six categories of data is achieved through automated collection logic deployed in the client-side front-end SDK layer: 1) Page lifecycle monitoring, automatically recording millisecond-level timestamps of page entry and exit and calculating dwell time; 2) User interaction capture, tracking interactive events such as button clicks and list scrolling and recording the screen position of interactive elements; 3) Resource loading monitoring, monitoring the success / failure status and time consumption of loading resources such as images and scripts, and associating them with corresponding page elements; 4) API performance tracking, recording the start and end time of interface calls, response status codes, and data volume; 5) Automatic error capture, intercepting JavaScript runtime errors and recording details such as HTTP status codes for resource loading failures; and 6) Performance metric collection, collecting key rendering metrics such as FCP and LCP, and synchronously recording memory usage change curves.

[0022] Step S102: Analyze the front-end embedded data to obtain analysis results of the driver client's operation and usage characteristics from different dimensions.

[0023] In this embodiment of the application, the front-end embedded data is analyzed to obtain analysis results of the driver client's operation and usage characteristics from different dimensions, such as... Figure 2 As shown, it includes: Step A1: Construct a complete user operation path with timestamps based on front-end event tracking data, and output a structured dataset.

[0024] First, the full set of front-end event tracking data collected from the ride-hailing driver's client is retrieved, covering various event data such as page lifecycle, user interaction, and resource loading. Second, using device ID and session ID as association identifiers, the discrete event tracking data is sorted sequentially by millisecond-level timestamps. Then, all operations within a single driver's session are chained together, such as "entering the homepage → clicking the order list → selecting to accept an order → jumping to the navigation page," forming a complete user operation path. Finally, attributes such as event type, operation duration, and page identifier are added to each path, and the data is organized in the format of "session ID-timestamp-operation behavior-associated page" to output a standardized structured dataset.

[0025] Step A2: Perform multi-dimensional indicator calculations on the dataset according to the preset time window to obtain basic indicator data.

[0026] First, define preset time windows, including different granularities such as 10-second real-time windows, 1-minute scrolling windows, and 1-hour statistical windows. Second, based on structured datasets, divide the data into three major indicator dimensions: page usage, interaction behavior, and performance. Clarify the calculation criteria for each dimension; for example, page usage calculates access frequency and average dwell time, while performance calculates API response time and page load success rate. Then, aggregate and calculate the indicators within each time window, such as calculating the average dwell time on the order page and the click-through rate of the order acceptance button within one hour. Finally, integrate the calculation results from different windows and dimensions, labeling the corresponding statistical time period and the number of users covered by the indicators to form a systematic set of basic indicator data.

[0027] Step A3: Identify erroneous event tracking data from the front-end event tracking data.

[0028] First, retrieve the original front-end event tracking data and clarify the rules for judging error tracking points, covering types such as JavaScript runtime errors, resource loading failures (e.g., HTTP 404 / 500 status codes), API call exceptions, and page rendering failures. Second, use a preset error feature matching algorithm to filter out event tracking data containing error identifier fields, while verifying the integrity of the data to ensure that error information includes key information such as error type, occurrence time, involved pages, and device model. Then, remove false error data caused by abnormal data collection, such as resource loading failure records falsely reported due to network fluctuations. Finally, classify and organize the filtered valid error data according to error type, and output a clearly labeled set of error tracking data.

[0029] Step A4: Analyze the differences between basic indicator data and error tracking data, and construct the analysis results of client operation and usage characteristics based on the differences.

[0030] Specifically, the analysis examines the differences between basic indicator data and error tracking data, and constructs analysis results on the driver-side client operation and usage characteristics based on these differences, including: Step A401: Associate the basic indicator data and the error tracking data according to preset entries to obtain an associated dataset, and perform event sequence matching on the data in the associated dataset to obtain the key event data corresponding to each error event.

[0031] The pre-defined association entries include three core identifiers: first, a globally unique TraceID, used to link all operations and error logs within a single session; second, a device ID, used to identify the terminal device to which the data belongs and eliminate cross-device data interference; and third, a millisecond-precision timestamp, used to calibrate the sequence of events. During the data association phase, basic metric data (covering normal operation benchmark data such as average page dwell time, API response time, and button click frequency) and error tracking data (including JavaScript runtime errors, HTTP status codes for resource loading failures, and API call exception logs) are retrieved. A mapping relationship is established through the pre-defined entries, invalid data lacking association identifiers is removed, and an association dataset containing four dimensions of information: "benchmark metric - error log - device - time". Subsequently, event sequence matching was performed on the associated dataset. Using timestamps as the sorting basis, the discrete basic indicator fluctuation data and error tracking data were time-series aligned to locate the indicator change nodes before, during, and after each error event. For example, the "404 error in map module resource loading" was bound to indicator data such as "sudden increase in page loading time" and "doubling of user dwell time". Finally, the key event data corresponding to each error event, which includes time sequence features and attribute features, was extracted.

[0032] Step A402: Based on the key event data of each error event, analyze and calculate the numerical and performance differences of the abnormal deviation feature dimensions, and output the differential features.

[0033] In the analysis and calculation phase, the basic indicators from the key event data are first used as the baseline for normal operation. For example, data such as "average order acceptance page loading time of 1.2 seconds," "average API response time of 200ms," and "button click success rate of 98%" are set as baseline thresholds. Subsequently, for the key data of each error event, the difference values ​​in three dimensions are calculated: In terms of indicator fluctuation, the deviation ratio between the real-time indicator at the time of the error event and the baseline threshold is compared. For example, the page loading time corresponding to an error event is 8.7 seconds, and the deviation from the baseline value is as high as 625%. In terms of operational behavior deviation, the trend of changes in user operational behavior before and after the error event is analyzed. For example, before the error occurred, users clicked the order acceptance button once every 3 seconds on average, and the operation frequency decreased by 50% after the error occurred. In terms of the scope of failure impact, the number of devices, user groups, and business scenario coverage that triggered similar error events are statistically analyzed. For example, a certain type of resource loading error affected 30% of low-configuration Android devices. At the same time, it will analyze the performance differences of each dimension, distinguish between "numerical excess anomalies" (such as the fluctuation range of indicators exceeding the threshold) and "behavioral mutation anomalies" (such as sudden changes in user operation behavior), and eliminate the influence of different indicator dimensions through data normalization processing. Finally, it will integrate the difference data of each dimension to form a differentiated feature that covers both numerical differences and performance differences, with error events as the unit.

[0034] Step A403: Based on the differentiated features, determine the triggering logic and the correlation between the impact links of each sub-feature in the operating feature domain of the client, and obtain the analysis results of the driver client's operating and usage features.

[0035] During the analysis, the differentiated features were first categorized into four sub-feature domains. For example, "page load time deviation" was categorized into the operational performance status sub-feature domain, "decreased user operation frequency" into the interaction behavior pattern sub-feature domain, and "resource loading error type" into the abnormal failure feature sub-feature domain. Subsequently, for the categorized differentiated features, the triggering logic and impact chains between each sub-feature were explored: on the one hand, correlation analysis was used to determine the triggering logic of the anomalies. For example, "404 error in map module resource loading" was identified as the direct cause of "sudden increase in page load time," and "page load timeout" further triggered the behavioral change of "user canceling operation." On the other hand, link analysis was used to clarify the scope of the anomalies' impact. For example, a certain type of API timeout error not only caused "order acceptance function failure" but also indirectly affected "delay in updating driver income statistics page data." Simultaneously, isolated and unrelated differentiated features were eliminated, and feature chains with causal relationships were integrated. For example, a complete impact chain of "resource loading failure → performance index deterioration → abnormal user behavior → business function obstruction" was constructed. Finally, the relationships between the various sub-feature domains are summarized to form an analysis result covering the driver client's operating and usage characteristics, including client usage behavior preferences, performance bottlenecks, and abnormal fault triggering mechanisms.

[0036] Step S103: Generate optimized operation strategies for the corresponding client of ride-hailing drivers based on the analysis results.

[0037] In this embodiment of the application, an optimized operation strategy for the corresponding client of the ride-hailing driver is generated based on the analysis results, such as... Figure 3 As shown, including: Step B1: Based on the analysis results, identify the problems of each sub-feature within the running feature domain and obtain a problem location list.

[0038] First, the analysis results of the driver client's operation and usage characteristics are retrieved and broken down into four sub-feature domains: page usage characteristics, interaction behavior patterns, operational performance status, and abnormal fault characteristics. From each sub-feature domain, abnormal behaviors that deviate significantly from the normal benchmark are extracted. For example, in the page usage characteristic domain, issues such as "the average daily dwell time on the wallet page is less than the benchmark value by 40%" and "the exit rate of the order details page is higher than the benchmark value by 60%" are identified. In the operational performance status domain, performance bottlenecks such as "the percentage of devices with the homepage LCP index exceeding 3 seconds reaches 35%" and "the average weekly API call timeout rate exceeds the standard by 20%" are screened. In the abnormal fault characteristic domain, fault points such as "the average daily occurrence of 404 errors in map module resource loading exceeds 600 times" and "the percentage of JavaScript runtime errors in Android versions below 8.0 reaches 75%" are extracted.

[0039] Secondly, each identified problem is labeled with multi-dimensional attributes to clarify the sub-feature domain to which the problem belongs, the core business scenario it affects, the device models and versions involved, the time distribution pattern and frequency range of the occurrence. For example, labeling "no response when clicking the order acceptance button" belongs to the interaction behavior pattern domain problem, which affects the core scenario of drivers accepting orders, occurs mainly during high-concurrency periods, and involves about 20% of low-end and mid-range Android devices.

[0040] Then, duplicate and isolated, unrelated issue entries are removed to avoid redundant records of similar issues. Finally, all issues are structured and integrated according to the dimensions of "issue level - scope of impact - frequency of occurrence" to generate a driver-side client issue location list with clear entries and well-defined attributes, providing precise issue guidance for subsequent strategy matching.

[0041] Step B2: For different types of problems in the problem location list, match the corresponding optimization operation rule library to obtain the initial operation strategy corresponding to different types of problems.

[0042] First, each problem in the problem identification list is categorized into four main types based on its attributes: performance optimization, interaction optimization, fault repair, and experience improvement. For example, "page loading time exceeds the limit" and "interface response delay" are categorized into performance optimization, "button click has no feedback" and "sluggish swiping operation" are categorized into interaction optimization, "resource loading failure" and "frequent runtime errors" are categorized into fault repair, and "low feature utilization" and "lengthy page jump path" are categorized into experience improvement.

[0043] Secondly, a pre-defined optimization and operation rule library is retrieved. This rule library is a standardized set of strategies built upon historical optimization cases, technical solutions, and business practice experience. It covers multiple mature solutions for each type of problem, and each solution is marked with its applicable scenarios, device compatibility range, and expected optimization effects. Then, precise matching is performed based on the attribute annotation information of the problem. For example, for "JavaScript runtime errors in Android versions below 8.0," targeted strategies such as "adaptation to old compatibility code" and "API interface version downgrade compatibility" are prioritized. For the "slow homepage loading" problem, combined strategies such as "lazy loading of non-core resources," "dynamic scheduling of CDN nodes," and "optimization of data caching strategies" are matched to ensure that the matched strategies are highly consistent with the problem scenario.

[0044] Finally, the matching strategies for each problem are compiled and summarized, and at least 1-2 initial operation strategies are generated for each problem, forming an initial strategy set covering various types of problems, providing sufficient strategy options for subsequent priority ranking.

[0045] Step B3: Prioritize the initial operation strategies based on the ride-hailing operation goals, and select the initial operation strategies with higher priority than the preset priority as candidate operation strategies.

[0046] First, the core operational objectives of the ride-hailing platform were clearly defined, encompassing four dimensions: improving driver operational efficiency, enhancing client-side stability, increasing core business conversion rates, and improving user experience satisfaction. Simultaneously, differentiated weighting coefficients were assigned to each dimension based on the platform's phased operational priorities. For example, the weighting for "improving core business conversion rates" was set at 40%, "enhancing client-side stability" at 30%, "improving driver operational efficiency" at 20%, and "improving user experience satisfaction" at 10%. Second, for each initial operational strategy, the technical and business teams jointly evaluated its contribution to the four operational objectives. For instance, the "map module fault repair strategy" contributed 95% to "enhancing client-side stability" and 80% to "improving core business conversion rates"; the "order acceptance button interaction optimization strategy" contributed 90% to "improving core business conversion rates" and 60% to "improving driver operational efficiency." The evaluation results were then quantified and assigned values.

[0047] Then, a weighted calculation model is used to obtain the priority score for each initial operational strategy. The calculation formula is: Strategy Priority Score = ∑(Target Contribution × Target Weight). For example, a strategy score = (95% × 30%) + (80% × 40%) + (50% × 20%) + (40% × 10%) = 76.5 points. Finally, a preset priority threshold is set, such as a score ≥ 60 points. Initial operational strategies with scores higher than the threshold are selected as candidate operational strategies to ensure that the selected strategies can best match the platform's core operational needs and avoid wasting resources on low-value optimization directions.

[0048] Step B4: Verify the feasibility of the candidate operational strategies, and optimize the operational strategies based on the verified candidate operational strategies.

[0049] First, conduct a technical feasibility verification. The technical team comprehensively evaluates whether candidate strategies conform to the current client-side technical architecture, development specifications, and hardware adaptation requirements. For example, for the "CDN node dynamic scheduling strategy," verify whether the existing architecture supports real-time monitoring and dynamic switching of multiple node statuses. For the "legacy compatibility adaptation strategy," verify whether the existing codebase has room for version adaptation. Strategies with excessively high technical implementation difficulty or conflicts with the existing architecture are eliminated, and alternative optimization solutions are proposed. Second, conduct a cost controllability verification. Assess whether the development cost, manpower cost, server resource cost, and time cost required for strategy implementation are within the preset budget. For example, compare the cost investment of the "full resource reconstruction" and "critical resource local optimization" strategies. The former has three times the manpower cost and twice the time cycle of the latter. Prioritize retaining the local optimization strategy with lower cost and shorter cycle.

[0050] Next, a risk foreseeability verification is conducted, analyzing potential technical and business risks that may arise after the strategy is implemented. For example, assessing whether "data caching strategy adjustment" will lead to data consistency issues, and whether "interaction logic modification" will cause risks of driver incompatibility. For high-risk strategies, supplementary and improved risk response plans are required, and high-risk strategies without effective plans are eliminated. Finally, an effectiveness quantifiability verification is conducted to confirm whether each candidate strategy has clear effectiveness evaluation indicators and judgment criteria. For example, the "page loading optimization strategy" needs to specify quantitative targets such as "FCP index reduction of 30%" and "page loading time shortened to within 1.5 seconds," and the "fault repair strategy" needs to specify evaluation criteria such as "error occurrence rate reduction of more than 90%." Strategies without clear quantitative indicators are eliminated. After comprehensive verification across four dimensions—technical feasibility, cost controllability, risk foreseeability, and effectiveness quantifiability—the candidate operational strategies for all verified items will be integrated to ultimately form a driver-side client optimization operational strategy that can be directly implemented, providing a clear action plan for subsequent version optimizations.

[0051] Step S104: Optimize the version of the client software for ride-hailing drivers based on the optimized operation strategy.

[0052] In this embodiment of the application, problems in each sub-feature within the running feature domain are determined based on the analysis results, resulting in a problem location list, such as... Figure 4 As shown, it includes: Step C1 involves breaking down the analysis results of the driver client's operation and usage characteristics to obtain a dimensional feature subset that includes page usage characteristics, interaction behavior patterns, operating performance status, and abnormal fault characteristics.

[0053] First, retrieve the complete analysis results of the driver client's operation and usage characteristics. These results cover multiple dimensions of information, including driver operation behavior, client performance, and fault occurrence. Perform data preprocessing to remove duplicate data, missing values, and outlier data to ensure the accuracy and completeness of the basic data.

[0054] Secondly, the dimensions and core categories of each dimension are clearly defined. The four core dimensions are page usage characteristics, interaction behavior patterns, runtime performance status, and abnormal fault characteristics. Page usage characteristics focus on usage-related data such as page access frequency, dwell time, and navigation path. Interaction behavior patterns focus on the frequency, timing, and abnormal behavior of operations such as button clicks and list scrolling. Runtime performance status covers performance data such as key rendering metrics, API response time, and resource loading efficiency. Abnormal fault characteristics include fault-related information such as error type, frequency of occurrence, and affected devices.

[0055] Then, according to the definition of each dimension, the preprocessed analysis results are categorized and split one by one. For example, "120 daily visits to the order list page, with an average stay of 8 seconds" is categorized as page usage features; "no response after 3 consecutive clicks of the order acceptance button" is categorized as interaction behavior patterns; "homepage LCP index of 2.8 seconds and API timeout rate of 5%" is categorized as operational performance status; and "300 daily 404 errors in map resource loading" is categorized as abnormal fault features. During the splitting process, feature label matching is used to ensure that the data classification is complete and without overlap. Finally, the data after splitting each dimension is structured and organized, and attribute information such as dimension identifier, data source, and statistical time window is added to each feature subset, forming four independent and standardized dimensional feature subsets.

[0056] Step C2 involves filtering based on a dimensional feature subset to obtain a set of abnormal feature items that exceed the baseline threshold range.

[0057] First, retrieve the four dimensional feature subsets and load the preset benchmark threshold system for each dimension. This system is based on historical operating data, industry standards, and business needs, and includes two types of thresholds: static thresholds and dynamic thresholds. Static thresholds include "average page dwell time 5-60 seconds" and "API response time ≤300ms". Dynamic thresholds include "core page access frequency fluctuation does not exceed ±20% of the daily average benchmark" and "error occurrence rate does not exceed 15% of the weekly average benchmark". The thresholds can be dynamically adjusted according to the platform's operational stage.

[0058] Secondly, threshold comparisons were conducted item by item for each dimensional feature subset. For example, within the page usage feature subset, data such as "average dwell time on the wallet page is only 3 seconds (below the static threshold of 5 seconds)" and "order details page visit frequency drops sharply by 35% (exceeding the dynamic threshold)" were selected; within the runtime performance status subset, data such as "homepage FCP index is 3.2 seconds (above the static threshold of 2.5 seconds)" and "resource loading success rate is 88% (below the static threshold of 95%)" were selected; within the abnormal fault feature subset, data such as "JavaScript runtime error rate is 8% (above the static threshold of 1%)" were selected. During the comparison process, the deviation between the actual value of the feature item and the benchmark threshold was recorded simultaneously.

[0059] Then, the initially screened suspected abnormal features are checked a second time to eliminate false anomalies caused by statistical errors or single, accidental fluctuations. For example, isolated data caused by abnormal operation of a single device is removed, while features that consistently exceed the threshold or appear in batches of devices are retained. Finally, the verified abnormal features are categorized and organized according to their respective dimensions, and key information such as the deviation ratio, occurrence time, and scope of each anomaly is labeled to form a structured set of abnormal features.

[0060] Step C3: Verify the set of abnormal features to obtain a valid set of abnormal features.

[0061] First, conduct anomaly verification by retrieving source data such as original front-end tracking data and client operation logs. For each item in the set of anomaly features, trace the data source to verify whether the collection process of the anomaly data is compliant and whether the statistical logic is accurate. For example, for the anomaly item "API response timeout rate exceeds the standard", check the original call log to confirm whether there is a timeout record, eliminate misjudgment anomalies caused by data collection vulnerabilities and statistical deviations, and directly remove anomalies that cannot be traced or are verified as false.

[0062] Secondly, perform anomaly correlation verification, analyze the temporal correlation, causal correlation, and scenario correlation between various anomaly features, such as verifying whether there is temporal synchronization between "page loading time exceeds the limit" and "resource loading failure", and determine whether the two are causally related. At the same time, investigate isolated and unrelated anomalies, such as occasional anomalies that occur once on a single device and have no impact on business. These anomalies are eliminated because they do not have universality or impact.

[0063] Next, the impact scope of the anomalies is verified. The number of devices, driver groups, and covered business scenarios involved in each anomaly feature are statistically analyzed. For example, if a "map module loading error" only affects 1% of low-end devices and does not impact core businesses such as order taking and navigation, it can be classified as a low-impact anomaly. If it affects more than 30% of devices and impacts core businesses, it is classified as a high-impact anomaly. Simultaneously, low-value anomalies with no actual business impact are removed. Finally, the anomaly features that pass the triple verification of authenticity, relevance, and impact scope are integrated, and the impact level, related anomaly chains, and affected businesses of each valid anomaly are labeled to form a standardized set of valid anomaly features.

[0064] Step C4 involves classifying the valid anomaly feature set to obtain a structured list of driver-side client-side problem locations with clear problem types and associated business scenarios.

[0065] In this embodiment, the client version of the ride-hailing driver is optimized based on the optimized operation strategy, including: breaking down and transforming the optimized operation strategy to obtain a structured list of version optimization requirements; designing a solution based on the list of version optimization requirements to obtain a version optimization document that includes the development scope, technology selection, and time nodes; developing and testing the version optimization document to obtain a client optimized version installation package that has passed the test; and conducting gray-scale iterations based on the client optimized version installation package to obtain a version optimization report.

[0066] Specifically, first, obtain the optimization operation strategy and clarify the optimization goals and expected effects in core areas such as performance improvement and fault repair; second, break down the strategy into concrete requirements according to the strategy category, such as breaking down "map loading optimization" into items such as resource path verification and compatibility adaptation; then, mark the category, associated scenario, priority and acceptance criteria of each requirement; finally, sort and organize them according to priority to generate a structured list containing core fields such as requirement ID and name, providing a clear basis for solution design.

[0067] Break down the requirements, define the core and secondary development scope, and clarify the compatible device models and versions; secondly, combine the technical architecture with the selection and adaptation scheme, such as using CDN scheduling + lazy loading combination for resource loading optimization; then, formulate phased time nodes, clarify the deliverables and responsible persons for each phase, and reserve buffer time to deal with unexpected problems; finally, integrate the development scope, technology selection and other contents to form a standardized document to guide the orderly progress of subsequent development and testing work.

[0068] The environment is set up according to the documentation, tasks are broken down and developed in a modular way, and code reviews are conducted regularly to ensure code quality. Secondly, a simulated test environment is built based on historical data, covering multiple networks, multiple devices, and multiple scenarios. Then, multi-dimensional tests such as functionality, performance, and compatibility are carried out, and unqualified issues are iteratively modified until the acceptance criteria are met. Finally, the qualified version is packaged, signature and security verification are completed, and an optimized client version installation package that can be installed and run normally is output.

[0069] A phased gray-scale strategy was developed, splitting the push groups by region, device type, and driver activity, and setting push intervals and monitoring nodes for each batch; secondly, after the push, the front-end SDK was used to collect operational data in real time to compare the differences in core indicators before and after optimization; then, if an anomaly occurred, the push was immediately paused and rolled back, the problem was investigated, iterative optimization was carried out, and the push was resumed; if the target was met, the full release was promoted; finally, the gray-scale data, optimization effects, problem records, etc., were summarized to form a version optimization report including effect evaluation, summary and suggestions.

[0070] In this embodiment of the application, the method further includes: performing data evaluation based on the running data of the optimized client version to obtain a quantitative evaluation result; and performing strategy iteration based on the quantitative evaluation result to obtain a personalized optimization sub-strategy for the driver segment.

[0071] First, retrieve the full runtime data of the optimized version collected by the front-end SDK, covering core dimensions such as performance indicators, error rate, and user operation behavior, and simultaneously extract historical benchmark data as a comparison reference; second, set up a quantitative evaluation indicator system, including core indicators such as performance improvement rate, failure rate reduction rate, and operational efficiency improvement, and clarify the calculation method and weight allocation of each indicator. Then, outliers are removed through data cleaning, and comparative analysis is used to calculate the differences in metrics between the optimized version and historical versions, such as the percentage reduction in homepage loading time and the decrease in map module error rate. Finally, the evaluation data from various dimensions are integrated to form a quantitative evaluation result that includes metric comparison, effect conclusions, and problem analysis, providing data support for subsequent strategy iterations.

[0072] The analysis and quantitative evaluation results pinpointed areas where optimization efforts fell short of expectations, while also identifying performance differences among different driver groups, such as the performance gap between high-end and low-end equipment, and the operational habits of high-frequency and low-frequency drivers. Secondly, the driver groups were segmented based on dimensions such as equipment configuration, order frequency, and geographical distribution to clarify the core pain points of each segment. Then, combining the evaluation results with the characteristics of each segment, targeted iterative optimization strategies were developed. For example, a lightweight resource loading sub-strategy was developed for low-end equipment users, and a quick-access interface sub-strategy was designed for high-frequency order-taking drivers. Finally, a feasibility pre-evaluation was conducted on the iterated sub-strategies to clarify the applicable groups, implementation paths, and expected effects of each sub-strategy, integrating them into a personalized optimization sub-strategy set for each driver segment.

[0073] This embodiment also provides an analysis device for embedded data, which is used to implement the above embodiments and preferred embodiments; details already described will not be repeated. As used below, the term "module" can refer to a combination of software and / or hardware that performs a predetermined function. Although the device described in the following embodiments is preferably implemented in software, hardware implementation, or a combination of software and hardware, is also possible and contemplated.

[0074] This embodiment provides a device for analyzing embedded data, such as... Figure 5 As shown, it includes: The data collection module 501 is used to collect front-end embedded data generated when ride-hailing drivers perform terminal operations on the corresponding client. Analysis module 502 is used to analyze front-end embedded data to obtain analysis results of the driver client's operation and usage characteristics from different dimensions; The generation module 503 is used to generate optimized operation strategies for the corresponding client of ride-hailing drivers based on the analysis results; Optimization module 504 is used to optimize the version of the client software for ride-hailing drivers based on the optimized operation strategy.

[0075] In this embodiment of the application, the analysis module 502 is used to construct a complete user operation path with timestamps based on the front-end tracking data and output a structured dataset; perform multi-dimensional indicator calculations on the dataset according to a preset time window to obtain basic indicator data; identify erroneous tracking data from the front-end tracking data; analyze the differential characteristics between the basic indicator data and the erroneous tracking data, and construct the analysis results of client operation and usage characteristics based on the differential characteristics.

[0076] In this embodiment, the analysis module 502 is used to associate basic indicator data and error tracking data according to preset entries to obtain an associated dataset, and to perform event sequence matching on the data in the associated dataset to obtain key event data corresponding to each error event; based on the key event data of each error event, to analyze and calculate the numerical differences and performance differences of the abnormal deviation feature dimension, and output differentiated features; based on the differentiated features, to determine the triggering logic and the correlation between the impact links of each sub-feature in the operating feature domain of the client, and to obtain the analysis results of the driver client's operating and usage features.

[0077] In this embodiment, the generation module 503 is used to determine the problems of each sub-feature within the operating feature domain based on the analysis results, and obtain a problem location list; for different types of problems in the problem location list, it matches the corresponding optimized operation rule base to obtain the initial operation strategy corresponding to different types of problems; it prioritizes the initial operation strategies using ride-hailing operation goals, and uses the initial operation strategies with higher priority than the preset priority as candidate operation strategies; it performs feasibility verification on the candidate operation strategies, and selects the candidate operation strategies that pass the verification as optimized operation strategies.

[0078] In this embodiment, the optimization module 504 is used to split the analysis results of the driver client's operation and usage characteristics to obtain a dimensional feature subset of page usage characteristics, interaction behavior patterns, operating performance status, and abnormal fault characteristics; to filter based on the dimensional feature subset to obtain a set of abnormal feature items that exceed the benchmark threshold range; to verify the set of abnormal feature items to obtain a set of valid abnormal features; and to classify based on the set of valid abnormal features to obtain a structured list of driver client problem locations with clear problem types and associated business scenarios.

[0079] In this embodiment, the optimization module 504 is used to decompose and transform the optimization operation strategy to obtain a structured version optimization requirement list; design a solution based on the version optimization requirement list to obtain a version optimization document that includes the development scope, technology selection, and time nodes; develop and test the version optimization document to obtain a client-side optimized version installation package that has passed the test; and perform gray-scale iteration based on the client-side optimized version installation package to obtain a version optimization report.

[0080] In this embodiment of the application, the device further includes: an evaluation module, used to perform data evaluation based on the operating data of the client-optimized version to obtain a quantitative evaluation result; and to perform strategy iteration based on the quantitative evaluation result to obtain a personalized optimization sub-strategy for the driver segment.

[0081] Please see Figure 6 , Figure 6 This is a schematic diagram of the structure of a computer device provided in an optional embodiment of the present invention, such as... Figure 6 As shown, the computer device includes one or more processors 10, memory 20, and interfaces for connecting the components, including high-speed interfaces and low-speed interfaces. The components communicate with each other via different buses and can be mounted on a common motherboard or otherwise installed as needed. The processors can process instructions executed within the computer device, including instructions stored in or on memory to display graphical information of a GUI on external input / output devices (such as display devices coupled to the interfaces). In some alternative implementations, multiple processors and / or multiple buses can be used with multiple memories and multiple memory modules, if desired. Similarly, multiple computer devices can be connected, each providing some of the necessary operations (e.g., as a server array, a group of blade servers, or a multiprocessor system).

[0082] Processor 10 may be a central processing unit, a network processor, or a combination thereof. Processor 10 may further include a hardware chip. The hardware chip may be an application-specific integrated circuit (ASIC), a programmable logic device (PLD), or a combination thereof. The programmable logic device may be a complex programmable logic device (CAMP), a field-programmable gate array (FPGA), a general-purpose array logic (GPA), or any combination thereof.

[0083] The memory 20 stores instructions executable by at least one processor 10 to cause the at least one processor 10 to perform the method shown in the above embodiments.

[0084] The memory 20 may include a program storage area and a data storage area. The program storage area may store the application program required for operation or at least one function; the data storage area may store data created by the use of the computer device based on the display of a mini-program landing page. Furthermore, the memory 20 may include high-speed random access memory and may also include non-transient memory, such as at least one disk storage device, flash memory device, or other non-transient solid-state storage device. In some alternative embodiments, the memory 20 may optionally include memory remotely located relative to the processor 10, which can be connected to the computer device via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.

[0085] The memory 20 may include volatile memory, such as random access memory; the memory may also include non-volatile memory, such as flash memory, hard disk or solid-state drive; the memory 20 may also include a combination of the above types of memory.

[0086] The computer device also includes a communication interface 30 for communicating with other devices or communication networks.

[0087] This invention also provides a computer-readable storage medium. The methods described above according to embodiments of the invention can be implemented in hardware or firmware, or implemented as computer code that can be recorded on a storage medium, or implemented as computer code downloaded via a network and originally stored on a remote storage medium or a non-transitory machine-readable storage medium and then stored on a local storage medium. Thus, the methods described herein can be processed by software stored on a storage medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware. The storage medium can be a magnetic disk, optical disk, read-only memory, random access memory, flash memory, hard disk, or solid-state drive, etc.; further, the storage medium can also include combinations of the above types of memory. It is understood that computers, processors, microprocessor controllers, or programmable hardware include storage components capable of storing or receiving software or computer code, which, when accessed and executed by the computer, processor, or hardware, implements the methods shown in the above embodiments.

[0088] Although embodiments of the invention have been described in conjunction with the accompanying drawings, those skilled in the art can make various modifications and variations without departing from the spirit and scope of the invention, and such modifications and variations all fall within the scope defined by the appended claims.

Claims

1. A method for analyzing embedded data, characterized in that, The method includes: Collect front-end embedded data generated when ride-hailing drivers perform terminal operations on the corresponding client; The front-end embedded data was analyzed to obtain analysis results on the operation and usage characteristics of the driver client from different dimensions; Based on the analysis results, an optimized operation strategy for the corresponding client of the ride-hailing driver is generated; Based on the optimized operation strategy, the client version corresponding to the ride-hailing driver is optimized.

2. The method according to claim 1, characterized in that, The analysis of the front-end embedded data yields analysis results on the driver client's operational and usage characteristics from different dimensions, including: Based on the aforementioned front-end data, a complete user operation path with timestamps is constructed, and a structured dataset is output. Perform multi-dimensional indicator calculations on the dataset according to a preset time window to obtain basic indicator data; Identify erroneous data points from the aforementioned front-end data points; The differences between the basic indicator data and the error tracking data are analyzed, and the analysis results of the client operation and usage characteristics are constructed based on the differences.

3. The method according to claim 2, characterized in that, The analysis of the differences between the basic indicator data and the error tracking data, and the construction of analysis results of the driver client's operation and usage characteristics based on the differences, includes: The basic indicator data and error tracking data are associated with each other according to preset items to obtain an associated dataset. Event sequence matching is performed on the data in the associated dataset to obtain the key event data corresponding to each error event. Based on the key event data for each error event, analyze and calculate the numerical and performance differences of the abnormal deviation feature dimensions, and output differentiated features; Based on the differentiated features, the triggering logic and the correlation between the impact links of each sub-feature in the operating feature domain of the client are determined, and the analysis results of the driver client's operating and usage features are obtained.

4. The method according to claim 1, characterized in that, The step of generating the optimized operation strategy for the ride-hailing driver's corresponding client based on the analysis results includes: Based on the analysis results, the problems of each sub-feature within the running feature domain are identified, and a problem location list is obtained; For different types of problems in the problem location list, the corresponding optimization operation rule library is matched to obtain the initial operation strategy corresponding to different types of problems; The initial operation strategies are prioritized based on the ride-hailing operation objectives, and the initial operation strategies with higher priority than the preset priority are selected as candidate operation strategies. The feasibility of the candidate operation strategies is verified, and the candidate operation strategies that pass the verification are then used as the optimization operation strategies.

5. The method according to claim 4, characterized in that, The problem identified based on the analysis results for each sub-feature within the operating feature domain is used to obtain a problem location list, including: The analysis results of the driver client's operation and usage characteristics are broken down to obtain a dimensional feature subset including page usage characteristics, interaction behavior patterns, operating performance status, and abnormal fault characteristics; Based on the dimensional feature subset, a set of abnormal feature items exceeding the baseline threshold range is obtained; The set of abnormal features is validated to obtain a valid set of abnormal features; Based on the effective anomaly feature set, a structured list of driver client problem locations with clear problem types and associated business scenarios is obtained.

6. The method according to claim 1, characterized in that, The optimization of the client version for the ride-hailing driver based on the optimized operation strategy includes: The optimized operation strategy was broken down and transformed to obtain a structured list of version optimization requirements; Based on the aforementioned version optimization requirement list, a solution is designed to obtain a version optimization document that includes the development scope, technology selection, and time nodes. The optimization document for this version was developed and tested to obtain a client-optimized version installation package that passed the tests. The optimized client-side installation package was used for canary rollout iterations to obtain a version optimization report.

7. The method according to claim 1, characterized in that, The method further includes: Data evaluation was conducted based on runtime data from the optimized client version to obtain quantitative evaluation results; Based on the quantitative evaluation results, the strategy is iterated to obtain personalized optimization sub-strategies for specific driver segments.

8. A device for analyzing embedded data, characterized in that, The device includes: The data collection module is used to collect front-end embedded data generated when ride-hailing drivers perform terminal operations on the corresponding client. The analysis module is used to analyze the front-end embedded data to obtain analysis results of the driver client's operation and usage characteristics from different dimensions; The generation module is used to generate optimized operation strategies for the corresponding client of the ride-hailing driver based on the analysis results; The optimization module is used to optimize the version of the client corresponding to the ride-hailing driver based on the optimized operation strategy.

9. A computer device, characterized in that, include: A memory and a processor, the memory and the processor being communicatively connected to each other, the memory storing computer instructions, the processor executing the computer instructions to perform the method of any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions for causing the computer to perform the method of any one of claims 1 to 7.