Data analysis method and device
By conducting horizontal comparison and vertical analysis of the traffic and sales data of the live broadcast room, we can identify false traffic and fake orders in live e-commerce, solve the identification problems in existing technologies, and achieve efficient and accurate anomaly detection and marketing optimization.
Patent Information
- Application Number
- CN202510902184.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-01
- Publication Date
- 2025-09-30
- Estimated Expiration
- 2045-07-01
AI Technical Summary
There is a problem in live e-commerce where fake traffic and fake orders are difficult to identify. Existing technical methods are prone to misjudgment or omission, and it is difficult to take into account the influence of multiple factors.
A two-stage data analysis method is adopted. First, the traffic and sales data of live broadcast rooms with the same category and similar popularity are compared horizontally to calculate the anomaly score; then, potential abnormal behavior is identified by vertically analyzing the voice information before the anchor is put on the shelf and the sales data in a short period of time after the shelf is put on the shelf.
It improves the accuracy and efficiency of anomaly detection, can promptly detect false traffic and fake order behaviors, maintain the health of the e-commerce ecosystem, optimize marketing strategies, and improve conversion rates and user satisfaction.
Smart Images

Figure CN120410573B_ABST
Abstract
Description
Technical Field
[0001] This specification relates to the field of e-commerce live streaming, and more specifically, this application relates to a data analysis method and device. Background Art
[0002] With the rapid rise of livestreaming e-commerce, more and more businesses, brands, and individual anchors are using livestreaming platforms to promote and sell their products. However, as livestreaming e-commerce rapidly develops, it also faces the following common problems:
[0003] 1. Fake traffic: Some live broadcast rooms artificially increase the popularity of live broadcasts by purchasing "fake fans" or using software to increase the number of views and likes, creating false high popularity and high interaction data, misleading platforms and consumers.
[0004] 2. Fake sales (fake orders): Some livestreaming platforms use fake orders and refunds to boost sales figures quickly. This behavior not only disrupts market order but also misleads real consumers about product sales and reviews.
[0005] 3. Difficulty in Monitoring and Identification: Traffic and sales data from livestreaming studios are often massive and fluctuate in real time, making it difficult for platforms to accurately distinguish real data from anomalous data using traditional single metrics (such as number of viewers or sales). Furthermore, livestreaming studios vary significantly in traffic sources, host styles, and promotional policies, making direct comparisons susceptible to interference from external factors.
[0006] Based on the above problems, related technologies usually adopt the following methods to detect anomalies:
[0007] 1. Single-dimensional statistical analysis: For example, judging fake traffic based solely on a sudden increase in the number of viewers or likes, or identifying fake orders based solely on abnormal order volume. However, such methods are prone to misjudgment or omission due to a lack of multiple verification methods.
[0008] 2. Model Prediction Comparison: Build a forecasting model to predict traffic and sales within a normal range, then compare the actual data with the predicted values. While this approach is effective, it relies heavily on model data and struggles to account for multiple factors such as the streamer's style and product type.
[0009] The above methods have limited effectiveness in practical applications and are difficult to accurately and quickly identify more subtle abnormal behaviors in live broadcast rooms. Therefore, it is necessary to propose a data analysis method and device to at least solve some of the above problems. Summary of the Invention
[0010] The Summary of the Invention introduces a series of simplified concepts that will be further described in the Detailed Description of the Invention. The Summary of the Invention of this application is not intended to limit the key features and essential technical features of the claimed technical solution, nor is it intended to determine the scope of protection of the claimed technical solution.
[0011] In a first aspect, the present application proposes a data analysis method, comprising:
[0012] Obtaining the first live broadcast traffic data of the first live broadcast room and the first sales data of the target product;
[0013] Obtaining second live broadcast traffic data of the second live broadcast room and second sales data of the target product, wherein the first live broadcast room and the second live broadcast room are live broadcast rooms selling the same product category, and the difference in live broadcast popularity between the first live broadcast room and the second live broadcast room is less than a preset threshold;
[0014] Determining whether the live broadcast sales data of the first live broadcast room and the second live broadcast room contain first abnormal information based on the first live broadcast traffic data, the first sales data, the second live broadcast traffic data, and the second sales data;
[0015] Obtaining the host's voice information in the first live broadcast room or the second live broadcast room before the product is put on the shelves and the sales data information in the first time period after the product is put on the shelves;
[0016] Based on the above-mentioned anchor voice information before the product is put on the shelves and the sales data information in the first time period after the product is put on the shelves, it is determined whether there is second abnormal information in the above-mentioned first live broadcast room or the above-mentioned second live broadcast room.
[0017] In a feasible implementation manner, the first live broadcast traffic data includes first user quantity data, first interaction quantity data, and first visit conversion data, and the second live broadcast traffic data includes second user quantity data, second interaction quantity data, and second visit conversion data;
[0018] The determining, based on the first live broadcast traffic data, the first sales data, the second live broadcast traffic data, and the second sales data, whether the live broadcast sales data of the first live broadcast room and the second live broadcast room contain first abnormal information includes:
[0019] Calculate the user quantity anomaly score based on the first user quantity data and the second user quantity data ;
[0020] Calculate the interaction quantity abnormality score based on the first interaction quantity data and the second interaction quantity data ;
[0021] Calculate the conversion rate anomaly score based on the first visit conversion data and the second visit conversion data ;
[0022] Calculate the sales difference anomaly score based on the first sales data and the second sales data ;
[0023] Calculate the traffic and sales correlation anomaly score based on the first user quantity data, the second user quantity data, the first sales data, and the second sales quantity data ;
[0024] The abnormal comprehensive score A is calculated according to the following formula:
[0025] + + + +
[0026] in, Score the above user number anomaly The corresponding first weight information, Score the abnormal number of interactions above The corresponding second weight information, Score the above conversion rate anomaly The corresponding third weight information, Score the above sales discrepancy anomaly The corresponding fourth weight information, Score the abnormal correlation between the above traffic and sales corresponding fifth weight information;
[0027] Based on the above-mentioned abnormal comprehensive score A and the preset threshold, it is determined whether the live broadcast sales data of the above-mentioned first live broadcast room and the above-mentioned second live broadcast room have the first abnormal information.
[0028] In a feasible implementation manner, the above method further includes:
[0029] Obtaining the traffic sensitivity and sales sensitivity of the target product, wherein the traffic sensitivity is determined based on the target product's launch time and product category information, and the sales sensitivity is determined based on the product price and promotion intensity;
[0030] Adjusting the first weight coefficient and the second weight coefficient according to the traffic sensitivity;
[0031] The third weight coefficient and the fourth weight coefficient are adjusted according to the sales sensitivity.
[0032] In a feasible implementation, it further includes:
[0033] Obtain the first live broadcast room type of the first live broadcast room and the second live broadcast room type of the second live broadcast room, wherein the live broadcast room types include brand live broadcast rooms, internet celebrity personal live broadcast rooms, and internet celebrity team live broadcast rooms;
[0034] In the case that the type of the first live broadcast room and the type of the second live broadcast room are different, a live broadcast room of the same type is matched for the first live broadcast room and / or the second live broadcast room, and a judgment is made as to whether there is any abnormality in the live broadcast sales data.
[0035] In a feasible implementation manner, the determining whether the first live broadcast room or the second live broadcast room has the second abnormal information based on the host voice information before the product is put on the shelves and the sales data information in the first time period after the product is put on the shelves includes:
[0036] Extracting voice features based on the host's voice information before the product is put on the shelves, wherein the voice features include emotional features, speech speed features, intonation features, and marketing word features;
[0037] Extracting sales data features of the first time period based on the sales data information of the first time period, wherein the sales data features include time series features and peak-valley features;
[0038] Based on the above-mentioned voice features and the above-mentioned sales data features, it is determined whether the above-mentioned first live broadcast room or the above-mentioned second live broadcast room has second abnormal information.
[0039] In a feasible implementation manner, the determining whether the first live broadcast room or the second live broadcast room has the second abnormal information based on the voice feature and the sales data feature includes:
[0040] Determine recommendation strength information based on the above-mentioned emotional characteristics and speech speed characteristics;
[0041] Calculating a first matching degree based on the recommendation strength information and the time series features;
[0042] Determine the key time nodes for recommendation based on the above-mentioned tone characteristics and marketing word characteristics;
[0043] Calculate a second matching degree based on the recommended key time nodes and the peak and valley characteristics;
[0044] Determine whether second abnormal information exists in the first live broadcast room or the second live broadcast room based on the first matching degree and the second matching degree.
[0045] In a feasible implementation manner, determining whether the first live broadcast room or the second live broadcast room has second abnormal information according to the first matching degree and the second matching degree includes:
[0046] When the first matching degree is less than the first preset matching degree and the second matching degree is less than the second preset matching degree, second abnormal information exists in the first live broadcast room or the second live broadcast room.
[0047] In a feasible implementation manner, determining whether the first live broadcast room or the second live broadcast room has second abnormal information according to the first matching degree and the second matching degree includes:
[0048] When the first matching degree is greater than or equal to the first preset matching degree and the second matching degree is less than the second preset matching degree, calculating first matching degree difference information according to the recommendation strength information and the time series feature;
[0049] Obtaining a time period corresponding to peak information of the first matching degree difference information;
[0050] Get the number of people entering the live broadcast room during the time period corresponding to the peak information;
[0051] When the number of people entering the live broadcast room is greater than the preset number, second abnormal information exists in the first live broadcast room or the second live broadcast room.
[0052] In a feasible implementation manner, determining whether the first live broadcast room or the second live broadcast room has second abnormal information according to the first matching degree and the second matching degree includes:
[0053] When the first matching degree is less than the first preset matching degree and the second matching degree is greater than or equal to the second preset matching degree, determining a hysteresis relationship between the peak-valley feature and the recommended key time node according to the recommended key time node and the peak-valley feature;
[0054] When the above-mentioned key time nodes are all after the time points corresponding to the above-mentioned peak and valley characteristics, second abnormal information exists in the above-mentioned first live broadcast room or the above-mentioned second live broadcast room.
[0055] In a second aspect, the present application proposes a data analysis device, comprising:
[0056] A first acquiring unit, configured to acquire first live broadcast traffic data of a first live broadcast room and first sales data of a target product;
[0057] A second acquiring unit is configured to acquire second live broadcast traffic data of a second live broadcast room and second sales data of the target product, wherein the first live broadcast room and the second live broadcast room are live broadcast rooms selling the same product category, and a difference in live broadcast popularity between the first live broadcast room and the second live broadcast room is less than a preset threshold;
[0058] a first judging unit, configured to judge whether the live broadcast sales data of the first live broadcast room and the second live broadcast room contain first abnormal information based on the first live broadcast traffic data, the first sales data, the second live broadcast traffic data, and the second sales data;
[0059] A third acquisition unit is configured to acquire the host's voice information in the first live broadcast room or the second live broadcast room before the product is put on the shelves and sales data information for a first time period after the product is put on the shelves;
[0060] The second judgment unit is used to judge whether the first live broadcast room or the second live broadcast room has second abnormal information based on the anchor voice information before the product is put on the shelves and the sales data information in the first time period after the product is put on the shelves.
[0061] In summary, in response to the difficulty in identifying fake traffic and fake orders faced by related technologies, the embodiment of the present application effectively improves the accuracy and efficiency of anomaly detection through a two-stage data analysis method of "first horizontal comparison, then vertical in-depth exploration". The method proposed in the embodiment of the present application first selects two live broadcast rooms of the same category and with a popularity difference less than a preset threshold for comparison. Since the product types are the same and the live broadcast popularity is not much different, the differences in external environment and user groups are also smaller, thus ensuring the comparability of data comparison and reducing the error rate. When comparing traffic data and sales data, not only conventional indicators such as the number of viewers, number of interactions, and transaction volume are included, but also the positive correlation between traffic and sales is fully considered. Potential anomalies can be discovered in a timely manner when traffic is significantly high or low while sales indicators are normal (or vice versa). Compared with single-dimensional or simple statistical analysis, it is easier to capture subtle abnormal performance and improve the accuracy of judgment. After confirming or suspecting that there are no (or yet to be confirmed) overall anomalies in the livestream, the present embodiment further collects pre-launch voice messages and sales data from the livestreamer for the first period after the livestream is launched. By analyzing the correlation between marketing language, speech speed, and intonation and real-time sales peaks and conversion rates, this method can deeply uncover more subtle abnormal behaviors. This in-depth analysis method, focusing on key nodes, is more capable of identifying short-term anomalies such as fake orders and instantaneous listings and sales, compared to traditional, coarse comparisons of full-field data, thereby improving detection sensitivity. This embodiment first uses a simple macro-level comparison (first anomaly information detection) to eliminate obvious fraudulent traffic or sales. If there are no significant anomalies or in-depth verification is required, a more detailed micro-level analysis (second anomaly information detection) is then performed. This hierarchical analysis process ensures both efficiency and accuracy, and can be flexibly applied to livestream events of varying scales and types. Based on the methods of the present embodiment, livestream platforms can not only promptly detect and address abnormal livestreams, maintaining a healthy e-commerce ecosystem, but also merchants and livestreamers can use the analysis results to optimize marketing language and enhance livestream performance, thereby increasing conversion rates and user satisfaction. The embodiment of the present application effectively overcomes the blind spots and limitations of existing technologies in identifying abnormal problems in live broadcasts through two major steps: macro comparison of highly similar live broadcast rooms of the same category and detailed analysis of key time periods. It can not only accurately detect potential abnormal behaviors such as brushing volume and brushing orders, but also provide multi-level and highly efficient data support for the comprehensive supervision of live broadcast platforms and merchant operation decisions. BRIEF DESCRIPTION OF THE DRAWINGS
[0062] Various other advantages and benefits will become apparent to those skilled in the art upon reading the detailed description of the preferred embodiment below. The accompanying drawings are for illustration purposes only and are not to be considered as limiting the present description. The same reference symbols are used throughout the drawings to represent the same components. In the drawings:
[0063] Figure 1A schematic diagram of a data analysis method provided in an embodiment of the present application;
[0064] Figure 2 A schematic diagram of the structure of a data analysis device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0065] The terms "first," "second," "third," "fourth," and so on (if any) in the specification and claims of this application and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that the numbers used in this way are interchangeable where appropriate so that the embodiments described herein can be implemented in an order other than that illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, apparatus, product, or device that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or are inherent to these processes, methods, products, or devices. The technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the embodiments described are only some of the embodiments of the present application, not all of the embodiments.
[0066] Figure 1 A schematic diagram of a data analysis method provided in an embodiment of the present application is provided. The method may specifically include:
[0067] S110: Obtain first live broadcast traffic data of a first live broadcast room and first sales data of a target product;
[0068] For example, the first live broadcast traffic data includes, but is not limited to, indicators such as the number of viewers, visits, UV / PV, interactive data (likes, bullet comments, comments, etc.), and viewing time of the live broadcast room. By obtaining the first live broadcast traffic data and first sales data of the first live broadcast room, the overall operating performance of the first live broadcast room during a certain time period or event and the sales status of the target product can be understood.
[0069] S120: Obtain second live broadcast traffic data of the second live broadcast room and second sales data of the target product, wherein the first live broadcast room and the second live broadcast room are live broadcast rooms selling the same product category, and the difference in live broadcast popularity between the first live broadcast room and the second live broadcast room is less than a preset threshold;
[0070] For example, the second live broadcast traffic data is similar to that of the first live broadcast room, including but not limited to the number of viewers, visits, interaction data, viewing time, etc. The first and second live broadcast rooms must belong to the same or similar product categories (such as beauty, clothing, digital products, etc.) to ensure comparability. By calculating the live broadcast popularity (for example, based on a comprehensive score such as the number of viewers, interaction volume, and reward amount), if the popularity difference between the two live broadcast rooms does not exceed a specified threshold, they can be considered to have similar popularity. This can eliminate the noise caused by "excessive popularity differences" and make comparative analysis more accurate. Obtain similar data corresponding to the second live broadcast room to provide a reference dimension for subsequent comparisons. By selecting live broadcast rooms with similar popularity and the same category, the interference caused by differences in the external environment can be minimized.
[0071] S130: Determine whether the live broadcast sales data of the first live broadcast room and the second live broadcast room contain first abnormal information based on the first live broadcast traffic data, the first sales data, the second live broadcast traffic data, and the second sales data.
[0072] For example, "first abnormal information" refers to whether a comparison of the overall traffic data and sales data of two live streaming rooms reveals any suspicious discrepancies. For example, abnormal number of viewers (suspected of fake traffic) and / or a significant mismatch between sales data and traffic (possibly fake orders or fraudulent transactions).
[0073] Compare the number of viewers, visits, interactions, and other indicators between the two live broadcast rooms. If the difference is far beyond a reasonable range and there is no reasonable explanation, there may be an anomaly. Compare the click-through rate, order volume, transaction volume, conversion rate, etc. of the target product in the two live broadcast rooms. If the difference in indicators is abnormal and cannot be explained by objective factors (such as the popularity of the anchor or the strength of the discount), there may also be an anomaly. Under normal circumstances, the greater the traffic of the live broadcast room, the higher the sales volume or conversion rate of the target product. If a live broadcast room is found to have extremely high traffic but extremely low sales (or vice versa), be alert to anomalies. If obvious anomalies in traffic or sales are confirmed (such as fake traffic or fake orders), it will be recorded as "the first anomaly information exists." If the comparison results show that the data performance of the two live broadcast rooms is generally normal, it can be considered that "the first anomaly information does not exist."
[0074] S140: Obtain the host's voice information from the first live broadcast room or the second live broadcast room before the product is put on the shelves and the sales data information for the first time period after the product is put on the shelves;
[0075] For example, when it is confirmed or assumed that there are no anomalies in the overall dimension, the live broadcast data before and after the product is put on the shelves can be analyzed more finely to determine whether there is a second abnormal information. Collect the host's voice clips in the period "before the product is put on the shelves". Analyze the speech content speed, tone, and emotional characteristics (such as excitement, dullness, urgency, and keyword recognition such as "flash sale" and "limited time offer"). At the same time, obtain sales data information in the first time period after the product is put on the shelves. For example, you can choose 1 to 5 minutes or less after the product is put on the shelves, because the audience's purchase feedback on the newly put on the shelves can best reflect the host's ability to bring goods and marketing effects. The host's voice information before the sales product is put on the shelves and the sales data information in the first time period after the product is put on the shelves can further identify potential abnormal behavior.
[0076] It should be noted that when judging whether there is a second abnormal information, the analysis is based on the host voice information before the listing in the first live broadcast room or the second live broadcast room and the sales data information in the first time period after the listing.
[0077] S150. Determine whether there is second abnormal information in the first live broadcast room or the second live broadcast room based on the host voice information before the product is put on the shelves and the sales data information in the first time period after the product is put on the shelves.
[0078] For example, the second anomaly refers to analyzing the host's voice characteristics and real-time sales data at the critical juncture before and after a product is released, revealing a mismatch between the host's promotional claims or recommendations and actual sales, or suspicious instances of fake orders or sudden spikes and drops. For example, under normal circumstances, a strong recommendation by a host at a specific moment often leads to a surge in sales within seconds or even a dozen seconds. If the voice characteristics indicate a strong sense of urgency or aggressive promotional rhetoric, yet sales data remain unchanged, this could be an anomaly (perhaps the host is putting on a show or the traffic data is inaccurate). A sudden, short-term surge in sales followed by a rapid decline could indicate concentrated fake orders. This can be considered an anomaly if there's no reasonable explanation to go along with the host's voice or promotional information. If the host frequently emphasizes "limited-time offers," "flash sales," or "only for the first 5 minutes," yet viewers don't follow up with any purchases, this indicates a significant disconnect between the "promotional marketing" and actual sales during the livestream, potentially indicating fake traffic or abnormal sales data.
[0079] If the aforementioned time series matching analysis and peak-valley detection confirm significant inconsistencies or suspicious fluctuations, this can be recorded as a secondary anomaly. If the voice and sales trends are consistent, it can be considered "no secondary anomaly" and the livestream promotion is normal in this respect.
[0080] In summary, in response to the difficulty in identifying fake traffic and fake orders faced by related technologies, the embodiment of the present application effectively improves the accuracy and efficiency of anomaly detection through a two-stage data analysis method of "first horizontal comparison, then vertical in-depth exploration". The method proposed in the embodiment of the present application first selects two live broadcast rooms of the same category and with a popularity difference less than a preset threshold for comparison. Since the product types are the same and the live broadcast popularity is not much different, the differences in external environment and user groups are also smaller, thus ensuring the comparability of data comparison and reducing the error rate. When comparing traffic data and sales data, not only conventional indicators such as the number of viewers, number of interactions, and transaction volume are included, but also the positive correlation between traffic and sales is fully considered. Potential anomalies can be discovered in a timely manner when traffic is significantly high or low while sales indicators are normal (or vice versa). Compared with single-dimensional or simple statistical analysis, it is easier to capture subtle abnormal performance and improve the accuracy of judgment. After confirming or suspecting that there are no (or yet to be confirmed) overall anomalies in the livestream, the present embodiment further collects pre-launch voice messages and sales data from the livestreamer for the first period after the livestream is launched. By analyzing the correlation between marketing language, speech speed, and intonation and real-time sales peaks and conversion rates, this method can deeply uncover more subtle abnormal behaviors. This in-depth analysis method, focusing on key nodes, is more capable of identifying short-term anomalies such as fake orders and instantaneous listings and sales, compared to traditional, coarse comparisons of full-field data, thereby improving detection sensitivity. This embodiment first uses a simple macro-level comparison (first anomaly information detection) to eliminate obvious fraudulent traffic or sales. If there are no significant anomalies or in-depth verification is required, a more detailed micro-level analysis (second anomaly information detection) is then performed. This hierarchical analysis process ensures both efficiency and accuracy, and can be flexibly applied to livestream events of varying scales and types. Based on the methods of the present embodiment, livestream platforms can not only promptly detect and address abnormal livestreams, maintaining a healthy e-commerce ecosystem, but also merchants and livestreamers can use the analysis results to optimize marketing language and enhance livestream performance, thereby increasing conversion rates and user satisfaction. The embodiment of the present application effectively overcomes the blind spots and limitations of existing technologies in identifying abnormal problems in live broadcasts through two major steps: macro comparison of highly similar live broadcast rooms of the same category and detailed analysis of key time periods. It can not only accurately detect potential abnormal behaviors such as brushing volume and brushing orders, but also provide multi-level and highly efficient data support for the comprehensive supervision of live broadcast platforms and merchant operation decisions.
[0081] In some examples, the first live broadcast traffic data includes first user quantity data, first interaction quantity data, and first visit conversion data, and the second live broadcast traffic data includes second user quantity data, second interaction quantity data, and second visit conversion data;
[0082] The determining, based on the first live broadcast traffic data, the first sales data, the second live broadcast traffic data, and the second sales data, whether the live broadcast sales data of the first live broadcast room and the second live broadcast room contain first abnormal information includes:
[0083] Calculate the user quantity anomaly score based on the first user quantity data and the second user quantity data ;
[0084] Calculate the interaction quantity abnormality score based on the first interaction quantity data and the second interaction quantity data ;
[0085] Calculate the conversion rate anomaly score based on the first visit conversion data and the second visit conversion data ;
[0086] Calculate the sales difference anomaly score based on the first sales data and the second sales data ;
[0087] Calculate the traffic and sales correlation anomaly score based on the first user quantity data, the second user quantity data, the first sales data, and the second sales quantity data ;
[0088] The abnormal comprehensive score A is calculated according to the following formula:
[0089] + + + + in, Score the above user number anomaly The corresponding first weight information, Score the abnormal number of interactions above The corresponding second weight information, Score the above conversion rate anomaly The corresponding third weight information, Score the above sales discrepancy anomaly The corresponding fourth weight information, Score the abnormal correlation between the above traffic and sales corresponding fifth weight information;
[0090] For example, the first user quantity data ( U1 ) is the number of independent users watching the first live broadcast room. The first interactive data ( I1 ) is the interactive data in the first live broadcast room (such as likes, comments, and total shares). The first visit conversion data ( C1) is the visit conversion data of the first live broadcast room, which can include conversion rate or actual transaction volume. S1 ) is the sales amount or sales volume of the target product in the first live broadcast room.
[0091] Second user quantity data ( U2 ) is the number of independent users watching the second live broadcast room. The second interactive data ( I2 ) is the interactive data in the second live broadcast room. The second visit conversion data ( C2 ) is the visit conversion data of the second live broadcast room. The second sales data ( S2 ) is the sales amount or sales volume of the target product in the second live broadcast room.
[0092] Calculate the user quantity anomaly score based on the first user quantity data and the second user quantity data :
[0093]
[0094] is the user quantity weight, reflecting the contribution of the user quantity to the overall anomaly index. .
[0095] Calculate the interaction quantity abnormality score based on the first interaction quantity data and the second interaction quantity data , the calculation formula is:
[0096]
[0097] is the weight of the interaction data. By comparing the interaction between the two live broadcast rooms, the interaction rate (the ratio of interaction volume to the number of users) is calculated:
[0098]
[0099] The difference in engagement rate can be expressed as:
[0100]
[0101] are the interaction rates of the first and second live broadcast rooms respectively. If the preset threshold is exceeded, there may be an abnormality in the live broadcast room.
[0102] Calculate the conversion rate anomaly score based on the first visit conversion data and the second visit conversion data , and its anomaly scoring formula is:
[0103]
[0104] is the conversion rate weight. is the amplification factor, which determines the sensitivity of the abnormality score.
[0105] Calculate the visit conversion rate of the target product:
[0106]
[0107] The difference in conversion rate can be expressed as:
[0108]
[0109] These are the conversion rates for the first and second live broadcast rooms respectively.
[0110] Calculate the sales difference anomaly score based on the first sales data and the second sales data , sales data is usually affected by both traffic and conversion rate. The anomaly scoring formula is:
[0111]
[0112] is the sales data weight. is the average of the sales variances in the historical data.
[0113] Compare the sales data of the target product in the two live broadcast rooms and calculate the sales rate (the ratio of sales to the number of users):
[0114]
[0115] The sales rate difference can be expressed as:
[0116]
[0117] Calculate the traffic and sales correlation anomaly score based on the first user quantity data, the second user quantity data, the first sales data, and the second sales quantity data :
[0118]
[0119] is the Pearson correlation coefficient (or other correlation indicator) between the number of users and sales. The weight associated with traffic and sales.
[0120] Compare the sales data of the target product in the two live broadcast rooms and calculate the sales rate (the ratio of sales to the number of users):
[0121]
[0122] The sales rate difference can be expressed as:
[0123]
[0124] In some examples, the method further includes:
[0125] Obtaining the traffic sensitivity and sales sensitivity of the target product, wherein the traffic sensitivity is determined based on the target product's launch time and product category information, and the sales sensitivity is determined based on the product price and promotion intensity;
[0126] Adjusting the first weight coefficient and the second weight coefficient according to the traffic sensitivity;
[0127] The third weight coefficient and the fourth weight coefficient are adjusted according to the sales sensitivity.
[0128] For example, traffic sensitivity is the degree to which a product relies on changes in live broadcast room traffic. The higher the traffic sensitivity value, the more a small increase in traffic can significantly boost sales. The main determinants of traffic sensitivity include launch time and product category. New products are more dependent on exposure in the early stages and are generally more traffic sensitive. Mature products, on the other hand, have relatively stable reputations and repurchase rates and are relatively less sensitive to single traffic shocks. Fast-moving consumer goods, high-repurchase products, or low-unit-price products are generally more traffic sensitive; high-end, professional, or low-frequency consumer goods generally have lower traffic sensitivity.
[0129] Sales sensitivity measures a product's reliance on sales tactics like promotions and price changes. Higher sales sensitivity indicates that even small adjustments to sales tactics or discounts can significantly impact sales. High-priced items are generally more sensitive to price changes and discounts, while low-priced items, due to their lower decision threshold, are not necessarily particularly sensitive to promotional changes. If the current promotion is strong, sales will be more sensitive to changes in promotion intensity. If the discount or promotion is smaller, sensitivity is less pronounced.
[0130] Traffic sensitivity is based on the time-to-market factor of the target product and category factors Determined. Time to market factor The shorter the time from the product to the market, the higher its dependence on traffic; the longer the time from the market (mature), the lower its sensitivity to traffic. For example: Indicates the number of days since the product was launched. Indicates a preset reference number of days (such as 180 days or 365 days), then:
[0131]
[0132] like ,but It is regarded as 0 (this means that the new product period has passed and the traffic pulling effect is relatively weak).
[0133] Category Factor Indicates the degree to which a product category relies on traffic. Fast-moving consumer goods, low-unit-price products, and impulse purchases are highly sensitive; specialized products with high unit prices or long decision cycles are less sensitive. For example, assigning empirical coefficients to different categories:
[0134]
[0135] Taking these two factors into consideration, the traffic sensitivity can be defined as follows:
[0136]
[0137] To adjust the weight, you can adjust it according to actual business needs. Common values are as follows Or set according to statistical results.
[0138] Sales Sensitivity ( ) is mainly determined by commodity price factors Promotion intensity factor Joint decision.
[0139] Price Factor Reflects the sensitivity of the product to price changes. The higher the price, the more consumers pay attention to the price difference caused by discounts or promotions. For example: Select a reference price (It can be the average price of similar products). If the current product price is :
[0140]
[0141] If the product price is higher than the reference average price, It can be regarded as 1 (the price is higher than the normal range); if it is significantly lower than the reference price, then Will decrease accordingly.
[0142] Promotion intensity factor Used to measure the strength of current product promotions. The greater the promotion, the more susceptible sales are to it and the higher the sales sensitivity. For example: if the promotion strength can be quantified as a discount rate or discount strength , and there is maximum promotion ,but:
[0143]
[0144] Similarly, you can also set multiple levels of intervals, such as discounts, direct discounts, gifts and other preferential forms, and then map them to the comprehensive scores. interval.
[0145] Combining these two factors, the example defines sales sensitivity as:
[0146]
[0147] To adjust the weight.
[0148] In this example, by obtaining the target product's traffic sensitivity and sales sensitivity and adjusting the first, second, third, and fourth weight coefficients in the analysis model accordingly, data analysis and anomaly detection can adapt to the characteristics of different products, significantly improving recognition accuracy and practicality. This approach not only effectively avoids the limitations of blindly measuring with fixed weights, but also provides an operational and scalable optimization approach for the supervision of live e-commerce platforms and merchant marketing strategies.
[0149] In some examples, this also includes:
[0150] Obtain the first live broadcast room type of the first live broadcast room and the second live broadcast room type of the second live broadcast room, wherein the live broadcast room types include brand live broadcast rooms, internet celebrity personal live broadcast rooms, and internet celebrity team live broadcast rooms;
[0151] In the case that the type of the first live broadcast room and the type of the second live broadcast room are different, a live broadcast room of the same type is matched for the first live broadcast room and / or the second live broadcast room, and a judgment is made as to whether there is any abnormality in the live broadcast sales data.
[0152] For example, in the rapid development of livestreaming e-commerce, different types of livestreaming rooms have a significant impact on traffic, sales results, and marketing strategies. The main types include brand livestreaming rooms, personal livestreaming rooms of influencers, and livestreaming rooms of influencer teams.
[0153] Brand livestreams are directly operated by the brand and often have a high degree of credibility and professionalism. Their traffic may come from loyal brand fans, but it's also common for traffic to come from the brand's own channels.
[0154] Internet celebrity livestreaming rooms are typically independently run by individual streamers, relying on their personal charisma and fan loyalty to attract traffic. While these rooms may be relatively small, they often offer a distinctive livestreaming style and strong fan engagement.
[0155] Livestreaming sessions hosted by influencers (KOLs) are led by professional operations, filming, and product selection teams. These hosts often possess strong fan bases and marketing planning capabilities. They have a wide range of traffic acquisition channels, a more professional livestreaming rhythm, and generally higher conversion rates.
[0156] Because different types of live streaming rooms differ significantly in terms of user profiles, operating models, fan loyalty, brand endorsements, etc., simply comparing them to determine whether there is fake traffic or fake sales may result in misjudgment or omission. Therefore, this embodiment proposes that when two live streaming rooms are detected to be of different types, one or both of them should be matched with a room of the same type before the subsequent anomaly detection process is carried out.
[0157] In some examples, determining whether the first live broadcast room or the second live broadcast room has second abnormal information based on the host's voice information before the product is put on the shelves and the sales data information in the first time period after the product is put on the shelves includes:
[0158] Extracting voice features based on the host's voice information before the product is put on the shelves, wherein the voice features include emotional features, speech speed features, intonation features, and marketing word features;
[0159] Extracting sales data features of the first time period based on the sales data information of the first time period, wherein the sales data features include time series features and peak-valley features;
[0160] Based on the above-mentioned voice features and the above-mentioned sales data features, it is determined whether the above-mentioned first live broadcast room or the above-mentioned second live broadcast room has second abnormal information.
[0161] For example, recorded live stream audio is processed and analyzed, and emotional features are identified using voice sentiment analysis models (such as deep learning or acoustic feature extraction) to identify the live streamer's emotion type (such as excitement, indifference, or urgency). Generally, the more "excited" or "urgent" the live streamer's emotions, the more likely viewers are to make impulse purchases after the product is released.
[0162] Speech rate is determined by counting the number of words or speech clips spoken by the host per unit time (e.g., words / second, phrases / second). Extremely fast speech rates are often used for limited-time promotions or flash sales to create a sense of urgency; excessively slow or choppy speech rates can reduce viewer interest.
[0163] Intonation is determined by analyzing the range of pitch fluctuations (e.g., fundamental frequency) of the host's voice, as well as the rise and fall of pitch at key moments. Sudden increases in pitch are often used to emphasize key points, such as in sentences like "Only ten left" or "The offer only lasts one minute."
[0164] Marketing keyword features are identified through keyword recognition, using core marketing terms frequently mentioned by live streamers before a product is put on the shelves (e.g., "special offer," "limited-time offer," "flash sale," "discount," "last chance," etc.). The frequent or intense use of marketing terms often indicates that the live streamer is strongly recommending the product.
[0165] After extracting speech features, these can be represented as multidimensional vectors, such as sentiment score, speech rate, intonation fluctuation, and marketing word frequency. These can be weighted or aggregated to form a quantitative description of recommendation strength or key recommendation time points.
[0166] Sales data features are extracted based on the sales data information of the first time period. This stage focuses on the sales performance within a short time window (for example, 1 to 5 minutes, or even shorter) after the product is put on the shelves, and extracts two core features: time series features and peak and valley features.
[0167] Time series features can be used to analyze the changing curves of sales, clicks, or conversion rates over a period of time. Whether there is a steady increase or a sudden surge (or drop) at a certain point in time. Sales metrics can be recorded in discrete time series (e.g., every second or every 10 seconds) and then curve fitting or differential analysis can be performed to determine whether sales correspond to the timing of the host's voice recommendations.
[0168] Peak-valley features are used to detect sales peaks or valleys within a time window. Peaks typically indicate a surge in sales over a very short period of time; valleys indicate stagnant sales or a significant drop. The peak-valley distribution of the sales curve is confirmed by finding local maxima and minima, and the time points at which these peaks and valleys occur are recorded.
[0169] These sales data features are also mapped to feature vectors, such as: (sales peak time, sales peak size, valley time, valley size, time series curve slope)
[0170] By matching the host's voice before a product is released with its immediate sales figures, we can determine if there are more subtle or deeper anomalies in the livestream. Compared to simply comparing overall traffic or sales, this embodiment performs a "voice-sales" linkage analysis at critical moments, which can uncover more subtle, short-term anomalies, such as fake orders or fake popularity that appear and disappear in seconds.
[0171] In some examples, determining whether the first live broadcast room or the second live broadcast room contains second abnormal information based on the voice feature and the sales data feature includes:
[0172] Determine recommendation strength information based on the above-mentioned emotional characteristics and speech speed characteristics;
[0173] Calculating a first matching degree based on the recommendation strength information and the time series features;
[0174] Determine the key time nodes for recommendation based on the above-mentioned tone characteristics and marketing word characteristics;
[0175] Calculate a second matching degree based on the recommended key time nodes and the peak and valley characteristics;
[0176] Determine whether second abnormal information exists in the first live broadcast room or the second live broadcast room based on the first matching degree and the second matching degree.
[0177] Example, recommendation strength information The calculation formula can be expressed as:
[0178] A
[0179] in,
[0180]
[0181] to is the weight coefficient, E is the emotional feature, S is the speech speed feature, To perform nonlinear mapping on the emotion feature E, its value is constrained to the range [−1, 1]. ke is used to control the magnitude of the nonlinearity. If ke is large, the emotion value is further amplified or compressed. is a Sigmoid function that performs a nonlinear transformation on the speech rate feature S, limiting its value to [0, 1]. s0 is the reference point for speech rate. When S = s0, the output of the Sigmoid function is 0.5. ks controls the slope of the Sigmoid function, determining the sensitivity of the output to changes in speech rate. is the square of the sentiment feature, which is used to capture the impact of nonlinear extreme values of the sentiment feature on the recommendation strength. is the nonlinear contribution term of the emotional feature, is the nonlinear contribution term of the speech rate feature, is the interaction term between emotion and speaking speed, is the quadratic term of the emotional feature, is the quadratic term of the speaking rate feature, is the quadratic interaction term between emotion and speaking rate.
[0182] The time series feature is the first time period after the product is put on the shelf (such as For example, if sales or conversion rates are counted every 10 seconds, a discrete time series is obtained. ). Although the recommended strength It is a comprehensive indicator, but it can be expanded into a time series , represents the host's recommendation strength in different time periods or segments. Measures the overall correlation or response between recommendation strength and time-series sales trends. For example:
[0183]
[0184] Calculate the possible time delay Under , the maximum correlation coefficient between the recommendation strength series and the sales volume series. If the value is high, it means that the host’s voice recommendation does increase sales; if If the number is extremely low, it means that the host has recommended the product with all his might but there is no sales movement (or the sudden increase or decrease in sales has nothing to do with the strength of the recommendation), which may be an anomaly.
[0185] Live streamers may intentionally raise their voice pitch and tone to emphasize or incite, for example, during countdowns or during the last few available spots. These "intonation peaks" can be located by analyzing the rise in audio fundamental frequency or energy during specific time periods. The presence of strong marketing terms such as "flash sale," "limited-time offer," or "hot item countdown" in live streams often indicates a potential key recommendation node. Keyword recognition or text analysis can be used to record the time these terms appear. Combining the intonation peaks with the time nodes of the marketing terms creates a set of key recommendation time nodes, Tc. These nodes are the moments most likely to influence viewers' decisions, such as "order now" or "end in 1 minute."
[0186] The sales curve in the first period after the product is put on the shelves often has one or more local peaks (sales surge points) and valleys (sales stagnation or decline points). The peak and valley time point set P can be obtained through time series analysis (such as local maximum / minimum detection). Under normal circumstances, the sales peak should appear immediately after the recommendation key node (with a certain delay), or overlap with it within a reasonable range. If the marketing words are very inciting but do not cause any peak response, or if there are bizarre sales peaks in the absence of marketing words (possibly fake orders), these are all abnormal signs. The second matching degree (M2) can be defined as "the alignment of the recommendation key node and the sales peak." The second matching degree (M2) formula can be expressed as:
[0187]
[0188] Indicates the number of recommended key nodes. For each key node , find the most recent sales peak And calculate the time difference between them, divided by a maximum tolerance time difference (like seconds). If all key nodes can match the sales peak within a reasonable delay range, then Higher; if there is no correspondence or the gap is huge, then Significantly reduced.
[0189] In some examples, determining whether the first live broadcast room or the second live broadcast room has second abnormal information based on the first matching degree and the second matching degree includes:
[0190] A: When the above first matching degree is less than the first preset matching degree and the above second matching degree is less than the second preset matching degree, there is a second abnormal information in the above first live broadcast room or the above second live broadcast room.
[0191] B: When the above first matching degree is greater than or equal to the above first preset matching degree and the above second matching degree is less than the above second preset matching degree, calculate the first matching degree difference information according to the above recommended intensity information and the above timing characteristics;
[0192] Obtain the time period corresponding to the peak information of the first matching degree difference information;
[0193] Obtain the number of people entering the live broadcast room information for the time period corresponding to the peak information;
[0194] When the above number of people entering the live broadcast room information is greater than the preset number of people, there is a second abnormal information in the above first live broadcast room or the above second live broadcast room.
[0195] C: When the above first matching degree is less than the above first preset matching degree and the above second matching degree is greater than or equal to the above second preset matching degree, determine the lag relationship between the above peak-valley characteristics and the above recommended key time nodes according to the above recommended key time nodes and the above peak-valley characteristics;
[0196] When the above key time nodes are all after the time points corresponding to the above peak-valley characteristics, there is a second abnormal information in the above first live broadcast room or the above second live broadcast room.
[0197] Exemplarily, condition A is that both matching degrees are lower than the preset values. The first matching degree M1 is used to measure the correlation between the host voice characteristics (such as emotion, speech rate) and the timing characteristics of sales data. The second matching degree M2 is used to measure the matching between the recommended key time nodes and the sales peak-valley characteristics. When the following conditions are met, directly determine the abnormality: M1 < T1 and M2 < T2. The preset threshold T1 of the first matching degree can be set to 0.7, and the preset threshold T2 of the second matching degree can be set to 0.6. If both the first matching degree and the second matching degree are lower than the threshold, it means that there is a serious mismatch between the host's voice recommendation content and the sales data, and the linkage relationship between the recommended key time points and the sales peak-valley moments is also very poor. In this case, directly determine that there is an abnormality in the live broadcast room without further analysis.
[0198] Condition B: The first matching degree is high, but the second matching degree is low. When the following conditions are met, it is necessary to further analyze whether there is an abnormality: M1 ≥ T1 and M2 < T2. The first matching degree is relatively high, and the overall correlation between the timing characteristics of the host voice recommendation and the sales data is good. The second matching degree is relatively low, that is, the recommended key time nodes do not match the sales peak-valley. In this case, focus on checking whether there is short-term traffic abnormality (such as brushing behavior), and the specific steps are as follows:
[0199] Calculate the first matching degree difference information through recommendation intensity information (such as voice emotion, speech rate, etc.) and timing characteristics, and capture the deviation points between voice characteristics and sales timing.
[0200] The difference information formula can be:
[0201] ΔM1 = ∣M1 predicted - M1 actual∣
[0202] M1 predicted is the theoretical matching degree calculated according to the recommendation intensity information. M1 actual is the actual timing characteristic. In the first matching degree difference curve, find the peak point and the corresponding time period, and mark the moments that may have abnormal traffic. Analyze the number of people entering the live broadcast room information during the peak time period, and check whether the number of newly entered viewers during this time period significantly exceeds the normal level. If the number of entrants N > N threshold (preset threshold, such as 500 people), it is determined as abnormal traffic. If the abnormal traffic condition is met, it is determined that there is a second abnormal information in the live broadcast room.
[0203] Condition C is that the first matching degree is low, but the second matching degree is high. That is, when the following condition holds, it is necessary to further analyze whether there is an abnormality: M1 < T1 and M2 ≥ T2. At this time, the first matching degree is relatively low, and the timing correlation between the voice recommendation content and the sales data is poor. The second matching degree is relatively high, and the recommended key time nodes match well with the sales peak and valley characteristics.
[0204] In this case, focus on checking the timing consistency between the sales data and the recommendation behavior. The specific steps are as follows:
[0205] Analyze the lag relationship between the peak and valley characteristics and the recommended key time nodes: Check whether the recommended key time nodes are always lagging behind the sales peak and valley characteristics. <00005�6>
[0206] The lag relationship can be:
[0207] ΔT = Tc - Tp
[0208] Tc is the recommended key time node (such as the time point when the host mentions "time-limited flash sale"). Tp is the sales peak and valley time point (such as the moment when the sales volume reaches the peak). If ΔT > 0 (that is, the recommended nodes are all lagging behind the peak and valley nodes), it means that the sales behavior occurs before the recommendation, and it may be that the sales volume is faked in advance to forge sales data. If all the recommended key time nodes are lagging behind the sales peak and valley time points, it is directly determined that there is a second abnormal information in the live broadcast room. The sales volume peak appears before the host's recommended voice, which may indicate that the sales data is false sales volume injected in advance, rather than the actual purchase behavior of the audience.
[0209] The method provided in this embodiment adopts targeted analysis logic based on different matching combinations to ensure comprehensive detection coverage. The method provided in this embodiment further captures short-term traffic anomalies and counterfeit sales behavior through auxiliary indicators such as difference information and lag relationships. The method provided in this embodiment can adjust the matching threshold, traffic number threshold, and lag time to adapt to different live broadcast scenarios. The method provided in this embodiment simplifies the logical path and can quickly classify and process abnormal live broadcast rooms.
[0210] Second, as Figure 2 FIG. 1 is a structural diagram of a data analysis device proposed in this application, comprising:
[0211] A first acquiring unit 21 is configured to acquire first live broadcast traffic data of a first live broadcast room and first sales data of a target product;
[0212] The second acquiring unit 22 is configured to acquire second live broadcast traffic data of a second live broadcast room and second sales data of the target product, wherein the first live broadcast room and the second live broadcast room are live broadcast rooms selling the same product category, and a difference in live broadcast popularity between the first live broadcast room and the second live broadcast room is less than a preset threshold;
[0213] A first judging unit 23 is configured to judge whether the live broadcast sales data of the first live broadcast room and the second live broadcast room have first abnormal information based on the first live broadcast traffic data, the first sales data, the second live broadcast traffic data, and the second sales data;
[0214] The third acquisition unit 24 is used to obtain the host voice information of the first live broadcast room or the second live broadcast room before the product is put on the shelves and the sales data information of the first time period after the product is put on the shelves;
[0215] The second judgment unit 25 is used to judge whether the first live broadcast room or the second live broadcast room has second abnormal information based on the host voice information before the product is put on the shelves and the sales data information in the first time period after the product is put on the shelves.
[0216] The above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application.
Claims
1. A data analysis method, characterized in that: include: Obtaining the first live broadcast traffic data of the first live broadcast room and the first sales data of the target product; Obtaining second live broadcast traffic data of a second live broadcast room and second sales data of the target product, wherein the first live broadcast room and the second live broadcast room are live broadcast rooms selling the same product category, and a difference in live broadcast popularity between the first live broadcast room and the second live broadcast room is less than a preset threshold; Determining whether the live broadcast sales data of the first live broadcast room and the second live broadcast room have first abnormal information based on the first live broadcast traffic data, the first sales data, the second live broadcast traffic data, and the second sales data; Obtaining the host's voice information in the first live broadcast room or the second live broadcast room before the product is put on the shelves and the sales data information in the first time period after the product is put on the shelves; Determining whether the first live broadcast room or the second live broadcast room has second abnormal information based on the host voice information before the product is put on the shelves and the sales data information in the first time period after the product is put on the shelves; The determining whether the first live broadcast room or the second live broadcast room has second abnormal information based on the host voice information before the product is put on the shelves and the sales data information in the first time period after the product is put on the shelves includes: Extracting voice features based on the host's voice information before the product is put on the shelves, wherein the voice features include emotional features, speech speed features, intonation features, and marketing word features; Extracting sales data features of the first time period based on the sales data information of the first time period, wherein the sales data features include time series features and peak-valley features; Determine whether second abnormal information exists in the first live broadcast room or the second live broadcast room based on the voice features and the sales data features.
2. The data analysis method according to claim 1, characterized in that The first live broadcast traffic data includes first user quantity data, first interaction quantity data and first visit conversion data, and the second live broadcast traffic data includes second user quantity data, second interaction quantity data and second visit conversion data; The determining, based on the first live broadcast traffic data, the first sales data, the second live broadcast traffic data, and the second sales data, whether the live broadcast sales data of the first live broadcast room and the second live broadcast room contain first abnormal information, includes: Calculate the user quantity anomaly score based on the first user quantity data and the second user quantity data ; Calculate an abnormal interaction score based on the first interaction data and the second interaction data ; Calculating a conversion rate anomaly score based on the first visit conversion data and the second visit conversion data ; Calculating a sales discrepancy anomaly score based on the first sales data and the second sales data ; Calculate the traffic and sales correlation anomaly score based on the first user quantity data, the second user quantity data, the first sales data, and the second sales data ; The abnormal comprehensive score A is calculated according to the following formula: + + + + ; in, Score the number of users as abnormal The corresponding first weight information, Score the number of interactions as abnormal The corresponding second weight information, Score the conversion rate anomaly The corresponding third weight information, Score the sales variance anomaly The corresponding fourth weight information, Score the anomaly of the traffic and sales correlation corresponding fifth weight information; Based on the abnormal comprehensive score A and the preset threshold, it is determined whether the live broadcast sales data of the first live broadcast room and the second live broadcast room contain first abnormal information.
3. The data analysis method according to claim 2, characterized in that: The method further comprises: Obtaining the traffic sensitivity and sales sensitivity of the target product, wherein the traffic sensitivity is determined based on the launch time and product category information of the target product, and the sales sensitivity is determined based on the product price and promotion intensity; adjusting the first weight information and the second weight information according to the traffic sensitivity; The third weight information and the fourth weight information are adjusted according to the sales sensitivity.
4. The data analysis method according to claim 2 or 3, characterized in that: The method comprises: Obtain a first live broadcast room type of the first live broadcast room and a second live broadcast room type of the second live broadcast room, wherein the live broadcast room types include brand live broadcast rooms, internet celebrity personal live broadcast rooms, and internet celebrity team live broadcast rooms; When the first live broadcast room type and the second live broadcast room type are different, a live broadcast room of the same live broadcast room type is matched for the first live broadcast room and / or the second live broadcast room, and a judgment is made as to whether there is any abnormality in the live broadcast sales data.
5. The data analysis method according to claim 1, wherein: The determining, based on the voice feature and the sales data feature, whether the first live broadcast room or the second live broadcast room has second abnormal information includes: Determining recommendation strength information based on the emotional characteristics and speech speed characteristics; Calculating a first matching degree according to the recommendation strength information and the time series feature; Determining a key time node for recommendation based on the intonation feature and the marketing word feature; Calculating a second matching degree according to the recommended key time node and the peak-valley features; Determine whether second abnormal information exists in the first live broadcast room or the second live broadcast room based on the first matching degree and the second matching degree.
6. The data analysis method according to claim 5, characterized in that: The determining whether second abnormal information exists in the first live broadcast room or the second live broadcast room according to the first matching degree and the second matching degree includes: When the first matching degree is less than a first preset matching degree and the second matching degree is less than a second preset matching degree, second abnormal information exists in the first live broadcast room or the second live broadcast room.
7. The data analysis method according to claim 6, characterized in that: The determining whether second abnormal information exists in the first live broadcast room or the second live broadcast room according to the first matching degree and the second matching degree includes: When the first matching degree is greater than or equal to the first preset matching degree and the second matching degree is less than the second preset matching degree, calculating first matching degree difference information according to the recommendation strength information and the time series feature; Obtaining a time period corresponding to peak information of the first matching degree difference information; Obtain information on the number of people entering the live broadcast room during the time period corresponding to the peak information; When the number of people entering the live broadcast room is greater than the preset number, second abnormal information exists in the first live broadcast room or the second live broadcast room.
8. The data analysis method according to claim 7, characterized in that: The determining whether second abnormal information exists in the first live broadcast room or the second live broadcast room according to the first matching degree and the second matching degree includes: When the first matching degree is less than the first preset matching degree and the second matching degree is greater than or equal to the second preset matching degree, determining a hysteresis relationship between the peak-valley feature and the recommended key time node according to the recommended key time node and the peak-valley feature; When the key time nodes are all after the time points corresponding to the peak and valley features, there is second abnormal information in the first live broadcast room or the second live broadcast room.
9. A data analysis device, characterized in that: include: A first acquiring unit, configured to acquire first live broadcast traffic data of a first live broadcast room and first sales data of a target product; A second acquiring unit is configured to acquire second live broadcast traffic data of a second live broadcast room and second sales data of the target product, wherein the first live broadcast room and the second live broadcast room are live broadcast rooms selling the same product category, and a difference in live broadcast popularity between the first live broadcast room and the second live broadcast room is less than a preset threshold; a first judging unit, configured to judge whether the live broadcast sales data of the first live broadcast room and the second live broadcast room contain first abnormal information based on the first live broadcast traffic data, the first sales data, the second live broadcast traffic data, and the second sales data; A third acquisition unit is configured to acquire the host's voice information in the first live broadcast room or the second live broadcast room before the product is put on the shelves and sales data information for a first time period after the product is put on the shelves; A second judgment unit is configured to judge whether second abnormal information exists in the first live broadcast room or the second live broadcast room based on the host voice information before the product is put on the shelves and the sales data information in the first time period after the product is put on the shelves; The determining whether the first live broadcast room or the second live broadcast room has second abnormal information based on the host voice information before the product is put on the shelves and the sales data information in the first time period after the product is put on the shelves includes: Extracting voice features based on the host's voice information before the product is put on the shelves, wherein the voice features include emotional features, speech speed features, intonation features, and marketing word features; Extracting sales data features of the first time period based on the sales data information of the first time period, wherein the sales data features include time series features and peak-valley features; Determine whether second abnormal information exists in the first live broadcast room or the second live broadcast room based on the voice features and the sales data features.
Citation Information
Patent Citations
Online live shopping platform data analysis processing method, system and device, and computer storage medium
CN113191845A
Non-operational interactive live broadcast data intelligent analysis system
CN118413708A