A low-cost high-accuracy e-commerce goods-carrying video analysis method and system
Patent Information
- Application Number
- CN202610307764.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-03-13
- Publication Date
- 2026-09-25
- Estimated Expiration
- 2046-03-13
AI Technical Summary
[0006]为了解决现有技术存在的现有方案均聚焦于直播带货的数据分析或合规监管层面,未针对带货视频本身的内容创意、画面设计、营销卖点等核心维度开展精细化、多模态的深度分析,无法解决当前行业对带货视频内容本身进行精准、低成本分析的核心需求,难以为带货视频的内容优化提供直接、有效的参考依据的技术问题,本发明实施例提供了一种低成本高准确度的电商带货视频分析方法及系统
本发明实施例提供的技术方案带来的有益效果至少包括:
Smart Images

Figure CN122269107B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data processing technology, and in particular to a low-cost, high-accuracy method and system for analyzing e-commerce live-streaming sales videos. Background Technology
[0002] With the popularization of video e-commerce, improving video content quality to drive product revenue growth has become a core focus of the industry. Currently, the industry's analysis of competitor's live-streaming sales videos mainly relies on two approaches: manual review and automated analysis using large-scale visual models. Manual analysis suffers from inherent limitations such as low efficiency and significant susceptibility to subjective factors, making it difficult to adapt to the analysis needs of massive amounts of live-streaming sales videos. Therefore, introducing large-scale visual models has become a key path to improve analysis efficiency and achieve standardization. However, existing large-scale visual models still face multiple technical bottlenecks in the practical application of e-commerce live-streaming sales video analysis: lightweight, small-parameter visual models (such as Qwen3-VL-8B-Instruct and Qwen3-VL-30B-Instruct) lack sufficient analysis accuracy, easily missing key details and core content in the video, and failing to accurately reproduce the design logic and selling points of the live-streaming sales video; while high-precision, large-parameter open-source visual models (such as Qwen3-VL-235B-Instruct) offer relatively better analysis quality, they suffer from extremely high deployment and usage costs. Stable operation requires 8 or more A100 or H100 graphics cards, and the output results are not stable when processing long and complex live-streaming e-commerce videos, and problems such as missing details and deviations in core descriptions are still likely to occur. The video analysis functions of mainstream closed-source large-parameter visual models in China are not yet mature and cannot meet the needs of refined analysis of live-streaming e-commerce videos. Meanwhile, closed-source models with video analysis capabilities abroad (such as Gemini) cannot be used stably in China due to geographical limitations, and these models also have problems with incomplete detail extraction and insufficient accuracy in the analysis of long and complex videos.
[0003] Video analytics, inherently complex due to its massive amounts of information and its involvement in both physical world content reconstruction and temporal logic analysis, is a technically challenging artificial intelligence task. E-commerce live-streaming videos, as a vertical video genre, not only face the inherent complexities of general video analytics but also demand higher precision and completeness in extracting core details such as video content creativity, character performance, scene setup, marketing rhetoric, and product selling points. This further increases the technical difficulty of achieving stable and accurate analysis. Currently, existing solutions for analyzing live-streaming e-commerce videos mainly fall into two categories: The first involves deploying large-parameter open-source visual models or calling closed-source visual models, guiding the model through refined prompts and multi-turn dialogue interactions, and summarizing complete analytical conclusions based on the results returned from multiple rounds. While this approach can improve the depth of analysis to some extent, it is limited by the model's context length, making it difficult to extract all the video content in detail. Furthermore, multi-turn interactions consume significant computing resources, further increasing the cost. The second approach involves extracting only the audio stream from the video and transcribing it into text, then having a large language model analyze it to obtain the general content framework of the video. Although this approach is lower in cost, the lack of visual dimension information results in a coarser analysis effect, only achieving basic content summaries and failing to meet the analytical needs of live-streaming e-commerce videos for details such as visual design and scene presentation.
[0004] Existing technologies also include some solutions for analyzing live-streaming e-commerce videos. For example, the invention patent announcement CN114897585B discloses a method and system for improving e-commerce efficiency based on multi-dimensional video analysis, which includes: acquiring the live-streaming room data; dividing the e-commerce video data into different e-commerce video segments and obtaining corresponding comment data for each segment; extracting relationship data between the e-commerce video segments and historical live-streaming room data from the corresponding comment data, and generating candidate e-commerce video segments; calculating the importance score of each candidate e-commerce video segment and synthesizing feature e-commerce video segments; and finally synthesizing the feature e-commerce video segments into the current feature live-streaming room data based on the generation time of each e-commerce video segment.
[0005] For example, the Chinese invention patent announcement CN115086721B discloses a data analysis-based ultra-high-definition live streaming system service supervision system, which includes: adopting an online intelligent supervision method when supervising the illegal sale of ultra-high-definition live streaming e-commerce videos; extracting live streaming sales information corresponding to the currently sold products from the live streaming ultra-high-definition live streaming e-commerce videos; evaluating the live streaming sales compliance index corresponding to the currently sold products; and having the live streaming platform perform targeted processing on the live streaming process of designated ultra-high-definition live streaming rooms to achieve effective supervision of live streaming e-commerce videos. Summary of the Invention
[0006] To address the shortcomings of existing technologies, which primarily focus on data analysis or compliance regulations for live-streaming e-commerce, neglecting in-depth, multimodal analysis of core dimensions such as content creativity, visual design, and marketing selling points, this invention provides a low-cost, high-accuracy method and system for analyzing e-commerce live-streaming videos. The technical solution is as follows: On the one hand, a low-cost and high-accuracy e-commerce live-streaming video analysis method is provided, including: performing multi-dimensional usability checks on the live-streaming video files to be analyzed; if the checks pass, the live-streaming video files to be analyzed are demultiplexed into video audio streams and video video streams; if the checks fail, the analysis process is terminated; performing streaming speech transcription on the video audio streams to generate text semantic sequences; performing real-time dynamic analysis on the video video streams to extract motion intensity coefficients; dynamically dividing the live-streaming video files to be analyzed into video segments based on the text semantic sequences and motion intensity coefficients, generating each semantic segment and its corresponding image frame set; and performing cross-segment summary aggregation based on the audio text corresponding to each semantic segment to generate an overall video summary; using the image frame sets, audio text, and overall video summary corresponding to each semantic segment as input, and combining them with a preset multi-dimensional analysis prompt vocabulary for parallel reasoning analysis to generate text analysis results under each preset analysis dimension; and structuring the analysis results text of each analysis dimension, and integrating and connecting them according to time order to generate a complete structured video description document.
[0007] On the other hand, a low-cost, high-accuracy e-commerce live-streaming video analysis system is provided. This system, applied to a low-cost, high-accuracy e-commerce live-streaming video analysis method, includes: a usability pre-detection and diversion module, an audio-visual segmentation processing module, a parallel inference analysis module, and an analysis result integration module. The usability pre-detection and diversion module performs multi-dimensional usability checks on the live-streaming video files to be analyzed. If the check passes, the video files are demultiplexed into video-audio streams and video-visual streams; if the check fails, the analysis process terminates. The audio-visual segmentation processing module performs streaming speech transcription on the video-audio streams to generate text semantic sequences and performs real-time dynamic analysis on the video-visual streams to extract motion. The intensity coefficient dynamically divides the video segments of the product promotion video file to be analyzed based on the text semantic sequence and the image motion intensity coefficient, generating each semantic segment and its corresponding image frame set. It also performs cross-segment summary aggregation based on the audio text corresponding to each semantic segment to generate an overall video summary. The parallel reasoning analysis module takes the image frame set, audio text, and overall video summary corresponding to each semantic segment as input, and performs parallel reasoning analysis in combination with a preset multi-dimensional analysis prompt vocabulary library to generate text analysis results under each preset analysis dimension. The analysis result integration module is used to structure the analysis result text of each analysis dimension, and integrate and connect them according to the time order to generate a complete structured description document of the video.
[0008] Beneficial effects The beneficial effects of the technical solutions provided in the embodiments of the present invention include at least the following: 1. This invention achieves intelligent and efficient e-commerce live-streaming video analysis through multi-dimensional usability compliance checks, dynamic video segmentation combining image motion intensity and text semantic features, adaptive frame extraction correction based on computational load, parallel inference analysis based on multi-modal input and multi-dimensional prompts, and confidence-driven dimension adjustment. This significantly reduces costs, improves analysis accuracy and detail completeness, and outputs structured description documents, providing detailed data support for subsequent applications.
[0009] 2. This invention matches the initial frame extraction frequency based on the video duration, performs a first-level adaptive correction based on the intensity of screen motion, and then introduces a second-level correction based on the computational load. It also accurately defines key frames through weighted scoring of motion and semantic features to differentiate frame extraction. At the same time, it optimizes the segmentation boundaries based on text semantic similarity and screen motion features, thereby realizing dynamic, intelligent, and computationally adapted video segmentation, improving resource utilization efficiency and segmentation accuracy.
[0010] 3. This invention calculates the semantic mutation coefficient of adjacent segments and dynamically adjusts the duration threshold by predefined semantic similarity and duration constraint thresholds to complete the segmentation and merging of segments. Then, compliance verification and iterative correction of the boundaries are performed to achieve fine adaptive adjustment of video semantic segmentation, so that the segmentation is highly consistent with the semantic features of the text, effectively avoiding the problem of being too fine or too coarse, and improving the detail and accuracy of subsequent analysis.
[0011] 4. This invention combines multimodal input data with a multidimensional analysis prompt lexicon for parallel inference, constructs a multi-index confidence evaluation system, dynamically adjusts the analysis dimensions (expansion, pruning, or minimum guarantee) based on the confidence score, and finally generates structured analysis results. This achieves intelligent adaptation of analysis dimensions to video content, improves the reliability and effectiveness of analysis results, and rationally allocates computing resources, balancing analysis depth and efficiency. Attached Figure Description
[0012] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0013] Figure 1 A flowchart illustrating the low-cost, high-accuracy e-commerce live-streaming video analysis method provided in this application embodiment; Figure 2 A flowchart illustrating the video segmentation analysis process of the low-cost, high-accuracy e-commerce live-streaming video analysis method provided in this application embodiment; Figure 3 This is a schematic diagram of the structure of the low-cost, high-accuracy e-commerce live-streaming video analysis system provided in the embodiments of this application. Detailed Implementation
[0014] Embodiments of the present disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of the present disclosure are shown in the drawings, it should be understood that embodiments of the present disclosure may be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of the present disclosure.
[0015] It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure. In the description of the embodiments of this disclosure, the term "comprising" and similar terms should be understood as open-ended inclusion, i.e., "including but not limited to". The term "based on" should be understood as "at least partially based on". The term "one embodiment" or "this embodiment" should be understood as "at least one embodiment". The terms "first", "second", etc., may refer to different or the same objects.
[0016] To make the technical problems, technical solutions and advantages of the present invention clearer, a detailed description will be given below in conjunction with the accompanying drawings and specific embodiments.
[0017] like Figure 1The diagram shows a flowchart of a low-cost, high-accuracy e-commerce live-streaming video analysis method provided in this application embodiment. The method includes the following steps: performing a multi-dimensional usability check on the live-streaming video file to be analyzed. If the check passes, the video file is demultiplexed into video audio streams and video video streams. If the check fails, the analysis process is terminated and a specific error reason code is returned to indicate a defect in the video file. The multi-dimensional usability check includes, but is not limited to, audio track existence checks, video frame rate compliance checks, video resolution compliance checks, and video bitrate compliance checks. A transcription tool, such as Whisper, is used to perform streaming speech transcription on the video audio stream to generate a text semantic sequence. Real-time dynamic analysis is performed on the video video stream to extract motion intensity coefficients. Based on the semantic integrity boundaries of the text semantic sequence and the scene switching points reflected by the motion intensity coefficients, the video file to be analyzed is dynamically divided into video segments. Multiple semantic segments with inherent consistency in both semantics and visuals, along with their corresponding image frame sets, are generated. Cross-segment summary aggregation is performed based on the audio text corresponding to each semantic segment to generate a video summary reflecting the overall theme of the video. The image frame sets corresponding to each semantic segment, ... The audio text and overall video summary serve as input, combined with a pre-set multi-dimensional analysis prompt word library for parallel reasoning analysis. This generates text analysis results for each pre-set analysis dimension. Core dimensions include, but are not limited to, environmental analysis, role and event analysis, visual design analysis, and marketing appeal analysis. The prompt words for the environmental analysis dimension guide the model to analyze the time, location, and set design factors of the video. The prompt words for the role and event analysis dimension guide the model to analyze the participants, role details, behaviors, and event development in the video. The prompt words for the visual design analysis dimension guide the model to analyze camera work, subtitles, animation effects, and post-production techniques. The prompt words for the marketing appeal analysis dimension guide the model to analyze the product-focused design and user attraction strategies. Each analysis dimension's prompt words contain semantic content guiding the model to reason about "why the video segments are designed this way," enhancing the depth and quality of the analysis. The analysis results text for each dimension are then structured, converted into a unified JSON format, and integrated and linked according to chronological order to generate a final, complete structured video description document containing timeline information.
[0018] In this embodiment, the present invention performs multi-dimensional availability checks on the audio track existence, frame rate, resolution, and bitrate of the product-selling video to ensure the effectiveness of the analysis material from the source. If the check fails, an error reason code is returned to terminate the process; if it passes, the video is demultiplexed into audio and video streams, laying the foundation for subsequent accurate analysis. The audio stream is transcribed into text semantic sequences using streaming speech transcription, and the video stream is dynamically analyzed in real time to extract motion intensity coefficients. Video segments are dynamically divided based on semantic integrity boundaries and scene transition points, generating semantic segments and corresponding image frame sets with inherent consistency in both semantics and visuals. Furthermore, cross-segment summaries of the audio text from each segment can be aggregated to form an overall video summary. This achieves scientific video segmentation while allowing the model to grasp the overall theme of the video, avoiding the one-sidedness of segmented analysis. Using the image frame sets, audio text, and overall summary as input, and combined with a pre-set multi-dimensional analysis prompt lexicon, parallel reasoning analysis is performed around four core elements: environment, roles and events, visual design, and marketing appeal. The analysis unfolds across multiple dimensions, with each dimension's prompts integrated into the semantic content guiding the model's reasoning regarding "the reasons for video segmentation design." This not only enhances analysis efficiency through parallel reasoning but also, through targeted prompt design and deep reasoning guidance, makes the analysis more aligned with the business needs of e-commerce live-streaming videos, significantly improving the depth, accuracy, and completeness of the analysis. Finally, the analysis results from each dimension are structured and converted into a unified JSON format, and then integrated and linked chronologically to generate a complete structured video description document with timeline information. This achieves standardized and structured presentation of the analysis results, preserving the detailed analysis content of each video segment while forming a complete video analysis framework through the timeline. This makes the analysis results easier to read, store, and subsequently re-analyze. At the same time, the volume of the structured data is much smaller than the original video, significantly reducing the cost and difficulty of subsequent storage and analysis. Overall, this achieves low-cost, high-accuracy, and high-stability full-dimensional analysis of e-commerce live-streaming videos, with significantly improved practicality and scalability of the analysis results.
[0019] Furthermore, the specific steps of the multi-dimensional usability check include: the multi-dimensional usability check rules include a first priority rule and a second priority rule; the first priority rule includes: detecting whether there is an audio stream in the video file to be analyzed. If not, the check is directly marked as failed; if there is, the first priority rule is marked as passed, and the second priority rule is executed sequentially: detecting whether the frame rate of the video file to be analyzed is higher than a preset frame rate threshold (e.g., 15fps), whether the resolution is higher than a preset resolution threshold (e.g., standard definition 480p), and whether the bitrate is higher than a preset bitrate threshold (e.g., 1000kbps). If yes, the corresponding indicator is marked as met; otherwise, the corresponding indicator is marked as not met; if any indicator of the above second priority rule is not met, the check is marked as failed, and the error code of the corresponding non-compliant item is recorded. All non-compliant items are summarized to generate a structured error reason code containing specific defect information; if all indicators of the above second priority rule are met, the check is marked as passed.
[0020] In this embodiment, the present invention establishes a logically clear and rigorously judged hierarchical inspection rule with first and second priorities, forming an admission mechanism for live-streaming e-commerce video materials, which has significant technical effects: First, the existence of audio stream is used as the first priority rule, directly judging videos without audio as failing the inspection and terminating the process, thus filtering out invalid materials that cannot meet the audio analysis requirements of live-streaming e-commerce videos from the core dimension, avoiding meaningless consumption of subsequent analysis resources; provided that the first priority rule passes, the second priority rule is executed in sequence to perform compliance checks on the frame rate, resolution, and bitrate of the video frame, accurately verifying the basic quality indicators of the video frame, and marking any failure of any indicator as a failure, and recording the error code for each failure separately. The code is ultimately compiled into a structured error cause code containing specific defect information, allowing users to accurately and intuitively understand the specific quality defects of the video file, providing clear guidance for the optimization or replacement of video materials. Only when the frame rate, resolution, and bit rate indicators of the second priority rule all meet the standards is the check marked as passed, ensuring that the e-commerce videos entering the subsequent analysis process have both a complete audio analysis foundation and picture quality that meets the requirements of image extraction and detailed analysis. This guarantees the effectiveness and accuracy of the materials in all subsequent analysis stages such as video segmentation, visual analysis, and audio transcription from the source, effectively avoiding problems such as analysis result deviations and omissions of details caused by material quality defects, and laying a solid material foundation for the high-quality advancement of the entire e-commerce e-commerce video analysis process.
[0021] Furthermore, the steps for dynamically dividing the video segments of the product promotion video file to be analyzed include: based on the total duration of the video segments of the product promotion video file to be analyzed, querying and obtaining the initial frame extraction frequency corresponding to the total duration of the video segment from a preset duration-frame extraction frequency mapping table. The duration-frame extraction frequency mapping table is used to determine the predefined reference relationship of the basic frame extraction frequency based on the total duration of the video. When the video duration is longer, the initial frame extraction frequency is reduced accordingly to control the total number of frames to balance the analysis granularity and computational overhead. For example, short videos (within 2 minutes): 3-5 images per second; medium and long videos (more than 2 minutes): 2-3 images per second. More images per second result in more detailed analysis, but may lead to some useless analysis in long videos; fewer images per second result in faster analysis, but may miss some details in short videos. The pixel differences between adjacent time frames in the video stream are calculated, and the results are used as the motion intensity coefficient. A pre-defined intensity-adjustment ratio mapping table is used to obtain the frame skipping frequency adjustment ratio corresponding to each time frame. This mapping table dynamically adjusts the frame skipping frequency according to a predefined relationship based on the motion intensity coefficient; the higher the motion intensity, the higher the frame skipping frequency adjustment ratio is to capture more detailed frames. The initial frame skipping frequency is then adaptively corrected using the frame skipping frequency adjustment ratio to obtain the dynamic frame skipping frequency for each time frame. The video stream is then skipped based on this dynamic frame skipping frequency. The process involves frame processing to generate image sets corresponding to each time frame; calculating the semantic similarity between adjacent time frames in the text semantic sequence; performing initial segmentation of the text semantic sequence based on semantic similarity; marking adjacent time frames as boundary points of different semantic segments when the semantic similarity between any two adjacent time frames is lower than a preset semantic boundary threshold; and marking adjacent time frames as belonging to the same semantic segment based on the marking results. The process also involves merging consecutive time frame intervals marked as belonging to the same semantic segment to generate initial semantic segments; and finally, boundary correction of the initial semantic segments based on the image motion intensity coefficient and semantic similarity to generate the final semantic segments and their corresponding image frame sets.
[0022] In this embodiment, the present invention achieves accurate and efficient segmentation of product-selling video clips in both semantic and visual dimensions through a multi-level adaptive design involving duration adaptation, dynamic image adjustment, semantic segmentation, and dual-dimensional boundary correction. This results in significant and diverse technical advantages: First, based on the total video duration, an initial frame extraction frequency is matched from a preset duration-frame extraction frequency mapping table, allowing the frame extraction density to dynamically adapt to the video duration. The longer the duration, the lower the frame extraction frequency. This avoids the surge in computational overhead and redundant analysis caused by excessive frame extraction in long videos, while also ensuring the precision of frame extraction in short videos, achieving both high granularity and high precision in analysis. The optimal balance of computational cost is achieved; secondly, the motion intensity coefficient of the image is obtained by calculating the pixel difference between adjacent frames, and the initial frame extraction frequency is adaptively corrected by combining the intensity-adjustment ratio mapping table, so that the frame extraction frequency of the image area with higher motion intensity is higher, which can accurately capture key detail frames such as image action and scene transition, effectively avoiding the omission of details caused by rapid dynamic changes in the image, and ensuring the complete restoration of the actual video content by the image frame set; at the same time, the semantic similarity of adjacent frames of the text semantic sequence is calculated, and the initial semantic segmentation is completed by using the preset semantic boundary threshold as the standard, and the semantic similarity is classified into... Positions with a degree below the threshold are marked as semantic boundary points. Consecutive frames with the same semantic meaning are merged to generate initial semantic segments, achieving natural and complete segmentation based on audio semantics. This ensures the inherent semantic consistency of each initial segment, aligning with the semantic logic of the sales pitch and selling points in the product promotion video. Finally, based on the scene switching features reflected by the motion intensity coefficient of the image and the semantic boundary features reflected by the audio semantic similarity, the initial semantic segments undergo dual-dimensional boundary correction. This ensures that the final generated semantic segments and their corresponding image frame sets not only conform to the semantic integrity boundaries at the audio level but also match the scene switching patterns at the image level, achieving dual inherent consistency between semantics and image. This fundamentally solves the problem of image and semantic disconnect caused by single-dimensional segmentation. It provides a clear-boundary, content-matching, and detailed analysis unit for subsequent parallel reasoning analysis combining image frame sets and audio text. This ensures both the accuracy and depth of subsequent segmentation analysis and controls the overall computational cost through adaptive adjustment of the frame extraction frequency. This allows the video segmentation process to combine analytical accuracy, content completeness, and execution efficiency, laying a crucial segmentation foundation for the high-quality advancement of the entire e-commerce product promotion video analysis process.
[0023] Furthermore, the determination of the dynamic frame extraction frequency also includes a secondary correction mechanism based on computing power load. Specific steps include: obtaining the current computing power utilization rate of the device; determining the current computing power load level by matching the computing power utilization rate of the device with the corresponding computing power load level range. The computing power load level is a priority hierarchy based on the current computing power utilization rate of the device, used to dynamically adjust the frame extraction strategy to balance analysis accuracy and system resource consumption. This includes low load, medium load, and high load. The low load level refers to the current device computing power utilization rate being below a preset light load threshold (e.g., 30%), indicating sufficient system resources. In this case, the frame extraction frequency after the first-level correction is maintained, and regular frame extraction is performed on all time frames to ensure maximum analysis granularity. For example, when processing video during server idle periods, with a CPU utilization rate of only 15%, the system uses high-frequency frame extraction for fast-moving segments and maintains basic frame extraction for static segments to fully preserve video details. The medium load level refers to the current device computing power utilization rate being in a preset medium range (e.g., 30%–70%), where system resources are competed for. At this point, a selective frequency reduction strategy is activated: only the sampling frequency of non-critical time frames is reduced, while critical time frames (such as sudden changes in the image or semantic turning points) maintain their original frequency, prioritizing the sampling density of important content under resource constraints. For example, when multiple tasks are running simultaneously on a local workstation and the computing power utilization rate reaches 55%, the system automatically lengthens the sampling interval of background images, but maintains a high sampling frequency for critical frames such as product close-ups and promotional script bursts. High load level refers to the current device computing power utilization rate exceeding the preset overload threshold (such as 70%), and system resources are approaching a bottleneck. At this point, an extreme minimum guarantee strategy is activated: only the sampling of critical time frames is retained, and non-critical frames are completely discarded, maintaining basic analysis capabilities with minimal resource overhead. For example, when processing live streams in real time on edge computing devices, if the CPU utilization rate has reached 85%, the system only extracts key transition frames such as when the host starts explaining, when the product is linked, and when the price is announced, discarding intermediate transition frames to ensure that the analysis process is not interrupted. If the computing load level is low, the frame extraction frequency after the first-level correction is maintained. If the computing load level is medium, the frame extraction frequency of non-critical time frames is reduced, while the frame extraction frequency of critical time frames is maintained. If the computing load level is high, only critical time frames are extracted, and non-critical time frames are ignored.
[0024] The steps for determining key and non-key time frames include: for each time frame, obtaining its corresponding motion intensity coefficient and semantic abruptness coefficient; normalizing the motion intensity coefficient and semantic abruptness coefficient of each time frame to obtain normalized motion intensity coefficient and normalized semantic abruptness coefficient of each time frame; performing a weighted summation operation on the normalized motion intensity coefficient and normalized semantic abruptness coefficient according to preset motion weight coefficient and semantic weight coefficient to obtain the importance score of each time frame; if the importance score of any time frame reaches the preset key frame threshold, then the time frame is marked as a key time frame; if the importance score of any time frame does not reach the preset key frame threshold, then the time frame is marked as a non-key time frame.
[0025] In this embodiment, the present invention adds a secondary correction mechanism based on computing power load on the basis of the first-level adaptive correction. Combined with the scientific judgment rules of critical and non-critical time frames, it forms an intelligent frame extraction strategy that is computing power adapted, prioritizes key components, and dynamically balances accuracy and resources, achieving multi-dimensional technical optimization effects: First, by obtaining the real-time computing power utilization rate of the device and matching the corresponding computing power load level, the computing power status is divided into three levels: low, medium, and high, and a differentiated frame extraction strategy is formulated. This allows the frame extraction frequency to be deeply adapted to the system resource status. Under low load, the frame extraction frequency after the first-level correction is maintained, making full use of ample resources to maximize the analysis granularity, completely preserving video images and semantic details, and ensuring the analysis. The system ensures high accuracy and completeness; under medium load, a selective frequency reduction strategy is activated, reducing the sampling frequency of non-critical time frames while maintaining the original frequency of critical frames. In the event of resource contention, priority is given to ensuring the sampling density of core content, effectively controlling system resource consumption and avoiding the omission of key information; under high load, an extreme minimum guarantee strategy is activated, retaining only the sampling of critical time frames and discarding non-critical frames to maintain basic analysis capabilities with minimal resource overhead. This ensures that the analysis process is not interrupted when system resources are on the verge of bottlenecks, significantly improving the system robustness and environmental adaptability of the frame extraction stage and even the entire video analysis process. It can flexibly cope with different computing power scenarios such as server idleness, multi-task parallelism, and real-time processing of edge devices. Meanwhile, the mechanism normalizes the motion intensity coefficient and semantic change coefficient of each time frame, and obtains an importance score by weighting and summing them with preset weight coefficients. It scientifically determines key and non-key time frames based on the key frame threshold, so that the identification of key frames takes into account both the dynamic features of the screen and the semantic features of the audio. It accurately locks the core analysis nodes of the sales video, such as screen change points, semantic turning points, product display close-ups, and promotional speech bursts. This gives the execution of the differentiated frame extraction strategy a clear and objective basis for judgment, avoiding the omission of key frames or invalid sampling of non-key frames caused by subjective division. Overall, this secondary correction mechanism allows the dynamic frame extraction frequency to not only adapt to the visual and semantic features of the video itself, but also to respond in real time to changes in device computing power load. While ensuring the accuracy of core content analysis in e-commerce videos, it achieves a dynamic and intelligent balance between analysis granularity and system resource consumption. This solves the problem of either wasting resources or having insufficient analysis accuracy when the frame extraction frequency is fixed under different computing power scenarios, and also avoids the loss of key information caused by indiscriminate frequency reduction. The resulting image frame set not only has the ability to completely restore the core content of the video, but also adapts to different computing power operating environments. This provides a foundation of high-quality, resource-adaptable, and detailed visual materials for subsequent video segmentation and parallel inference analysis, while further optimizing the computational efficiency and resource utilization of the entire analysis process.
[0026] Further, the steps for boundary correction of the initial semantic segmentation include: S61: Predefine the semantic similarity benchmark threshold and the segmentation duration constraint threshold, the segmentation duration constraint threshold includes the minimum segmentation duration threshold and the maximum segmentation duration threshold; S62: Calculate the vector similarity of the text semantics of each adjacent initial semantic segment to obtain the semantic mutation coefficient of each adjacent initial semantic segment; S63: If the semantic mutation coefficient of any adjacent initial semantic segment is greater than the preset mutation upper limit threshold, the difference between the semantic mutation coefficient and the preset mutation upper limit threshold is recorded as the segmentation urgency. According to the segmentation urgency, the corresponding minimum segmentation duration reduction coefficient is retrieved from the preset urgency-adjustment coefficient mapping table. Based on the minimum segmentation duration reduction coefficient, the minimum segmentation duration threshold is dynamically reduced, and the adjacent initial semantic segment is segmented from the current boundary point to generate two independent semantic segments; S64: If the semantic mutation of any adjacent initial semantic segment is greater than the preset mutation upper limit threshold threshold, the minimum segmentation duration constraint threshold is reduced, and the adjacent initial semantic segment is segmented from the current boundary point to generate two independent semantic segments; If the coefficient is within the range of the preset smoothing lower threshold and the preset mutation upper threshold, no additional processing is performed; S65: If the semantic mutation coefficient of any adjacent initial semantic segment is less than the preset smoothing lower threshold, the difference between the preset smoothing lower threshold and the semantic mutation coefficient is recorded as the segment continuity. Based on the segment continuity, the corresponding segment maximum duration increase coefficient is obtained from the preset continuity-adjustment coefficient mapping table. The segment maximum duration threshold is dynamically increased based on the segment maximum duration increase coefficient, and the adjacent initial semantic segments are merged into one semantic segment; S66: Based on the duration constraint range formed by the reduced segment minimum duration threshold and the increased segment maximum duration threshold, boundary compliance verification processing is performed on each semantic segment; S67: Steps S61 to S66 are repeated until all adjacent initial semantic segments are traversed and no new segmentation or merging operations occur. The final semantic segment is generated based on the finally adjusted segment boundary.
[0027] The steps for performing boundary compliance verification on each semantic segment include: if the duration of any semantic segment is lower than the reduced minimum duration threshold, then the semantic segment is merged with the semantic segment with the smallest semantic mutation coefficient among the adjacent semantic segments, and the semantic mutation coefficient of the merged adjacent segments is recalculated; if the duration of any semantic segment is between the reduced minimum duration threshold and the increased maximum duration threshold, then the semantic segment remains unchanged; if the duration of any semantic segment is higher than the increased maximum duration threshold, then the boundary point with the largest semantic mutation coefficient is found within the semantic segment, and the semantic segment is split into two independent semantic segments based on the boundary point, and the semantic mutation coefficient of the split adjacent segments is recalculated.
[0028] In this embodiment, the present invention constructs an initial semantic segmentation boundary correction step, combined with a dedicated boundary compliance verification process, to build a fully closed-loop semantic segmentation optimization mechanism that includes quantitative judgment, dynamic adjustment, compliance verification, and iterative optimization. This achieves refined, intelligent, and standardized correction of the semantic segmentation boundaries of live-streaming e-commerce videos, with significant technical effects and a layered, interconnected approach: First, by predefining a semantic similarity benchmark threshold and a segmentation duration constraint threshold containing minimum and maximum durations, a basic judgment framework is established for segmentation correction, avoiding the analysis failure problem caused by overly fragmented or long segments from the source; then, by calculating vector similarity, the semantic correlation between adjacent initial semantic segments is transformed into a quantifiable semantic mutation coefficient, freeing the judgment of semantic boundaries from subjectivity and providing objective and accurate numerical basis, thus providing scientific support for subsequent differentiated correction. Simultaneously, precise dynamic correction strategies are formulated for different ranges of semantic mutation coefficients. When the mutation coefficient exceeds the upper limit, the minimum duration threshold is dynamically reduced by matching the segmentation urgency coefficient and segmentation is performed. This accurately identifies key nodes of drastic semantic mutations, and effective segmentation can be achieved even if the original segment is close to the minimum duration, ensuring the segmentation independence of content with strong semantic differences. When the mutation coefficient is in the smooth range, no processing is performed, which fits the actual scenario of natural semantic transition, reduces invalid calculations, and improves correction efficiency. When the mutation coefficient is below the lower limit, the maximum duration threshold is dynamically increased by matching the paragraph continuity coefficient and merging is performed. This accurately identifies content with highly coherent semantics, breaks through the original duration limit to achieve reasonable merging, avoids the incorrect segmentation of content with weak semantic differences, and makes the segmentation highly consistent with the natural semantic logic of the sales video's explanation of sales pitches and introduction of selling points. The accompanying boundary compliance verification process further standardizes the dynamically corrected semantic segments, ensuring compliance in both duration and semantics. For segments shorter than the minimum corrected duration, they are merged with the adjacent segment with the smallest semantic mutation coefficient, and the coefficient is recalculated to avoid excessively short fragmented segments and ensure the segment's analytical value. Segments with durations within the compliance range remain unchanged, maintaining the reasonable corrected boundaries. For segments with durations exceeding the maximum corrected duration, the boundary point with the largest semantic mutation coefficient is found within them, and the segments are split and the coefficient is recalculated to avoid excessively long redundant segments and ensure that the segment granularity adapts to subsequent analysis needs. Finally, by cyclically executing the correction and verification steps until no new segmentation or merging operations are performed, a closed-loop optimization without blind spots is formed. This ensures that all final generated semantic segments strictly conform to the dynamically adjusted duration constraint range and that each segment boundary highly matches the semantic features of the video, completely resolving core issues that may exist in the initial semantic segmentation, such as boundary deviations, duration non-compliance, and semantic connection and segment boundary disconnection.Overall, this technical solution ensures that the generated semantic segments not only closely adhere to the natural semantic logic of the e-commerce video, achieving strong inherent consistency at the semantic level, but also balances the rationality of segment granularity, the standardization of duration, and the completeness of content through dynamic duration threshold adjustment and standardized compliance verification. This avoids fragmented and redundant segmentation issues, providing high-quality analysis units with clear boundaries, semantic coherence, duration compliance, and strong analytical capabilities for subsequent parallel reasoning analysis that combines semantic segments with corresponding image frame sets. This fundamentally improves the accuracy, effectiveness, and efficiency of subsequent segmentation analysis, further solidifying the core foundation of the entire e-commerce video analysis process. It allows subsequent multi-dimensional analyses such as environment, roles, visual design, and marketing appeal to be conducted based on high-quality segmentation units, ensuring the accuracy and depth of the overall analysis results.
[0029] Figure 2 The flowchart of the video segmentation analysis method for low-cost and high-accuracy e-commerce live-streaming video analysis provided in this application embodiment includes the following steps: combining the image frame sets, audio text, and overall summary text of each semantic segment as multimodal input data; extracting key image frames for each semantic segment (e.g., uniform sampling or keyframe detection) and converting them into feature vectors through a visual encoder; simultaneously, extracting semantic features from the corresponding audio text (e.g., speech recognition results) and the overall video summary text through a text encoder. Then, integrating these features into a unified representation through feature concatenation, cross-attention mechanisms, or multimodal fusion layers (e.g., cross-modal interaction of Transformer) to ensure that images, audio text, and global context information can mutually enhance each other. Finally, inputting the fused multimodal features into a large model enables the model to perform collaborative reasoning based on local images, local audio, and global context, thereby outputting more accurate analysis results.
[0030] By employing concurrent invocation mechanisms (such as asynchronous requests or multithreading), multimodal input data and corresponding prompt words from a multidimensional analysis prompt word library for each core dimension are input into a multimodal large-scale model for parallel inference analysis. This generates preliminary text analysis results for each core dimension. A multimodal large-scale model refers to a deep learning model capable of simultaneously processing and understanding multiple data types such as text, images, audio, and video. Through cross-modal alignment and fusion techniques, it maps information from different modalities to a unified semantic space, thereby achieving a more comprehensive and human-like intelligent understanding and generation. These models integrate visual, auditory, and linguistic information for reasoning in scenarios such as video analysis, graphic creation, and intelligent customer service, significantly improving the accuracy and richness of tasks. Common multimodal large-scale models include: OpenAI GPT-4V, which supports mixed image and text input; Google Gemini, which has native multimodal inference capabilities; DeepMind Flamingo, which excels at video and text; OpenAI CLIP, which performs image-text contrastive learning; the DALL·E series, which generates images from text; and domestic models such as Qwen-VL and InternVL. These models have demonstrated powerful cross-modal capabilities in areas such as image description, visual question answering, and video content analysis.
[0031] Based on the generation probability distribution, semantic similarity from multiple samplings, and cross-modal feature consistency scores, the confidence scores of the preliminary analysis results for each core dimension are dynamically evaluated. If the confidence score of the preliminary analysis result for any core dimension is higher than the preset high confidence threshold, an expansion mechanism is triggered. From the preset extended analysis dimension prompt word library, extended dimensions that meet the preset conditions of relevance to the current multimodal input data content are selected and added to the current analysis task. The extended analysis dimension prompt word library is a preset, structured set of prompt words used to trigger the expansion mechanism to mine deeper commercial value information in the video when the confidence of the core dimension analysis is high. Each dimension in this word library is equipped with a specially designed prompt word template to guide the multimodal large model to conduct secondary mining analysis of video content from specific perspectives (such as audience psychology, competitor comparison, compliance risks, etc.), thereby enriching the dimensional breadth of the analysis results while ensuring the accuracy of the analysis. For example, it includes prompt words for audience emotional resonance, competitor comparison, and compliance risk dimensions. To determine whether an expanded dimension meets preset conditions, a multimodal encoder can be used to map the input data into feature vectors, and then cosine similarity calculations can be performed between these vectors and the prompt word vectors for each expanded dimension. Dimensions with similarity scores higher than a threshold (e.g., 0.7) are considered to meet the conditions. Alternatively, zero-shot evaluation can be performed using the large model itself, requiring the model to determine whether the expanded dimension is highly relevant to the current video content and output "yes / no" along with a confidence level, thus achieving dynamic expansion. This mechanism ensures that newly added analysis dimensions always focus on the core video content, avoiding waste of computational resources.
[0032] If the confidence score of any core dimension's preliminary analysis result falls within the preset low-confidence threshold and high-confidence threshold range, no additional processing is performed. If the confidence score of any core dimension's preliminary analysis result is below the low-confidence threshold, a pruning mechanism is triggered, and that core dimension is removed from the current analysis task. If the confidence scores of all core dimensions are below the low-confidence threshold, a safety net analysis mode is triggered, directly generating a global video content overview based on the multimodal input data as a structured analysis result. Representative keyframes are extracted from the image frame sets of each semantic segment and concatenated with the corresponding audio text and overall video summary to form a simplified multimodal input, which is then input into the preset multimodal summarization generation model. Generating a global video content overview as a structured analysis result, the multimodal summarization generation model is a lightweight model specifically designed for video content condensation, activated in baseline analysis mode. This model can receive and fuse input data from different modalities (including key visual features of image frames, semantic features of audio transcription text, and global information summarizing the overall video). Through cross-modal alignment and information compression techniques, it generates a coherent, concise, and comprehensive global text overview covering the core points of the video, ensuring that the system can still output meaningful analysis results even if core dimension analysis fails. Currently, the main existing models include: MMSFT (Multilingual Video Summary). Multimodal Summarization by Fine-Tuning Transformers: A Transformer model specifically fine-tuned for multilingual multimodal summarization tasks, capable of handling text and image inputs and generating aligned multimodal summaries; QUBVIS: A query-based multimodal summarization system based on CLIP and Vision Language Model, using a Transformer architecture for video summarization and employing a GPT-2 decoder to generate subtitles for the summarized videos; V2Xum-LLM (V2Xum-LLaMA): The first framework to unify different video summarization tasks into a large language model text decoder, achieving task-controllable video summarization and general-purpose multimodal large models such as the Gemini series and GPT-4V / 5 through time prompts and task instructions; Based on an adjusted set of analysis dimensions, it generates structured analysis results for each time segment under each analysis dimension.
[0033] The confidence scores for the preliminary analysis results of each core dimension are obtained as follows: The generation probability distribution corresponding to the preliminary analysis results of each core dimension is obtained, and its average log probability is calculated as the first confidence index. The generation probability distribution is obtained by normalizing the raw scores output by the decoding layer of the multimodal large model. When the model generates each word, the decoder calculates a raw score for all possible words in the vocabulary. These scores are then processed by a normalized exponential function to convert them into probability values that sum to 1, representing the likelihood of each word being selected. Developers can obtain these probability values by calling specific parameters of the model interface. For example, enabling the log probability option in the request will return the log probability of each generated word, which can then be used to obtain the actual probability distribution through exponential calculation. These probability distributions reflect the model's certainty about each generated word. Multiple inferences are performed on the same multimodal input data using different random sampling parameters at a preset number of iterations, resulting in multiple candidate analysis results. The semantic similarity between these candidate results is calculated as a second confidence index. Random sampling parameters are modulating variables used to control the randomness and diversity of text generated by large models. They influence the richness of the output results by adjusting the probability distribution or limiting the range of candidate words. Common parameters include: a temperature coefficient, which affects randomness by adjusting the smoothness of the probability distribution; a lower temperature coefficient makes high-probability words stand out more, resulting in more stable generation results; and a higher temperature coefficient makes the probabilities of each word tend to be more even, resulting in more diverse generation results. The first K samples limit the model to randomly select only from the top K words with the highest probabilities, excluding long-tail low-probability words. Probability threshold sampling dynamically selects the smallest set of words whose cumulative probability reaches a set threshold for sampling, achieving adaptive filtering. By combining and adjusting these parameters, multiple different analysis results can be obtained for the same input. Single-modal features are extracted from image frame sets, audio text, and the overall summary. These single-modal features are then compared with multimodal fusion features across modalities, and the intermodal consistency score is calculated as the third confidence index. The intermodal consistency score is calculated by comparing the degree of matching between each single-modal feature and the fusion feature. First, a dedicated encoder is used to extract single-modal feature vectors from image frame sets, audio text, and the overall summary, while a joint feature representation incorporating all information is obtained through a multimodal fusion module. Then, methods such as vector similarity calculation, information correlation analysis, or contribution assessment are used to quantify the correspondence between each single-modal feature and the fusion feature. A higher score indicates that the information of that single modality is effectively preserved during the fusion process, with mutual corroboration between modalities; a lower score suggests potential modal information conflicts or insufficient fusion. Based on preset confidence weight coefficients, the first, second, and third confidence indices are weighted and summed to obtain the confidence scores for the preliminary analysis results of each core dimension.
[0034] In this embodiment, the present invention constructs a full-link intelligent analysis system that integrates deep fusion of multimodal features, parallel reasoning across multiple dimensions, dynamic confidence assessment, intelligent optimization of analysis dimensions, and a safety net mechanism. This system achieves a comprehensive improvement in the accuracy, depth, breadth, and robustness of e-commerce live-streaming video analysis. The technical effects are progressive and synergistic, specifically manifested in the following ways: First, key frames are extracted from the image frame sets of each semantic segment and converted into visual feature vectors. Simultaneously, audio text and overall video summary text are extracted as semantic features. Then, multimodal feature fusion is completed, allowing the visual details of the images, the semantic information of the audio, and the overall plot background to mutually enhance each other and form a unified representation. This enables the large model to conduct collaborative reasoning based on local visuals, local audio, and the overall plot, fundamentally solving the problems of one-sided information in single-modal analysis and the disconnect between cross-modal information. This significantly improves the accuracy and comprehensiveness of the analysis results, aligning with the composite content characteristics of live-streaming videos: "visuals, dialogue, and overall selling points." Secondly, by employing a concurrent invocation mechanism, multimodal input data and multi-dimensional core prompts are input into the multimodal large-scale model inference in parallel. Leveraging the cross-modal alignment and fusion capabilities of the multimodal large-scale model, synchronous analysis of core dimensions such as environment, roles and events, visual design, and marketing appeal is achieved, significantly improving analysis efficiency and avoiding the resource waste of single-dimensional serial inference. Simultaneously, it ensures that each dimension's analysis is based on unified multimodal features, guaranteeing the consistency of analysis results. Furthermore, by quantifying indicators across three dimensions—probability distribution, multiple sampling semantic similarity, and cross-modal feature consistency score—and combining them with pre-defined weighted summations, a confidence score for the preliminary analysis results of each core dimension is obtained. This transforms the reliability of the model analysis results into quantifiable values, providing an objective and accurate basis for judging the analysis effect and overcoming the limitations of subjective judgment. Specifically, the average logarithmic probability reflects the determinism of model generation, multiple sampling semantic similarity reflects the stability of the results, and the cross-modal consistency score reflects the effectiveness of multimodal information fusion. The combination of these three factors enables a comprehensive, multi-dimensional evaluation of confidence, ensuring the scientific rigor of the judgment results. Based on this, an intelligent dimension optimization strategy of expansion, maintenance, and pruning is formulated based on the confidence score. When the confidence score is higher than the high threshold, the expansion mechanism is triggered. Dimensions highly related to the video content (such as audience psychology, competitor comparison, and compliance risks) are selected from the expanded dimension lexicon through cosine similarity or zero-shot evaluation of a large model and added to the analysis. This allows for the discovery of deeper commercial value of the video while ensuring accuracy, enriching the breadth of the analysis results. The newly added dimensions always focus on the core content, avoiding waste of computing resources. When the confidence score is between the high and low thresholds, the original analysis dimensions are maintained to ensure the stability of the analysis process. When the confidence score is lower than the low threshold, the core dimension is pruned and removed to prevent low-quality analysis results from dragging down the overall analysis effect. This achieves dynamic adaptation of analysis dimensions and precise allocation of resources.Simultaneously, when the confidence scores of all core dimensions fall below the low threshold, a safety net analysis mode is triggered, activating a lightweight multimodal summarization generation model. This model concatenates keyframes, audio text, and the overall summary into a concise multimodal input, generating a global video content overview. Even if core dimension analysis fails, the system can still output meaningful results covering the core points, completely resolving the issue of analysis process interruptions or lack of output due to partial dimension failures. This significantly improves the robustness and fault tolerance of the entire analysis system. Finally, based on the adjusted set of analysis dimensions, structured analysis results for each time segment under the corresponding dimensions are generated. This ensures that the final output not only matches the content characteristics of each semantic segment but also dynamically adjusts the breadth of dimensions based on the analysis results, balancing the accuracy, depth, breadth, and standardization of the analysis. Overall, this parallel inference analysis step fully leverages the cross-modal understanding capabilities of the multimodal large model. Deep fusion of multimodal features lays the foundation for accurate analysis, parallel inference improves analysis efficiency, dynamic confidence assessment enables scientific judgment of analysis effectiveness, intelligent dimensional optimization allows for dynamic adaptation of analysis depth and breadth, and a safety net mechanism ensures system robustness. Ultimately, it generates more accurate, comprehensive, and commercially valuable structured analysis results for e-commerce videos, while achieving optimal allocation of analysis resources. This gives the entire analysis process high precision, high efficiency, high adaptability, and high fault tolerance, providing high-quality dimensional analysis support for the subsequent generation of complete video structured description documents.
[0035] like Figure 3The diagram shown is a structural schematic of the low-cost, high-accuracy e-commerce live-streaming video analysis system provided in this application embodiment. It adopts a distributed architecture of control nodes and service nodes, including: an availability pre-detection and diversion module, an audio-visual segmentation processing module, a parallel inference analysis module, and an analysis result integration module. The control node is a standard server (cloud service monthly rental approximately 300-500 RMB), deployed with an availability pre-detection and distribution module, an audio-visual segmentation processing module, and an analysis result integration module. It is responsible for receiving video tasks, performing multi-dimensional availability checks, demultiplexing the video into audio and video streams, generating semantic segments and their corresponding image frame sets and overall video summaries through audio-visual segmentation processing, and monitoring the GPU status of each service node, distributing segmented analysis tasks to the service nodes as needed. The service nodes are one or more servers equipped with consumer-grade graphics cards (such as NVIDIA 3090 / 4090) (monthly rental approximately 2500-3000 RMB per server), deployed with a parallel inference analysis module. This module receives segmented tasks, performs parallel inference analysis using a pre-set multi-dimensional analysis prompt dictionary, generates text analysis results for each dimension, and reports the remaining GPU memory status every second. The control node collects the results from all service nodes, and the analysis result integration module performs structured integration and concatenation in chronological order to generate a complete structured video description document. The minimum total deployment cost of this system is approximately 3,000 yuan per month, and its processing speed can be improved by horizontally scaling up by adding service nodes. The system comprises several modules: Availability Pre-detection and Triage Module: This module performs multi-dimensional availability checks on the video files to be analyzed. If the check passes, the video file is demultiplexed into audio and video streams; otherwise, the analysis process terminates. Audio-visual Segmentation Module: This module performs streaming speech transcription on the audio and video streams to generate semantic text sequences. It also performs real-time dynamic analysis on the video streams, extracting motion intensity coefficients. Based on the semantic text sequences and motion intensity coefficients, it dynamically segments the video files to be analyzed, generating each semantic segment and its corresponding image frame set. Furthermore, it performs cross-segment summary aggregation based on the audio text corresponding to each semantic segment, generating an overall video summary. Parallel Inference Analysis Module: This module takes the image frame sets, audio text, and overall video summary corresponding to each semantic segment as input, and performs parallel inference analysis using a pre-defined multi-dimensional analysis prompt dictionary to generate text analysis results for each pre-defined analysis dimension. Analysis Result Integration Module: This module structures the analysis result text from each analysis dimension and integrates and connects them chronologically to generate a complete structured video description document.
[0036] Through the above description of the implementation methods, those skilled in the art can clearly understand that, for the sake of convenience and brevity, only the division of the above functional modules is used as an example. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the above functions can be divided into different functional modules to complete all or part of the functions described above.
[0037] In the several embodiments provided in this application, it should be understood that the disclosed systems and methods can be implemented in other ways. For example, the embodiments described above are merely illustrative; for instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another device, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces, or indirect coupling or communication connection between devices or units, and may be electrical, mechanical, or other forms.
[0038] The units described as separate components may or may not be physically separate. A component shown as a unit can be one or more physical units, located in one place or distributed in multiple different locations. Some or all of the units can be selected to achieve the purpose of this embodiment, depending on actual needs.
[0039] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0040] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions within the technical scope disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. A low-cost, high-accuracy e-commerce sales video analysis method, characterized in that, Includes the following steps: A multi-dimensional usability check is performed on the product livestreaming video file to be analyzed. If the check passes, the product livestreaming video file to be analyzed is demultiplexed into video audio stream and video video stream. If the check fails, the analysis process is terminated. The multi-dimensional usability check includes audio track existence check, video frame rate compliance check, video resolution compliance check, and video bitrate compliance check. The video audio stream is transcribed into a text semantic sequence. The video frame stream is dynamically analyzed in real time to extract the motion intensity coefficient. Based on the text semantic sequence and the motion intensity coefficient, the video file to be analyzed is dynamically divided into video segments. Each semantic segment and its corresponding image frame set are generated. Cross-segment summary aggregation is performed based on the audio text corresponding to each semantic segment to generate an overall video summary. The image frame set, audio text and video summary corresponding to each semantic segment are used as input. Parallel reasoning analysis is performed in combination with a preset multi-dimensional analysis prompt word library to generate text analysis results under each preset analysis dimension. The analysis dimensions include environmental analysis dimension, role and event analysis dimension, screen design analysis dimension and marketing attractiveness analysis dimension. The text analysis results from each analysis dimension are structured and integrated and linked according to time order to generate a complete structured video description document.
2. The low-cost, high-accuracy e-commerce sales video analysis method as described in claim 1, characterized in that: The specific steps of the multi-dimensional availability check include: The multi-dimensional availability check rules include first priority rules and second priority rules; The first priority rule includes: detecting whether there is an audio stream in the video file to be analyzed; if not, the check is marked as failed. If it exists, it is marked as the first priority rule passed, and the second priority rule is executed in sequence: check whether the frame rate of the video frame of the video file to be analyzed is higher than the preset frame rate threshold, whether the resolution is higher than the preset resolution threshold, and whether the bit rate is higher than the preset bit rate threshold. If it is, it is marked as the corresponding indicator meets the standard; otherwise, it is marked as the corresponding indicator does not meet the standard. If any indicator of the above second priority rule fails to meet the standard, the check is marked as failed, and the error code of the corresponding failure item is recorded; If all the indicators of the above second priority rule are met, the check is marked as passed.
3. The low-cost, high-accuracy e-commerce sales video analysis method as described in claim 1, characterized in that: The steps for dynamically dividing the video segments of the product-selling video file to be analyzed include: Based on the total duration of the video segments in the product promotion video file to be analyzed, the initial frame extraction frequency corresponding to the total duration of the video segment is retrieved from the preset duration-frame extraction frequency mapping table. The pixel difference between adjacent time frames in the video stream is calculated, and the calculation result is used as the motion intensity coefficient of the image. The frame skipping frequency adjustment ratio corresponding to each time frame is obtained from the preset intensity-adjustment ratio mapping table. The initial frame-sampling frequency is adaptively corrected using the frame-sampling frequency adjustment ratio to obtain the dynamic frame-sampling frequency corresponding to each time frame. Based on the dynamic frame extraction frequency, the video stream is processed by frame extraction to generate an image set corresponding to each time frame. Calculate the semantic similarity between adjacent time frames in the text semantic sequence, and perform initial segmentation of the text semantic sequence based on the semantic similarity to generate initial semantic segments; Based on the motion intensity coefficient and the semantic similarity, the initial semantic segmentation is boundary-corrected to generate the final semantic segmentation and its corresponding image frame set.
4. The low-cost, high-accuracy e-commerce sales video analysis method as described in claim 3, characterized in that: The determination of the dynamic frame-sampling frequency also includes a secondary correction mechanism based on computing load, the specific steps of which include: Obtain the current computing power utilization rate of the device and determine the current computing power load level, which includes low load, medium load and high load. If the computing load level is low, the frame extraction frequency after the first-level adaptive correction will be maintained. If the computing load level is medium, reduce the frame extraction frequency of non-critical time frames and maintain the frame extraction frequency of critical time frames. If the computing load level is high, only critical time frames will be extracted, and non-critical time frames will be ignored.
5. The low-cost, high-accuracy e-commerce sales video analysis method as described in claim 4, characterized in that: The method for determining the key time frames and non-key time frames is as follows: For each time frame, obtain its corresponding motion intensity coefficient and semantic change coefficient; The motion intensity coefficient and semantic mutation coefficient of each time frame are normalized to obtain the normalized motion intensity coefficient and normalized semantic mutation coefficient of each time frame. Based on the preset motion weight coefficients and semantic weight coefficients, the normalized motion intensity coefficients and normalized semantic mutation coefficients are weighted and summed to obtain the importance score of each time frame. If the importance score of any time frame reaches the preset key frame threshold, then that time frame is marked as a key time frame. If the importance score of any time frame does not reach the preset key frame threshold, then the time frame is marked as a non-key time frame.
6. The low-cost, high-accuracy e-commerce sales video analysis method as described in claim 3, characterized in that: The step of boundary correction for the initial semantic segmentation includes: S61: Predefined semantic similarity benchmark threshold and segment duration constraint threshold, wherein the segment duration constraint threshold includes the minimum segment duration threshold and the maximum segment duration threshold; S62: Calculate the vector similarity of the text semantics of each adjacent initial semantic segment to obtain the semantic mutation coefficient of each adjacent initial semantic segment; S63: If the semantic mutation coefficient of any adjacent initial semantic segment is greater than the preset mutation upper limit threshold, the difference between the semantic mutation coefficient and the preset mutation upper limit threshold is recorded as the segmentation urgency. According to the segmentation urgency, the corresponding segment minimum duration reduction coefficient is obtained from the preset urgency-adjustment coefficient mapping table. Based on the segment minimum duration reduction coefficient, the segment minimum duration threshold is dynamically reduced, and the adjacent initial semantic segment is divided from the current boundary point to generate two independent semantic segments. S64: If the semantic mutation coefficient of any adjacent initial semantic segment is within the range of the preset smoothing lower threshold and the preset mutation upper threshold, no additional processing is performed. S65: If the semantic mutation coefficient of any adjacent initial semantic segment is less than the preset smoothing lower limit threshold, the difference between the preset smoothing lower limit threshold and the semantic mutation coefficient is recorded as the segment continuity. The corresponding segment maximum duration increase coefficient is obtained from the preset continuity-adjustment coefficient mapping table according to the segment continuity. The segment maximum duration threshold is dynamically increased based on the segment maximum duration increase coefficient, and the adjacent initial semantic segments are merged into one semantic segment. S66: Based on the duration constraint interval formed by the reduced minimum duration threshold and the increased maximum duration threshold of the segment, perform boundary compliance verification on each semantic segment; S67: Repeat steps S61 to S66 until all adjacent initial semantic segments are traversed and no new segmentation or merging operations occur. Generate the final semantic segments based on the final adjusted segment boundaries.
7. The low-cost, high-accuracy e-commerce sales video analysis method as described in claim 6, characterized in that: The steps for performing boundary compliance verification on each semantic segment include: If the duration of any semantic segment is lower than the reduced minimum duration threshold, then the semantic segment is merged with the semantic segment with the smallest semantic mutation coefficient among the adjacent semantic segments, and the semantic mutation coefficient of the merged adjacent segments is recalculated. If the duration of any semantic segment is between the reduced minimum duration threshold and the increased maximum duration threshold, then the semantic segment remains unchanged. If the duration of any semantic segment is higher than the maximum duration threshold of the segment after the increase, then the boundary point with the largest semantic mutation coefficient is found within the semantic segment. Based on the boundary point, the semantic segment is split into two independent semantic segments, and the semantic mutation coefficient of the adjacent segments after splitting is recalculated.
8. The low-cost, high-accuracy e-commerce sales video analysis method as described in claim 1, characterized in that: The steps of the parallel inference analysis include: The image frame sets and audio text of each semantic segment are combined with the overall summary text as multimodal input data; The multimodal input data and the prompt words corresponding to each analysis dimension in the multidimensional analysis prompt word library are input into the multimodal large model for parallel reasoning analysis to generate preliminary text analysis results under each analysis dimension. Based on the generation probability distribution, semantic similarity from multiple samplings, and cross-modal feature consistency scores, the confidence scores of the preliminary analysis results for each analysis dimension are dynamically evaluated. If the confidence score of the preliminary analysis result of any analysis dimension is higher than the preset high confidence threshold, the expansion mechanism is triggered. From the preset extended analysis dimension prompt word library, an extended dimension that meets the preset conditions and is related to the current multimodal input data content is selected and added to the current analysis task. If the confidence score of the preliminary analysis result of any analysis dimension is within the preset low confidence threshold and high confidence threshold range, no additional processing will be performed; If the confidence score of the preliminary analysis result of any analysis dimension is lower than the low confidence threshold, the pruning mechanism is triggered to remove the analysis dimension from the current analysis task. If the confidence level of all analysis dimensions is lower than the low confidence threshold, the baseline analysis mode is triggered, and a global video content overview is directly generated based on the multimodal input data as the text analysis result. Based on the adjusted set of analysis dimensions, text analysis results for each time segment under each analysis dimension are generated.
9. The low-cost, high-accuracy e-commerce sales video analysis method as described in claim 8, characterized in that: The confidence scores of the preliminary analysis results for each analysis dimension are obtained as follows: Obtain the generation probability distribution corresponding to the preliminary analysis results of each analysis dimension, and calculate its average log probability as the first confidence index; Multiple inferences are performed on the same multimodal input data using different random sampling parameters at a preset number of times to obtain multiple candidate analysis results, and the semantic similarity between each candidate result is calculated as the second confidence index. Single-modal features are extracted from image frame sets, audio text, and overall summary respectively. The single-modal features are compared with multimodal fusion features across modalities, and the consistency score between modalities is calculated as the third confidence index. Based on the preset confidence weight coefficients, the first confidence index, the second confidence index, and the third confidence index are weighted and summed to obtain the confidence scores of the preliminary analysis results for each analysis dimension.
10. A system applying the low-cost, high-accuracy e-commerce sales video analysis method as described in any one of claims 1-9, characterized in that, include: Availability pre-detection and triage module, audio-visual segmentation processing module, parallel inference analysis module, and analysis result integration module; The availability pre-detection and diversion module is used to perform multi-dimensional availability checks on the product livestreaming video files to be analyzed. If the check passes, the product livestreaming video files to be analyzed are demultiplexed into video audio streams and video video streams. If the check fails, the analysis process is terminated. The audio-visual segmentation processing module is used to perform streaming speech transcription on the video audio stream to generate a text semantic sequence, perform real-time dynamic analysis on the video video stream to extract the motion intensity coefficient, dynamically divide the video file to be analyzed into video segments based on the text semantic sequence and the motion intensity coefficient, generate each semantic segment and its corresponding image frame set, and perform cross-segment summary aggregation based on the audio text corresponding to each semantic segment to generate an overall video summary. The parallel reasoning analysis module is used to take the image frame set, audio text and video summary corresponding to each semantic segment as input, and perform parallel reasoning analysis in combination with the preset multi-dimensional analysis prompt word library to generate text analysis results under each preset analysis dimension. The analysis result integration module is used to structure the text analysis results of each analysis dimension, and integrate and connect them according to the time sequence to generate a complete video structured description document.
Citation Information
Patent Citations
A method and system for improving sales efficiency based on multi-dimensional video analysis
CN114897585B
A data analysis-based ultra-high-definition live streaming system service supervision system
CN115086721B
Method and device for splitting video and electronic equipment
CN120769112A
Video Segmentation and Content Classification Automation Apparatus
KR102715517B1