Video stream processing method and apparatus therefor
By collecting video streams and motion state data streams, a motion intensity index is generated for image stabilization and encoding optimization. This solves the problem of the inability to optimally balance image quality and smoothness in video stream processing methods, and achieves coordinated optimization of stability and quality of video streams in dynamic network environments.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- VIVO MOBILE COMM (SHENZHEN) CO LTD
- Filing Date
- 2026-04-17
- Publication Date
- 2026-06-16
AI Technical Summary
In existing technologies, video stream processing methods operate independently at each stage, making it impossible to strike an optimal balance between image quality assurance and smoothness, resulting in issues such as blurry or stuttering images in professional live streaming scenarios.
By acquiring video streams and motion state data streams, a motion intensity index is generated, and image stabilization and encoding optimization are performed. The transmission mode is determined by combining network state parameters, and the video stream is reconstructed by weighted fusion at the receiving end.
It achieves coordinated optimization of video stream stability and quality in dynamic network environments, meeting the dual requirements of high robustness and high quality for professional live streaming scenarios.
Smart Images

Figure CN122226962A_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of electronic equipment technology, and specifically relates to a video stream processing method and apparatus. Background Technology
[0002] With the widespread adoption of 5G technology, live video streaming using smartphones and other electronic devices has become an important application scenario. At the same time, users are increasingly demanding higher video quality, expecting both clear and stable images and smooth, uninterrupted playback. However, hand tremors at the capture end and network fluctuations at the transmission end are two core factors affecting the final presentation. Therefore, comprehensive end-to-end processing of the video stream, while simultaneously addressing motion interference and network variations, has become crucial for improving live streaming quality.
[0003] Various video stream processing methods have been proposed in related technologies, but the processing of each stage is independent of each other. This makes it impossible for electronic devices to make the optimal trade-off between image quality and smoothness when facing network fluctuations. This results in either stuttering or interruption due to blindly pursuing high image quality, or excessive compression of the image to ensure smoothness, causing key details to become blurry or pixelated. This makes it difficult to meet the dual requirements of high robustness and high quality in professional live streaming scenarios. Summary of the Invention
[0004] The purpose of this application is to provide a video stream processing method and apparatus that can solve the problem that the optimal balance between image quality and smoothness cannot be achieved in live streaming scenarios due to the independent processing of each stage of electronic equipment.
[0005] In a first aspect, embodiments of this application provide a video stream processing method applied to an electronic device, the method comprising: Acquire video stream and simultaneously acquire motion state data stream of the electronic device; wherein, the video stream includes N original video frames, and the motion state data stream includes the i-th motion state data corresponding to the i-th original video frame, 1≤i≤N, where N is a positive integer greater than 1; Based on the i-th motion state data, an i-th motion intensity index is generated; wherein, the i-th motion intensity index is used to characterize the intensity of motion of the electronic device within the time window corresponding to the i-th original video frame; Based on the i-th motion intensity index, the i-th original video frame is subjected to anti-shake processing to obtain the i-th stable video frame; Based on the i-th motion intensity index and the content characteristics of the i-th stable video frame, the i-th stable video frame is encoded to obtain the i-th encoded data; Based on the network status parameters, the i-th motion intensity index, and the content characteristics of the i-th stable video frame, the i-th transmission mode corresponding to the i-th stable video frame is determined; According to the i-th transmission mode, the i-th encoded data is streamed.
[0006] Secondly, embodiments of this application provide a video stream processing method applied to a receiving device, the method comprising: Receive data packets from a video stream from an electronic device; wherein the data packets include: core feature data, secondary feature data, redundancy check data, motion state data, transmission mode identifier, and encoding priority score; When the transmission mode identifier indicates a high robustness mode, the core feature data is subjected to forward error correction decoding recovery processing based on the redundant check data to obtain the recovered core feature data. The fusion weights are determined based on the completeness of the restoration of the core feature data and the encoding priority score; wherein, the fusion weights are used to characterize the contribution ratio of the core feature data, secondary feature data and motion state data respectively during video reconstruction. Based on the fusion weights, the recovered core feature data, secondary feature data, and motion state data are weighted and fused to reconstruct the video stream.
[0007] Thirdly, embodiments of this application provide a video stream processing apparatus applied to an electronic device, the apparatus comprising: The acquisition module is used to acquire video streams and synchronously acquire motion state data streams of the electronic device; wherein, the video stream includes N original video frames, and the motion state data stream includes the i-th motion state data corresponding to the i-th original video frame, 1≤i≤N, where N is a positive integer greater than 1; The generation module is used to generate an i-th motion intensity index based on the i-th motion state data; wherein the i-th motion intensity index is used to characterize the intensity of motion of the electronic device within the time window corresponding to the i-th original video frame; The image stabilization module is used to perform image stabilization processing on the i-th original video frame according to the i-th motion intensity index to obtain the i-th stable video frame; The encoding processing module is used to encode the i-th stable video frame according to the i-th motion intensity index and the content features of the i-th stable video frame to obtain the i-th encoded data; The first determining module is used to determine the i-th transmission mode corresponding to the i-th stable video frame based on network status parameters, the i-th motion intensity index, and the content characteristics of the i-th stable video frame. The transmission module is used to stream the i-th encoded data according to the i-th transmission mode.
[0008] Fourthly, embodiments of this application provide a video stream processing apparatus applied to a receiving device, the apparatus comprising: A receiving module is used to receive data packets from a video stream from an electronic device; wherein the data packets include: core feature data, secondary feature data, redundancy check data, motion state data, transmission mode identifier, and encoding priority score; The decoding processing module is used to perform forward error correction decoding recovery processing on the core feature data based on the redundant check data when the transmission mode identifier indicates a high robustness mode, so as to obtain the recovered core feature data. The second determining module is used to determine the fusion weight based on the restoration completeness of the core feature data and the encoding priority score; wherein, the fusion weight is used to characterize the contribution ratio of the core feature data, secondary feature data and motion state data respectively during video reconstruction; The reconstruction module is used to perform weighted fusion of the recovered core feature data, the secondary feature data, and the motion state data according to the fusion weights, and reconstruct the video stream.
[0009] Fifthly, embodiments of this application provide an electronic device, the electronic device including a processor and a memory, the memory storing a program or instructions executable on the processor, the program or instructions being executed by the processor to implement the steps of the video stream processing method as described in the first or second aspect.
[0010] In a sixth aspect, embodiments of this application provide a computer-readable storage medium storing a program or instructions that, when executed by a processor, implement the steps of the video stream processing method as described in the first or second aspect.
[0011] In a seventh aspect, embodiments of this application provide a chip, the chip including a processor and a communication interface, the communication interface being coupled to the processor, the processor being used to run programs or instructions to implement the video stream processing method as described in the first or second aspect.
[0012] Eighthly, embodiments of this application provide a computer program product stored in a storage medium, which is executed by at least one processor to implement the video stream processing method as described in the first or second aspect.
[0013] In this embodiment, a video stream is acquired, and a motion state data stream of the electronic device is acquired simultaneously. The video stream includes N original video frames, and the motion state data stream includes the i-th motion state data corresponding to the i-th original video frame, where 1 ≤ i ≤ N, and N is a positive integer greater than 1. Based on the i-th motion state data, an i-th motion intensity index is generated. This i-th motion intensity index characterizes the intensity of motion of the electronic device within the time window corresponding to the i-th original video frame. Based on the i-th motion intensity index, the i-th original video frame is subjected to image stabilization processing to obtain an i-th stable video frame. Based on the i-th motion intensity index and the content characteristics of the i-th stable video frame, the i-th stable video frame is encoded to obtain i-th encoded data. Based on network state parameters, the i-th motion intensity index, and the content characteristics of the i-th stable video frame, the i-th transmission mode corresponding to the i-th stable video frame is determined. Based on the i-th transmission mode, the i-th encoded data is streamed.
[0014] As can be seen, in this embodiment, by converting the motion state data of the electronic device into a motion intensity index and reusing it throughout the entire process of image stabilization, encoding, and transmission decision-making, the collaborative utilization of motion perception information across multiple stages is achieved. This solves the problem of the inability to optimally balance image quality and smoothness in live streaming scenarios due to the independent processing of each stage of the electronic device. Image stabilization based on the motion intensity index improves the stability of video frames. Encoding based on the motion intensity index and video content characteristics enables content-adaptive encoding. Determining the transmission mode based on network state parameters and video content characteristics enables joint decision-making between network supply and content demand. Adapting streaming transmission according to the determined transmission mode improves the transmission efficiency of video streams in dynamic network environments. This embodiment can achieve collaborative optimization of motion perception, content characteristics, and network state throughout the entire link of video stream acquisition, processing, encoding, and transmission in a dynamically changing network environment. While ensuring video stability, it improves transmission robustness and image quality in live streaming scenarios, meeting the dual requirements of high robustness and high quality in professional live streaming scenarios. Attached Figure Description
[0015] Figure 1 This is one of the flowcharts of a video stream processing method provided in some embodiments of this application; Figure 2 This is a flowchart of one implementation of S103 provided in some embodiments of this application; Figure 3 This is a flowchart of one implementation of S104 provided in some embodiments of this application; Figure 4 This is a flowchart of one implementation of S105 provided in some embodiments of this application; Figure 5This is a second flowchart of a video stream processing method provided in some embodiments of this application; Figure 6 This is one of the structural block diagrams of a video stream processing apparatus provided in some embodiments of this application; Figure 7 This is a second structural block diagram of a video stream processing device provided in some embodiments of this application; Figure 8 This is a schematic diagram of the structure of an electronic device provided in some embodiments of this application; Figure 9 This is a schematic diagram of the hardware structure of an electronic device that implements the various embodiments of this application. Detailed Implementation
[0016] The technical solutions of the embodiments of this application will be clearly described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application are within the scope of protection of this application.
[0017] The terms "first," "second," etc., used in the specification and claims of this application are used to distinguish similar objects and not to describe a specific order or sequence. It should be understood that such terms can be used interchangeably where appropriate so that embodiments of this application can be implemented in orders other than those illustrated or described herein, and the objects distinguished by "first," "second," etc., are generally of the same class and the number of objects is not limited; for example, a first object can be one or more. Furthermore, in the specification and claims, "and / or" indicates at least one of the connected objects, and the character " / " generally indicates that the preceding and following objects are in an "or" relationship.
[0018] With the widespread adoption of mobile internet and 5G mobile communication technology, live video streaming using electronic devices has become a significant application scenario. Users have increasingly higher demands for live stream quality, expecting both clear and stable images and smooth, uninterrupted playback. However, hand tremors at the capture end and network fluctuations at the transmission end are two core factors affecting the final presentation. Hand tremors cause blurring, while network fluctuations can lead to stuttering, delays, or even image loss. Therefore, how to perform end-to-end comprehensive processing of the video stream while simultaneously addressing motion interference and network variations has become a key technical challenge for improving live stream quality.
[0019] Various video stream processing methods have been proposed in related technologies, but most only optimize a single stage. In image stabilization, the main approach is to eliminate hand shake by fusing optical and electronic image stabilization, using inertial measurement unit data to drive lens compensation or image correction. However, the generated motion data is discarded after stabilization and cannot be used for subsequent encoding and transmission. In video encoding, fixed or simple adaptive bitrate control is typically used, failing to finely differentiate based on the dynamic changes in video content. Regarding network transmission, although transmission strategies can be dynamically adjusted based on parameters such as bandwidth and latency, the decision-making is limited to the network status itself and does not incorporate the valuable information of the video content.
[0020] The fragmented processing of these various stages prevents electronic devices from making globally optimal decisions based on the actual importance of the video content when faced with real-time changes in network conditions. This is especially true in professional-grade multi-band live streaming scenarios, where image optimization on the terminal side and transmission decisions on the network side are disconnected. Motion data from the image stabilization process is discarded after compensation, network transmission strategies fail to utilize pre-calculated motion intensity information, and the encoding process fails to jointly optimize based on network capacity and content importance. This disconnect in the "perception-processing-transmission" chain prevents electronic devices from intelligently determining whether to prioritize the clarity and continuity of key moving subjects or maintain overall high resolution in complex network environments with congestion or frequent signal switching. The result is often either frequent video stream interruptions due to blindly pursuing high bitrates, or excessive compression of image quality to ensure smoothness, leading to blurry or pixelated keyframes, failing to meet the dual requirements of high robustness and high quality in professional live streaming scenarios.
[0021] To address the aforementioned technical problems, this application provides a video stream processing method and apparatus.
[0022] The video stream processing method provided in the embodiments of this application will be described in detail below with reference to the accompanying drawings.
[0023] It should be noted that the video stream processing method provided in this application is applicable to scenarios that require real-time video transmission, including but not limited to: professional-grade multi-band live streaming, mobile real-time live streaming, video calls, remote conferencing, drone aerial photography, action camera recording, real-time cloud gaming interaction, augmented reality (AR) and virtual reality (VR) immersive experiences, etc. This application does not limit these scenarios.
[0024] Figure 1 This is one of the flowcharts of a video stream processing method provided in some embodiments of this application. This method is applied to electronic devices, such as... Figure 1As shown, the method may include the following steps: S101, S102, S103, S104, S105 and S106.
[0025] In S101, a video stream is acquired, and a motion state data stream of the electronic device is acquired simultaneously; wherein, the video stream includes N original video frames, and the motion state data stream includes the i-th motion state data corresponding to the i-th original video frame, 1≤i≤N, where N is a positive integer greater than 1.
[0026] In this embodiment, the electronic device can drive the camera component to capture raw video streams through its built-in high-end mobile platform. For example, the mobile processing platform can be a Qualcomm Snapdragon series, Apple A series, or MediaTek Dimensity series chip, featuring a professional-grade image signal processor (ISP) and hardware acceleration unit. Furthermore, the video stream can be pre-processed in real-time using the hardware accelerators of the mobile processing platform, such as the image signal processor, video processing unit, and vector processor. Taking the Qualcomm Snapdragon series chip as an example, its integrated image signal processor supports high-resolution video capture, including 8K resolution, and can operate at high frame rates of 30 frames per second or higher, such as supporting 30fps capture at 8K resolution or 120fps capture at 4K resolution. Combined with the vector processor and the video processing unit in the graphics processor, real-time noise reduction, high dynamic range synthesis, and low-latency encoding can be achieved, thereby supporting professional-grade parameter settings including high frame rates of 60 frames per second or higher, wide dynamic range, and low-latency encoding.
[0027] In this embodiment of the application, when acquiring video streams, the electronic device acquires motion state data streams at a sampling frequency higher than the video frame rate, which can provide continuous, high-resolution motion trajectory information for each original video frame in the video stream, thereby supporting microsecond-level timestamp alignment and more accurate image stabilization compensation.
[0028] For example, the video frame rate is 60fps, and the electronic device synchronously acquires motion state data streams at a frequency of 200Hz through an inertial measurement unit (IMU); wherein, the inertial measurement unit typically includes a gyroscope and an accelerometer.
[0029] In this embodiment, the motion state data may include angular velocity data collected by a gyroscope and linear acceleration data collected by an accelerometer, used to capture the high-precision motion state of the electronic device in three-dimensional space. Each original video frame in the video stream is aligned with the motion state data within the corresponding time window via a high-speed bus to establish a microsecond-level timestamp, forming a correspondence between the i-th original video frame and the i-th motion state data, where 1≤i≤N, and N is the total number of original video frames in the video stream.
[0030] For example, taking a video frame captured at 60fps as an example, the exposure time of each frame is about 16.7 milliseconds. The inertial measurement unit can collect 3 to 4 motion data sampling points within this time window, and establish a precise correspondence by aligning the timestamps.
[0031] In this embodiment, the precise alignment of high-frequency motion state data with video frames provides an accurate basis for motion compensation in subsequent image stabilization processing. At the same time, the synchronously acquired motion state data provides reusable motion information for the encoding and transmission stages, avoiding repeated analysis of the original sensor data in each stage.
[0032] In S102, the i-th motion intensity index is generated based on the i-th motion state data.
[0033] In this embodiment, the i-th motion intensity index is used to characterize the intensity of motion of the electronic device within the time window corresponding to the i-th original video frame. Its physical meaning is to characterize the total jitter energy generated by the electronic device due to the composite motion during that time period.
[0034] In this embodiment, the motion intensity index fuses the raw motion state data collected by the inertial measurement unit in real time through an algorithm, transforming low-level sensor readings into a unified quantitative feature that characterizes the energy of image jitter and the complexity of scene motion. The magnitude of the motion intensity index reflects the intensity of motion; for example, 0.85 represents severe jitter, and 0.12 represents slight shaking. Its dynamic trend reflects the dynamic changes in motion patterns and can serve as a priori input for subsequent encoding and transmission stages to determine the complexity of scene motion and adjust encoding strategies and transmission priorities accordingly.
[0035] In this embodiment, the raw sensor data can be transformed into unified quantitative features, which serve as a common input for stabilization, encoding, and transmission decisions. This enables each stage to perform collaborative optimization based on the same motion perception benchmark. On the one hand, it realizes the transformation of motion perception from isolated front-end processing to driving collaborative optimization throughout the entire process. On the other hand, it avoids the waste of motion information being discarded after stabilization, providing a precise compensation basis and decision-making foundation for subsequent implementation of optical and electronic image stabilization.
[0036] In some embodiments, when the motion state data includes angular velocity data collected by a gyroscope and linear acceleration data collected by an accelerometer, the above S102 may specifically include the following step: S1021.
[0037] In S1021, the i-th motion intensity index is generated based on the angular velocity data and linear acceleration data in the i-th motion state data.
[0038] Specifically, the inertial measurement unit built into the electronic device captures its angular velocity and linear acceleration data in three-dimensional space in real time at a sampling frequency higher than the video frame rate, forming a high-precision motion state data stream. To ensure that the subsequent collaborative processing of image stabilization and encoding has an accurate time reference, the raw data from the gyroscope and accelerometer are first timestamped and noise filtered. Then, a motion intensity index is calculated in real time using a motion state fusion and characterization algorithm. This algorithm obtains the i-th motion intensity index I by integrating the product of the angular velocity amplitude and the linear acceleration amplitude within a specific time window. m (t), calculated using the following formula (1): (1) Where ω(τ) is the angular velocity vector at time point τ, a(τ) is the corresponding linear acceleration vector, and ΔT is a short-time integral window associated with the video frame period.
[0039] For example, when shooting a static landscape, the amplitudes of angular velocity and acceleration are relatively small, and the value of the motion intensity index of the integral result is close to 0.1; when shooting a moving scene, the amplitudes of angular velocity and acceleration are relatively large, and the value of the motion intensity index of the integral result can reach 0.8 or more.
[0040] As can be seen, in this embodiment, the motion intensity index generated based on angular velocity and linear acceleration can simultaneously capture the rotational and translational motions of the electronic device, thereby comprehensively characterizing the total jitter energy generated by the composite motion. Compared to relying solely on data from a single sensor, fusing angular velocity and linear acceleration can more accurately reflect the true amplitude and directional changes of hand tremors, providing a more precise motion perception basis for image stabilization compensation, encoding strategies, and transmission decisions.
[0041] In S103, the i-th original video frame is subjected to anti-shake processing according to the i-th motion intensity index to obtain the i-th stable video frame.
[0042] In this embodiment, the i-th motion intensity index is positively correlated with the intensity of the image stabilization process: the higher the motion intensity index, the more violent the electronic device shakes, and the stronger the image stabilization process becomes; the lower the motion intensity index, the more stable the electronic device is, and the weaker the image stabilization process becomes.
[0043] In this embodiment, the original video frame is stabilized according to the motion intensity index, so that the stabilization intensity is adaptively matched with the actual shaking degree. When there is severe shaking, high intensity compensation is used to quickly stabilize the picture, and when there is slight shaking, low intensity fine compensation is used to filter high frequency tremors. This avoids the problem of insufficient compensation when there is severe shaking and overcompensation when there is slight shaking, which is a problem of fixed intensity stabilization. Thus, while ensuring the stability of the picture, unnecessary image correction is reduced, and the overall quality and efficiency of stabilization processing are improved.
[0044] In some embodiments, image stabilization is divided into two stages: optical image stabilization and electronic image stabilization. Correspondingly, such as... Figure 2 As shown, the above S103 may specifically include the following steps: S1031 and S1032.
[0045] In S1031, optical image stabilization is performed on the i-th original video frame according to the i-th motion intensity index to obtain the i-th image frame after optical image stabilization; wherein, the compensation intensity of the optical image stabilization is positively correlated with the i-th motion intensity index.
[0046] In this embodiment, when the value of the i-th motion intensity index is high, a higher compensation intensity of optical image stabilization is used to cope with severe shaking; when the value of the i-th motion intensity index is low, a finer compensation intensity of optical image stabilization is used to filter high-frequency vibrations, thereby ensuring accurate physical compensation during image exposure.
[0047] In some embodiments, the above S1031 may specifically include the following steps: S10311; In S10311, the servo gain coefficient is determined based on the i-th motion intensity index; wherein the servo gain coefficient is positively correlated with the i-th motion intensity index; the drive displacement is determined based on the servo gain coefficient and the angular velocity data in the i-th motion state data; based on the drive displacement, the lens or image sensor of the electronic device is controlled to move in the opposite direction of the device shaking direction to obtain the i-th image frame after optical image stabilization.
[0048] Specifically, according to the i-th exercise intensity index I m (t), dynamically adjust the compensation response curve of the optical image stabilization component. During the exposure of the i-th original video frame, determine the servo gain coefficient based on the i-th motion intensity index. The drive displacement D of the optical image stabilization. ois It is determined by the following formula (2): (2) Wherein, the negative sign indicates that the driving direction is opposite to the device vibration direction, k(I m (t) represents the servo gain coefficient. When the motion intensity index is high, k(I) m (t) takes a larger value to cope with severe shaking. When the value of the motion intensity index is low, k(I) m (t) is taken to a smaller value to filter high-frequency vibration; During the exposure period T exp The integral of the angular velocity within the exposure time represents the total angular offset of the electronic device during the exposure process; t exp This marks the moment the exposure begins.
[0049] Based on the driving displacement D oisThe optical image stabilization system controls the lens or image sensor to move in the opposite direction of the device's shaking. That is, when the electronic device shakes upwards, the lens moves downwards; when the electronic device shakes to the right, the lens moves to the left. This counteracts image shift caused by hand shake, ensuring precise physical compensation during image exposure. After the exposure of the i-th original video frame, the optical image stabilization component immediately resets the lens or image sensor to the center position, preparing for compensation in the next frame.
[0050] For example, when shooting intense motion, the value of the i-th motion intensity index is 0.85, the servo gain coefficient is set to a large value, and the optical image stabilization uses high gain and large displacement for rapid compensation; when shooting still images, the value of the i-th motion intensity index is 0.12, the servo gain coefficient is set to a small value, and the optical image stabilization switches to fine micro-stepping mode to filter high-frequency vibrations.
[0051] In S1032, based on the i-th motion intensity index and the compensation data of optical image stabilization, electronic image stabilization is applied to the i-th image frame that has undergone optical image stabilization to obtain the i-th stable video frame.
[0052] Considering the residual geometric distortion that may still exist within the image frame after optical image stabilization is reset, in this embodiment, electronic image stabilization is performed on the i-th image frame after optical image stabilization based on the i-th motion intensity index and the compensation data of optical image stabilization. This can fuse the deterministic compensation completed by optical image stabilization with the residual motion information that has not been fully compensated. The correction amplitude is adaptively adjusted using the motion intensity index, so that the correction strength of electronic image stabilization matches the current level of shaking: the correction strength is increased to eliminate residual deformation when there is severe shaking, and the correction strength is reduced to avoid overcorrection when there is slight shaking. Thus, while eliminating residual geometric distortion, the natural look of the image is maintained. This forms a deep complement to optical image stabilization in terms of timing and performance, and together outputs a highly stable video frame.
[0053] In some embodiments, S1032 may specifically include the following step: S10321.
[0054] In step S10321, an optical transformation matrix is determined based on the compensation data from the optical image stabilization (OIS) processing. This OIS characterizes the geometric deformation of the image caused by OIS. An additional angular offset is determined based on the angular velocity data in the i-th motion state data. This additional angular offset characterizes the residual rotation angle that was not fully compensated by OIS from the start of exposure to the current processing time. An optical-motion coupling correction matrix is determined based on the OIS, the additional angular offset, and the i-th motion intensity index. This optical-motion coupling correction matrix characterizes the transformation matrix used to perform comprehensive geometric correction on the image after fusing optical compensation information and motion sensing information. Based on the optical-motion coupling correction matrix, geometric correction processing is performed on the i-th image frame after OIS processing to obtain the i-th stable video frame.
[0055] Specifically, to address the residual geometric distortion that may still exist within the image frame after optical image stabilization is reset, the electronic image stabilization component executes a pixel-level collaborative correction algorithm, introducing an optical and motion coupling correction matrix C. frame The matrix is calculated using the following formula (3): (3) Among them, R ois The optical transformation matrix, representing the actual drive displacement vectors performed by the optical image stabilization components within the current frame exposure period, characterizes the geometric deformation of the image caused by optical image stabilization; integral term. Characterizing from the start time t of exposure exp up to the image processing time t of this frame exp The additional angular offset between +Δt is calculated from gyroscope data, representing the residual rotation angle not fully compensated by optical image stabilization; α and β are weighting coefficients for optical compensation data and motion sensing data, respectively, and their values are dynamically determined based on the physical calibration parameters of the current camera module; softmax(I m The function (t) normalizes the exercise intensity index into a weighting factor. For example, the sigmoid function can be used to normalize I. m (t) is mapped to the [0,1] interval, so that the amplitude of the entire correction matrix is adaptively matched with the severity of the current jitter.
[0056] The core purpose of the above formula (3) is not simply to superimpose the two types of data, but to fuse the deterministic vector of optical compensation and the random drift of motion perception into an optimal geometric transformation by using normalized weights based on motion intensity. This transformation is directly applied to the coordinate mapping of image pixels, thereby eliminating pixel misalignment and deformation related to physical motion and generating geometrically highly stable video frames.
[0057] For example, when shooting intense motion, the value of the i-th motion intensity index is 0.85, and the electronic image stabilization performs fine correction based on the compensation data of optical image stabilization and the motion intensity index; when shooting still images, the value of the i-th motion intensity index is 0.12, and the electronic image stabilization only needs slight adjustment.
[0058] As can be seen, in this embodiment, by collaboratively utilizing data from all available sensors and actuators, a deep complementarity between optical image stabilization (OIS) and electronic image stabilization (EIS) in terms of timing and performance is achieved: OIS is responsible for rapid physical compensation during exposure, while EIS is responsible for fine residual correction after exposure. Together, they output highly stable video frames. The resulting pixel-level stable image frames reduce the data burden caused by motion redundancy in subsequent encoding stages, providing a data foundation for intelligent encoding steps geared towards transmission.
[0059] In S104, the i-th stable video frame is encoded based on the i-th motion intensity index and the content characteristics of the i-th stable video frame to obtain the i-th encoded data.
[0060] In this embodiment, the content features of the i-th stable video frame are used to characterize the spatial detail distribution of the i-th stable video frame, including but not limited to: texture complexity, edge intensity, color distribution, etc. Among them, texture complexity is an important indicator for measuring the richness of image detail. For example, the texture complexity of the face area is high, while the texture complexity of the sky or wall area is low.
[0061] In this embodiment, video frames are encoded by combining the motion intensity index and the content features of the video frames: the motion intensity index reflects the degree of dynamic change between frames, and the content features reflect the spatial detail distribution within the frame. Both characterize the importance of the video frame from the temporal and spatial domains, respectively. Frames with rapid motion require more bits to preserve the continuity of the motion trajectory, frames with complex textures require more bits to preserve the integrity of spatial details, and frames that possess both characteristics (such as close-ups of faces in moving scenes) have the highest requirements for encoding quality.
[0062] In this embodiment, the electronic device can dynamically allocate coding resources for each frame based on a comprehensive evaluation of "intra-frame detail richness" and "inter-frame motion intensity." For frames with intense motion or complex textures, the compression intensity is reduced to retain more details; for frames with gentle motion and simple textures, the compression intensity is increased to reduce bitstream overhead. Thus, visually critical information is prioritized for preservation within limited bandwidth, ensuring that the bitstream's information density matches the content value. Simultaneously, the output coding priority score can be used for subsequent transmission decisions, achieving coordinated optimization of coding and transmission.
[0063] In some embodiments, such as Figure 3As shown, the above S104 may specifically include the following steps: S1041 and S1042.
[0064] In S1041, an i-th coding priority score is generated based on the i-th motion intensity index and the content features of the i-th stable video frame; wherein, the i-th coding priority score is used to characterize the content importance of the i-th stable video frame.
[0065] Specifically, the electronic device receives the geometrically stabilized video frame sequence output from the previous collaborative stabilization step via a pre-trained intelligent coding module, and simultaneously reads in the motion intensity index precisely corresponding to each frame. This intelligent coding module executes a content-adaptive feature importance mapping algorithm, coupling the global motion energy remaining after stabilization with the richness of local spatial details within the frame for evaluation, calculating a global coding priority score for each video frame. This score couples the importance of the content with the potential requirements for transmission reliability, making the video stream no longer a homogeneous data packet, but rather endowed with an intrinsic value label, providing precise and quantifiable content-side requirement input for subsequent network transmission strategies.
[0066] This intelligent encoding module is implemented based on a deep learning architecture, employing lightweight convolutional encoder-decoder networks such as MobileNet or ShuffleNet as the encoder backbone. During training, the module learns from broadcast-grade video datasets, such as high-bitrate video footage encompassing various live scenarios like sports events, news broadcasts, and concerts. Its optimization objective is not solely reconstruction accuracy, but rather a balance between reconstruction quality and bitrate overhead. Specifically, its training loss function L... total The following formula (4): (4) Where, x i Represents the video frame before encoding, y i D represents the frame that has been encoded and then decoded and reconstructed. R( ) is a function that measures reconstruction distortion, typically using a combination of multi-scale structural similarity loss and mean squared error; ) is the bitrate estimation function that estimates the actual number of bits generated after encoding the current frame; λ is a key Lagrange multiplier used to dynamically adjust the distortion D during training. ) and bitrate R( A trade-off between these two factors is struck. By minimizing this joint loss, the model is driven to learn a content-adaptive feature representation, which allocates more bits to keyframes with complex textures or intense motion to preserve information, while efficiently compressing static or smooth regions.
[0067] Ultimately, the feature maps learned by the encoder portion of the intelligent encoding module, with their activation intensities across different channels, contain information about the importance of the content. During deployment, based on this characteristic, texture variance is extracted in real-time from the intermediate layer feature maps of the encoder. Equal statistical measures, combined with the exercise intensity index I m (t), calculate the coding priority score P enc This completes the mapping from content importance to quantitative scores.
[0068] In some embodiments, when the content features include texture complexity, the above S1041 may specifically include the following steps: S10411.
[0069] In S10411, the motion intensity index of the i-th frame is normalized to obtain the normalized motion intensity; the normalized motion intensity is used to characterize the intensity of motion of the current frame relative to the historical period; multi-scale gradient operation is performed on the i-th stable video frame to obtain the texture complexity variance; the texture complexity variance is used to characterize the spatial detail richness of the i-th stable video frame; the coding priority score of the i-th frame is determined based on the normalized motion intensity and the texture complexity variance.
[0070] Specifically, the coding priority score P enc The calculation formula is as follows: Formula (5): (5) Among them, I m (t) represents the motion intensity index corresponding to the current frame, max(I) m ) represents the motion intensity index I within a short sliding window. m The historical maximum value of (t) is used to normalize the assessment of the current intensity of the exercise; δ is a very small normal number to prevent the denominator from being zero and to ensure the quality of the baseline coding. The texture complexity variance is obtained by performing multi-scale gradient operations on the stable frame. For example, the Sobel operator is used to calculate the gradient magnitude at 3×3 and 5×5 scales to obtain the texture complexity variance. Formula (5) couples the global motion energy that still exists after image stabilization with the richness of local spatial details within the frame to calculate P. enc The value directly determines the intensity allocation of the intelligent encoding module for feature extraction and compression of the frame.
[0071] In S1042, the i-th stable video frame is encoded according to the i-th encoding priority score to obtain the i-th encoded data.
[0072] Specifically, the intelligent coding module calculates the coding priority score P. enc Adaptive encoding is performed on video frames. When Penc A higher value indicates significant motion or rich texture in the current frame, and the intelligent encoding module will automatically reduce the compression intensity and retain a more complete multi-level feature map; when P... enc When the value is low, a high compression mode is activated, retaining only the most basic spatial structure features and low-resolution motion vectors. This dynamic process ensures that the final output video feature stream maintains a match between its information density and content value while reducing the amount of data. This process generates two outputs: low-volume encoded data and a score P for each frame. enc These findings collectively reflect the video content's demand for transmission resources, providing a quantitative basis for subsequent link quality assessment and strategy decision-making steps before transmission.
[0073] In some embodiments, S1042 may specifically include the following step: S10421.
[0074] In S10421, if the i-th encoding priority score is greater than or equal to a preset score threshold, the i-th stable video frame is subjected to a first encoding process to obtain the i-th encoded data; wherein, the first encoding process uses a first compression intensity, and the encoded data retains multi-level feature maps, including deep motion trajectory features and fine texture features; if the i-th encoding priority score is less than the preset score threshold, the i-th stable video frame is subjected to a second encoding process to obtain the i-th encoded data; wherein, the second encoding process uses a second compression intensity, and the encoded data retains only basic spatial structure features and low-resolution motion vectors; the first compression intensity is lower than the second compression intensity.
[0075] For example, taking a close-up frame of a person's face as an example, the motion intensity index is 0.78, the texture complexity variance is 0.92, and the encoding priority score is 0.72. The intelligent encoding module uses a low compression intensity to retain complete facial details and deep motion trajectory features. Taking a background frame as an example, the motion intensity index is 0.12, the texture complexity variance is 0.15, and the encoding priority score is 0.02. The intelligent encoding module uses a high compression intensity to retain only the basic outline and basic spatial structure features.
[0076] As can be seen, in this embodiment, bit resources are dynamically allocated according to the importance of the content, so that the information density of the video bitstream matches the value of the content. Under limited bandwidth, key visual information is reserved first, realizing the intelligent association between the encoding output and the transmission guarantee requirements, and achieving the effect of transforming from "best-effort" transmission to "on-demand guarantee" transmission.
[0077] In S105, the i-th transmission mode corresponding to the i-th stable video frame is determined based on the network state parameters, the i-th motion intensity index, and the content characteristics of the i-th stable video frame.
[0078] In this embodiment, network state parameters are used to characterize the real-time communication quality of the transmission link, including available bandwidth, average latency, and instantaneous packet loss rate. Available bandwidth reflects the data transmission capacity of the link, average latency reflects the transmission time of data packets from the sender to the receiver, and instantaneous packet loss rate reflects the stability of the link.
[0079] Considering that adjusting transmission strategies solely based on network status parameters cannot distinguish between important and ordinary frames, and selecting transmission methods solely based on video content may lead to transmission failures due to ignoring actual network conditions, this application embodiment determines the transmission mode by combining network status parameters, motion intensity index, and video frame content characteristics: network status parameters reflect "how much can be transmitted, how fast, and how stable," while motion intensity index and content characteristics reflect "how important the current frame is." This enables the transmission mode to simultaneously respond to network fluctuations and changes in content value, achieving adaptive adjustment of the transmission strategy in a dynamic network environment.
[0080] In this embodiment, an adaptive selection of the transmission mode is achieved by jointly evaluating network supply capacity and the importance of video content: when network conditions are good and the current frame content is of low importance, a high throughput mode is selected to fully utilize bandwidth resources; when network conditions are poor or the current frame content is of high importance, a high robustness mode is selected to prioritize the reliable transmission of critical data. Thus, in a dynamically changing network environment, electronic devices can make a reasonable trade-off between throughput and reliability, avoiding blindly pursuing high bitrates that lead to stuttering, and also avoiding excessive compression that causes blurring of critical images.
[0081] In some embodiments, such as Figure 4 As shown, the above S105 may specifically include the following steps: S1051, S1052, S1053 and S1054.
[0082] In S1051, an i-th coding priority score is generated based on the i-th motion intensity index and the content features of the i-th stable video frame; wherein, the i-th coding priority score is used to characterize the content importance of the i-th stable video frame.
[0083] In this embodiment, the i-th motion intensity index and the content features of the i-th stable video frame are fused into the i-th coding priority score, which is used to quantify the content transmission requirements of the current frame. For details on the specific generation method of the coding priority score, please refer to the relevant descriptions in S1041 and S10411, which will not be repeated here.
[0084] In S1052, the transmission strategy decision value is calculated based on the network state parameters and the i-th encoding priority score; wherein, the transmission strategy decision value is used to characterize the degree of matching between network supply capacity and content transmission demand.
[0085] In this embodiment, the electronic device can continuously monitor the transmission quality parameters of all currently available network interfaces and simultaneously obtain the current video frame encoding priority score from the intelligent encoding module. To make optimal transmission decisions in complex network environments, a joint evaluation and decision function is introduced. By fusing network quality parameters and video content priority scores, a real-time transmission strategy decision value is calculated.
[0086] In some embodiments, when the network state parameters include available bandwidth, average latency and instantaneous packet loss rate, the above S1052 may specifically include the following steps: S10521.
[0087] In S10521, bandwidth adequacy is determined based on available bandwidth and a preset baseline bandwidth; bandwidth adequacy characterizes the sufficiency of the current available bandwidth relative to the baseline bandwidth; link reliability attenuation factor is determined based on average latency and instantaneous packet loss rate; link reliability attenuation factor characterizes the impact of network latency and packet loss on transmission reliability; network supply capacity value is determined based on bandwidth adequacy and link reliability attenuation factor; network supply capacity value characterizes the overall transmission capacity that the current network can provide; and transmission strategy decision value is determined based on network supply capacity value and the i-th coding priority score.
[0088] Specifically, the transmission strategy decision value Q decision The calculation formula is as follows: Formula (6): (6) Among them, B avail For the currently available bandwidth, B req This is the preset baseline bandwidth based on the basic rate requirements of the video feature bitstream. It quantifies the sufficiency of network bandwidth; D avg For the average delay, L loss The instantaneous packet loss rate A link reliability evaluation factor was constructed that rapidly decreases with increasing network latency and packet loss; the product of these two factors reflects the overall network supply capacity. enc Let be the priority score for the i-th encoding, and γ be a positive coefficient used to adjust the weight of content importance. This constitutes the content delivery demand item. Subtracting the content delivery demand item from the overall network supply capacity allows the decision value to accurately represent the dynamic difference between the current network supply and content delivery demand.
[0089] For example, taking a static landscape shot with a network bandwidth of 100Mbps, a latency of 15ms, a packet loss rate close to 0, and an encoding priority score of 0.05, the bandwidth adequacy is relatively high, the link reliability attenuation factor is close to 1, and the transmission strategy decision value is positive. Taking a close-up shot with a network bandwidth of 8Mbps, a latency of 180ms, a packet loss rate of 12%, and an encoding priority score of 0.78, the bandwidth adequacy is relatively low, the link reliability attenuation factor is relatively small, and the transmission strategy decision value is negative.
[0090] In S1053, if the transmission strategy decision value is greater than or equal to the preset decision threshold, the i-th transmission mode is determined to be the high throughput mode; wherein, the high throughput mode is used to aggregate multiple network links to transmit complete encoded data.
[0091] In this embodiment, when the transmission strategy decision value is consistently higher than a preset decision threshold, it indicates that the network condition is excellent and sufficient to support high bitrate demands. The decision module then selects a high throughput mode, scheduling multi-band aggregation links such as 5G-A to transmit the complete video feature bitstream. This mode prioritizes utilizing the high-bandwidth links of multi-band aggregation, pursuing high throughput and low latency.
[0092] In S1054, if the transmission strategy decision value is less than the preset decision threshold, the i-th transmission mode is determined to be a high robustness mode; wherein, the high robustness mode is used to perform hierarchical transmission and redundancy protection of the encoded data.
[0093] In this embodiment, when the transmission strategy decision value falls below a preset decision threshold, it indicates that the network quality is unstable or the reliability requirements of the content exceed the current network's guarantee capacity. The decision module automatically triggers and enables a high-robustness mode. This mode is used for hierarchical transmission and redundancy protection of encoded data, prioritizing the reliable transmission of critical content.
[0094] As can be seen, in this embodiment, by constructing a decision function that integrates real-time network quality parameters and video content priority scores, abstract network parameters and specific video content values are modeled in a unified manner, realizing a shift from passive link monitoring to proactive intelligent decision-making. The electronic device can autonomously switch between a high-throughput mode that prioritizes throughput and a high-robustness mode that prioritizes reliability, based on the principle of global optimization, dynamically optimizing the balance between throughput and reliability in complex network environments.
[0095] In S106, the i-th encoded data is streamed according to the i-th transmission mode.
[0096] In this embodiment of the application, the encoded data of each video frame constitutes a continuous video stream in chronological order and is sent to the receiving device in a streaming manner.
[0097] In some embodiments, when the i-th transmission mode is a high-throughput mode, the electronic device aggregates all available network links such as 5G-A, 5G, Wi-Fi, 4G, etc., and distributes the complete encoded data evenly to each link to improve bandwidth utilization.
[0098] In some embodiments, when the i-th transmission mode is a high robustness mode, the above S106 may specifically include the following steps: S1061, S1062 and S1063.
[0099] In S1061, if the i-th encoding priority score is greater than or equal to the preset classification threshold, the i-th encoded data is determined as the core feature data, and forward error correction redundancy protection is added to the i-th encoded data before streaming transmission through the main network link; wherein, the main network link is a stable link selected from multiple available network links for transmitting the core feature data.
[0100] In this embodiment, under high robustness mode, the electronic device executes a dynamic redundancy allocation and data classification algorithm based on the negative magnitude of the real-time transmission strategy decision value calculated in the previous stage and the encoding priority score for each frame provided by the intelligent encoding module. This algorithm first performs critical quantification on the encoded data corresponding to each video frame in the video feature bitstream: when the encoding priority score of the i-th frame is greater than or equal to a preset classification threshold, the encoded data of that frame is determined as core feature data. The core feature data contains key motion and structural information. After the electronic device adds forward error correction redundancy protection to it, it reliably transmits it through the main network link with the least latency jitter selected from multiple available network links.
[0101] In some embodiments, S1061 may specifically include the following step: S10611.
[0102] In S10611, the number of redundant packets K is calculated based on the i-th encoding priority score and network state parameters. i Fountain code encoding is applied to the i-th encoded data to generate K. i Redundant check data; combine the i-th encoded data and K... i After the redundant verification data is merged, it is streamed through the main network link.
[0103] In this embodiment, the electronic device calculates the basic redundancy based on the i-th encoding priority score, the average latency of the current main network link, and the packet loss rate; it then adaptively adjusts the basic redundancy based on the historical transmission success rate and rounds it up to obtain the number of redundant packets K. i .
[0104] Specifically, the number of redundant packets K i The calculation is adaptive and is dynamically determined by the following formula (7): (7) Among them, P enc It is the encoding priority score corresponding to that video frame, D avg With L loss η is the latest average latency and packet loss rate assessment value of the currently selected stable link used for transmitting core data, which together characterize the severity of the link; η is a strength coefficient that is adjusted online based on historical transmission success rate, used to control the overall redundancy tendency of the system; the function ceil ensures that the number of redundant packets generated in the end is an integer. The physical meaning of this formula (7) is: for data content of higher importance, more redundant packets are allocated for protection when the network conditions are assessed as worse. Denominator It ensures a sensitive response to the degree of network degradation, while avoiding unbounded expansion of redundancy when latency is too high through square root operation.
[0105] The specific generation process is as follows: the encoded data of the i-th stable video frame is divided into R... i The original data packet will contain R i The core feature data block of each raw data packet is input into the fountain code encoder, for example, using RaptorQ code. The encoder calculates K... i Values, generate R i +K i A encoded data packet. This K i The additional packets are redundant check data, which are mathematically equivalent to the original data packets. Any R i Combinations of individual packets can be used to decode and recover the original data blocks, thus providing fault tolerance in the event of network packet loss.
[0106] For example, taking a human face frame with an encoding priority score of 0.78 as an example, the number of redundant packets is calculated to be 8. After adding redundancy protection to the core feature data, it is streamed through the 5G main link.
[0107] In S1062, if the i-th encoding priority score is less than the preset classification threshold, the i-th encoded data is determined as secondary feature data, and the i-th encoded data is streamed through other available links; wherein, other available links are links other than the main network link among multiple available network links.
[0108] In this embodiment, when the i-th encoding priority score is less than a preset classification threshold, the encoded data of that frame is determined as secondary feature data. Secondary feature data does not enjoy strong error correction protection, only undergoes basic encapsulation, and is transmitted through other available links in a best-effort manner.
[0109] For example, taking a background frame with an encoding priority score of 0.02 as an example, the electronic device classifies it as secondary feature data and streams it through a 4G link.
[0110] In S1063, the i-th motion state data is synchronously transmitted to the receiving device as auxiliary information.
[0111] In this embodiment, the electronic device continuously transmits the acquired high-precision motion state data stream synchronously to the receiving device as an independent, low-bandwidth auxiliary channel. This motion state data can be used in the video reconstruction process at the receiving end, providing an additional information source for decoding.
[0112] As can be seen, in this embodiment, a hierarchical survivability protection system based on content importance is constructed through a dynamic classification and adaptive redundancy protection mechanism driven by both real-time network quality and content value, under harsh transmission environments with limited and unstable network resources. Hierarchical transmission ensures that critical content receives the highest priority protection when the network is limited; the redundancy mechanism enables the receiving end to fully recover the original data from partially lost data packets; and the synchronous transmission of motion state data provides an additional information source for the receiving device's reconstruction, ensuring that the receiving end can recover the key motion and structural information necessary for video reconstruction with a high probability of lossless recovery.
[0113] As can be seen from the above embodiments, this embodiment achieves the coordinated utilization of motion perception information across multiple stages by converting the motion state data of electronic devices into a motion intensity index and reusing it throughout the entire process of image stabilization, encoding, and transmission decision-making. This solves the problem of the inability to optimally balance image quality and smoothness in live streaming scenarios due to the independent processing of each stage of electronic devices. Image stabilization based on the motion intensity index improves the stability of video frames. Encoding based on the motion intensity index and video content characteristics enables content-adaptive encoding. Determining the transmission mode based on network state parameters and video content characteristics enables joint decision-making between network supply and content demand. Adapting streaming transmission to the determined transmission mode improves the transmission efficiency of video streams in dynamic network environments. This embodiment can achieve coordinated optimization of motion perception, content characteristics, and network state throughout the entire link of video stream acquisition, processing, encoding, and transmission in a dynamically changing network environment. While ensuring video stability, it improves transmission robustness and image quality in live streaming scenarios, meeting the dual requirements of high robustness and high quality in professional live streaming scenarios.
[0114] Figure 5 This is a second flowchart of a video stream processing method provided in some embodiments of this application. This method is applied to a receiving device, such as... Figure 5As shown, the method may include the following steps: S501, S502, S503 and S504.
[0115] In S501, data packets of video stream from an electronic device are received; wherein the data packets include: core feature data, secondary feature data, redundancy check data, motion state data, transmission mode identifier, and encoding priority score.
[0116] In this embodiment, the cloud or a professional workstation can act as the receiving device to receive the data packets sent through the aforementioned transmission steps. The receiving device first identifies the transmission mode identifier of the data packet based on the transmission protocol header information to determine the transmission mode corresponding to the current data packet. The data packet contains the following: core feature data, i.e., encoded data classified as high priority by the sender according to the encoding priority score; secondary feature data, i.e., encoded data classified as low priority by the sender according to the encoding priority score; redundancy check data, i.e., redundant information generated by the sender after adding forward error correction protection to the core feature data; motion state data, i.e., motion state data stream of the electronic device synchronously collected by the sender; transmission mode identifier, used to indicate whether the current data packet corresponds to a high-throughput mode or a high-robust mode; and encoding priority score, used to characterize the importance of the video frame content.
[0117] In S502, when the transmission mode identifier indicates a high robustness mode, the core feature data is recovered by forward error correction decoding based on the redundancy check data.
[0118] In this embodiment, when the transmission mode identifier indicates a high robustness mode, the receiving device initiates a robust decoding process guided by content importance. The receiving device utilizes the received redundant check data to perform forward error correction decoding and recovery of potentially lost or damaged core feature data packets. Because the sending end uses fountain codes such as RaptorQ codes to redundantly encode the core feature data, the receiving device can recover the original core feature data block as long as it receives a sufficient number of encoded data packets.
[0119] In S503, the fusion weights are determined based on the restoration completeness and coding priority score of the core feature data; the fusion weights are used to characterize the contribution ratio of the core feature data, secondary feature data and motion state data during video reconstruction.
[0120] Specifically, the receiving device executes a motion-aware feature-weighted reconstruction algorithm to calculate fused features for video reconstruction. This algorithm introduces dynamic adaptive coefficients v and μ, the values of which are determined by the recovery completeness ratio ρ of the core feature data and the current frame coding priority score P. enc The specific relationship is determined jointly by formulas (8) and (9) as follows: (8) (9) Where ρ is the proportion of completeness of the restored core feature data, 0≤ρ≤1, P enc This represents the encoding priority score for the current frame. The fusion weight for core feature data is v, the fusion weight for motion state data is μ, and the fusion weight for secondary feature data is 1-v. When the completeness ρ of core feature recovery is high, the contribution weight of core feature data is larger; when ρ decreases due to network deterioration, the contribution weights of secondary feature data and motion state data increase accordingly.
[0121] In S504, the recovered core feature data, secondary feature data, and motion state data are weighted and fused according to the fusion weights to reconstruct the video stream.
[0122] Specifically, the receiving device calculates the final fusion feature F used for video reconstruction using the following formula (10). final : (10) Among them, F core J( represents the complete core feature data recovered by the receiving device through forward error correction decoding using redundant check data.) ) is the core feature decoder; F secondary It is the successfully received secondary feature data, U( ) is a secondary feature upsampling and compensation network; I m(t) It is a synchronously received time series of motion intensity index of electronic devices, W( ) is a motion feature weighted network, which is based on I m(t) The temporal changes dynamically generate a set of feature weight maps to enhance the spatiotemporal coherence of the reconstructed frames.
[0123] The motion feature weighted network W( It employs a lightweight temporal-spatial convolutional neural network structure, whose input is the synchronously received motion intensity exponent time series I. m(t) The output is a feature weight map W that matches the spatial dimensions of the reconstructed frame. t It is used to weight the reliability of different spatial locations during feature fusion.
[0124] Optionally, the network first preprocesses the input time series to extract stable trends, for a given time series with the current frame time t... c Centered sliding window T w The weighted mean kinetic energy E is calculated using the following formula (11). m : (11) Where, σ t Control the rate of time decay. Afterwards, the network will... m With the normalized current sequence segment [I] m (t c-L ),..., I m (t c The concatenated layers are fed into an encoder consisting of a one-dimensional convolutional layer and a fully connected layer, mapping it into a low-dimensional latent vector. This latent vector is then copied and reassembled as a conditional input to a two-dimensional spatial transformer, ultimately generating a spatial weight map W. t The training objective of the entire network is to maximize the spatiotemporal coherence gain it brings in assisting video reconstruction, which is measured by comparing the optical flow consistency between the fused output frame and the real frame in a multi-frame consecutive window.
[0125] As can be seen, the above formula (10) constructs a multi-level, compensatory feature reconstruction framework: when the completeness ρ of the core feature recovery is high, the decoder trusts and mainly uses the high-quality content reconstructed from the core features, and the motion state data is only used as a fine-tuning aid; when ρ decreases due to network deterioration, the receiving device automatically increases the contribution weight of secondary feature data and motion state data, and uses the spatial details of secondary features and the temporal dynamic priors contained in the motion state information to fill and correct feature gaps and ambiguities caused by the loss of core data. After completing the weighted fusion, the receiving device uses a decoder that matches the intelligent encoding module of the sending end to reconstruct a high-quality stable video stream. If the transmission mode identifier indicates a high throughput mode, the receiving device directly uses the completely received video feature bitstream for decoding and reconstruction.
[0126] As can be seen from the above embodiments, in this embodiment, the video stream finally reconstructed by the receiving device is not only visually high quality and rich in detail, but also maintains a high degree of spatiotemporal stability and smoothness consistent with the original image stabilization processing of the sending device.
[0127] The video stream processing method provided in this application can be executed by a video stream processing device. This application uses a video stream processing device executing the video stream processing method as an example to illustrate the video stream processing device provided in this application.
[0128] Figure 6 This is one of the structural block diagrams of a video stream processing device provided in some embodiments of this application, applied to electronic devices, such as... Figure 6 As shown, the video stream processing device 600 may include: an acquisition module 601, a generation module 602, a stabilization module 603, an encoding module 604, a first determination module 605, and a transmission module 606; The acquisition module 601 is used to acquire video streams and synchronously acquire motion state data streams of the electronic device; wherein, the video stream includes N original video frames, and the motion state data stream includes the i-th motion state data corresponding to the i-th original video frame, 1≤i≤N, where N is a positive integer greater than 1; The generation module 602 is used to generate an i-th motion intensity index based on the i-th motion state data; wherein, the i-th motion intensity index is used to characterize the intensity of motion of the electronic device within the time window corresponding to the i-th original video frame; The image stabilization module 603 is used to perform image stabilization processing on the i-th original video frame according to the i-th motion intensity index to obtain the i-th stable video frame; The encoding processing module 604 is used to encode the i-th stable video frame according to the i-th motion intensity index and the content characteristics of the i-th stable video frame to obtain the i-th encoded data; The first determining module 605 is used to determine the i-th transmission mode corresponding to the i-th stable video frame based on network status parameters, the i-th motion intensity index and the content characteristics of the i-th stable video frame; The transmission module 606 is used to stream the i-th encoded data according to the i-th transmission mode.
[0129] As can be seen from the above embodiments, this embodiment achieves the coordinated utilization of motion perception information across multiple stages by converting the motion state data of electronic devices into a motion intensity index and reusing it throughout the entire process of image stabilization, encoding, and transmission decision-making. This solves the problem of the inability to optimally balance image quality and smoothness in live streaming scenarios due to the independent processing of each stage of electronic devices. Image stabilization based on the motion intensity index improves the stability of video frames. Encoding based on the motion intensity index and video content characteristics enables content-adaptive encoding. Determining the transmission mode based on network state parameters and video content characteristics enables joint decision-making between network supply and content demand. Adapting streaming transmission to the determined transmission mode improves the transmission efficiency of video streams in dynamic network environments. This embodiment can achieve coordinated optimization of motion perception, content characteristics, and network state throughout the entire link of video stream acquisition, processing, encoding, and transmission in a dynamically changing network environment. While ensuring video stability, it improves transmission robustness and image quality in live streaming scenarios, meeting the dual requirements of high robustness and high quality in professional live streaming scenarios.
[0130] Optionally, as an embodiment, the image stabilization module 603 is specifically used to perform optical image stabilization on the i-th original video frame according to the i-th motion intensity index to obtain the i-th image frame after optical image stabilization; wherein, the compensation intensity of the optical image stabilization is positively correlated with the i-th motion intensity index; and to perform electronic image stabilization on the i-th image frame after optical image stabilization according to the i-th motion intensity index and the compensation data of the optical image stabilization to obtain the i-th stable video frame.
[0131] Optionally, as an embodiment, the encoding processing module 604 is specifically used to generate an i-th encoding priority score based on the i-th motion intensity index and the content features of the i-th stable video frame; wherein, the i-th encoding priority score is used to characterize the content importance of the i-th stable video frame; and to encode the i-th stable video frame according to the i-th encoding priority score to obtain the i-th encoded data.
[0132] Optionally, as an embodiment, the first determining module 605 is specifically configured to generate an i-th encoding priority score based on the i-th motion intensity index and the content characteristics of the i-th stable video frame; wherein the i-th encoding priority score is used to characterize the content importance of the i-th stable video frame; calculate a transmission strategy decision value based on network state parameters and the i-th encoding priority score; wherein the transmission strategy decision value is used to characterize the degree of matching between network supply capacity and content transmission demand; determine the i-th transmission mode as a high throughput mode when the transmission strategy decision value is greater than or equal to a preset decision threshold; wherein the high throughput mode is used to aggregate multiple network links to transmit complete encoded data; determine the i-th transmission mode as a high robustness mode when the transmission strategy decision value is less than the preset decision threshold; wherein the high robustness mode is used to perform hierarchical transmission and redundancy protection for encoded data.
[0133] Figure 7 This is a second structural block diagram of a video stream processing apparatus provided in some embodiments of this application, applied to a receiving end device, such as... Figure 7 As shown, the video stream processing device 700 may include: a receiving module 701, a decoding processing module 702, a second determining module 703, and a reconstruction module 704; The receiving module 701 is used to receive data packets of video stream from an electronic device; wherein the data packets include: core feature data, secondary feature data, redundancy check data, motion state data, transmission mode identifier, and encoding priority score; The decoding processing module 702 is used to perform forward error correction decoding recovery processing on the core feature data based on the redundant check data when the transmission mode identifier indicates a high robustness mode, so as to obtain the recovered core feature data. The second determining module 703 is used to determine the fusion weight based on the restoration completeness of the core feature data and the encoding priority score; wherein the fusion weight is used to characterize the contribution ratio of the core feature data, secondary feature data and motion state data respectively during video reconstruction; The reconstruction module 704 is used to perform weighted fusion of the recovered core feature data, the secondary feature data and the motion state data according to the fusion weight, and reconstruct the video stream.
[0134] As can be seen from the above embodiments, in this embodiment, the video stream finally reconstructed by the receiving device is not only visually high quality and rich in detail, but also maintains a high degree of spatiotemporal stability and smoothness consistent with the original image stabilization processing of the sending device.
[0135] The video stream processing device in this application embodiment can be an electronic device or a component within an electronic device, such as an integrated circuit or a chip. The electronic device can be a terminal or other devices besides a terminal. For example, the electronic device can be a mobile phone, tablet computer, laptop computer, PDA, in-vehicle electronic device, mobile internet device (MID), augmented reality (AR) / virtual reality (VR) device, robot, wearable device, ultra-mobile personal computer (UMPC), netbook, or personal digital assistant (PDA), etc. It can also be a server, network attached storage (NAS), personal computer (PC), television (TV), ATM, or self-service machine, etc. This application embodiment does not specifically limit the device.
[0136] The video stream processing device in this application embodiment can be a device with an operating system. This operating system can be Android, iOS, or other possible operating systems; this application embodiment does not specifically limit the specific operating system used.
[0137] The video stream processing device provided in this application embodiment can achieve the above-mentioned... Figure 1 or Figure 5 To avoid repetition, the various processes implemented in the method embodiment shown will not be described again here.
[0138] Optionally, such as Figure 8 As shown, this application embodiment also provides an electronic device 800, including a processor 801 and a memory 802. The memory 802 stores a program or instructions that can run on the processor 801. When the program or instructions are executed by the processor 801, they implement the various steps of the above-described video stream processing method embodiment and can achieve the same technical effect. To avoid repetition, they will not be described again here.
[0139] It should be noted that the electronic devices in the embodiments of this application include the mobile electronic devices and non-mobile electronic devices described above.
[0140] Figure 9 This is a schematic diagram of the hardware structure of an electronic device that implements the various embodiments of this application.
[0141] The electronic device 900 includes, but is not limited to, components such as: radio frequency unit 901, network module 902, audio output unit 903, input unit 904, sensor 905, display unit 906, user input unit 907, interface unit 908, memory 909, and processor 910.
[0142] Those skilled in the art will understand that the electronic device 900 may also include a power supply (such as a battery) for supplying power to various components. The power supply may be logically connected to the processor 910 through a power management system, thereby enabling functions such as managing charging, discharging, and power consumption through the power management system. Figure 9 The electronic device structure shown does not constitute a limitation on the electronic device. The electronic device may include more or fewer components than shown, or combine certain components, or have different component arrangements, which will not be elaborated here.
[0143] In some embodiments, the processor 910 is configured to acquire a video stream and simultaneously acquire a motion state data stream of the electronic device; wherein the video stream includes N original video frames, and the motion state data stream includes i-th motion state data corresponding to the i-th original video frame, 1≤i≤N, where N is a positive integer greater than 1; generate an i-th motion intensity index based on the i-th motion state data; wherein the i-th motion intensity index is used to characterize the intensity of motion of the electronic device within a time window corresponding to the i-th original video frame; perform image stabilization processing on the i-th original video frame based on the i-th motion intensity index to obtain an i-th stable video frame; encode the i-th stable video frame based on the i-th motion intensity index and the content characteristics of the i-th stable video frame to obtain i-th encoded data; determine the i-th transmission mode corresponding to the i-th stable video frame based on network state parameters, the i-th motion intensity index, and the content characteristics of the i-th stable video frame; and stream the i-th encoded data according to the i-th transmission mode.
[0144] As can be seen, in the embodiments of this application, motion perception, content features and network status can be coordinated and optimized in the entire link of video stream acquisition, processing, encoding and transmission in a dynamically changing network environment. While ensuring video stability, the transmission robustness and picture quality in live streaming scenarios are improved, meeting the dual requirements of high robustness and high quality in professional live streaming scenarios.
[0145] Optionally, as an embodiment, the processor 910 is specifically configured to perform optical image stabilization on the i-th original video frame according to the i-th motion intensity index to obtain an i-th image frame after optical image stabilization; wherein the compensation intensity of the optical image stabilization is positively correlated with the i-th motion intensity index; and to perform electronic image stabilization on the i-th image frame after optical image stabilization according to the i-th motion intensity index and the compensation data of the optical image stabilization to obtain an i-th stable video frame.
[0146] Optionally, as an embodiment, the processor 910 is specifically configured to generate an i-th encoding priority score based on the i-th motion intensity index and the content features of the i-th stable video frame; wherein the i-th encoding priority score is used to characterize the content importance of the i-th stable video frame; and to encode the i-th stable video frame according to the i-th encoding priority score to obtain i-th encoded data.
[0147] Optionally, as an embodiment, the processor 910 is specifically configured to generate an i-th encoding priority score based on the i-th motion intensity index and the content characteristics of the i-th stable video frame; wherein the i-th encoding priority score is used to characterize the content importance of the i-th stable video frame; calculate a transmission strategy decision value based on network state parameters and the i-th encoding priority score; wherein the transmission strategy decision value is used to characterize the degree of matching between network supply capacity and content transmission demand; determine the i-th transmission mode as a high throughput mode when the transmission strategy decision value is greater than or equal to a preset decision threshold; wherein the high throughput mode is used to aggregate multiple network links to transmit complete encoded data; determine the i-th transmission mode as a high robustness mode when the transmission strategy decision value is less than the preset decision threshold; wherein the high robustness mode is used to perform hierarchical transmission and redundancy protection for encoded data.
[0148] In some embodiments, the processor 910 is configured to receive data packets from a video stream from an electronic device; wherein the data packets include: core feature data, secondary feature data, redundancy check data, motion state data, transmission mode identifier, and encoding priority score; when the transmission mode identifier indicates a high robustness mode, forward error correction decoding recovery processing is performed on the core feature data based on the redundancy check data to obtain recovered core feature data; a fusion weight is determined based on the recovery completeness of the core feature data and the encoding priority score; wherein the fusion weight is used to characterize the contribution ratio of the core feature data, secondary feature data, and motion state data during video reconstruction; and the recovered core feature data, secondary feature data, and motion state data are weighted and fused according to the fusion weight to reconstruct the video stream.
[0149] It should be understood that, in this embodiment, the input unit 904 may include a graphics processing unit (GPU) 9041 and a microphone 9042. The GPU 9041 processes image data of still images or videos obtained by an image capture device (such as a camera) in video capture mode or image capture mode. The display unit 906 may include a display panel 9061, which may be configured in the form of a liquid crystal display, an organic light-emitting diode, or the like. The user input unit 907 includes at least one of a touch panel 9071 and other input devices 9072. The touch panel 9071 is also called a touch screen. The touch panel 9071 may include a touch detection device and a touch controller. Other input devices 9072 may include, but are not limited to, physical keyboards, function keys (such as volume control buttons, power buttons, etc.), trackballs, mice, and joysticks, which will not be described in detail here.
[0150] The memory 909 can be used to store software programs and various data. The memory 909 may primarily include a first storage area for storing programs or instructions and a second storage area for storing data. The first storage area may store the operating system, application programs or instructions required for at least one function (such as sound playback, image playback, etc.). Furthermore, the memory 909 may include volatile memory or non-volatile memory, or both. The non-volatile memory may be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. Volatile memory can be random access memory (RAM), static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDRSDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous linked dynamic random access memory (Synchlink DRAM, SLDRAM), and direct memory bus RAM (DRRAM). The memory 909 in the embodiments of this application includes, but is not limited to, these and any other suitable types of memory.
[0151] Processor 910 may include one or more processing units; optionally, processor 910 integrates an application processor and a modem processor, wherein the application processor mainly handles operations involving the operating system, user interface, and applications, and the modem processor mainly handles wireless communication signals, such as a baseband processor. It is understood that the aforementioned modem processor may also not be integrated into processor 910.
[0152] This application also provides a readable storage medium storing a program or instructions. When the program or instructions are executed by a processor, they implement the various processes of the above-described video stream processing method embodiments and achieve the same technical effect. To avoid repetition, they will not be described again here.
[0153] The processor is the processor in the electronic device described in the above embodiments. The readable storage medium includes computer-readable storage media, such as computer read-only memory (ROM), random access memory (RAM), magnetic disk, or optical disk.
[0154] This application embodiment also provides a chip, which includes a processor and a communication interface. The communication interface is coupled to the processor. The processor is used to run programs or instructions to implement the various processes of the above-described video stream processing method embodiments and can achieve the same technical effect. To avoid repetition, it will not be described again here.
[0155] It should be understood that the chip mentioned in the embodiments of this application may also be referred to as a system-on-a-chip, system chip, chip system, or system-on-a-chip, etc.
[0156] This application provides a computer program product, which is stored in a storage medium and executed by at least one processor to implement the various processes of the video stream processing method embodiments described above, and can achieve the same technical effect. To avoid repetition, it will not be described again here.
[0157] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element. Furthermore, it should be noted that the scope of the methods and apparatuses in the embodiments of this application is not limited to performing functions in the order shown or discussed, but may also include performing functions substantially simultaneously or in the reverse order, depending on the functions involved. For example, the described methods may be performed in a different order than described, and various steps may be added, omitted, or combined. Additionally, features described with reference to certain examples may be combined in other examples.
[0158] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a computer software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal (which may be a mobile phone, computer, server, or network device, etc.) to execute the methods described in the various embodiments of this application.
[0159] The embodiments of this application have been described above with reference to the accompanying drawings. However, this application is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of this application without departing from the spirit and scope of the claims, and all of these forms are within the protection scope of this application.
Claims
1. A video stream processing method, applied to electronic devices, characterized in that, The method includes: Acquire video stream and simultaneously acquire motion state data stream of the electronic device; wherein, the video stream includes N original video frames, and the motion state data stream includes the i-th motion state data corresponding to the i-th original video frame, 1≤i≤N, where N is a positive integer greater than 1; Based on the i-th motion state data, an i-th motion intensity index is generated; wherein, the i-th motion intensity index is used to characterize the intensity of motion of the electronic device within the time window corresponding to the i-th original video frame; Based on the i-th motion intensity index, the i-th original video frame is subjected to anti-shake processing to obtain the i-th stable video frame; Based on the i-th motion intensity index and the content characteristics of the i-th stable video frame, the i-th stable video frame is encoded to obtain the i-th encoded data; Based on the network status parameters, the i-th motion intensity index, and the content characteristics of the i-th stable video frame, the i-th transmission mode corresponding to the i-th stable video frame is determined; According to the i-th transmission mode, the i-th encoded data is streamed.
2. The method according to claim 1, characterized in that, The step of performing image stabilization processing on the i-th original video frame according to the i-th motion intensity index to obtain the i-th stable video frame includes: Based on the i-th motion intensity index, optical image stabilization is applied to the i-th original video frame to obtain the i-th image frame after optical image stabilization; wherein, the compensation intensity of the optical image stabilization is positively correlated with the i-th motion intensity index; Based on the i-th motion intensity index and the compensation data of the optical image stabilization process, electronic image stabilization is applied to the i-th image frame that has undergone optical image stabilization to obtain the i-th stable video frame.
3. The method according to claim 1, characterized in that, The step of encoding the i-th stable video frame based on the i-th motion intensity index and the content features of the i-th stable video frame to obtain the i-th encoded data includes: Based on the i-th motion intensity index and the content features of the i-th stable video frame, an i-th encoding priority score is generated; wherein, the i-th encoding priority score is used to characterize the content importance of the i-th stable video frame; Based on the i-th encoding priority score, the i-th stable video frame is encoded to obtain the i-th encoded data.
4. The method according to claim 1, characterized in that, The step of determining the i-th transmission mode corresponding to the i-th stable video frame based on network state parameters, the i-th motion intensity index, and the content characteristics of the i-th stable video frame includes: Based on the i-th motion intensity index and the content features of the i-th stable video frame, an i-th encoding priority score is generated; wherein, the i-th encoding priority score is used to characterize the content importance of the i-th stable video frame; Based on the network state parameters and the i-th encoding priority score, a transmission strategy decision value is calculated; wherein, the transmission strategy decision value is used to characterize the degree of matching between network supply capacity and content transmission demand; If the transmission strategy decision value is greater than or equal to a preset decision threshold, the i-th transmission mode is determined to be a high throughput mode; wherein, the high throughput mode is used to aggregate multiple network links to transmit complete encoded data; If the transmission strategy decision value is less than the preset decision threshold, the i-th transmission mode is determined to be a high robustness mode; wherein, the high robustness mode is used to perform hierarchical transmission and redundancy protection of encoded data.
5. A video stream processing method, applied to a receiving device, characterized in that, The method includes: Receive data packets from a video stream from an electronic device; wherein the data packets include: core feature data, secondary feature data, redundancy check data, motion state data, transmission mode identifier, and encoding priority score; When the transmission mode identifier indicates a high robustness mode, the core feature data is subjected to forward error correction decoding recovery processing based on the redundant check data to obtain the recovered core feature data. The fusion weights are determined based on the completeness of the restoration of the core feature data and the encoding priority score; wherein, the fusion weights are used to characterize the contribution ratio of the core feature data, secondary feature data and motion state data respectively during video reconstruction. Based on the fusion weights, the recovered core feature data, secondary feature data, and motion state data are weighted and fused to reconstruct the video stream.
6. A video stream processing device, applied to electronic devices, characterized in that, The device includes: The acquisition module is used to acquire video streams and synchronously acquire motion state data streams of the electronic device; wherein, the video stream includes N original video frames, and the motion state data stream includes the i-th motion state data corresponding to the i-th original video frame, 1≤i≤N, where N is a positive integer greater than 1; The generation module is used to generate an i-th motion intensity index based on the i-th motion state data; wherein the i-th motion intensity index is used to characterize the intensity of motion of the electronic device within the time window corresponding to the i-th original video frame; The image stabilization module is used to perform image stabilization processing on the i-th original video frame according to the i-th motion intensity index to obtain the i-th stable video frame; The encoding processing module is used to encode the i-th stable video frame according to the i-th motion intensity index and the content features of the i-th stable video frame to obtain the i-th encoded data; The first determining module is used to determine the i-th transmission mode corresponding to the i-th stable video frame based on network status parameters, the i-th motion intensity index, and the content characteristics of the i-th stable video frame. The transmission module is used to stream the i-th encoded data according to the i-th transmission mode.
7. The apparatus according to claim 6, characterized in that, The image stabilization module is specifically used to perform optical image stabilization on the i-th original video frame according to the i-th motion intensity index to obtain the i-th image frame after optical image stabilization; wherein, the compensation intensity of the optical image stabilization is positively correlated with the i-th motion intensity index; and to perform electronic image stabilization on the i-th image frame after optical image stabilization according to the i-th motion intensity index and the compensation data of the optical image stabilization to obtain the i-th stable video frame.
8. The apparatus according to claim 6, characterized in that, The encoding processing module is specifically used to generate an i-th encoding priority score based on the i-th motion intensity index and the content features of the i-th stable video frame; wherein, the i-th encoding priority score is used to characterize the content importance of the i-th stable video frame; and to encode the i-th stable video frame according to the i-th encoding priority score to obtain the i-th encoded data.
9. The apparatus according to claim 6, characterized in that, The first determining module is specifically configured to generate an i-th encoding priority score based on the i-th motion intensity index and the content characteristics of the i-th stable video frame; wherein the i-th encoding priority score is used to characterize the content importance of the i-th stable video frame; calculate a transmission strategy decision value based on network state parameters and the i-th encoding priority score; wherein the transmission strategy decision value is used to characterize the degree of matching between network supply capacity and content transmission demand; determine the i-th transmission mode as a high throughput mode if the transmission strategy decision value is greater than or equal to a preset decision threshold; wherein the high throughput mode is used to aggregate multiple network links to transmit complete encoded data; determine the i-th transmission mode as a high robustness mode if the transmission strategy decision value is less than the preset decision threshold; wherein the high robustness mode is used to perform hierarchical transmission and redundancy protection for encoded data.
10. A video stream processing apparatus, applied to a receiving end device, characterized in that, The device includes: A receiving module is used to receive data packets from a video stream from an electronic device; wherein the data packets include: core feature data, secondary feature data, redundancy check data, motion state data, transmission mode identifier, and encoding priority score; The decoding processing module is used to perform forward error correction decoding recovery processing on the core feature data based on the redundant check data when the transmission mode identifier indicates a high robustness mode, so as to obtain the recovered core feature data. The second determining module is used to determine the fusion weight based on the restoration completeness of the core feature data and the encoding priority score; wherein, the fusion weight is used to characterize the contribution ratio of the core feature data, secondary feature data and motion state data respectively during video reconstruction; The reconstruction module is used to perform weighted fusion of the recovered core feature data, the secondary feature data, and the motion state data according to the fusion weights, and reconstruct the video stream.