Video workflow data processing system

By designing a layered image framework, the video quality is dynamically adjusted according to the communication link status, which solves the problem of unstable video data transmission, achieves the integrity and smoothness of video transmission, and improves the user experience.

CN119996737BActive Publication Date: 2025-12-26SHANGHAI YUNTI INFORMATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510483413.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-17
Publication Date
2025-12-26
Estimated Expiration
2045-04-17

AI Technical Summary

Technical Problem

Video data transmission is characterized by large data volumes, high real-time requirements, and network instability due to network complexity, which affects video playback quality and user experience.

Method used

A layered image frame design is adopted, in which each workflow image frame contains at least one layer of image frame. The resolution of the layered image frame is related to the communication link status, and the video quality is dynamically adjusted to adapt to network conditions.

Benefits of technology

It ensures the integrity and smoothness of video transmission, reduces image synthesis time, adapts to the image processing and display needs of terminal devices, effectively utilizes network resources, and improves user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119996737B_ABST
    Figure CN119996737B_ABST
Patent Text Reader

Abstract

The application provides a video workflow data processing system. The system sends the collected video data to the server through the corresponding communication link by each terminal device in the terminal device cluster to generate a video data set, and generates video workflow data according to the link state of each communication link and the video data set, wherein each frame of workflow image includes at least one layer of image frame, and the resolution of each layered image frame is associated with the corresponding link state, so that the system can automatically adjust the video quality according to the communication link state between the server and the terminal device, and ensure the integrity and smoothness of video transmission between the server and each terminal device.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of data processing, and particularly relates to a video workflow data processing system. BACKGROUND

[0002] With the rapid development of information technology, video data is increasingly widely used in various fields, such as online meetings, security monitoring, online education, remote medical treatment, intelligent transportation, etc. Real-time transmission and processing of video data has become one of the core technologies supporting these applications. However, in actual applications, the transmission of video data faces many challenges.

[0003] Video data has the characteristics of large data volume and high real-time requirements. High-resolution, high-frame-rate video data requires a large amount of network bandwidth, and the complexity and uncertainty of the network environment (such as network congestion, packet loss, delay, etc.) often lead to instability of video data transmission, thereby affecting the video playback quality and user experience.

[0004] In view of the challenges in video data transmission, existing technical solutions mainly optimize from the following aspects:

[0005] Compression encoding technology: by adopting efficient video compression encoding algorithm, the size of video data is reduced, and the demand for network bandwidth is reduced. However, compression encoding technology can improve the transmission efficiency of video data, but it will also reduce the video quality to a certain extent, especially in the case of high compression ratio.

[0006] Adaptive bit rate adjustment technology: according to the real-time change of network bandwidth, the bit rate of video data is dynamically adjusted to adapt to different network environments. However, this technology usually needs frequent communication and negotiation between the server and the terminal device, which increases the complexity and delay of the system. SUMMARY

[0007] The present application provides a video workflow data processing system, which can automatically adjust the video quality according to the communication link state between the server and the terminal device, ensuring the integrity and smoothness of video transmission between the server and each terminal device.

[0008] In a first aspect, the present application provides a video workflow data processing system, comprising: a terminal device cluster and a server;

[0009] Each terminal device in the terminal device cluster sends the collected video data to the server through the corresponding communication link to generate a video data set, and the video data set includes image data sequences at each time node;

[0010] The server generates video workflow data according to the link states of the respective communication links and the video data set, each frame of workflow image in the video workflow data comprising at least one layer of image frame, and the resolution of each sub-layer image frame being associated with the corresponding link state;

[0011] The server transmits the corresponding workflow image in the video workflow data to the corresponding terminal device through the respective communication links.

[0012] In the above scheme, each terminal device in the terminal device cluster transmits the collected video data to the server through the corresponding communication link to generate a video data set, and generates video workflow data according to the link states of the respective communication links and the video data set, wherein each frame of workflow image comprises at least one layer of image frame, and the resolution of each sub-layer image frame is associated with the corresponding link state, so that the system can automatically adjust the video quality according to the communication link state between the server and the terminal device, ensuring the integrity and smoothness of video transmission between the server and each terminal device. It is worth mentioning that the video workflow data generated by the above-mentioned layered image frame mode can reduce the image synthesis time compared with the traditional server directly transmitting the final image, and can better adapt to the image processing and image display requirements of the corresponding terminal device.

[0013] Optionally, the data structure of the video workflow data comprises a time guide part and an image data part;

[0014] The time guide part is used to accommodate each time node arranged in time;

[0015] The image data part comprises a workflow image under each time node, and each sub-layer image frame of the workflow image has a different resolution, and each sub-layer image frame is used to accommodate at least one image data in the image data sequence.

[0016] In the above scheme, by designing the layered image frame, the system can dynamically adjust the resolution of each sub-layer image frame in the workflow image according to the link state of the communication link. For example, in the case of good link state, the system can select a high-resolution image frame for transmission to provide a higher-quality video experience; while in the case of poor link state, a low-resolution image frame can be selected for transmission to ensure the smoothness of the video. And each sub-layer image frame is used to accommodate at least one image data in the image data sequence, which can realize different resolution processing for different image data to ensure the quality of the transmitted workflow image while adapting to the specific link state to ensure the smoothness of the transmission.

[0017] Optionally, the number of layers of the layered images included in the workflow image is associated with the corresponding link state.

[0018] In the above scheme, by associating the number of layers of the layered images in the workflow image with the link state, the system can perceive and adapt to different network conditions in real time, where the link state, such as bandwidth, delay and packet loss rate, is a key factor affecting the quality of video transmission.

[0019] Optionally, the number of layers of the layered images included in the workflow image dynamically changes with the change of the link state.

[0020] In the above scheme, by dynamically adjusting the number of layers of the layered images, the system can more effectively utilize network resources. When the link bandwidth is sufficient, the system can select high-resolution layered images for transmission to provide a better video experience.

[0021] Optionally, each layer image frame in the workflow image includes a sequence of accommodation areas corresponding to the number of terminal devices in the terminal device cluster, and each accommodation area in the sequence of accommodation areas is used to correspond to one of the image data in the sequence of image data.

[0022] In the above scheme, by designing each layer image frame in the workflow image to contain a sequence of accommodation areas corresponding to the number of terminal devices, the system realizes the mapping between video data and terminal devices. Each accommodation area corresponds to the image data of one terminal device, ensuring the accuracy and consistency of video data in the process of processing, transmission and display, not only simplifying the data processing process, but also improving the efficiency and accuracy of data processing.

[0023] Optionally, after the server sends the corresponding workflow image in the video workflow data to the corresponding terminal device through each communication link, it further includes:

[0024] If the number of image frames in the workflow image is not unique, the terminal device aligns each image frame in the workflow image to generate and display the corresponding target image.

[0025] In the above scheme, by position alignment processing, the terminal device can accurately superimpose multiple image frames to form a complete target image. This step is crucial to ensure the integrity and accuracy of the video content. In the video workflow, different levels of image frames may contain image data of different resolutions or different perspectives. Through position alignment, these image data can be accurately combined to generate a high-quality fused image. It is worth mentioning that each image frame in the above workflow image can be partially filled with corresponding image data, and the unfilled area is set to empty.

[0026] Optionally, the link state is a link state of downlink transmission of the server to the corresponding terminal device.

[0027] In the above scheme, by focusing on the link state of the downlink transmission of the server to the terminal device, the system can monitor the key indicators such as the bandwidth, delay and packet loss rate of the downlink in real time. This real-time monitoring mechanism enables the system to quickly perceive the changes in the link state and respond in a timely manner. For example, when detecting insufficient downlink bandwidth, the system can dynamically adjust the resolution of the image frames in the video stream to ensure smooth transmission of the video content. Based on the real-time monitoring of the downlink state, the system can optimize the transmission strategy of the video stream. In the case of good link state, the system can choose higher video code rate and resolution image frames to provide higher quality video experience. While in the case of poor link state, the system can choose to reduce the code rate or resolution of the image frames to reduce the transmission delay and packet loss rate, ensuring the continuity and stability of the video stream. This dynamic transmission strategy adjustment based on the link state can significantly improve the efficiency and quality of video transmission.

[0028] In a second aspect, the present application provides a video workflow data processing method applied to a video workflow data processing system, the system comprising: a terminal device cluster and a server; the method comprising:

[0029] Each terminal device in the terminal device cluster sends the collected video data to the server through the corresponding communication link to generate a video data set, the video data set comprising a sequence of image data at each time node;

[0030] The server generates video workflow data according to the link state of each communication link and the video data set, each frame of workflow image in the video workflow data comprising at least one layer of image frame, the resolution of each layered image frame being associated with the corresponding link state;

[0031] The server sends the corresponding workflow image in the video workflow data to the corresponding terminal device through each communication link.

[0032] In a third aspect, the present application provides an electronic device comprising:

[0033] a processor; and,

[0034] a memory for storing executable instructions of the processor;

[0035] wherein the processor is configured to execute any one of the possible methods described in the first aspect by executing the executable instructions.

[0036] In a fourth aspect, the present application provides a computer readable storage medium, wherein computer execution instructions are stored in the computer readable storage medium, and the computer execution instructions are used to implement any possible method in the first aspect when executed by a processor.

[0037] The video workflow data processing system provided by the present application can automatically adjust the video quality according to the communication link state between the server and the terminal device, and ensure the integrity and fluency of the video transmission between the server and the terminal device. BRIEF DESCRIPTION OF DRAWINGS

[0038] The accompanying drawings, which are incorporated herein and constitute part of the specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application.

[0039] Figure 1 is a structural schematic diagram of a video workflow data processing system according to an example embodiment of the present application;

[0040] Figure 2 is a flow schematic diagram of a video workflow data processing method according to an example embodiment of the present application;

[0041] Figure 3 is a mapping schematic diagram of a layered image frame and an image data sequence according to an example embodiment of the present application;

[0042] Figure 4 is a structural schematic diagram of a layered image frame according to an example embodiment of the present application;

[0043] Figure 5 is a structural schematic diagram of a layered image frame according to an example embodiment of the present application;

[0044] Figure 6 is a structural schematic diagram of a layered image frame according to an example embodiment of the present application;

[0045] Figure 7 is a structural schematic diagram of a layered image frame according to an example embodiment of the present application;

[0046] Figure 8 is a structural schematic diagram of an electronic device according to an example embodiment of the present application.

[0047] The specific embodiments of the application have been shown and described in the above drawings and specification, it is to be understood that the application is not limited to the embodiments disclosed, but rather can be practiced with modification and alteration within the scope of the appended claims. Accordingly, the specification and drawings are to be regarded as illustrative in nature and not as restrictive. DETAILED DESCRIPTION

[0048] The illustrative embodiments described herein will be made with reference to exemplary embodiments illustrated in the drawings. Like numbers refer to like elements throughout. The following description is made in connection with the illustrated embodiments, but is not intended to be limited to the illustrative embodiments. Rather, the scope of the application is limited only by the appended claims.

[0049] To solve the above problems, the embodiments provided by the application propose a video workflow data processing system, which mainly includes two components of terminal device cluster and server. The terminal device cluster is composed of multiple devices with video data acquisition function, such as cameras, video recording devices, etc. They are connected and data transmitted with the server through their respective communication links. The server as the center of data processing and transmission is responsible for receiving video data from the terminal device cluster, and generating video workflow data according to the link state of the communication link. The specific technical concept is as follows:

[0050] Hierarchical image frame design: the system adopts hierarchical image frame design, and each frame of workflow image contains at least one layer of image frame, and the resolution of each hierarchical image frame is associated with the corresponding link state. Specifically, the server dynamically adjusts the resolution of each hierarchical image frame in the workflow image according to the real-time monitored communication link state (such as bandwidth, delay, packet loss rate, etc.). When the link state is good, high-resolution image frame is selected to provide high-quality video experience; when the link state is poor, low-resolution image frame is selected to ensure the smoothness of the video.

[0051] Data differentiated transmission: the system dynamically adjusts the structure and resolution of the generated video workflow image according to the link state and display capability of the terminal device. For terminal devices with good link state, high-quality integrated images are provided; for terminal devices with poor link state, layered display images are provided to reduce the transmission and display burden. This data differentiated transmission strategy not only improves the user experience, but also effectively avoids resource waste.

[0052] Efficient data processing and transmission: By adopting a hierarchical image frame design and a data differential transmission strategy, the system can more effectively utilize network resources, reducing unnecessary data transmission and processing burden. At the same time, the system also has strong data processing capability, which can analyze, process and optimize the received video data in real time, ensuring the integrity and smoothness of video transmission.

[0053] Audio signal integrated processing: The system also integrates audio signal detection and processing functions. When audio signals are detected in the target video data, the corresponding image data is configured to the image frame with the highest resolution to ensure clear display of the content. This design is particularly important in scenarios such as multi-person online meetings, significantly improving user communication efficiency and experience.

[0054] Figure 1 is a structural diagram of a video workflow data processing system according to an example embodiment. As shown in Figure 1 the embodiment provides a video workflow data processing system 100, which includes a terminal device cluster 110 and a server 120.

[0055] Among them, the terminal device cluster 110 is composed of multiple terminal devices, each of which has video data acquisition function, such as camera, video recording device, etc. These terminal devices are connected and data transmitted with the server 120 through their respective communication links (such as wired network, Wi-Fi, 4G / 5G, etc.).

[0056] And the server 120 as the center of data processing and transmission, responsible for receiving video data from the terminal device cluster 110, and generating video workflow data according to the link state of the communication link. The above-mentioned server 120 has strong data processing capability, which can analyze, process and optimize the received video data in real time.

[0057] In order to further illustrate the architecture and working principle of the video workflow data processing system in the present application, the following provides several possible application scenarios:

[0058] In one possible application scenario, multi-person online video conference is carried out, and each participant sends the collected video data to the server through the camera on the terminal device and the respective communication link. Each terminal device (camera) in the terminal device cluster continuously collects video data. These video data are transmitted to the server in real time through the respective communication link (such as Wi-Fi, 4G / 5G, etc.).

[0059] The server receives video data from the terminal device cluster and organizes and stores these data according to time nodes to form a video data set, which contains image data sequences at each time node.

[0060] The server generates video workflow data according to the link states (such as bandwidth, delay, packet loss rate, etc.) of each communication link and the video data set. Each frame of workflow image in the video workflow data includes at least one image frame, and the resolution of each sub-layer image frame is associated with the corresponding link state. For example, when the communication link state of a certain terminal device is good, the server generates a high-resolution image frame for the terminal device; and when the link state is poor, a low-resolution image frame is generated.

[0061] The server sends the corresponding workflow images in the video workflow data to the corresponding terminal devices through each communication link. The terminal devices receive and display these workflow images, thereby realizing smooth online video conference.

[0062] In addition, the video workflow data processing system in this application can also be applied to large-scale activities such as sports events, concerts, and exhibitions, which need to monitor and record every corner of the activities in real time. By deploying multiple terminal devices (such as cameras) and connecting them to the server, comprehensive monitoring and recording of the activity site can be achieved.

[0063] Figure 2 is a flowchart of a video workflow data processing method according to an example embodiment of the present application. As shown in Figure 2 The video workflow data processing method provided by the embodiment includes:

[0064] S201, each terminal device in the terminal device cluster sends the collected video data to the server through the corresponding communication link to generate a video data set.

[0065] In this step, each terminal device in the terminal device cluster sends the collected video data to the server through the corresponding communication link to generate a video data set, and the video data set includes image data sequences at each time node.

[0066] Specifically, the terminal device cluster includes multiple terminal devices, which can be deployed in different geographical locations or application scenarios. Each terminal device is equipped with a video acquisition module, such as a camera, for real-time capture of video pictures. The collected video data exists in the form of continuous image frames. Each terminal device establishes a connection with the server through its exclusive communication link. These communication links can be wired networks (such as Ethernet) or wireless networks (such as Wi-Fi, 4G / 5G mobile networks), etc. The terminal device encodes and packages the collected image data into data packets in real time, and then sends them to the server through its communication link. The server receives video data packets from each terminal device and organizes them into a video data set according to the time stamp and other information.

[0067] S202, the server generates video workflow data according to the link state of each communication link and the video data set.

[0068] In this step, the server generates video workflow data according to the link state of each communication link and the video data set. Each frame of workflow image in the video workflow data includes at least one image frame. The resolution of each sub-layer image frame is associated with the corresponding link state. The resolution of the image data contained in each layer image frame is consistent with the resolution of the sub-layer image frame. It can be understood that the resolution of the sub-layer image frame is used to constrain the resolution of the image data contained therein. It is worth noting that the link state is the link state of the server downloading to the corresponding terminal device.

[0069] Specifically, the video data set is an ordered data structure that contains image data sequences at each time node. The time node can be a fixed time interval or a dynamic time node based on event triggering. The image data sequence is a set of video frames collected from all terminal devices at these time nodes.

[0070] The server monitors the communication link state between each terminal device in real time. The link state includes bandwidth, delay, packet loss rate and other key indicators, which directly affect the transmission quality and efficiency of video data. The server obtains link state information by sending probe packets regularly or using existing network protocols. According to the monitored link state and the video data set, the server dynamically generates video workflow data. Each frame of workflow image in the video workflow data is designed to contain at least one image frame. The resolution of these image frames is associated with the corresponding link state: when the link state is good, the image data is configured to a high-resolution image frame; when the link state is poor, the image data is configured to a low-resolution image frame. The server also dynamically adjusts the number of image frames according to the stability of the link state to adapt to the configuration of different resolution image data under different link states.

[0071] Optionally, the data structure of the above-mentioned video workflow data includes a time guide part and an image data part. The time guide part is used to contain each time node arranged in time.

[0072] The image data part includes workflow images at each time node. Each sub-layer image frame of the workflow image has a different resolution. Each sub-layer image frame is used to contain at least one image data in the image data sequence.

[0073] Specifically, the main role of the time guide part is to accommodate each time node arranged according to time. These time nodes represent the key time points in the video data, which are used to organize and index the video data in chronological order. Through the time guide part, the system can quickly locate the video data at any time point, thereby realizing accurate playing and control of the video. In specific implementation, the time guide part can be implemented by using data structures such as linked list, array or database. These data structures can efficiently store and retrieve time node information, ensuring that the system still maintains good performance when processing a large amount of video data. The image data part contains the workflow images under each time node. These workflow images are generated by the server according to the received video data and the link state of the communication link. Each workflow image is composed of at least one image frame, and these image frames have different resolutions for accommodating at least one image data in the image data sequence.

[0074] Each workflow image is composed of multiple layers of image frames, and the resolutions of these image frames are associated with the corresponding link state. When the communication link state is good, the system can select high-resolution image frames to transmit high-quality video data; when the link state is poor, low-resolution image frames can be selected to ensure the smoothness of the video.

[0075] In order to realize the design of layered image frames, the system needs to maintain an image frame level table. This table records the number of layers of workflow images under each time node, the resolutions of each layer, and the corresponding link state information. When the server receives new video data, it will select appropriate image frames from the image frame level table according to the current link state to accommodate these data.

[0076] After determining the layered image frames, the system needs to accommodate the received image data into these frames. Since each image frame has different resolutions, the system needs to scale, crop and other image processing according to the original resolution of the image data and the resolution of the target frame, to ensure that the image data can be correctly filled into the frame. In specific implementation, the system can use image processing algorithms to realize the scaling and cropping of image data. These algorithms can efficiently process a large amount of image data and generate image frames that meet the requirements. At the same time, the system can also perform other processing on the image data according to actual needs, such as enhancement, filtering, etc., to improve the quality and visual effect of the video.

[0077] wherein, Figure 3 is a mapping diagram of layered image frames and image data sequences according to an example embodiment of the present application, as shown in Figure 3As shown, each layer image frame in the workflow image includes a sequence of accommodation zones corresponding to the number of terminal devices in the terminal device cluster 110, and each accommodation zone in the sequence of accommodation zones is used to correspond to one of the image data in the sequence of image data.

[0078] By designing each layer image frame in the workflow image to contain a sequence of accommodation zones corresponding to the number of terminal devices, the system achieves the mapping between video data and terminal devices. Each accommodation zone corresponds to the image data of one terminal device, ensuring the accuracy and consistency of video data in the process of processing, transmission and display, not only simplifying the process of data processing, but also improving the efficiency and accuracy of data processing. Moreover, since each accommodation zone corresponds to the image data of one terminal device, the system can process the video data of multiple terminal devices in parallel. This parallel processing mechanism greatly improves the processing speed of video data, enabling the system to respond more quickly to requests from terminal devices. In addition, each layer image frame in the workflow image contains a sequence of accommodation zones corresponding to the number of terminal devices, allowing the system to flexibly adapt to different numbers and types of terminal devices. Regardless of the number of terminal devices, the system can adapt to new requirements by adjusting the number and layout of accommodation zones. In addition, in terms of video data display, the system can select appropriate image data frames for display according to the resolution and display capabilities of the terminal devices. Furthermore, each accommodation zone corresponds to the image data of one terminal device, which ensures the integrity and consistency of video data. During video data processing, the system can monitor the data state of each accommodation zone in real time, and timely detect and handle data loss or damage. At the same time, during video data transmission and display, the system can ensure the integrity and consistency of data through the verification and validation mechanism, avoiding video quality problems caused by data errors.

[0079] Finally, designing each layer image frame in the workflow image to contain a sequence of accommodation zones corresponding to the number of terminal devices also improves the scalability and maintainability of the system. As the number of terminal devices increases or decreases, the system can easily adapt to new requirements by adjusting the number and layout of accommodation zones without the need for large-scale modification or upgrade of the system. In addition, in terms of system maintenance, since each accommodation zone corresponds to a specific terminal device, the system can more easily troubleshoot and repair work, improving the stability and reliability of the system.

[0080] Optionally, the number of layers of the layered images included in the workflow images is associated with the corresponding link status. By associating the number of layers of the layered images in the workflow images with the link status, the system can perceive and adapt to different network conditions in real time, where the link status, such as bandwidth, delay, and packet loss rate, is a key factor affecting the quality of video transmission. When the link status is good, the system can configure the image data in the video content to a higher resolution image frame, and when the link status is poor, it can configure part of the image data with higher priority to a high-resolution image frame, and part of the image data with lower priority to a low-resolution image frame, thereby reducing the bandwidth demand of video transmission and ensuring smooth display of the video. The design of the number of layers of the layered images associated with the link status also realizes the dynamic adjustment and optimization of the video resolution. During video transmission, the system can adjust the resolution of each layered image in the workflow image in real time according to the change of the link status.

[0081] Further, the number of layers of the layered images included in the workflow images dynamically changes with the change of the link status. By dynamically adjusting the number of layers of the layered images, the system can more effectively utilize network resources. When the link bandwidth is sufficient, the system can select high-resolution layered images for transmission to provide a more high-definition video experience. When the link bandwidth is limited, the system can select low-resolution layered images to reduce the amount of data transmitted and ensure continuous playback of the video. This dynamic resolution adjustment mechanism enables the system to provide adaptive video transmission effects under various network conditions. In addition, the design of associating the number of layers of the layered images with the link status also makes the system more flexible and scalable. The system can flexibly adjust the number of layers and resolution of the layered images according to actual needs to adapt to different application scenarios and terminal device requirements.

[0082] S203, the server sends the corresponding workflow images in the video workflow data to the corresponding terminal devices through each communication link.

[0083] In this step, the server sends the corresponding workflow images in the video workflow data to the corresponding terminal devices through each communication link.

[0084] Specifically, the server encapsulates the generated video workflow data into data packets and sends them to each terminal device through the corresponding communication link. During transmission, the server dynamically adjusts the transmission rate and priority of the data packets according to the link status to ensure the smoothness and integrity of video transmission.

[0085] After the terminal device receives the video workflow data packet from the server, it performs decoding and recombination operations to recover the original workflow image. If the workflow image contains multiple image frames (i.e., the number of layers is greater than 1), the terminal device will perform alignment processing operations. Alignment processing refers to precisely aligning multiple image frames in spatial position so as to combine them into a complete target image for display. Alignment processing can involve image transformation techniques such as translation, rotation, scaling, and interpolation algorithms. The terminal device selects appropriate workflow images for display or storage based on its display capabilities and user needs.

[0086] If the number of image frames in the workflow image is not unique, the terminal device will perform alignment processing on each image frame in the workflow image to generate and display the corresponding target image. Through position alignment processing, the terminal device can accurately superimpose multiple image frames to form a complete target image. This step is crucial for ensuring the integrity and accuracy of video content. In video workflows, different levels of image frames can contain image data of different resolutions or different perspectives. Through position alignment, these image data can be accurately combined to generate high-quality fused images. It is worth noting that each image frame in the above workflow image can be partially filled with corresponding image data, and the unfilled areas are set to empty.

[0087] In the above scheme, each terminal device in the terminal device cluster sends the collected video data to the server through the corresponding communication link to generate a video data set and generates a video workflow data based on the link state of each communication link and the video data set, wherein each frame of workflow image includes at least one layer of image frame, and the resolution of each layered image frame is associated with the corresponding link state, so that the system can automatically adjust the video quality according to the communication link state between the server and the terminal device, ensuring the integrity and smoothness of video transmission between the server and each terminal device.

[0088] Based on the above embodiment, in one possible implementation, the first terminal device in the terminal device cluster sends the collected first video data to the server through the first communication link, and the second terminal device sends the collected second video data to the server through the second communication link. The video data set includes a target data sequence at a target time node, and the target data sequence includes first image data of the first video data at the target time node and second image data of the second video data at the target time node.

[0089] Figure 4 is a structural diagram of a layer of image frame according to an example embodiment of the present application. As shown in FIG. 1, the image frame includes a plurality of image data, and each image data is associated with a corresponding time node. The image data can be filled with image data of different resolutions or different perspectives, and the unfilled areas are set to empty. Figure 4As shown, if the first communication link and the second communication link are in the same state, each frame of the video workflow data only includes one layer of image frame, i.e., the first layer of image frame 310, and the first image data and the second image data are accommodated in the one layer of image frame (as shown in the filled part).

[0090] The system identifies whether the first communication link and the second communication link are in the same state by monitoring and evaluating the link states of the first communication link and the second communication link in real time. When it is determined that the states are the same, the system automatically triggers specific video workflow data processing logic. This technical solution effectively avoids the confusion or errors of processing logic caused by the difference in link states, and ensures the consistency and predictability of the system behavior in different link environments. By uniformly processing the case where the link states are the same, the system can more efficiently perform subsequent video workflow data processing tasks. Among them, in the case where the link states of the first communication link and the second communication link are good, the first layer of image frame 310 is configured with high resolution, and in the case where the link states of the first communication link and the second communication link are poor, the first layer of image frame 310 is configured with low resolution.

[0091] In the single-layer image frame, the system ingeniously accommodates the first image data and the second image data. Through advanced data integration technology, the system can seamlessly integrate data from different links into the same image frame while maintaining the integrity and accuracy of the data. Efficient integration and presentation of multi-source data are achieved, providing users with a more comprehensive and comprehensive view of video workflow data. Users can observe data information from different links in the same image frame at the same time.

[0092] If the first communication link and the second communication link are in different states, and the link state of the first communication link is better than that of the second communication link, the first workflow image in the video workflow data includes the first layer of image frame 310 for accommodating the first image data and the second image data (as shown in the filled part). Figure 4 Figure 5 is a structural schematic diagram of a multi-layer image frame according to an example embodiment of the present application. As shown, Figure 5 The second workflow image includes the first layer of image frame 310 for accommodating the first image data and the second layer of image frame 320 for accommodating the second image data, wherein the resolution of the first layer of image frame 310 is higher than that of the second layer of image data, the first workflow image is used to be sent to the first terminal device, and the second workflow image is used to be sent to the second terminal device.

[0093] ​In view of the link state difference, the system adopts a hierarchical image framework design strategy. In the first workflow image, a first-layer image framework is designed to accommodate both the first image data and the second image data; while in the second workflow image, the first-layer image framework is designed to accommodate the first image data, and an additional second-layer image framework is designed to accommodate the second image data. At the same time, the resolution of the first-layer image framework is ensured to be higher than that of the second-layer image framework.

[0094] The hierarchical image framework design realizes differentiated presentation and efficient transmission of data. The high-resolution first-layer image framework ensures the clarity and integrity of key data, while the low-resolution second-layer image framework reduces the occupation of bandwidth resources. This design not only meets the demand for high-quality data transmission, but also takes into account the effective use of system resources.

[0095] According to the link state difference, the system generates first workflow images and second workflow images with different data accommodation structures and resolution characteristics. The first workflow images are directed to the first terminal devices with better link states, providing integrated high-quality data views; while the second workflow images are directed to the second terminal devices with relatively poor link states, providing hierarchical display of data views. The data differentiation presentation technology ensures that terminal devices under different link states can receive data that best suits their transmission capabilities and display needs. This not only improves user experience, but also avoids wasting resources or user dissatisfaction due to excessive data transmission or poor display effect.

[0096] In other words, the system dynamically adjusts the structure and resolution of the generated video workflow images according to the link state and display capability of the terminal device. For the first terminal device with a better link state, an integrated high-quality image is provided; while for the second terminal device with a relatively poor link state, a hierarchical display image is provided to reduce the transmission and display burden. The terminal device adaptability optimization technology significantly improves the compatibility and support capability of the system for different terminal devices. No matter what network environment or display capability the terminal device is in, the system can provide an adaptive data transmission and display scheme, thereby ensuring user experience.

[0097] In addition, through the hierarchical image framework design and data differentiation presentation technology, the system can dynamically adjust the data transmission volume and display quality according to the link state and terminal device requirements. On the premise of ensuring data integrity and clarity, the occupation of system resources is minimized. The system resource efficient utilization technology significantly improves the overall performance and stability of the system. By reducing unnecessary data transmission and display burden, the system can more efficiently utilize computing resources, bandwidth resources and storage resources, thereby improving the response speed and processing capacity of the system. At the same time, this also reduces the operation and maintenance cost and energy consumption level of the system.

[0098] Figure 6 is a structural diagram of a layer image frame according to an example embodiment of the present application. As shown, the third terminal device in the terminal device cluster sends the collected third video data to the server through the third communication link, and the target data sequence includes third image data of the third video data at the target time node. Figure 6 If the first communication link, the second communication link and the third communication link are in the same state, each frame of the workflow image in the video workflow data includes only one layer of image frame, i.e., the first layer image frame 410, and the first image data, the second image data and the third image data are accommodated in the one layer image frame.

[0099] If the first communication link and the second communication link are in the same state, and the link state of the first communication link and the second communication link is better than that of the third communication link, the first workflow image in the video workflow data includes the first layer image frame 410 for accommodating the first image data, the second image data and the third image data (as shown). Figure 7 is a structural diagram of a multi-layer image frame according to an example embodiment of the present application. As shown, Figure 7 the second workflow image includes the first layer image frame 410 for accommodating the first image data and the second image data, and the second layer image frame 420 for accommodating the third image data, wherein the resolution of the first layer image frame 410 is higher than that of the second layer image frame 420, the first workflow image is used to be sent to the first terminal device and the second terminal device, and the second workflow image is used to be sent to the third terminal device. It is worth mentioning that if there are other terminal devices, a third layer image frame 430 can be further set to accommodate corresponding data images.

[0100] The system forms a link state evaluation model by monitoring the link states (such as bandwidth, delay, packet loss rate and the like) of the first communication link, the second communication link and the third communication link in real time. When the link states of the first communication link and the second communication link are consistent and are better than that of the third communication link, the system automatically triggers the differentiated workflow image generation logic based on the link state. In view of the difference in the link states, the system uses the hierarchical image frame technology to design the first workflow image as a single-layer high-resolution frame (the first layer image frame) for accommodating the first image data, the second image data and the third image data, and the second workflow image as a double-layer frame structure, wherein the first layer image frame (high resolution) accommodates the first image data and the second image data, and the second layer image frame (low resolution) accommodates the third image data. This design ensures that when the link state is limited, the continuity of the overall data transmission can still be ensured by reducing the transmission quality (resolution) of part of the image data.

[0101] The first layer image frame adopts a high resolution design, ensuring that the image quality is maintained at a high level when transmitted to the first terminal device and the second terminal device, thereby meeting the high-precision visual processing requirements. For the third communication link with poor link state, the system uses a low-resolution second layer image frame to accommodate the third image data, thereby reducing the data transmission volume by reducing the image resolution, and thereby realizing stable transmission of the third image data under limited link bandwidth.

[0102] The first workflow image contains all the image data (the first image data, the second image data, and the third image data) and adopts a high-resolution first layer image frame, which is suitable for the first terminal device and the second terminal device with good link state. This strategy ensures that the terminal device can receive complete and high-resolution video workflow data under a high-quality link environment. The second workflow image is designed for the third terminal device with limited link state, and realizes differentiated transmission of image data through a hierarchical frame structure. The high-resolution first layer image frame ensures the transmission quality of the first image data and the second image data, while the low-resolution second layer image frame ensures stable transmission of the third image data under limited bandwidth. This strategy effectively balances the data transmission requirements under different link states and improves the overall data transmission efficiency of the system.

[0103] If the first communication link and the second communication link are in the same state, and the link state of the third communication link is better than the link states of the first communication link and the second communication link, then the first workflow image in the video workflow data includes a first layer image frame 410 for accommodating the third image data and a second layer image frame 420 for accommodating the first image data and the second image data, and the second workflow image includes a first layer image frame 410 for accommodating the first image data, the second image data, and the third image data, wherein the resolution of the first layer image frame 410 is higher than that of the second layer image frame 420, the first workflow image is used to be issued to the first terminal device and the second terminal device, and the second workflow image is used to be issued to the third terminal device.

[0104] Through the differentiated workflow image generation and distribution strategy based on the link state, the system can dynamically adjust the transmission mode of image data according to the actual state of different links, thereby maximizing the use of limited link bandwidth resources while ensuring data transmission continuity. The differentiated workflow image design fully considers the performance differences of different terminal devices, ensuring that the terminal device can receive video workflow data suitable for its processing capacity under various link states. This design avoids resource waste and performance bottlenecks, improving the overall system efficiency.

[0105] It is worth noting that the hierarchical image frame technology and the differentiated workflow image generation logic adopt a modular design concept, facilitating the expansion and upgrade of system functions. In the future, more levels of image frames or adjustment of the generation rules of workflow images can be added according to actual needs to adapt to the changing link status and terminal device requirements. Through standardized interface and protocol design, the system can be compatible with various types of terminal devices and communication links, ensuring stable and efficient data transmission in different network environments. This design enhances the cross-platform compatibility and interoperability of the system, laying a solid foundation for its widespread application.

[0106] On the basis of the above-mentioned embodiments, it is worth noting that if the workflow image includes multiple image frames, and it is determined that there is an audio signal in the target video data corresponding to the target terminal device in the video data set, each target image data corresponding to the target terminal device is configured in the target image frame in the multiple image frames, wherein the target image frame is the image frame with the highest resolution in the multiple image frames.

[0107] The system integrates audio signal detection to analyze the audio stream in the target video data in real time and identify whether there is an audio signal. This mechanism can be based on an audio feature extraction algorithm (such as Mel Frequency Cepstral Coefficient (MFCC) analysis) to ensure high-precision judgment of the existence of an audio signal and avoid false positives or false negatives.

[0108] When it is detected that there is an audio signal in the target video data, the system automatically triggers the image frame allocation logic to uniformly configure all target image data (including but not limited to video frames, key frames, etc.) corresponding to the target terminal device to the target image frame with the highest resolution in the multiple image frames. This process is achieved through an association mapping table of image data and audio signals to ensure the accuracy and real-time performance of the allocation. It is worth noting that in a multi-person online conference scenario, if it is detected that there is an audio signal in the target video data corresponding to the target terminal device, it means that the user corresponding to the target terminal device is making a voice input, and there may be a situation of displaying corresponding content. Therefore, by configuring the corresponding image data to the target image frame with the highest resolution, the clear display of the display content can be ensured.

[0109] Figure 8 is a structural schematic diagram of an electronic device according to an example embodiment. As shown in Figure 8 The electronic device 500 provided by the embodiment includes a processor 501 and a memory 502. The processor 501 and the memory 502 are connected through a bus.

[0110] The memory 502 is used for storing a computer program, and the memory can also be a flash memory.

[0111] The processor 501 is configured to execute the instructions stored in the memory to implement the steps of the above method. Details can be referred to the related description in the above method embodiments.

[0112] Optionally, the memory 502 can be independent or integrated with the processor 501.

[0113] When the memory 502 is independent of the processor 501, the electronic device 500 further includes:

[0114] The bus 503 is configured to connect the memory 502 and the processor 501.

[0115] The embodiment further provides a readable storage medium, and the readable storage medium stores the computer program. When at least one processor of the electronic device executes the computer program, the electronic device executes the method provided in the various embodiments.

[0116] The embodiment further provides a program product, and the program product includes the computer program stored in the readable storage medium. At least one processor of the electronic device can read the computer program from the readable storage medium, and the at least one processor executes the computer program to enable the electronic device to implement the method provided in the various embodiments.

[0117] Other embodiments of the application will be apparent to those skilled in the art from consideration of the specification and practice of the application disclosed herein. For example, the methods and apparatuses of the present application can be used in conjunction with any number of computer systems or devices. It is intended that the specification and examples be considered as exemplary only, with the true scope and spirit of the application being indicated by the following claims.

[0118] It should be understood that the application is not limited to the precise construction that has been described above and illustrated in the accompanying drawings, and that various modifications and changes can be made by those skilled in the art without departing from the scope of the application. The scope of the application is limited only by the appended claims.

Claims

1. A video workflow data processing system, characterized by, include: Terminal device clusters and servers; Each terminal device in the terminal device cluster sends the collected video data to the server through a corresponding communication link to generate a video data set. The video data set includes image data sequences at various time points. The first terminal device in the terminal device cluster sends the collected first video data to the server through a first communication link, and the second terminal device sends the collected second video data to the server through a second communication link. The video data set includes a target data sequence at a target time point. The target data sequence includes the first image data of the first video data at the target time point and the second image data of the second video data at the target time point. The server generates video workflow data based on the link status of each communication link and the video data set. Each workflow image in the video workflow data includes at least one layer of image frames. The resolution of each layer of image frames is associated with the corresponding link status. The resolution of each layer of image frames is used to constrain the resolution of the image data contained therein. The server sends the corresponding workflow image in the video workflow data to the corresponding terminal device through various communication links; The server generates video workflow data based on the link status of each communication link and the video data set, including: If the first communication link and the second communication link are in the same state, then each frame of the workflow image in the video workflow data includes only one layer of image frame, and the first image data and the second image data are contained in the one layer of image frame; If the first communication link and the second communication link are in different states, and the link state of the first communication link is better than that of the second communication link, then the first workflow image in the video workflow data includes a first layer image frame for accommodating the first image data and the second image data, and the second workflow image includes a first layer image frame for accommodating the first image data and a second layer image frame for accommodating the second image data, wherein the resolution of the first layer image frame is higher than the resolution of the second layer image frame, the first workflow image is used to be sent to the first terminal device, and the second workflow image is used to be sent to the second terminal device; If the workflow image includes a multi-layer image frame, then if it is determined that there is an audio signal in the target video data corresponding to the target terminal device in the video data set, then each target image data corresponding to the target terminal device is configured in the target image frame in the multi-layer image frame, wherein the target image frame is the image frame with the highest resolution in the multi-layer image frame.

2. The video workflow data processing system of claim 1, wherein, The data structure of the video workflow data includes: a time guidance section and an image data section; The time guide section is used to accommodate the various time nodes arranged in chronological order. The image data part includes workflow images at respective time nodes, each layered image frame of the workflow images having a different resolution, and each layered image frame being configured to accommodate at least one image data in the image data sequence.

3. The video workflow data processing system of claim 2, wherein, The number of layers of the layered images included in the workflow images is associated with corresponding link states.

4. The video workflow data processing system of claim 3, wherein, The number of layers of the layered images included in the workflow images dynamically changes with the link states.

5. The video workflow data processing system of any of claims 1-4, wherein, Each layered image frame in the workflow images includes a sequence of accommodation areas corresponding to the number of terminal devices in the terminal device cluster, and each accommodation area in the sequence of accommodation areas is configured to correspond to one of the image data in the image data sequence.

6. The video workflow data processing system of claim 5, wherein, After the server transmits the corresponding workflow image in the video workflow data to the corresponding terminal device through each communication link, the method further includes: If the number of image frames in the workflow image is not unique, the terminal device aligns each image frame in the workflow image to generate and display a corresponding target image.

7. The video workflow data processing system of any of claims 1-4, wherein, The link state is a link state of the server in downlink transmission to the corresponding terminal device.

8. A method of video workflow data processing, the method comprising: The method is applied to a video workflow data processing system, and the system includes a terminal device cluster and a server. Each terminal device in the terminal device cluster transmits collected video data to the server through a corresponding communication link to generate a video data set, the video data set including an image data sequence at each time node, a first terminal device in the terminal device cluster transmits collected first video data to the server through a first communication link, and a second terminal device transmits collected second video data to the server through a second communication link; the video data set includes a target data sequence at a target time node, the target data sequence including first image data of the first video data at the target time node and second image data of the second video data at the target time node; The server generates video workflow data according to the link states of each communication link and the video data set, each frame of workflow image in the video workflow data including at least one layered image frame, the resolution of each layered image frame being associated with a corresponding link state, and the resolution of each layered image frame being a resolution for constraining the image data accommodated therein; The server transmits the corresponding workflow image in the video workflow data to the corresponding terminal device through each communication link. The server generates video workflow data according to the link states of each communication link and the video data set, including: If the first communication link and the second communication link are in the same state, each frame of workflow image in the video workflow data includes only one layered image frame, and the first image data and the second image data are accommodated in the one layered image frame; If the first communication link and the second communication link are in different states, and the link state of the first communication link is better than the link state of the second communication link, the first workflow image in the video workflow data comprises a first layer image frame for accommodating the first image data and the second image data, and the second workflow image comprises the first layer image frame for accommodating the first image data and a second layer image frame for accommodating the second image data, wherein the resolution of the first layer image frame is higher than the resolution of the second layer image frame, the first workflow image is used for being sent to the first terminal device, and the second workflow image is used for being sent to the second terminal device. If the workflow image comprises a plurality of layer image frames, and it is determined that there is audio signal in the target video data corresponding to the target terminal device in the video data set, each target image data corresponding to the target terminal device is configured in a target image frame in the plurality of layer image frames, wherein the target image frame is the image frame with the highest resolution in the plurality of layer image frames.

9. An electronic device, comprising: Comprise: a processor; and a memory for storing executable instructions of the processor; wherein the processor is configured to execute the method of claim 8 by executing the executable instructions.

10. A computer-readable storage medium, characterized in that, The computer readable storage medium stores computer execution instructions, and the computer execution instructions are executed by the processor to implement the method of claim 8. The computer readable storage medium stores computer execution instructions, and the computer execution instructions are executed by the processor to implement the method of claim 8.

Citation Information

Patent Citations

  • Multi-layer video processing method and system and readable storage medium

    CN112235606A

  • Image processing method and system and electronic equipment

    CN117395459A