Multimodal Video Data Compression and Transmission Method

The method addresses inefficient compression and transmission of multiple-modal video data by prioritizing features and dynamically managing transmission paths, enhancing efficiency and reliability.

CN120091137BActive Publication Date: 2025-07-15GUANGDONG IDATATECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510576133.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-06
Publication Date
2025-07-15
Estimated Expiration
2045-05-06

AI Technical Summary

Technical Problem

The existing multimodal video data processing technology fails to differentiate the characteristics of different modes, resulting in the compression strategy not meeting actual business needs and failing to effectively solve the problem of efficient transmission of compressed data in complex network environments.

Method used

By extracting features such as the pixel proportion of the subject object in the image mode and the number of mouse clicks, the proportion of effective speech duration of the audio mode and the number of clip marks, the keyword proportion and copy times of the text mode, multi-dimensional fusion calculation is used to determine the importance of the feature, and quantitative evaluation of the value of the mode, and layered compression and dynamic selection and switching of the transmission path are performed based on feature priorities.

Benefits of technology

It improves the effectiveness and transmission quality of compression strategies, enhances user experience, optimizes overall transmission efficiency, improves the real-time adaptability and reliability of the system, and ensures the quality of retention of important information and the real-time transmission of transmission.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120091137B_ABST
    Figure CN120091137B_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of video data compression and transmission, and specifically discloses a method for compressing and transmitting multi-modal video data. The method first extracts the feature set of multi-modal video data, dynamically evaluates the feature priorities based on the feature set, then adopts a hierarchical compression strategy to differentially compress different priority modalities, and finally screens available paths through historical transmission data, combines the modality priorities to match path intervals, and monitors the network status in real time, dynamically determines the transmission status, and quickly switches to the sub-optimal path in case of anomalies; through dynamic feature priority evaluation, hierarchical compression strategy and real-time transmission path monitoring and switching, the present invention solves the deficiencies of traditional methods in terms of compression efficiency, transmission quality and real-time adaptability, effectively avoids the impact of network congestion or failures on the transmission quality, and significantly improves the compression efficiency, transmission quality and user experience.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of video data compression and transmission, and more specifically, to a method for compressing and transmitting multi-modal video data. Background Art

[0002] At present, with the rapid development of information technology, multi-modal video data is increasingly widely used in fields such as telemedicine, online education, and video conferencing. The efficient compression and reliable transmission of multi-modal video data are the core technologies to ensure its application experience. Existing multi-modal video processing technologies usually adopt a unified compression strategy and transmission path allocation method, without differentiating the importance of different modal features. Therefore, there is an urgent need for a hierarchical compression and intelligent transmission path allocation method based on modal feature priorities to solve the collaborative optimization problem of quality, efficiency, and reliability in multi-modal video data processing.

[0003] For example, the patent with the Chinese patent application publication number CN116152362A discloses a data compression method, system, and electronic device. Among them, the data compression method includes: obtaining original high-channel line data, inputting the original high-channel line data into a trained encoder model for encoding processing and outputting a latent feature variable, where the encoder model is iteratively trained at least according to the line data to be trained, training parameters, and a preset generator model, and using a preset entropy encoder to perform entropy encoding processing on the latent feature variable and generating a compression result of the high-channel line data.

[0004] The following problems also exist in the above prior art: 1. It only targets single-modal high-channel line data and performs unified compression through an encoder model, without differentiating the feature differences and importance of different modalities for hierarchical compression. At the same time, it relies on a deep learning model for end-to-end compression, and the compression strategy is automatically learned by the model, without explicitly introducing features related to the business scenario, making it difficult to dynamically adjust the compression strategy according to actual business needs.

[0005] 2. It only focuses on the data compression link, does not involve the problem of data transmission after compression, and does not perform transmission path selection, network status monitoring, and dynamic switching mechanism according to different data modalities, and cannot solve the problems of efficient and stable transmission of compressed data in a complex network environment. Summary of the Invention

[0006] In view of this, to solve the problems raised in the above background art, a method for compressing and transmitting multi-modal video data is proposed.

[0007] The object of the present invention can be achieved by the following technical solutions: The present invention provides a multi-modal video data compression and transmission method, including the following steps: S1. Extract features from multi-modal video data, and respectively obtain the feature sets of each modality in the video, where each modality includes an image modality, an audio modality, and a text modality.

[0008] S2. According to the feature sets, determine the feature priorities of each modality based on the modality feature importance evaluation rule.

[0009] S3. According to the feature priorities of each modality, adopt a hierarchical compression strategy to compress the video data of each modality, and obtain the compressed video data of each modality.

[0010] S4. Extract the historical transmission data of each transmission path, screen the available transmission paths corresponding to the multi-modal video data accordingly, and allocate the most suitable transmission path for the compressed video data of each modality in combination with the feature priorities.

[0011] S5. Real-time monitor the network condition parameters of the most suitable transmission path of the video data of each modality, and determine the transmission state of the most suitable transmission path based on the transmission path switching rule. When the determination result is abnormal, perform transmission path switching.

[0012] Compared with the prior art, the embodiments of the present invention have at least the following advantages or beneficial effects: (1) By extracting features such as the proportion of the main object pixels in the image modality and the number of mouse clicks, the proportion of the effective speech duration in the audio modality and the number of segment markers, and the keyword ratio and the number of copies in the text modality, the present invention uses multi-dimensional fusion calculation to determine the feature importance, realizes the quantitative evaluation of the modality value, makes the feature priority evaluation more in line with the actual scenario requirements, provides a basis for hierarchical compression, and improves the effectiveness of the compression strategy.

[0013] (2) By dividing the modalities into three levels: high, medium, and low according to the feature priorities, and adopting a hierarchical compression strategy according to the feature priorities, the present invention can significantly improve the compression efficiency and transmission quality. This differential compression strategy not only improves the retention quality of important information, but also optimizes the overall transmission efficiency by reasonably allocating compression resources, enhances the user experience, and improves the real-time adaptability of the system.

[0014] (3) By screening available paths through historical transmission data, matching path intervals in combination with modality priorities, and real-time monitoring the network state, the present invention can dynamically determine the transmission state and quickly switch to the sub-optimal path in case of abnormality, effectively avoiding the impact of network congestion or faults on the transmission quality, and improving the real-time performance and reliability of the transmission. Description of the Drawings

[0015] To more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0016] Figure 1 It is a schematic flowchart of the method steps of the present invention.

[0017] Figure 2 It is a schematic flowchart of the generation process of the feature importance of the image modality of the present invention.

[0018] Figure 3 It is a flowchart for judging the switching of the transmission path of the present invention. Specific embodiments

[0019] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the protection scope of the present invention.

[0020] Please refer to Figure 1 As shown, the present invention provides a multi-modal video data compression and transmission method, including: S1. Extract features from the multi-modal video data to obtain the feature sets of each modality in the video respectively, where each modality includes an image modality, an audio modality, and a text modality.

[0021] In the specific embodiments of the present invention, the feature sets of each modality are as follows: the image modality includes the proportion of the pixel area of the main object and the number of times the image is clicked by the mouse, the audio modality includes the proportion of the effective speech duration and the number of times the audio segment is marked, and the text modality includes the ratio of the number of keywords to the total number of words and the number of times the text is copied.

[0022] It should be noted that the method for obtaining the proportion of the pixel area of the main object is as follows: 1) Preprocessing: Convert the video frame to the RGB format and adjust the resolution to the input requirement of the algorithm. 2) Target detection: Input the target detection model to obtain the bounding boxes and class labels of all detected objects. 3) Main object screening: Filter out irrelevant objects according to the main object categories preset for the video service scenario (such as "face" and "license plate" preset for security videos, and "lectern" and "teacher" preset for education videos). 4) Area calculation: Calculate the sum of the pixel areas of the screened objects and divide it by the total pixel area of the video frame to obtain the proportion of the main object.

[0023] It should be noted that the method for obtaining the proportion of the effective speech duration is as follows: 1) Audio framing: Divide the audio signal into frames of a fixed duration. 2) VAD detection: Calculate the energy value and zero-crossing rate for each frame, and input them into the VAD classifier to determine whether it is a speech frame. 3) Speech segment merging: Merge consecutive speech frames into speech segments, and filter out noise segments shorter than the set threshold. 4) Activity calculation: Divide the total duration of the effective speech segments by the total duration of the audio to obtain the speech activity.

[0024] It should be noted that the method for obtaining the ratio of the number of keywords to the total number of words is as follows: 1) Text preprocessing: Perform word segmentation, stop word removal, and part-of-speech tagging on the text (captions, OCR recognition results). 2) Keyword extraction: Extract entity words through a text named entity recognition model, and combine with the preset business dictionary of the video (such as "disease" and "drug" in medical videos, and "knowledge points" and "concepts" in educational videos) to screen for professional terms. 3) Density calculation: Divide the number of keywords by the total number of words in the text to obtain the ratio of the number of keywords to the total number of words.

[0025] It should be further noted that the number of mouse clicks on the image is captured from the client click coordinates, the number of audio segment markings is obtained by extracting from the audio clip operation log, and the number of text copies is obtained from the text interaction event monitoring log.

[0026] S2. According to the feature set, determine the feature priorities of each modality based on the modality feature importance evaluation rule.

[0027] In a specific embodiment of the present invention, the specific implementation process of determining the feature priorities of each modality based on the modality feature importance evaluation rule is as follows: Perform multi-dimensional fusion calculation based on the proportion of the subject object pixel area and the number of mouse clicks on the image to obtain the feature importance of the image modality.

[0028] Please refer to Figure 2 As shown, in a specific embodiment of the present invention, the specific analysis process of the feature importance of the image modality is as follows: Perform dynamic difference calculation on the proportion of the subject object pixel area and the number of mouse clicks on the image respectively with a preset value.

[0029] It should be noted that the dynamic difference calculation specifically refers to: taking the difference between the proportion of the subject object pixel area and the preset proportion of the subject object pixel area, and taking the difference between the number of mouse clicks on the image and the preset number of mouse clicks on the image.

[0030] Perform normalization processing on each dynamic difference to obtain the content contribution degree and user attention degree of the image modality.

[0031] It should be noted that the specific process of obtaining the content contribution degree and user attention degree of the image modality is as follows: The ratio of the difference between the proportion of the pixel area of the main object and the preset proportion of the pixel area of the main object to the preset proportion of the pixel area of the main object is used as the content contribution degree, and the ratio of the difference between the number of times the image is clicked by the mouse and the preset number of times the image is clicked by the mouse to the preset number of times the image is clicked by the mouse is used as the user attention degree.

[0032] The content contribution degree and the user attention degree are weighted and calculated with the corresponding proportion weights respectively, and then summed to generate the feature importance degree of the image modality.

[0033] In a specific embodiment of the present invention, when generating the feature importance degree of the image modality, the importance of the content contribution degree and the user attention degree needs to be dynamically determined in combination with the business scenario: In the scenario centered on content value, the content contribution degree is more important because it directly determines the effectiveness and accuracy of information. In the scenario oriented to user experience, the weight of user attention degree is higher, and the quality of the areas with high-frequency user attention needs to be guaranteed first. In general scenarios, both need to be considered balancedly to ensure that the feature importance degree reflects both the essential value of the image content and the actual focus of user attention.

[0034] The feature importance degree of the audio modality is obtained by performing importance evaluation based on the proportion of the effective voice duration and the number of times the audio segment is marked.

[0035] It should be noted that the specific process of obtaining the feature importance degree of the audio modality is as follows: The ratios of the differences between the proportion of the effective voice duration and the number of times the audio segment is marked and the corresponding set values to the corresponding set values are calculated respectively, and the results of the two ratios are summed to obtain the feature importance degree of the audio modality.

[0036] The proportion of the number of keywords to the total number of words and the number of times the text is copied are comprehensively processed to obtain the feature importance degree of the text modality.

[0037] It should be noted that the specific process of obtaining the feature importance degree of the text modality is as follows: The ratio of the difference between the ratio of the number of keywords to the total number of words and the ratio of the number of keywords to the total number of words in the set reference to the ratio of the number of keywords to the total number of words in the set reference is calculated, and the ratio of the difference between the number of times the text is copied and the number of times the text is copied in the set reference to the number of times the text is copied in the set reference is calculated. Then the results of the two ratios are summed to obtain the feature importance degree of the text modality.

[0038] The image modality, audio modality, and text modality are sorted in descending order of feature importance degree to generate a feature priority queue, thereby obtaining the feature priorities of each modality.

[0039] In the embodiments of the present invention, by extracting features such as the proportion of the main object pixels in the image modality, the number of mouse clicks, the proportion of the effective speech duration in the audio modality, the number of segment markers, the keyword ratio in the text modality, and the number of copies, multi-dimensional fusion calculation is used to determine the feature importance, realizing the quantitative evaluation of the modality value, making the feature priority evaluation more in line with the actual scenario requirements, providing a basis for hierarchical compression, and improving the effectiveness of the compression strategy.

[0040] S3. According to the feature priorities of each modality, a hierarchical compression strategy is adopted to compress the video data of each modality, and the compressed video data of each modality is obtained.

[0041] In a specific embodiment of the present invention, the specific method of using the hierarchical compression strategy to compress the video data of each modality is as follows: The modalities with the highest feature priority, medium feature priority, and lowest feature priority are respectively denoted as the high-priority modality, medium-priority modality, and low-priority modality, and they are matched with the compression strategies corresponding to each priority modality stored in the database, so as to obtain the compression strategies corresponding to the high-priority modality, medium-priority modality, and low-priority modality respectively.

[0042] In a specific embodiment of the present invention, the compression strategies corresponding to each priority modality stored in the database include but are not limited to the high-quality compression strategy with an extremely low compression ratio corresponding to the high-priority modality, the strategy of balancing compression efficiency and quality corresponding to the medium-priority modality, and the high-efficiency compression strategy with a high compression ratio corresponding to the low-priority modality.

[0043] It should be further noted that different priority modalities correspond to different compression strategies. The basis lies in the differences in the business requirements and quality requirements of each modality, as well as the need to adapt to the network resource status. The advantage of this is that it can not only guarantee the quality of key services of high-priority modalities, but also optimize the resource utilization of medium-priority modalities through the strategy of balancing compression efficiency and quality, and can also improve the transmission efficiency of low-priority modalities by virtue of the high compression ratio strategy, thereby reducing the impact on high-priority modalities and enhancing the overall performance of the system.

[0044] In the embodiments of the present invention, by classifying modalities into high, medium, and low levels according to the feature priorities and adopting a hierarchical compression strategy based on the feature priorities, the compression efficiency and transmission quality can be significantly improved. This differential compression strategy not only improves the retention quality of important information, but also optimizes the overall transmission efficiency by reasonably allocating compression resources, enhances the user experience, and improves the real-time adaptability of the system.

[0045] S4. Extract the historical transmission data of each transmission path, screen the available transmission paths corresponding to the multi-modal video data accordingly, and allocate the most suitable transmission path for the compressed video data of each modality in combination with the feature priorities.

[0046] In a specific embodiment of the present invention, the specific process of screening the available transmission paths corresponding to the multi-modal video data is as follows: based on the historical transmission data of each transmission path, the number of packets sent recorded by the sender and the number of packets received feedback by the receiver, as well as the sending time point of the sender and the receiving time point of the receiver for each transmission, are respectively subjected to coupling processing to obtain the packet loss rate and the delay change rate of each transmission path.

[0047] It should be noted that the methods for obtaining the number of packets sent recorded by the sender and the sending time point of the sender corresponding to each transmission are as follows: a monitoring module is embedded in the network protocol stack of the sender. Whenever a data packet is sent, the number of packets sent is accumulated and recorded through a counter, and the sending timestamp of the data packet is recorded using the high-precision clock of the system. At the same time, a unique identifier is assigned to each data packet for subsequent matching. The methods for obtaining the number of packets received feedback by the receiver and the receiving time point of the receiver are as follows: after receiving the data packet, the receiver also parses the unique identifier of the data packet through the monitoring module, accumulates the number of packets received through the counter, records the receiving timestamp, and transmits the packet reception information back to the sender through a feedback mechanism.

[0048] In a specific embodiment of the present invention, the specific process of obtaining the packet loss rate and the delay change rate of each transmission path is as follows: the difference between the number of packets sent recorded by the sender and the number of packets received feedback by the receiver corresponding to each transmission of each transmission path is used as the number of lost packets corresponding to each transmission, and the ratio of the number of lost packets to the number of packets sent recorded by the sender corresponding to each transmission is used as the packet loss rate corresponding to each transmission. The average calculation result of the packet loss rates corresponding to each transmission is used as the packet loss rate of each transmission path.

[0049] The time interval between the sending time point of the sender and the receiving time point of the receiver corresponding to each transmission of each transmission path is used as the delay duration corresponding to each transmission. The absolute value of the deviation of the delay duration corresponding to adjacent transmissions is obtained. The sum of the absolute values of the deviations of the delay durations corresponding to all adjacent transmissions is divided by the number of transmissions to obtain the delay change rate of each transmission path.

[0050] The packet loss rate and the delay change rate of each transmission path are respectively compared with the packet loss rate critical value and the delay change rate critical value stored in the database. If the packet loss rate of a certain transmission path is less than the packet loss rate critical value and the delay change rate is less than the delay change rate critical value, then this transmission path is recorded as an available transmission path. If the packet loss rate of a certain transmission path is greater than the packet loss rate critical value or the delay change rate is greater than the delay change rate critical value, then this transmission path is recorded as an unavailable transmission path. Thus, the available transmission paths corresponding to the multi-modal video data are obtained.

[0051] It should be noted that the setting of the packet loss rate critical value and the delay change rate critical value stored in the database is mainly based on the requirements of the service scenario and the industry quality standard. In a specific embodiment of the present invention, the packet loss rate critical value can be set to 25%, and the delay change rate critical value can be set to 60%.

[0052] In a specific embodiment of the present invention, the specific process of the combined feature priority to allocate the most suitable transmission path for each modal video data after compression is as follows: extract the packet loss rate and the delay change rate of each available transmission path from the packet loss rate and the delay change rate of each transmission path, and match the packet loss rate and the delay change rate of each available transmission path with the packet loss rate interval and the delay change rate interval corresponding to each priority mode stored in the database. If the packet loss rate of a certain available transmission path is within the packet loss rate interval corresponding to a certain priority mode and the delay change rate is also within the delay change rate interval corresponding to this priority mode, then this available transmission path is used as the transmission path corresponding to this priority mode, and thus the transmission paths corresponding to each priority mode are obtained.

[0053] It should be noted that the packet loss rate interval and the delay change rate interval corresponding to each priority mode stored in the database are specified by industry standards and protocol specifications. In a specific embodiment of the present invention, the packet loss rate interval of the high-priority mode is set to 0% - 5%, the delay change rate interval is set to 0% - 20%, the packet loss rate interval of the medium-priority mode is set to 5% - 15%, the delay change rate interval is set to 20% - 40%, the packet loss rate interval of the low-priority mode is set to 15% - 25%, and the delay change rate interval is set to 40% - 60%.

[0054] Based on the packet loss rate and the delay change rate of each transmission path corresponding to each priority mode, the reliability of each transmission path is analyzed to obtain the reliability of each transmission path, and the transmission path corresponding to the maximum reliability is selected as the most suitable transmission path for the corresponding priority mode.

[0055] It should be noted that the specific method for obtaining the reliability of each transmission path is as follows: divide the difference between the packet loss rate and the delay change rate of each transmission path corresponding to each priority mode and the corresponding set reference value by the corresponding set reference value and then sum them up, and substitute the summation result into the exponential function with as the base to obtain the reliability of each transmission path.

[0056] S5. Real-time monitor the network condition parameters of the most suitable transmission path of each modal video data, and judge the transmission state of the most suitable transmission path based on the transmission path switching rule. When the judgment result is abnormal, perform transmission path switching.

[0057] In a specific embodiment of the present invention, the specific process of determining the transmission state of the optimal transmission path based on the transmission path switching rule is as follows: extract the total number of packets sent by the sender, the number of acknowledged packets received by the receiver, the time points of each probe packet sent by the sender, and the time points of each probe packet received by the receiver from the network condition parameters of the optimal transmission path of each modality video data.

[0058] It should be noted that the ways to obtain the total number of packets sent by the sender and the number of acknowledged packets received by the receiver are as follows: the sender adds a unique sequence number to each transmitted data packet, and after the receiver receives it, it feeds back the received sequence numbers through an acknowledgment message, and confirms the number of received packets by comparing the sequence numbers. The ways to obtain the time points of each probe packet sent by the sender and the time points of each probe packet received by the receiver are as follows: the sender generates a probe packet and embeds the sending timestamp to obtain the time points of each probe packet sent, and after the receiver receives it, it carries this timestamp in the response packet and adds the receiving timestamp as the time points of each probe packet received.

[0059] Take the ratio of the difference between the total number of packets sent by the sender and the number of acknowledged packets received by the receiver to the total number of packets sent as the real-time packet loss rate of the optimal transmission path of each modality video data.

[0060] Perform a change rate analysis based on the time points of each probe packet sent by the sender and the time points of each probe packet received by the receiver to obtain the real-time delay change rate of the optimal transmission path of each modality video data.

[0061] It should be noted that the specific process of obtaining the real-time delay change rate of the optimal transmission path of each modality video data is as follows: take the interval duration between the time points of each probe packet sent by the sender and the time points of each probe packet received by the receiver as the delay duration corresponding to each sending by the sender, obtain the absolute value of the deviation of the delay durations corresponding to adjacent sendings, sum up the absolute values of the deviation of the delay durations corresponding to all adjacent sendings and then divide by the number of sendings to obtain the real-time delay change rate of the optimal transmission path of each modality video data.

[0062] If the real-time packet loss rate of the optimal transmission path of a certain modality video data is not within the packet loss rate interval corresponding to the priority modality of this modality video data or the real-time delay change rate is not within the delay change rate interval corresponding to the priority modality of this modality video data, then it is determined that there is an abnormality in the optimal transmission path of this modality video data. If both the real-time packet loss rate and the real-time delay change rate of the optimal transmission path of a certain modality video data are within the packet loss rate interval and the delay change rate interval corresponding to the priority modality of this modality video data, then it is determined that the optimal transmission path of this modality video data is normal and there is no need to switch the transmission path.

[0063] Please refer to Figure 3As shown in the figure, in a specific embodiment of the present invention, the specific method for performing transmission path switching is as follows: Step 1: Select the transmission path with the second highest reliability from the reliabilities of the transmission paths corresponding to each priority mode as the sub-optimal path.

[0064] Step 2: Monitor the network condition parameters of the sub-optimal path of the video data of each mode in real time, and determine the transmission state of the sub-optimal path in the same way as the determination method of the transmission state of the optimal transmission path.

[0065] It should be noted that the network condition parameters of the sub-optimal path of the video data of each mode also include the total number of packets sent by the sending end, the number of acknowledged packets received by the receiving end, and the time points of each probe packet sent by the sending end and the time points of each probe packet received by the receiving end.

[0066] Step 3: When the determination result of the transmission state of the sub-optimal path is abnormal, re-execute Steps 1 to 2. When the determination result of the transmission state of the sub-optimal path is normal, stop the iteration.

[0067] The embodiment of the present invention screens available paths through historical transmission data, combines the mode priority to match the path interval, and monitors the network state in real time, can dynamically determine the transmission state, and quickly switch to the sub-optimal path in case of abnormality, effectively avoiding the impact of network congestion or failure on the transmission quality, and improving the real-time performance and reliability of the transmission.

[0068] The above content is only an example and explanation of the concept of the present invention. Those skilled in the art of the present technology can make various modifications or supplements to the described specific embodiments or use similar methods to replace them, as long as they do not deviate from the concept of the invention or exceed the scope defined by the present invention, they should fall within the protection scope of the present invention.

Claims

1. A multi-modal video data compression and transmission method, characterized in that It includes the following steps: S1. Extract features from the multi-modal video data to obtain the feature sets of each modality in the video respectively, where each modality includes an image modality, an audio modality, and a text modality; The feature sets of each modality are as follows: the image modality includes the proportion of the pixel area of the main object and the number of times the image is clicked by the mouse, the audio modality includes the proportion of the effective speech duration and the number of times the audio segment is marked, and the text modality includes the ratio of the number of keywords to the total number of words and the number of times the text is copied; S2. According to the feature sets, determine the feature priorities of each modality based on the modality feature importance evaluation rules; The specific implementation process of determining the feature priorities of each modality based on the modality feature importance evaluation rules is as follows: Perform multi-dimensional fusion calculation based on the proportion of the pixel area of the main object and the number of times the image is clicked by the mouse to obtain the feature importance of the image modality; Evaluate the feature importance of the audio modality based on the proportion of the effective speech duration and the number of times the audio segment is marked; Perform comprehensive processing on the ratio of the number of keywords to the total number of words and the number of times the text is copied to obtain the feature importance of the text modality; Arrange the image modality, audio modality, and text modality in descending order of feature importance to generate a feature priority queue, thereby obtaining the feature priorities of each modality; S3. According to the feature priorities of each modality, adopt a hierarchical compression strategy to compress the video data of each modality to obtain the compressed video data of each modality; S4. Extract the historical transmission data of each transmission path, screen the available transmission paths corresponding to the multi-modal video data accordingly, and allocate the most suitable transmission path for the compressed video data of each modality in combination with the feature priorities; S5. Real-time monitor the network status parameters of the most suitable transmission path of the video data of each modality, and determine the transmission status of the most suitable transmission path based on the transmission path switching rules. When the determination result is abnormal, perform transmission path switching.

2. The multimodal video data compression and transmission method according to claim 1, characterized in that: The specific analysis process of the feature importance of the image modality is as follows: Perform dynamic difference calculation on the proportion of the pixel area of the main object and the number of times the image is clicked by the mouse respectively with a preset value; Perform normalization processing on each dynamic difference to obtain the content contribution degree and user attention degree of the image modality; Perform weighted calculation on the content contribution degree and user attention degree respectively with the corresponding proportion weights and then sum them to generate the feature importance of the image modality.

3. The multimodal video data compression and transmission method according to claim 1, characterized in that: The specific method of adopting a hierarchical compression strategy to compress the video data of each modality is as follows: Denote the modality with the highest feature priority, the medium feature priority, and the lowest feature priority as the high-priority modality, the medium-priority modality, and the low-priority modality respectively, and match them with the compression strategies corresponding to each priority modality stored in the database, so as to obtain the compression strategies corresponding to the high-priority modality, the medium-priority modality, and the low-priority modality respectively.

4. The multimodal video data compression and transmission method according to claim 3, characterized in that: The specific process of screening the available transmission paths corresponding to the multi-modal video data is as follows: Based on the historical transmission data of each transmission path, perform coupling processing on the number of packets sent recorded by the sender and the number of packets received feedback by the receiver and the sending time point of the sender and the receiving time point of the receiver corresponding to each transmission respectively to obtain the packet loss rate and delay change rate of each transmission path; Compare the packet loss rate and the delay change rate of each transmission path with the packet loss rate threshold and the delay change rate threshold stored in the database respectively. If the packet loss rate of a certain transmission path is less than the packet loss rate threshold and the delay change rate is less than the delay change rate threshold, then mark this transmission path as an available transmission path. If the packet loss rate of a certain transmission path is greater than the packet loss rate threshold or the delay change rate is greater than the delay change rate threshold, then mark this transmission path as an unavailable transmission path. Thus, each available transmission path corresponding to the multi-modal video data is obtained.

5. The multimodal video data compression and transmission method according to claim 4, characterized in that: The specific process of obtaining the packet loss rate and the delay change rate of each transmission path is as follows: Take the difference between the number of packets sent recorded by the sender and the number of packets received feedback by the receiver for each transmission corresponding to each transmission path as the number of lost packets for each transmission, and take the ratio of it to the number of packets sent recorded by the sender for each transmission as the packet loss rate for each transmission. Take the calculation result of the average value of the packet loss rates for each transmission as the packet loss rate of each transmission path; Take the interval duration between the sending time point of the sender and the receiving time point of the receiver for each transmission corresponding to each transmission path as the delay duration for each transmission. Obtain the absolute value of the deviation of the delay duration between adjacent transmissions. Sum up the absolute values of the deviation of the delay duration between all adjacent transmissions and then divide by the number of transmissions to obtain the delay change rate of each transmission path.

6. The multimodal video data compression and transmission method according to claim 4, characterized in that: The specific process of allocating the most suitable transmission path for each compressed modal video data in combination with the feature priority is as follows: Extract the packet loss rate and the delay change rate of each available transmission path from the packet loss rate and the delay change rate of each transmission path, and match the packet loss rate and the delay change rate of each available transmission path with the packet loss rate interval and the delay change rate interval corresponding to each priority modal stored in the database respectively. If the packet loss rate of a certain available transmission path is within the packet loss rate interval corresponding to a certain priority modal and the delay change rate is also within the delay change rate interval corresponding to this priority modal, then take this available transmission path as the transmission path corresponding to this priority modal. Thus, each transmission path corresponding to each priority modal is obtained; Conduct reliability analysis based on the packet loss rate and the delay change rate of each transmission path corresponding to each priority modal to obtain the reliability of each transmission path, and select the transmission path corresponding to the maximum reliability as the most suitable transmission path for the corresponding priority modal.

7. The multimodal video data compression and transmission method according to claim 6, wherein: The specific process of determining the transmission state of the most suitable transmission path based on the transmission path switching rule is as follows: Extract the total number of packets sent by the sender, the number of acknowledged packets received by the receiver, and the time points of each sending of the probe packet by the sender and the time points of each receiving of the probe packet by the receiver from the network status parameters of the most suitable transmission path of each modal video data; Take the ratio of the difference between the total number of packets sent by the sender and the number of acknowledged packets received by the receiver to the total number of packets sent as the real-time packet loss rate of the most suitable transmission path of each modal video data; Conduct change rate analysis based on the time points of each sending of the probe packet by the sender and the time points of each receiving of the probe packet by the receiver to obtain the real-time delay change rate of the most suitable transmission path of each modal video data; If the real-time packet loss rate of the optimal transmission path of a certain modal video data is not within the packet loss rate interval corresponding to the priority mode of the modal video data or the real-time delay change rate is not within the delay change rate interval corresponding to the priority mode of the modal video data, it is determined that there is an abnormality in the optimal transmission path of the modal video data. If both the real-time packet loss rate and the real-time delay change rate of the optimal transmission path of a certain modal video data are within the packet loss rate interval and the delay change rate interval corresponding to the priority mode of the modal video data, it is determined that the optimal transmission path of the modal video data is normal and there is no need to switch the transmission path.

8. The multimodal video data compression and transmission method according to claim 7, characterized in that: The specific method for switching the transmission path is as follows: Step 1: Select the transmission path with the second highest reliability from the reliabilities of the transmission paths corresponding to each priority mode as the sub-optimal path; Step 2: Real-time monitor the network condition parameters of the sub-optimal path of each modal video data, and similarly determine the transmission state of the sub-optimal path according to the transmission state determination method of the optimal transmission path; Step 3: When the determination result of the transmission state of the sub-optimal path is abnormal, re-execute Steps 1 to 2. When the determination result of the transmission state of the sub-optimal path is normal, stop the iteration.

Citation Information

Patent Citations

  • Data compression method and system and electronic device

    CN116152362A

  • Information transmission method and device and storage medium

    CN119094369A

  • Voice scheduling exchange system with dual-network integration

    CN119363643A