Video data transmission method and head-mounted display equipment

By segmenting video data and adding personalized identification information, the problem of video transmission failure in unstable network environments for head-mounted display devices was solved, enabling fast transmission and high-quality display of video data and improving user experience.

CN121967633APending Publication Date: 2026-05-01GEER TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
GEER TECH CO LTD
Filing Date
2024-10-31
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

Video data transmission failures in head-mounted display devices under unstable network conditions cause buffering and reduce the user experience.

Method used

The acquired video data is divided into packets to generate multiple small video data packets. Personalized identification information is added to each data packet to ensure that the video processing module can restore the original video data based on this identification information.

Benefits of technology

By splitting video data packets for transmission, situations where video data cannot be sent to the video processing module are avoided, ensuring fast video data transmission and improving the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121967633A_ABST
    Figure CN121967633A_ABST
Patent Text Reader

Abstract

The invention discloses a video data transmission method and a head-mounted display device, and relates to the technical field of wearable device.The video data transmission method comprises the steps that obtained first video data are subpackaged to obtain multiple pieces of second video data, and the data volume of the second video data is smaller than that of the first video data; determining personalized identification information corresponding to the plurality of second video data according to the first video data; writing each piece of personalized identification information into respective matched second video data to obtain each target video data packet; and transmitting the target video data packets to a video processing module, so that the video processing module restores the target video data packets to obtain the first video data. According to the invention, the technical effect that the head-mounted display device can quickly transmit the video data to the video processing module is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Video data transmission methods and head-mounted display devices Technical Field

[0001] This application relates to the field of wearable device technology, and more particularly to a method for transmitting video data and a head-mounted display device. Background Technology

[0002] AR (Augmented Reality) devices or VR (Virtual Reality) devices are head-mounted display devices that are currently developing and becoming increasingly widespread. Head-mounted display devices typically generate video data of sufficiently high quality to meet the immersive experience required by users when wearing AR / VR devices.

[0003] In related technologies, after generating high-quality video data, the video generation module in a head-mounted display device typically sends the video data to the video processing module of the head-mounted display device, which then processes the video data to display high-quality video data to the user.

[0004] However, since high-quality video data usually has a large data volume, technical problems such as video data transmission failure are prone to occur when the head-mounted display device is in an unstable network environment. This results in users watching stuttering video content while wearing the head-mounted display device, which greatly reduces the user experience. Summary of the Invention

[0005] The main objective of this application is to provide a method for transmitting video data and a head-mounted display device, aiming to solve the technical problem of video data transmission failure in related technologies.

[0006] To achieve the above objectives, this application proposes a method for transmitting video data, the method comprising:

[0007] The acquired first video data is divided into multiple second video data, wherein the data volume of the second video data is smaller than that of the first video data;

[0008] Based on the first video data, determine the personalized identification information corresponding to each of the multiple second video data;

[0009] Each personalized identifier is written into its corresponding matching second video data to obtain each target video data packet;

[0010] Each of the target video data packets is transmitted to the video processing module so that the video processing module can reconstruct the first video data based on each of the target video data packets.

[0011] In one embodiment, the personalized identification information includes image frame identification information;

[0012] The step of determining the personalized identification information corresponding to each of the multiple second video data based on the first video data includes:

[0013] The first video data contains a plurality of first image frames, and the plurality of second video data each contains a plurality of second image frames;

[0014] A plurality of first image frames and a plurality of second image frames are compared to determine the image frame mapping relationship between the plurality of first image frames and the plurality of second image frames;

[0015] The frame generation time corresponding to each of the multiple first image frames is determined, and based on the image frame mapping relationship and the frame generation time, the multiple image frame identification information corresponding to each of the multiple second video data is determined.

[0016] In one embodiment, the personalized identification information further includes data packet identification information;

[0017] After the step of determining the image frame mapping relationship between the plurality of first image frames and the plurality of second image frames, the method further includes:

[0018] Based on the image frame mapping relationship, sorting is performed on multiple second video data to obtain a sorting result;

[0019] Based on the sorting result, the data packet identification information corresponding to each of the multiple second video data is determined.

[0020] In one embodiment, the personalized identification information further includes content identification information;

[0021] The step of determining the personalized identification information corresponding to each of the multiple second video data based on the first video data further includes:

[0022] Determine the frame content data contained in each of the multiple first image frames;

[0023] Based on the content data of each frame and the mapping relationship between the image frames, the content identification information corresponding to each of the multiple second video data is determined.

[0024] In one embodiment, the step of writing each of the personalized identifiers into the corresponding matched second video data to obtain each target video data packet includes:

[0025] Each personalized identifier is written into the tail storage space of the corresponding second video data to obtain each target video data packet.

[0026] In one embodiment, the step of processing the acquired first video data into multiple second video data packets includes:

[0027] Obtain the first video data;

[0028] The first video data is encoded using a preset Android encoder to obtain the third video data, wherein the data volume of the third video data is smaller than that of the first video data and larger than that of the second video data.

[0029] The target video data length is determined, and the third video data is divided into multiple second video data according to the target video data length.

[0030] In one embodiment, the step of determining the length of the target video data includes:

[0031] Obtain the real-time data transmission rate;

[0032] Determine multiple preset data transmission rates and preset video data lengths that match each of the multiple preset data transmission rates;

[0033] Based on the real-time data transmission rate, multiple preset data transmission rates are filtered to determine the target data transmission rate;

[0034] The preset video data length corresponding to the target data transmission rate is determined as the target video data length.

[0035] In one embodiment, the step of determining the length of the target video data further includes:

[0036] Determine the length of the actual video data corresponding to the third video data, and obtain the preset number of data packets;

[0037] The target video data length is determined based on the actual video data length and the preset number of data packets.

[0038] In one embodiment, the step of dividing the third video data into packets according to the target video data length to obtain multiple second video data includes:

[0039] The target image frames contained in the third video data are determined based on the length of the target video data.

[0040] The third video data is divided into packets based on each target image frame to obtain multiple second video data.

[0041] In addition, to achieve the above objectives, this application also proposes a head-mounted display device, the device comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, the computer program being configured to implement the steps of the video data transmission method described above.

[0042] The video data transmission method provided in this application embodiment involves dividing the acquired first video data into packets to obtain multiple second video data, wherein the data volume of the second video data is smaller than that of the first video data; determining personalized identification information corresponding to each of the multiple second video data based on the first video data; writing each personalized identification information into its corresponding matching second video data to obtain each target video data packet; and transmitting each target video data packet to a video processing module so that the video processing module can reconstruct the first video data based on each target video data packet.

[0043] In this embodiment, during operation, the head-mounted display device first performs packet processing on the acquired first video data containing multiple frames of images to obtain multiple second video data with reduced data size. Then, the head-mounted display device determines the personalized identification information corresponding to each of the multiple second video data based on the first video data. Next, the head-mounted display device writes each personalized identification information into its corresponding matching second video data to obtain multiple target video data packets containing different identification information. Finally, the head-mounted display device transmits each target video data packet to a preset video processing module, so that the video processing module can perform restoration processing on each target video data packet based on the personalized identification information contained in each target video data packet to obtain the first video data.

[0044] Thus, this application solves the technical problem of video data transmission failure in related technologies. Specifically, by splitting the originally large high-quality video data into multiple smaller video data packets for transmission, this application avoids the situation where video data cannot be sent to the video processing module when the video data is large. It also ensures that the video processing module can restore the original content of the video data based on the personalized identification information contained in each of the multiple video data packets. This achieves the technical effect of enabling the head-mounted display device to quickly transmit video data to the video processing module, further enhancing the user experience when wearing the head-mounted display device. Attached Figure Description

[0045] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.

[0046] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0047] Figure 1 is a flowchart illustrating the video data transmission method according to Embodiment 1 of this application.

[0048] Figure 2 is a schematic diagram of the packet processing flow involved in an embodiment of the video data transmission method of this application;

[0049] Figure 3 is a schematic diagram of the header of the target data packet involved in an embodiment of the video data transmission method of this application;

[0050] Figure 4 is a schematic diagram of the tail storage space of the target data packet involved in an embodiment of the video data transmission method of this application;

[0051] Figure 5 is a simplified flowchart of the video data transmission method of this application;

[0052] Figure 6 is a schematic diagram of the module structure of the video data transmission device of this application;

[0053] Figure 7 is a schematic diagram of the hardware operating environment involved in the video data transmission method in this application embodiment.

[0054] The purpose, features, and advantages of this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation

[0055] It should be understood that the specific embodiments described herein are merely illustrative of the technical solutions of this application and are not intended to limit this application.

[0056] To better understand the technical solution of this application, a detailed description will be provided below in conjunction with the accompanying drawings and specific implementation methods.

[0057] In this embodiment, for ease of description, the following description uses a head-mounted display device equipped with an Android encoder, or a mobile terminal, data storage control terminal, PC, or other terminal connected to an electronic control unit associated with the head-mounted display device, as the execution subject. The head-mounted display device includes, but is not limited to, Mixed Reality (MR) devices (e.g., MR glasses or MR helmets), Augmented Reality (AR) devices (e.g., AR glasses or AR helmets), Virtual Reality (VR) devices (e.g., VR glasses or VR helmets), Extended Reality (XR) devices, or some combination thereof.

[0058] Based on the aforementioned head-mounted display device, the overall concept of the video data transmission method of this application is proposed.

[0059] AR (Augmented Reality) devices or VR (Virtual Reality) devices are head-mounted display devices that are currently experiencing rapid development and widespread adoption. Head-mounted displays typically generate video data of sufficient high quality to meet the immersive experience required by users wearing AR / VR devices. In related technologies, after generating high-quality video data, the video generation module within the head-mounted display device usually sends the video data to the device's video processing module. The video processing module then processes the video data to present it to the user in a high-quality manner. However, because high-quality video data is usually quite large, technical issues such as video data transmission failures can easily occur when the head-mounted display device is in an unstable network environment. This can lead to choppy video content being viewed by the user while wearing the head-mounted display, significantly reducing the user experience.

[0060] To address the above issues, this application provides a video data transmission method, comprising: packetizing acquired first video data to obtain multiple second video data, wherein the data volume of the second video data is smaller than that of the first video data; determining personalized identification information corresponding to each of the multiple second video data based on the first video data; writing each personalized identification information into its corresponding matching second video data to obtain each target video data packet; and transmitting each target video data packet to a video processing module for the video processing module to reconstruct the first video data based on each target video data packet.

[0061] Thus, this application solves the technical problem of video data transmission failure in related technologies. Specifically, by splitting the originally large high-quality video data into multiple smaller video data packets for transmission, this application avoids the situation where video data cannot be sent to the video processing module when the video data is large. It also ensures that the video processing module can restore the original content of the video data based on the personalized identification information contained in each of the multiple video data packets. This achieves the technical effect of enabling the head-mounted display device to quickly transmit video data to the video processing module, further enhancing the user experience when wearing the head-mounted display device.

[0062] Based on the overall concept of the video data transmission method of this application, this application provides a video data transmission method. Referring to Figure 1, Figure 1 is a flowchart illustrating the first embodiment of the video data transmission method of this application. In this embodiment, the video data transmission method includes steps S10 to S40:

[0063] Step S10: The acquired first video data is divided into multiple second video data, wherein the data volume of the second video data is smaller than that of the first video data;

[0064] It should be noted that the first video data is video data composed of multiple image frames captured by the head-mounted display device. Understandably, due to the packet size limitation of UDP (User Datagram Protocol), if the first video data exceeds 65507 bytes, it cannot be transmitted to the video processing module within the head-mounted display device. Furthermore, the first video data can specifically be RGB format video data.

[0065] In this embodiment, when the head-mounted display device is running, the video generation module in the head-mounted display device first obtains the first video data that needs to be sent to the video processing module of the head-mounted display device. The video generation module inputs the first video data into the Android encoder configured in the head-mounted display device so that the first video data can be processed by the Android encoder to obtain multiple second video data with smaller data volume.

[0066] For example, when the head-mounted display device worn by the user is AR glasses, the AR glasses first call the video generation module that it communicates with to obtain first RGB video data containing multiple image frames in RGB format. The video generation module then inputs the first RGB video data into the Android encoder configured in the head-mounted display device. The Android encoder then performs packet processing on the first RGB video data to obtain multiple second AVC video data with smaller data size and AVC format.

[0067] In this way, when the head-mounted display device acquires a first video data with a large data volume, it can process the first video data through its own configured Android encoder to obtain multiple second video data with smaller data volumes.

[0068] In one feasible implementation, step S10 may specifically include steps S101 to S103:

[0069] Step S101: Obtain the first video data;

[0070] Step S102: Encode the first video data using a preset Android encoder to obtain third video data, wherein the data volume of the third video data is smaller than that of the first video data and larger than that of the second video data.

[0071] Step S103: Determine the target video data length, and perform packet processing on the third video data according to the target video data length to obtain multiple second video data.

[0072] It should be noted that the third video data, after encoding, has a smaller file size than the first video data but larger than the second video data, and is in YUV format. Furthermore, the Android encoder is a pre-installed encoder within the head-mounted display device; it is understood that there are many ways to install the Android encoder, and this application does not impose any restrictions on this. Additionally, the target video data length is the same as the length of the split second video data.

[0073] In this embodiment, when the head-mounted display device is running, the video generation module within the head-mounted display device first acquires first video data containing multiple image frames. Then, the video generation module sends the first video data to a preset Android encoder within the head-mounted display device, whereby the Android encoder performs a format conversion operation to compress the size of the first video data, resulting in smaller third video data. Finally, the Android encoder determines the target video data length for the second video data to be packetized, and splits the third video data according to the target video data length to obtain multiple second video data sets with a smaller data size than the third video data set.

[0074] For example, please refer to Figure 2, which is a schematic diagram of the packet processing flow involved in an embodiment of the video data transmission method of this application. When the head-mounted display device worn by the user is AR glasses, the AR glasses first call its own configured video generation module to obtain first RGB video data containing multiple image frames in RGB format. Then, as shown in Figure 2, the video generation module inputs the obtained first RGB video data into the Android encoder configured in the AR glasses. The Android encoder performs C2D conversion processing on the first RGB video data to compress the volume of the first RGB video data, thereby obtaining third YUV video data with a smaller volume and a YUV format. Finally, the Android encoder determines the target video data length of the second AVC video data to be split out, and splits the third YUV video data according to the target video data length to obtain multiple second AVC video data with a data volume smaller than the third YUV video data.

[0075] In this way, by using an Android encoder to process the first video data to obtain a smaller third video data, and then splitting the third video data to obtain multiple smaller second video data, the head-mounted display device can further improve the processing speed of video data and accelerate the transmission efficiency of the video data transmission process.

[0076] In one feasible implementation, the step of "determining the target video data length" in step S103 above may specifically include steps S1031 to S1034:

[0077] Step S1031: Obtain the real-time data transmission rate;

[0078] Step S1032: Determine multiple preset data transmission rates and preset video data lengths that match each of the multiple preset data transmission rates;

[0079] Step S1033: Based on the real-time data transmission rate, filter multiple preset data transmission rates to determine the target data transmission rate;

[0080] Step S1034: Determine the preset video data length corresponding to the target data transmission rate as the target video data length.

[0081] In this embodiment, after obtaining the third video data, the Android encoder can first determine the real-time data transmission rate between the video generation module and the video processing module. Then, the Android encoder reads the storage module configured in the head-mounted display device to obtain multiple preset data transmission rates and preset video data lengths that match each of the multiple preset data transmission rates. Next, the Android encoder filters the multiple preset data transmission rates based on the obtained real-time data transmission rate to determine the target data transmission rate that matches the real-time data transmission rate. Finally, the Android encoder determines the preset video data length corresponding to the target data transmission rate as the target video data length corresponding to the split second video data.

[0082] For example, after obtaining the third YUV video data, the Android encoder can first detect the real-time data transmission rate between the video generation module and the video processing module in the AR glasses. Then, the Android encoder reads the storage module configured in the AR glasses to obtain a parameter mapping table containing multiple preset data transmission rates and preset video data lengths that match each of the preset data transmission rates. Next, the Android encoder queries the parameter mapping table based on the obtained real-time data transmission rate to compare the real-time data transmission rate with the multiple preset data transmission rates contained in the parameter mapping table, thereby determining the target data transmission rate that matches the real-time data transmission rate among the multiple preset data transmission rates. Finally, the Android encoder determines the preset video data length that matches the target data transmission rate in the parameter mapping table as the target video data length corresponding to the split second AVC video data.

[0083] In this way, the head-mounted display device can filter out the target video data length that is more compatible with the second video data based on the real-time data transmission rate between the video generation module and the video processing module, and then split the third video data into second video data of appropriate size according to the target video data length.

[0084] In one feasible implementation, the step of "determining the target video data length" in step S103 above may further include steps S1035 to S1036:

[0085] Step S1035: Determine the actual video data length corresponding to the third video data, and obtain the preset number of data packets;

[0086] Step S1036: Determine the target video data length based on the actual video data length and the preset number of data packets.

[0087] It should be noted that the preset number of data packets refers to the number of data packets for the second video data obtained after processing the third video data into packets. It is understood that the specific value of this preset number of data packets can be set by technical personnel according to actual needs, and this application does not impose any restrictions on it.

[0088] In this embodiment, after obtaining the third video data, the Android encoder can determine the target video data length based on the real-time data transmission rate, and can also first read the actual video data length of the third video data. At the same time, the Android encoder reads the aforementioned storage module to obtain the preset number of data packets. Then, the Android encoder calculates the target video data length that each second video data to be separated matches based on the actual video data length and the preset number of data packets.

[0089] For example, after obtaining the aforementioned third YUV video data, the Android encoder can, in addition to determining the target video data length based on the aforementioned parameter mapping table and real-time data transmission rate, first identify the third YUV video data to determine its actual video data length as A. Simultaneously, the Android encoder reads the aforementioned storage module to obtain the preset data packet quantity stored by the technician as 3. The Android encoder then determines that the third YUV video data needs to be divided into 3 second AVC video data. Afterward, based on the actual video data length A and the preset data packet quantity 3, the Android encoder calculates the video data length of each second AVC video data as A / 3. The Android encoder then determines the video data length A / 3 as the target video data length.

[0090] In this way, the head-mounted display device can determine the target video data length that matches the second video data based on the video data length and the preset number of data packets, and then split the third video data into second video data of appropriate size according to the target video data length.

[0091] In one feasible implementation, the step S103 above, "splitting the third video data according to the target video data length to obtain multiple second video data", may specifically include steps S1037 to S1038:

[0092] Step S1037: Determine each target image frame contained in the third video data based on the length of the target video data;

[0093] Step S1038: The third video data is divided into packets according to each target image frame to obtain multiple second video data.

[0094] In this embodiment, after determining the target video data length, the Android encoder further reads each third image frame contained in the third video data. The Android encoder filters each third image frame based on the target video data length to determine each target image frame that matches the target video data length. Finally, the Android encoder uses each target image frame as a packet node to perform packet processing on the third video data, thereby obtaining multiple second video data that meet the target video data length.

[0095] For example, after determining the target video data length, the Android encoder further reads each third image frame contained in the aforementioned third YUV video data. The Android encoder then determines the image frame number of each third image frame and filters each third image frame based on the target video data length and each image frame number to determine the splitting nodes contained in the third YUV video data and to determine each target image frame located at the splitting node. Finally, the Android encoder splits the third YUV video data according to the splitting node where each target image frame is located to obtain multiple second AVC video data sets that match the target video data length.

[0096] In this way, the head-mounted display device uses an Android encoder to split the third video data, thereby obtaining multiple second video data sets with smaller data volumes.

[0097] Step S20: Determine personalized identification information corresponding to each of the multiple second video data based on the first video data;

[0098] It should be noted that the personalized identification information is identification information that can distinguish the second video data; specifically, the personalized identification information may include frameid, which can indicate the image frame number, sliceid, which can indicate the position of the second video data within the original first video data, timestamp, which can indicate the original video content, and pose, which can indicate the pose of the AR glasses.

[0099] In this embodiment, after the Android encoder obtains multiple second video data, it inputs the multiple second video data and the aforementioned first video data into the data processing unit configured in the head-mounted display device. The data processing unit then determines the personalized identification information matched by each second video data based on the first video data.

[0100] For example, as shown in Figure 2, after obtaining multiple second AVC video data, the Android encoder inputs the multiple second AVC video data and the first RGB video data into the video frame data processing unit configured in the head-mounted display device. The video frame processing module determines the image frame identifier information frameid, data packet identifier information sliceid, timestamp information timestamp, and pose information that match each of the multiple second AVC video data according to the original first RGB video data.

[0101] In this way, the head-mounted display device can obtain different personalized identification information, and then use the different personalized identification information to reflect the differences between multiple second video data, so that the video processing module can restore multiple second video data based on the different personalized identification information.

[0102] In one feasible implementation, the personalized identification information includes image frame identification information; step S20 above may specifically include steps S201 to S203:

[0103] Step S201: Determine the plurality of first image frames contained in the first video data, and determine the plurality of second image frames contained in each of the plurality of second video data;

[0104] Step S202: Compare the plurality of first image frames and the plurality of second image frames to determine the image frame mapping relationship between the plurality of first image frames and the plurality of second image frames;

[0105] Step S203: Determine the frame generation time corresponding to each of the multiple first image frames, and based on the image frame mapping relationship and the frame generation time, determine the multiple image frame identification information corresponding to each of the multiple second video data.

[0106] In this embodiment, after obtaining multiple second video data, the Android encoder first sends the multiple second video data and the aforementioned first video data to the aforementioned data processing unit. At this time, the data processing unit reads the first video data to determine each first image frame contained in the first video data. Simultaneously, the data processing unit reads the multiple second video data to determine the multiple second image frames contained in each of the multiple second video data. Then, the data processing unit compares each first image frame and each second image frame to determine the image frame mapping relationship between each first image frame and each second image frame. Finally, the data processing unit reads the frame generation time corresponding to each first image frame and, based on the frame generation time and the image frame mapping relationship, determines the frame generation time corresponding to each second image frame. The data processing unit generates image frame identification information based on the frame generation time and determines the image frame identification information corresponding to each second image frame.

[0107] For example, after obtaining multiple second AVC video data, the Android encoder first inputs the multiple second AVC video data and the aforementioned first RGB video data into the aforementioned data processing unit. The data processing unit first reads each first image frame contained in the first RGB video data and each second image frame contained in each of the multiple second AVC video data. Then, the data processing unit compares each first image frame and each second image frame to determine the second image frame that matches each first image frame, and then determines the image frame mapping relationship generated between each first image frame and each second image frame. Finally, the data processing unit determines the frame generation time corresponding to each first image frame, and determines the frame generation time corresponding to each second image frame based on the frame generation time and the image frame mapping relationship. The data processing unit then generates the image frame identification information frameid that matches each second image frame based on the frame generation time.

[0108] In this way, the head-mounted display device can obtain different personalized identification information, and then use the different personalized identification information to reflect the differences between multiple second video data, so that the video processing module can restore multiple second video data based on the different personalized identification information.

[0109] In one feasible implementation, the personalized identification information further includes data packet identification information; after step S202 above, the video data transmission method of this application may further include steps S204 to S205:

[0110] Step S204: Sort the multiple second video data based on the image frame mapping relationship to obtain a sorting result;

[0111] Step S205: Determine the data packet identification information corresponding to each of the multiple second video data according to the sorting result.

[0112] In this embodiment, after obtaining the above-mentioned image frame mapping relationship, the data processing unit can also determine the frame generation time of each first image frame, and sort each second video data according to the frame generation time and the image frame mapping relationship to obtain a second sorting result. Then, the data processing unit generates data packet identification information corresponding to each of the multiple second video data according to the second sorting result.

[0113] For example, after obtaining the above-mentioned image frame mapping relationship, the data processing unit may first determine the frame generation time corresponding to each first image frame, and sort each second image frame according to the frame generation time and image frame mapping relationship to obtain a sorting result. Then, the data processing unit determines the positional relationship between multiple second AVC video data based on the sorting result, and generates multiple data packet identification information sliceid that matches the multiple second AVC video data respectively and can indicate the positional relationship.

[0114] In this way, the head-mounted display device can obtain different personalized identification information, and then use the different personalized identification information to reflect the differences between multiple second video data, so that the video processing module can restore multiple second video data based on the different personalized identification information.

[0115] In one feasible implementation, the personalized identification information further includes content identification information; step S20 above may also include steps S206 to S207:

[0116] Step S206: Determine the frame content data contained in each of the multiple first image frames;

[0117] Step S207: Based on the mapping relationship between the content data of each frame and the image frame, determine the content identification information corresponding to each of the multiple second video data.

[0118] It should be noted that the content identification information is the identification information that can reflect the image content such as timestamp and pose of the second video data, specifically including the timestamp information and pose information mentioned above.

[0119] In this embodiment, the Android encoder, after obtaining multiple second video data and inputting the multiple second video data and the first video data to the data processing unit so that the data processing unit can obtain the image frame identification information and data packet identification information, can also first read the multiple first image frames contained in the first video data and determine the timestamp data corresponding to each of the multiple first image frames. Then, the data processing unit reads the multiple second video data and determines the second image frames contained in each of the multiple second video data, thereby determining the image frame mapping relationship between each first image frame and each second image frame. Based on the image frame mapping relationship, the data processing unit determines the frame content data that matches each of the second image frames. The data processing unit generates multiple timestamp information based on each frame content data and determines the timestamp information that matches each of the multiple second video data. Similarly, the data processing unit can also first determine the pose data corresponding to each of the first image frames and determine the pose data that matches each of the second image frames according to the image frame mapping relationship. The data processing unit then generates multiple pose information based on each pose data and determines the pose information that matches each of the multiple second video data.

[0120] For example, after the Android encoder generates multiple second AVC video data, it inputs the multiple second AVC video data and the first RGB video data to the aforementioned data processing unit. The data processing unit first reads each first image frame contained in the first RGB video data and determines the frame content data contained in each of the multiple first image frames. Then, the data processing unit reads multiple second AVC video data and determines the second image frames contained in each of the multiple second AVC video data. The data processing unit then compares each first image frame and each second image frame to determine the image frame mapping relationship between each first image frame and each second image frame. The data processing unit then determines the frame content data corresponding to each of the second image frames based on the image frame mapping relationship. The data processing unit generates multiple timestamp information based on each frame content data and determines the timestamp information corresponding to each of the multiple second AVC video data.

[0121] Similarly, in addition to reading the frame content data contained in the first RGB video data to generate timestamp information, the data processing unit can also first read the pose data contained in each first image frame. Then, the data processing unit reads multiple second AVC video data and determines the second image frames contained in each of the multiple second AVC video data, thereby determining the image frame mapping relationship between each second image frame and each first image frame. After that, the data processing unit determines the pose data matched by each second image frame based on the image frame mapping relationship. The data processing unit then generates multiple pose information based on each pose data and determines the pose information corresponding to each second AVC video data.

[0122] In this way, the head-mounted display device can obtain different personalized identification information, and then use the different personalized identification information to reflect the differences between multiple second video data, so that the video processing module can restore multiple second video data based on the different personalized identification information.

[0123] Step S30: Write each of the personalized identification information into the corresponding matching second video data to obtain each target video data packet;

[0124] In this embodiment, after obtaining each personalized identifier information, the data processing unit reads the data storage space contained in each second video data and writes each personalized identifier information into the data storage space of the corresponding second video data to obtain multiple target video data packets containing different personalized identifier information.

[0125] For example, after obtaining the frameid, sliceid, timestamp, and pose information of each matching second AVC video data, the data processing unit further writes the frameid, sliceid, timestamp, and pose information of each matching second AVC video data into the data storage space of each matching second AVC video data, thereby obtaining multiple target video data packets containing different frameid, sliceid, timestamp, and pose information.

[0126] In this way, the head-mounted display device can write different personalized identification information into the multiple split second video data to determine whether the video processing module can restore the original content of the video data based on the multiple personalized identification information.

[0127] In one feasible implementation, the personalized identification information further includes content identification information; step S30 above may also include step S301:

[0128] Step S301: Write each of the personalized identification information into the tail storage space of the corresponding matching second video data to obtain each target video data packet.

[0129] It should be noted that, referring to Figure 3, which is a schematic diagram of the header of the target data packet involved in an embodiment of the video data transmission method of this application, in related technologies, head-mounted display devices typically need to request additional storage space from memory when adding personalized identification information to video data. The personalized identification information is then written into this storage space, and the storage space containing the personalized identification information is combined with the video data as the header shown in Figure 3 to obtain the video data packet. However, the process of requesting additional storage space increases memory copy latency, thereby reducing the transmission efficiency of video data and further degrading the user experience. Based on the encoding characteristics of the Android encoder, there is a space at the end of the data frame of the video data output by the Android encoder that is much larger than the storage space for the personalized identification information. Therefore, writing the personalized information data into this end storage space can avoid the head-mounted display device performing memory request operations, thereby reducing memory copy latency and improving the transmission efficiency of video data.

[0130] In this embodiment, after obtaining the personalized information corresponding to each second video data, the data processing unit first determines the tail storage space contained in each second video data. The data processing unit writes each personalized identification information into the tail storage space of the matching second video data to obtain target video data packets containing different personalized identification information.

[0131] For example, please refer to Figure 4, which is a schematic diagram of the tail storage space of the target data packet according to an embodiment of the video data transmission method of this application. After obtaining the image frame identifier information frameid, data packet identifier information sliceid, timestamp information timestamp, and pose information matching each of the second AVC video data, the data processing unit first reads the tail storage space contained in each of the second AVC video data. Then, without re-allocating memory, the data processing unit inputs the obtained image frame identifier information frameid, data packet identifier information sliceid, timestamp information timestamp, and pose information pose to the corresponding tail storage space of each matching second AVC video data. As shown in Figure 4, multiple target video data packets containing different personalized identifier information are obtained.

[0132] In this way, the head-mounted display device can write different personalized identification information into the multiple split second video data to determine whether the video processing module can restore the original content of the video data based on the multiple personalized identification information.

[0133] Step S40: Transmit each of the target video data packets to the video processing module so that the video processing module can reconstruct the first video data based on each of the target video data packets;

[0134] In this embodiment, after obtaining each target video data packet, the data processing unit sends each target video data packet to the video processing module in the head-mounted display device. The video processing module sorts each target video data packet based on the personalized identification information contained in each target video data packet, thereby restoring the first video data.

[0135] For example, after the data processing unit obtains each target video data packet, the video generation module inputs each target video data packet into the video processing module configured in the AR glasses. The video processing module sorts each target video data packet based on the sliceid information contained in each target video data packet, thereby restoring the aforementioned first video data. At the same time, the video processing module can also determine whether there are any frame drops in each target video data packet based on the frameid information contained in each target video data packet. In addition, the video processing module can also perform corresponding video processing operations based on other auxiliary information such as timestamp information and pose information contained in each target video data packet.

[0136] In this embodiment, when the head-mounted display device is running, the video generation module within the head-mounted display device first acquires the first video data that needs to be sent to the video processing module of the head-mounted display device. The video generation module inputs the first video data into the Android encoder configured within the head-mounted display device, so that the Android encoder processes the first video data to obtain multiple second video data with smaller data volumes. Then, the Android encoder inputs the multiple second video data and the aforementioned first video data into the data processing unit configured within the head-mounted display device. The data processing unit determines the personalized identification information that matches each second video data based on the first video data. Then, the data processing unit reads the data storage space contained in each second video data and writes each personalized identification information into the data storage space of the corresponding second video data to obtain multiple target video data packets containing different personalized identification information. Finally, the data processing unit sends each target video data packet to the video processing module within the head-mounted display device. The video processing module sorts the target video data packets based on the personalized identification information contained in each target video data packet to reconstruct the first video data.

[0137] Thus, this application solves the technical problem of video data transmission failure in related technologies. Specifically, by splitting the originally large high-quality video data into multiple smaller video data packets for transmission, this application avoids the situation where video data cannot be sent to the video processing module when the video data is large. It also ensures that the video processing module can restore the original content of the video data based on the personalized identification information contained in each of the multiple video data packets. This achieves the technical effect of enabling the head-mounted display device to quickly transmit video data to the video processing module, further enhancing the user experience when wearing the head-mounted display device.

[0138] For example, to help understand the implementation flow of the video data transmission method obtained by combining the above embodiments, please refer to Figure 5. Figure 5 is a simplified flowchart of the video data transmission method of this application. Specifically:

[0139] In this embodiment, when a user wears AR glasses, the AR glasses first call the video generation module connected to itself to obtain first video data containing multiple first image frames. The video generation module inputs the obtained first video data into the Android encoder configured in the AR glasses. The Android encoder performs format conversion on the first video data to obtain third video data with a smaller size than the first video data. At the same time, the Android encoder determines the real-time data transmission rate between the video generation module and the video processing module in the AR glasses, and determines the target video data length based on the real-time data transmission rate. The Android encoder then identifies the third video data according to the target video length to determine each packet node in the third video data and each target image frame on the packet node. The Android encoder performs packet processing on the third video data based on each target image frame to obtain multiple second video data with a smaller data size than the third video data.

[0140] Subsequently, the Android encoder inputs the generated second video data and first video data into the data processing unit configured in the video generation module. The data processing unit first reads the first image frames contained in the first video data and reads the second image frames contained in each of the second video data to determine the image frame mapping relationship between each first image frame and each second image frame. The data processing unit then reads the frame generation time, timestamp data, and pose data contained in each first image frame. Based on the frame generation time, timestamp data, pose data, and image frame mapping relationship, the data processing unit determines the personalized identification information such as image frame identification information, data packet identification information, timestamp information, and pose information corresponding to each of the second video data.

[0141] Then, the data processing unit writes the image frame identification information, data packet identification information, timestamp information and pose information of each image frame into the tail storage space of their respective second video data, so as to form multiple target video data packets containing different image frame identification information, data packet identification information, timestamp information and pose information.

[0142] Finally, the video generation module sends the generated target video data packets to the video processing unit inside the AR glasses. The video processing unit reads the data packet identification information contained in each target video data packet, and then restores each target video data packet according to the data packet identification information to obtain the first video data. At the same time, the video processing unit also reads the image frame identification information contained in each target video data packet to determine whether there are any frame drops in each target video data packet. In addition, the video processing unit also reads the timestamp information and pose information contained in each target video data packet to perform video rendering operations according to the timestamp information and pose information.

[0143] It should be noted that the above examples are only for understanding this application and do not constitute a limitation on the video data transmission method of this application. Any simple modifications based on this technical concept are within the protection scope of this application.

[0144] This application also provides a video data transmission device. Referring to FIG6, the video data transmission device includes:

[0145] The data splitting module 10 is used to split the acquired first video data into multiple second video data, wherein the data volume of the second video data is smaller than that of the first video data.

[0146] The identifier generation module 20 is used to determine personalized identifier information corresponding to each of the multiple second video data based on the first video data;

[0147] The identifier adding module 30 is used to write each of the personalized identifier information into the corresponding matching second video data to obtain each target video data packet;

[0148] The data transmission module 40 is used to transmit each of the target video data packets to the video processing module, so that the video processing module can restore the first video data based on each of the target video data packets.

[0149] In one feasible implementation, the personalized identification information includes image frame identification information; the identification generation module 20 is further configured to:

[0150] The first video data contains a plurality of first image frames, and the plurality of second video data each contains a plurality of second image frames;

[0151] A plurality of first image frames and a plurality of second image frames are compared to determine the image frame mapping relationship between the plurality of first image frames and the plurality of second image frames;

[0152] The frame generation time corresponding to each of the multiple first image frames is determined, and based on the image frame mapping relationship and the frame generation time, the multiple image frame identification information corresponding to each of the multiple second video data is determined.

[0153] In one feasible implementation, the personalized identification information further includes data packet identification information; the identification generation module 20 is further configured to:

[0154] Based on the image frame mapping relationship, sorting is performed on multiple second video data to obtain a sorting result;

[0155] Based on the sorting result, the data packet identification information corresponding to each of the multiple second video data is determined.

[0156] In one feasible implementation, the personalized identification information further includes content identification information; the identification generation module 20 is further configured to:

[0157] Determine the frame content data contained in each of the multiple first image frames;

[0158] Based on the content data of each frame and the mapping relationship between the image frames, the content identification information corresponding to each of the multiple second video data is determined.

[0159] In one feasible implementation, the aforementioned identifier adding module 30 is further configured to:

[0160] Each personalized identifier is written into the tail storage space of the corresponding second video data to obtain each target video data packet.

[0161] In one feasible implementation, the data splitting module 10 is further configured to:

[0162] Obtain the first video data;

[0163] The first video data is encoded using a preset Android encoder to obtain the third video data, wherein the data volume of the third video data is smaller than that of the first video data and larger than that of the second video data.

[0164] The target video data length is determined, and the third video data is divided into multiple second video data according to the target video data length.

[0165] In one feasible implementation, the data splitting module 10 is further configured to:

[0166] Obtain the real-time data transmission rate;

[0167] Determine multiple preset data transmission rates and preset video data lengths that match each of the multiple preset data transmission rates;

[0168] Based on the real-time data transmission rate, multiple preset data transmission rates are filtered to determine the target data transmission rate;

[0169] The preset video data length corresponding to the target data transmission rate is determined as the target video data length.

[0170] In one feasible implementation, the data splitting module 10 is further configured to:

[0171] Determine the length of the actual video data corresponding to the third video data, and obtain the preset number of data packets;

[0172] The target video data length is determined based on the actual video data length and the preset number of data packets.

[0173] In one feasible implementation, the data splitting module 10 is further configured to:

[0174] The target image frames contained in the third video data are determined based on the length of the target video data.

[0175] The third video data is divided into packets based on each target image frame to obtain multiple second video data.

[0176] The video data transmission apparatus provided in this application, employing the video data transmission method described in the above embodiments, can solve the technical problem of video data transmission failure in related technologies. Compared with the prior art, the beneficial effects of the video data transmission apparatus provided in this application are the same as those of the video data transmission method provided in the above embodiments, and other technical features in the video data transmission apparatus are the same as those disclosed in the methods of the above embodiments, and will not be repeated here.

[0177] This application provides a head-mounted display device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, which are executed by the at least one processor to enable the at least one processor to perform the video data transmission method in Embodiment 1 above.

[0178] Referring to Figure 7 below, a schematic diagram of a head-mounted display device suitable for implementing embodiments of this application is shown. The head-mounted display device in the embodiments of this application may include, but is not limited to, a head-mounted display device with an Android encoder, or a mobile terminal, data storage control terminal, PC, or other terminal connected to an electronic control unit associated with the head-mounted display device. The head-mounted display device shown in Figure 7 is merely an example and should not impose any limitations on the functionality and scope of use of the embodiments of this application.

[0179] As shown in Figure 7, the head-mounted display device may include a processing unit 1001 (e.g., a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 1002 or a program loaded from a storage device 1003 into a random access memory (RAM) 1004. The RAM 1004 also stores various programs and data required for the operation of the head-mounted display device. The processing unit 1001, ROM 1002, and RAM 1004 are interconnected via a bus 1005. An input / output (I / O) interface 1006 is also connected to the bus. Typically, the following systems can be connected to I / O interface 1006: input devices 1007 including, for example, touchscreens, touchpads, keyboards, mice, image sensors, microphones, accelerometers, gyroscopes, etc.; output devices 1008 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 1003 including, for example, magnetic tapes, hard disks, etc.; and communication devices 1009. Communication device 1009 allows the head-mounted display to communicate wirelessly or wiredly with other devices to exchange data. Although head-mounted display devices with various systems are shown in the figures, it should be understood that it is not required to implement or possess all the systems shown. More or fewer systems may be implemented alternatively.

[0180] Specifically, according to the embodiments disclosed in this application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments disclosed in this application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device, or installed from storage device 1003, or installed from ROM 1002. When the computer program is executed by processing device 1001, it performs the functions defined in the methods of the embodiments disclosed in this application.

[0181] The head-mounted display device provided in this application, employing the video data transmission method described in the above embodiments, can solve the technical problem of video data transmission failure in related technologies. Compared with the prior art, the beneficial effects of the head-mounted display device provided in this application are the same as those of the video data transmission method provided in the above embodiments, and other technical features of this head-mounted display device are the same as those disclosed in the method of the previous embodiment, and will not be repeated here.

[0182] It should be understood that the various parts disclosed in this application can be implemented using hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in any suitable manner in one or more embodiments or examples.

[0183] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

[0184] This application provides a computer-readable storage medium having computer-readable program instructions (i.e., a computer program) stored thereon, the computer-readable program instructions being used to execute the video data transmission method described in the above embodiments.

[0185] The computer-readable storage medium provided in this application may be, for example, a USB flash drive, but is not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems or devices, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to: electrical connections having one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this embodiment, the computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system or device. The program code contained on the computer-readable storage medium may be transmitted using any suitable medium, including but not limited to: wires, optical cables, RF (Radio Frequency), etc., or any suitable combination thereof.

[0186] The aforementioned computer-readable storage medium may be included in the head-mounted display device; or it may exist independently and not assembled into the head-mounted display device.

[0187] The aforementioned computer-readable storage medium carries one or more programs that, when executed by a head-mounted display device, cause the head-mounted display device to: perform packet processing on the acquired first video data to obtain multiple second video data, wherein the data volume of the second video data is smaller than that of the first video data; determine personalized identification information corresponding to each of the multiple second video data based on the first video data; write each of the personalized identification information into its corresponding matching second video data to obtain each target video data packet; and transmit each of the target video data packets to a video processing module so that the video processing module can reconstruct the first video data based on each of the target video data packets.

[0188] Computer program code for performing the operations of this application can be written in one or more programming languages ​​or a combination thereof, including object-oriented programming languages ​​such as Java, Smalltalk, and C++, and conventional procedural programming languages ​​such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a Local Area Network (LAN) or a Wide Area Network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0189] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0190] The modules described in the embodiments of this application can be implemented in software or hardware. The names of the modules do not necessarily limit the functionality of the unit itself.

[0191] The readable storage medium provided in this application is a computer-readable storage medium that stores computer-readable program instructions (i.e., a computer program) for executing the above-described video data transmission method, and can solve the technical problem of video data transmission failure in related technologies. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided in this application are the same as the beneficial effects of the video data transmission method provided in the above embodiments, and will not be repeated here.

[0192] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the video data transmission method described above.

[0193] The computer program product provided in this application can solve the technical problem of video data transmission failure in related technologies. Compared with the prior art, the beneficial effects of the computer program product provided in this application are the same as the beneficial effects of the video data transmission method provided in the above embodiments, and will not be repeated here.

[0194] The above description is only a part of the embodiments of this application and does not limit the patent scope of this application. All equivalent structural transformations made under the technical concept of this application and using the contents of the specification and drawings of this application, or direct / indirect applications in other related technical fields, are included in the patent protection scope of this application.

Claims

1. A method for transmitting video data, characterized in that, The video data transmission method includes: splitting the acquired first video data into packets to obtain multiple second video data, wherein the data volume of the second video data is smaller than that of the first video data; determining personalized identification information corresponding to each of the multiple second video data based on the first video data; writing each personalized identification information into its corresponding matching second video data to obtain each target video data packet; and transmitting each target video data packet to a video processing module so that the video processing module can reconstruct the first video data based on each target video data packet.

2. The video data transmission method as described in claim 1, characterized in that, The personalized identification information includes image frame identification information; the step of determining the personalized identification information corresponding to each of the multiple second video data based on the first video data includes: determining multiple first image frames contained in the first video data, and determining multiple second image frames contained in each of the multiple second video data; comparing the multiple first image frames and the multiple second image frames to determine the image frame mapping relationship between the multiple first image frames and the multiple second image frames; determining the frame generation time corresponding to each of the multiple first image frames, and determining the multiple image frame identification information corresponding to each of the multiple second video data based on the image frame mapping relationship and each frame generation time.

3. The video data transmission method as described in claim 2, characterized in that, The personalized identification information also includes data packet identification information; After determining the image frame mapping relationship between the plurality of first image frames and the plurality of second image frames, the method further includes: sorting the plurality of second video data based on the image frame mapping relationship to obtain a sorting result; Based on the sorting result, the data packet identification information corresponding to each of the multiple second video data is determined.

4. The video data transmission method as described in claim 2, characterized in that, The personalized identification information also includes content identification information; the step of determining the personalized identification information corresponding to each of the multiple second video data based on the first video data further includes: determining the frame content data contained in each of the multiple first image frames; and determining the content identification information corresponding to each of the multiple second video data based on the mapping relationship between each frame content data and the image frame.

5. The video data transmission method as described in claim 1, characterized in that, The step of writing each of the personalized identifiers into the corresponding matching second video data to obtain each target video data packet includes: writing each of the personalized identifiers into the tail storage space of the corresponding matching second video data to obtain each target video data packet.

6. The video data transmission method as described in claim 1, characterized in that, The step of subpackaging the acquired first video data to obtain multiple second video data includes: acquiring the first video data; encoding the first video data using a preset Android encoder to obtain third video data, wherein the data volume of the third video data is smaller than that of the first video data and larger than that of the second video data; determining the target video data length, and subpackaging the third video data according to the target video data length to obtain multiple second video data.

7. The video data transmission method as described in claim 6, characterized in that, The step of determining the target video data length includes: acquiring a real-time data transmission rate; determining multiple preset data transmission rates and preset video data lengths that match each of the multiple preset data transmission rates; filtering the multiple preset data transmission rates based on the real-time data transmission rate to determine a target data transmission rate; and determining the preset video data length corresponding to the target data transmission rate as the target video data length.

8. The video data transmission method as described in claim 6, characterized in that, The step of determining the target video data length further includes: determining the corresponding actual video data length of the third video data and obtaining a preset number of data packets; and determining the target video data length based on the actual video data length and the preset number of data packets.

9. The video data transmission method as described in claim 6, characterized in that, The step of dividing the third video data into packets based on the target video data length to obtain multiple second video data includes: determining each target image frame contained in the third video data based on the target video data length; and dividing the third video data into packets based on each target image frame to obtain multiple second video data.

10. A head-mounted display device, characterized in that, The device includes: a memory, a processor, and a computer program stored in the memory and executable on the processor, the computer program being configured to implement the steps of the video data transmission method as described in any one of claims 1 to 9.