Method and apparatus for merging data across encoding formats
Patent Information
- Application Number
- CN202611217098.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-08-12
- Publication Date
- 2026-09-11
AI Technical Summary
[0005]本申请提供了一种跨编码格式的数据合并方法及装置,旨在解决目前由于不同编码格式的编码结构差异导致代码分散重复、新旧参数冲突,使得数据传输效率低下的技术问题
[0008]本申请提供一种跨编码格式的数据合并方法及装置,本申请方法通过接收来自摄像头图像信号处理器的原始图像数据,按照预设编码参数进行编码输出NAL单元,并通过读取NAL单元的头部信息识别编码格式类型以及NAL类型,将NAL单元封装为统一描述符,将不同编码格式类型对应的NAL类型映射为统一格式的统一NAL类型。通过统一NAL类型将不同编码格式中语义相同但原始枚举值不同的NAL类型映射为同一枚举值,使上层处理逻辑与编码格式无关,从编码格式识别层面消除因格式差异导致的重复处理开销。在统一NAL类型为参数集类型时,从统一描述符中提取当前参数集数据,将当前参数集数据输入参数集版本池进行版本管理并确定版本号信息,同时生成对应的参数集数据包,通过版本号追踪参数集变更,从而避免重复传输未变化的参数集,降低参数集重复传输率。在统一NAL类型为访问单元分隔符或补充增强信息时,直接生成对应的非关键帧数据包,无需经过参数集版本池管理和关键帧合并流程,减少不必要的处理开销,提高非关键帧的传输效率。在统一NAL类型为关键帧类型时,从统一描述符中提取标识信息,从参数集版本池中提取与标识信息相对应的至少一个目标参数集数据以及各目标参数集数据对应的当前版本号信息,通过对比历史版本号信息和当前版本号信息,确定目标参数集数据的版本变更状态,进而根据版本变更状态确定数据合并模式,使得数据合并模式与实际需求相适配,减少了无效数据拼接带来的带宽浪费。基于确定的数据合并模式将统一描述符中的关键帧数据与参数集版本池中对应的当前版本号信息的各目标参数集数据进行合并,获得关键帧数据包,使参数集与关键帧随行发送,客户端一次性获取完整解码参数和关键帧数据,无需等待额外参数集即可开始解码,降低解码等待延迟。将关键帧数据包、参数集数据包和非关键帧数据包分别缓存至对应优先级的目标优先级队列,作为优先级队列中的缓存数据包,实现数据包的分级缓存,基于预设的优先级调度策略,按照预设的多级优先级顺序,将缓存于优先级队列中的各缓存数据包依次发送至客户端,通过轮询机制确保高优先级的缓存数据包优先发送,避免关键帧被非关键帧阻塞,保障关键数据的传输实时性,提高数据传输效率。
Smart Images

Figure CN122741501A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of communication technology, and in particular to a method and apparatus for merging data across encoding formats. Background Technology
[0002] With the rapid development of IoT and smart home technologies, embedded real-time communication (RTC) technology has been widely used in video surveillance, smart doorbells, and video access control. The RTC SDK (Software Development Kit) acts as a bridge connecting embedded devices and clients, responsible for transmitting video data collected by the device in real time to mobile apps and other clients for decoding and display.
[0003] In embedded real-time communication scenarios, the transmission efficiency and real-time performance of video data directly impact user experience. Video data captured by a camera, after being compressed by an encoder, needs to be transmitted to the client via a network. Encoders typically use the H.264 or H.265 video encoding standards, which are currently the mainstream video encoding standards. However, the NAL header structures of the two encoding formats are completely different, and the number of parameter set types also differs (H.264 has two types, H.265 has three). Traditional video transmission systems require separate processing code for each encoding format, resulting in fragmented and repetitive NAL processing code for H.264 and H.265, conflicts between old and new parameters, significant bandwidth waste, and low data transmission efficiency.
[0004] Therefore, how to improve data transmission efficiency is a technical problem that urgently needs to be solved in this field. Summary of the Invention
[0005] This application provides a method and apparatus for merging data across encoding formats, aiming to solve the technical problem of low data transmission efficiency caused by code fragmentation and duplication and conflicts between old and new parameters due to differences in the encoding structure of different encoding formats.
[0006] In a first aspect, this application provides a method for merging data across encoding formats, the method comprising the following steps: The system receives raw image data from the image signal processor of the camera, encodes the raw image data according to preset encoding parameters, and outputs at least one NAL unit. By reading the header information of the NAL unit, the encoding format type and NAL type of the NAL unit are identified, and the NAL unit is encapsulated into a unified descriptor of a unified format, so that the NAL types of different encoding formats are mapped to a unified NAL type. The unified NAL type includes parameter set type, keyframe type, non-keyframe type, access unit separator and supplementary enhancement information. When the unified NAL type is a parameter set type, the current parameter set data is extracted from the unified descriptor and input into the parameter set version pool for version management to determine the version number information of the current parameter set data; and a parameter set data package corresponding to the current parameter set data is generated. When the unified NAL type is a non-critical frame type, access unit separator, or supplementary enhancement information, the corresponding non-critical frame data packet is directly generated. When the unified NAL type is a keyframe type, identification information is extracted from the unified descriptor, and at least one target parameter set data corresponding to the identification information and the current version number information corresponding to each target parameter set data are extracted from the parameter set version pool. Obtain the historical version number information corresponding to the historical parameter set data of the NAL unit of the key frame type that was previously merged, and determine the version change status of the target parameter set data by comparing the historical version number information with the current version number information; Determine the data merging mode based on the version change status; Based on the data merging mode, the keyframe data in the unified descriptor is merged with the target parameter set data corresponding to the current version number information in the parameter set version pool to obtain the keyframe data packet; The keyframe data packets, the parameter set data packets, and the non-keyframe data packets are respectively cached in priority queues of corresponding priorities, and are used as cached data packets in the priority queues; Based on a preset priority scheduling strategy, each cached data packet in the priority queue is sent to the client in a preset multi-level priority order so that the client can decode and display the cached data packets.
[0007] Secondly, this application also provides a cross-encoding format data merging apparatus, the cross-encoding format data merging apparatus comprising: The data encoding module is used to receive raw image data from the image signal processor of the camera, encode the raw image data according to preset encoding parameters, and output at least one NAL unit. The encoding type identification module is used to identify the encoding format type and NAL type of the NAL unit by reading the header information of the NAL unit, and encapsulate the NAL unit into a unified descriptor of a unified format, so that the NAL types of different encoding formats are mapped to a unified NAL type. The unified NAL type includes parameter set type, keyframe type, non-keyframe type, access unit separator and supplementary enhancement information. The version management module is used to extract the current parameter set data from the unified descriptor when the unified NAL type is a parameter set type, and input the current parameter set data into the parameter set version pool for version management, and determine the version number information of the current parameter set data; The non-critical frame management module is used to directly generate the corresponding non-critical frame data packet when the unified NAL type is a non-critical frame type, access unit separator, or supplementary enhancement information. The version number information extraction module is used to extract identification information from the unified descriptor when the unified NAL type is a keyframe type, and to extract at least one target parameter set data corresponding to the identification information and the current version number information corresponding to each target parameter set data from the parameter set version pool. The version change status determination module is used to obtain the historical version number information corresponding to the historical parameter set data of the NAL unit of the key frame type that was merged last time, and to determine the version change status of the target parameter set data by comparing the historical version number information and the current version number information. The merge mode determination module is used to determine the data merge mode based on the version change status. The data merging module is used to merge the keyframe data in the unified descriptor with the target parameter set data corresponding to the current version number information in the parameter set version pool based on the data merging mode, so as to obtain a keyframe data packet. The data caching module is used to cache the key frame data packets, the parameter set data packets, and the non-key frame data packets into priority queues of corresponding priorities, as cached data packets in the priority queues; The data transmission module is used to send each cached data packet in the priority queue to the client in a preset multi-level priority order based on a preset priority scheduling strategy, so that the client can decode and display the cached data packets.
[0008] This application provides a method and apparatus for data merging across encoding formats. The method receives raw image data from a camera image signal processor, encodes it according to preset encoding parameters, and outputs NAL units. It identifies the encoding format type and NAL type by reading the header information of the NAL units, encapsulates the NAL units into a unified descriptor, and maps NAL types corresponding to different encoding format types to a unified NAL type. By using a unified NAL type, NAL types with the same semantics but different original enumeration values in different encoding formats are mapped to the same enumeration value, making the upper-layer processing logic independent of the encoding format and eliminating the overhead of redundant processing caused by format differences at the encoding format identification level. When the unified NAL type is a parameter set type, the current parameter set data is extracted from the unified descriptor, input into a parameter set version pool for version management, and version number information is determined. Simultaneously, a corresponding parameter set data packet is generated. The version number tracks parameter set changes, thereby avoiding the repeated transmission of unchanged parameter sets and reducing the parameter set duplication rate. When the unified NAL type is an access unit separator or supplementary enhancement information, the corresponding non-critical frame data packets are directly generated without going through parameter set version pool management and keyframe merging processes, reducing unnecessary processing overhead and improving the transmission efficiency of non-critical frames. When the unified NAL type is a keyframe type, identification information is extracted from the unified descriptor, and at least one target parameter set data corresponding to the identification information and the current version number information corresponding to each target parameter set data are extracted from the parameter set version pool. By comparing the historical version number information and the current version number information, the version change status of the target parameter set data is determined, and then the data merging mode is determined according to the version change status, so that the data merging mode is adapted to the actual needs and reduces the bandwidth waste caused by invalid data splicing. Based on the determined data merging mode, the keyframe data in the unified descriptor is merged with the target parameter set data of each target parameter set with the corresponding current version number information in the parameter set version pool to obtain the keyframe data packet. The parameter set and keyframe are sent together, and the client obtains the complete decoding parameters and keyframe data at one time, without waiting for additional parameter sets to start decoding, reducing decoding waiting latency. Keyframe data packets, parameter set data packets, and non-keyframe data packets are cached in target priority queues of corresponding priorities, serving as cached data packets in the priority queues. This achieves hierarchical caching of data packets. Based on a preset priority scheduling strategy, each cached data packet in the priority queue is sent to the client sequentially according to a preset multi-level priority order. A polling mechanism ensures that high-priority cached data packets are sent first, preventing keyframes from being blocked by non-keyframes, ensuring the real-time transmission of critical data, and improving data transmission efficiency. Attached Figure Description
[0009] To more clearly illustrate the technical solutions of the embodiments of this application, the drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0010] Figure 1 A flowchart illustrating an embodiment of a cross-encoding format data merging method provided in this application; Figure 2 A schematic flowchart of an extended embodiment of the cross-encoding format data merging method provided in this application, specifically step S103; Figure 3 This application provides a state transition diagram for parameter set entries in the parameter set version pool. Figure 4 A schematic diagram illustrating the execution process of the priority scheduling strategy polling mechanism provided in this application embodiment; Figure 5 A flowchart illustrating a specific embodiment of the cross-encoding format data merging method provided in this application; Figure 6 This is a schematic diagram of the structure of a first embodiment of a cross-encoding format data merging device provided in this application; Figure 7 This is a schematic block diagram of the structure of a computer device provided in an embodiment of this application.
[0011] The realization of the purpose, functional features and advantages of this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation
[0012] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0013] The flowchart shown in the attached diagram is for illustrative purposes only and does not necessarily include all content and operations / steps, nor does it necessarily have to be performed in the order described. For example, some operations / steps can be broken down, combined, or partially merged, so the actual execution order may change depending on the actual situation.
[0014] The following detailed description of some embodiments of this application is provided in conjunction with the accompanying drawings. Unless otherwise specified, the following embodiments and features can be combined with each other.
[0015] Please refer to Figure 1 , Figure 1 This is a flowchart illustrating an embodiment of a cross-encoding format data merging method provided in this application.
[0016] like Figure 1 As shown, the cross-encoding format data merging method includes steps S101 to S109.
[0017] S101. Receive raw image data from the image signal processor of the camera, encode the raw image data according to preset encoding parameters, and output at least one NAL unit.
[0018] In embedded real-time communication scenarios, video data originates from the light signal acquisition of the camera module. Taking an IPC (IP Camera, network camera) as an example, the image sensor inside the camera converts the light signal into an electrical signal, outputting raw image data. This raw image data is then sent to an image signal processor (ISP) for processing, including noise reduction, white balance, color correction, and other operations, outputting standard YUV format raw image frames.
[0019] The raw YUV image frames are fed into the Video Encoding Subsystem (VENC) for compression encoding. The encoder encodes the YUV image frames into a video stream in H.264 or H.265 format according to preset encoding parameters (such as resolution, frame rate, bit rate, etc.). The encoder can be a hardware encoder (such as the VENC module inside the Hisilicon Hi3516 / Hi3518 series chip) or a software encoder.
[0020] For each frame processed, the encoder outputs one or more NAL (Network Abstraction Layer) units. For example, for H.265 encoded keyframes (IDR (Instantaneous Decoding Refresh) frames), the encoder typically outputs four NAL units: Video Parameter Set (VPS), Sequence Parameter Set (SPS), Picture Parameter Set (PPS), and the keyframe data itself. For ordinary frames (P-frames or B-frames), the encoder typically outputs only one NAL unit.
[0021] Among them, the Video Parameter Set (VPS) is a parameter set type unique to the H.265 encoding format, carrying global decoding parameters at the video level, such as quality level and grade. The Sequence Parameter Set (SPS) is a parameter set type shared by H.264 and H.265, carrying sequence-level decoding parameters such as resolution, frame rate, and bitrate. The Picture Parameter Set (PPS) is a parameter set type shared by H.264 and H.265, carrying picture-level decoding parameters such as entropy coding mode and slice grouping.
[0022] S102. By reading the header information of the NAL unit, the encoding format type and NAL type of the NAL unit are identified, and the NAL unit is encapsulated into a unified descriptor of a unified format, so that the NAL types of different encoding formats are mapped to a unified NAL type. The unified NAL type includes parameter set type, keyframe type, non-keyframe type, access unit separator and supplementary enhancement information.
[0023] NAL is a data encapsulation format defined in the H.264 / H.265 standard. Each NAL unit consists of three parts: a start code, a NAL header, and a NAL payload. The start code marks the beginning of a NAL; the NAL header identifies the type of this NAL unit (whether it is a parameter set, a keyframe, or a normal frame); and the NAL payload is the actual encoded video binary data.
[0024] The video stream output by the encoder consists of continuous binary data, with NAL units separated by start codes. The continuous stream is divided into independent NAL units by detecting the position of the start code.
[0025] The start code can have two lengths: 3 bytes (0x000001) or 4 bytes (0x00000001). Therefore, the start code can be determined by scanning the bitstream data: a consecutive byte sequence of 0x00 0x00 0x01 indicates a 3-byte start code; a consecutive byte sequence of 0x00 0x00 0x00 0x01 indicates a 4-byte start code. Then, using the position information of the start code, the bitstream is divided into multiple NAL units.
[0026] Generally, the structure of each NAL unit can be represented as: start code (3 or 4 bytes) + NAL header (1 byte for H.264, 2 bytes for H.265) + NAL payload (actual video encoded data). Therefore, the NAL header structures of H.264 and H.265 are completely different.
[0027] Further, the target byte in the header information of the NAL unit is read; when the byte value of the target byte is within a first value range, the encoding format type of the NAL unit is determined to be a first encoding format; when the byte value of the target byte is within a second value range, the encoding format type of the NAL unit is determined to be a second encoding format, wherein the first value range and the second value range do not overlap.
[0028] In one embodiment, the encoding format type of the NAL unit may include H.264 and H.265, or other encoding format types.
[0029] The first encoding format corresponds to a first value range, and the second encoding format corresponds to a second value range. The first value range (1 to 12) and the second value range (32 to 35) do not overlap, therefore the same NAL unit will not simultaneously satisfy both format conditions, ensuring the uniqueness and determinism of the recognition result.
[0030] Specifically, the `nal_unit_type` field (NAL unit type field) in H.264 occupies 5 bits (bits 4 to 0), with a maximum value of 31; the `nal_unit_type` field in H.265 occupies 6 bits (bits 6 to 1), with a minimum value of 0, but the parameter set type value starts from 32. Therefore, when the first encoding format is H.264, the first value range is the value of bits 4 to 0 of the target byte between 1 and 12; when the second encoding format is H.265, the second value range is the value of bits 6 to 1 of the target byte between 32 and 35. This allows the main NAL type value range (1 to 12) of H.264 and the parameter set type value range (32 to 35) of H.265 to be completely separated in numerical space. The encoding format type can be uniquely determined by parsing the two different bit fields of the same target byte, without the need for an additional format identifier field.
[0031] For other encoding format types, their corresponding value ranges and bit field resolution methods can be defined similarly and added to the mapping table of the unified descriptor to achieve support for new encoding formats without modifying the upper-level processing logic.
[0032] In one specific embodiment, the first byte after the start code in the NAL unit is read as the target byte. This target byte is the first byte of the NAL header, and both H.264 and H.265 NAL headers begin from this byte.
[0033] Bit-field parsing is performed on the target byte. Specifically, bits 4 to 0 of the target byte can be extracted to obtain the first candidate type value. This bit field corresponds to the value position of the nal_unit_type field in the H.264 format. If the value range of the first candidate type value is between 1 and 12, the encoding format type of the NAL unit is determined to be H.264 format. This value range covers the main NAL types defined in the H.264 standard, including non-IDR frames (type 1), IDR frames (type 5), SPS (type 7), PPS (type 8), etc. Simultaneously, bits 6 to 1 of the target byte can be extracted to obtain the second candidate type value. This bit field corresponds to the value position of the nal_unit_type field in the H.265 format. If the value range of the second candidate type value is between 32 and 35, the encoding format type of the NAL unit is determined to be H.265 format. This range of values covers the main parameter set types defined in the H.265 standard, including VPS (type 32), SPS (type 33), PPS (type 34), etc.
[0034] In existing technologies, implementing NAL parsing, parameter set management, and IDR merging for H.264 requires approximately 1500 lines of C++ code, while the same functionality in H.265 requires approximately 1700 lines. Although both share similar logic, their NAL header structures and the number of parameter set types differ (H.264 has two parameter sets, while H.265 has three), necessitating independent implementations. Each new feature addition or bug fix requires simultaneous modifications to both sets of code, doubling the manpower cost. To address this technical issue, this application encapsulates the original NAL unit into a unified descriptor with a standardized format. This allows subsequent steps to be processed based on the unified descriptor, eliminating the need to access the format-related fields of the original NAL unit and thus disregarding underlying format differences.
[0035] Specifically, after the encoding format is identified, the original NAL units are encapsulated into a unified descriptor with a uniform format. The unified descriptor is a data structure used to convert NAL units with different encoding formats into a format-independent standardized representation, so that subsequent processing steps do not need to concern themselves with the differences in the underlying encoding formats.
[0036] The core design idea of the unified descriptor is to mark the original format source through the encoding format type field (codec_type field), and to map different formats of the same type of NAL type to the same enumeration value through the unified NAL type field (nal_type field). The specific mapping relationship is shown in Table 1 below: Table 1. Uniform Descriptor Mapping Relationship Table In one embodiment, the unified descriptor may include at least the following core fields: encoding format type field, unified NAL type field, stream identifier field, parameter set identifier field, etc. Depending on actual needs, the unified descriptor may also adaptively include other fields, such as raw NAL type field, time-domain level field, data pointer field, data length field, display timestamp field, decoding timestamp field, etc.
[0037] The encoding format type field is used to identify the source format of the original NAL unit. For example, 0 represents H.264, 1 represents H.265, and other values can be extended to represent other encoding format types. The stream identifier field is used to identify the video stream to which the NAL unit belongs. For example, 0 represents the main stream, and 1 represents the sub-stream. The parameter set identifier field is used to identify the ID of the parameter set, such as sps_id for sequence parameter sets, pps_id for image parameter sets, etc.
[0038] The unified NAL type field is used to map semantically identical but different original enumeration values in H.264 and H.265 NAL types to the same enumeration value. Each NAL type corresponds one-to-one with the unified NAL type, and both include Video Parameter Set (VPS), Sequence Parameter Set (SPS), Picture Parameter Set (PPS), Keyframe IDR, Non-Keyframe (P / B frame), Access Unit Separator (AUD), and Supplemental Enhancement Information (SEI). However, the enumeration values corresponding to the same NAL type are different in different encoding format types (such as H.264 and H.265). This application maps the different enumeration values corresponding to the same NAL type in different encoding format types to the same enumeration value through the mapping relationship shown in Table 1 above, so that the NAL types (different representations) corresponding to the same type in different encoding format types are represented by the same unified NAL type. For example, the NAL type of the sequence parameter set (SPS) is represented as nal_unit_type=7 in the H.264 encoding format type and as nal_unit_type=33 in the H.265 encoding format type. The two have the same semantics but different enumeration values. After being mapped to the unified descriptor, they are both represented by the unified NAL type (33).
[0039] Specifically, as shown in Table 1 above, the specific mapping relationship is as follows: The Video Parameter Set (VPS) is a unique NAL type in the H.265 encoding format, and there is no corresponding NAL type in the H.264 encoding format. The VPS carries global decoding parameters of the video layer, such as quality and level, and is used to initialize the video layer configuration of the decoder. The H.265 Video Parameter Set (VPS) (nal_unit_type=32) is mapped to a uniform value of 32.
[0040] The Sequence Parameter Set (SPS) is a NAL type common to both H.264 and H.265. The SPS carries sequence-level decoding parameters, such as resolution, frame rate, bit rate, and image format. In H.264, nal_unit_type=7 represents SPS, while in H.265, nal_unit_type=33 represents SPS. The two have the same semantics but different original enumeration values, and are uniformly mapped to 33.
[0041] The Image Parameter Set (PPS) is a NAL type common to both H.264 and H.265. The PPS carries image-level decoding parameters, such as entropy coding mode, slice grouping, and deblocking filtering. In H.264, nal_unit_type=8 represents the PPS, while in H.265, nal_unit_type=34 represents the PPS; both are uniformly mapped to 34.
[0042] A keyframe IDR (Instantaneous Decoding Refresh) is a special type of keyframe. Upon receiving an IDR frame, the decoder can immediately refresh the reference frame buffer and decode independently without relying on previous frames. In H.264, nal_unit_type=5 indicates an IDR frame, while in H.265, nal_unit_type=19 or 20 indicates an IDR frame (19 is IDR_W_RADL, 20 is IDR_N_LP), and is uniformly mapped to 19.
[0043] Non-key frames, also known as P / B frames (inter-frame coded frames), are used where P frames refer to the previous frame for predictive coding, and B frames refer to both the previous and next frames for bidirectional predictive coding. In both H.264 and H.265, nal_unit_type=1 indicates a non-IDR frame, and the original enumeration values are the same, uniformly mapped to 1.
[0044] The Access Unit Delimiter (AUD) is used to mark the starting position of an access unit. An access unit is a group of NAL units that the decoder can decode independently. In H.264, nal_unit_type=9 represents AUD, and in H.265, nal_unit_type=35 represents AUD. The mapping is uniformly 35.
[0045] Supplemental Enhancement Information (SEI) is used to carry auxiliary information outside the decoding process, such as timecode, user data, HDR metadata, etc. In H.264, nal_unit_type=6 represents SEI, and in H.265, nal_unit_type=39 represents SEI; the unified mapping is 39. For other unified NAL types, new mapping relationships can be added by extending the mapping table.
[0046] NAL types with the same semantics (such as SPS, PPS, IDR, non-IDR, AUD, SEI) in different encoding formats (such as H.264 and H.265) are mapped to the same unified value. The mapped unified descriptor can be used as the standard data interface for subsequent steps. After format recognition and encapsulation are completed, subsequent steps are all processed based on the unified descriptor. There is no need to access the format-related fields of the original NAL unit. Therefore, the underlying format differences are no longer a concern, making the upper-level processing logic completely independent of the encoding format. This eliminates the redundant processing overhead caused by format differences from the encoding format recognition level.
[0047] After the unified descriptor obtained from the aforementioned steps enters the subsequent steps, the unified NAL type field in the unified descriptor is first read to determine whether the NAL unit is a parameter set type.
[0048] The parameter set types include the following three: Video Parameter Set (VPS, unified NAL type value 32): Specific to the H.265 encoding format, carrying global decoding parameters at the video level, such as quality level and grade. This type is not available in the H.264 encoding format; Sequence Parameter Set (SPS, unified NAL type value 33): Shared by both H.264 and H.265, carrying sequence-level decoding parameters, such as resolution, frame rate, bit rate, and image format; Image Parameter Set (PPS, unified NAL type value 34): Shared by both H.264 and H.265, carrying image-level decoding parameters, such as entropy coding mode, slice grouping, and deblocking filtering.
[0049] S103. When the unified NAL type is a parameter set type, extract the current parameter set data from the unified descriptor, input the current parameter set data into the parameter set version pool for version management, determine the version number information of the current parameter set data, and generate the parameter set data package corresponding to the current parameter set data.
[0050] When the value of the unified NAL type field is 32, 33, or 34, the NAL unit is determined to be a parameter set type and enters the parameter set version pool management process. Otherwise, the NAL unit skips step S102 and proceeds directly to subsequent processing.
[0051] In existing technologies, when an IPC camera dynamically changes its encoding parameters (e.g., switching from 1080p to 720p), the encoder outputs a new SPS / PPS. The traditional approach is to directly forward the new parameter set, but the client may still be using the old parameter set for decoding, leading to decoding failures or screen artifacts. More seriously, when parameter set versions change frequently, old and new parameters conflict, resulting in approximately 18% to 25% of bandwidth being wasted on repeatedly transmitting unchanged parameter sets. To address this, this application implements unified management of parameter set data through a parameter set version pool.
[0052] Furthermore, such as Figure 2 As shown, step S103 specifically includes steps S1031 to S1036.
[0053] S1031. Construct index data for the current parameter set data, wherein the index data includes stream identifier, NAL type, and parameter set identifier.
[0054] In one embodiment, for a unified descriptor determined to be a parameter set type, three fields are extracted from the unified descriptor as indexes to construct the query key for the parameter set version pool: stream identifier (stream_id), unified NAL type (nal_type), and parameter set identifier (param_set_id). The stream identifier ensures data isolation during multi-path concurrency, the unified NAL type eliminates encoding format differences, and the parameter set identifier supports switching between multiple configurations under the same type.
[0055] The stream identifier (stream_id) identifies the video stream to which this parameter set belongs. In multi-channel video transmission scenarios, the parameter sets of different video streams (such as the main stream and sub-streams) need to be managed independently. For example, the main stream identifier is 0, and the sub-stream identifier is 1.
[0056] The uniform NAL type (nal_type) identifies the type of the parameter set, namely VPS (32), SPS (33), or PPS (34). This field has been uniformly mapped and is no longer related to the original encoding format.
[0057] The parameter set identifier (param_set_id) identifies different instances within the same type of parameter set. For example, the sps_id field for SPS and the pps_id field for PPS. Different parameter set identifiers correspond to different decoding parameter configurations.
[0058] For example, the structure of an index key-value pair can be represented as (stream identifier, uniform NAL type, parameter set identifier), which uniquely identifies an entry in the parameter set version pool.
[0059] S1032. Query the historical parameter set entries corresponding to the index data in the parameter set version pool.
[0060] Using the extracted triples as keys, a query is performed in the parameter set version pool to determine if a corresponding historical parameter set entry exists.
[0061] The parameter set version pool is a hash table structure with triples as keys, where each entry corresponds to a unique parameter set instance. Query operations locate the target entry based on three fields: stream identifier, uniform NAL type, and parameter set identifier.
[0062] The query results fall into two categories: either there is no historical parameter set entry corresponding to the index data in the parameter set version pool, or a historical parameter set entry corresponding to the index data already exists in the parameter set version pool. Different operations are performed depending on the query result.
[0063] S1033. When there is no historical parameter set entry corresponding to the index data in the parameter set version pool, construct the current parameter set entry corresponding to the current parameter set data and determine the version number information of the current parameter set data.
[0064] In one embodiment, if the entry corresponding to the key value does not exist in the parameter set version pool, it indicates that the parameter set has appeared for the first time. In this case, a new entry for the current parameter set is created in the parameter set version pool.
[0065] Specifically, the current parameter set data is stored in a newly created current parameter set entry. This data may include a data pointer and data length, pointing to the binary payload of the original NAL unit. Then, the version number is initialized to 1, an auto-incrementing integer counting from 1. The entry status is set to valid, indicating that the parameter set entry is currently in use and can be merged and queried by subsequent keyframes. The current timestamp is recorded as the creation time. At this point, the version number of the current parameter set entry is determined to be 1, and the current parameter set data enters the management lifecycle of the parameter set version pool.
[0066] S1034. When a historical parameter set entry corresponding to the index data exists in the parameter set version pool, compare the historical parameter set data corresponding to the historical parameter set entry with the current parameter set data.
[0067] In one embodiment, if a historical parameter set entry corresponding to the key value already exists in the parameter set version pool, it is necessary to compare whether the historical parameter set data is consistent with the current parameter set data.
[0068] The comparison can be performed byte-by-byte or hash value comparison. Specifically, the data pointer and data length of the historical parameter set entries are extracted and compared with the data pointer and data length of the current parameter set data. If the binary content of the two is completely identical, they are considered the same; if there are any byte differences, they are considered different.
[0069] The comparison results fall into two categories: the historical parameter set data is the same as the current parameter set data, or the historical parameter set data is different from the current parameter set data.
[0070] S1035. When the historical parameter set data is the same as the current parameter set data, the historical version number information corresponding to the historical parameter set data is determined as the version number information of the current parameter set data.
[0071] S1036. When the historical parameter set data is different from the current parameter set data, the historical parameter set data is updated to the current parameter set data, and the version number information is updated based on the historical version number information corresponding to the historical parameter set data, and the updated version number information is determined to be the version number information of the current parameter set data.
[0072] If the historical parameter set data is completely identical to the current parameter set data, it indicates that the parameter set has not changed. In this case, no data update operation is performed, and the historical version number information corresponding to the historical parameter set data is directly determined as the version number information of the current parameter set data. At the same time, the status of the historical parameter set entries remains unchanged and remains valid, avoiding meaningless version increments for unchanged parameter sets and reducing the update overhead of the parameter set version pool.
[0073] If there are differences between the historical parameter set data and the current parameter set data, it indicates that the parameter set has changed. In this case, the parameter set data in the historical parameter set entry is updated to the current parameter set data. The updated data pointer points to the binary payload of the current parameter set data, and the data length is updated to the length of the current parameter set data. Simultaneously, the version number information is updated based on the historical version number information corresponding to the historical parameter set data. Specifically, the historical version number is automatically incremented by 1; for example, if the historical version number is v1, the updated version number is v2. The version number only increases and never decreases, exhibiting a monotonically increasing behavior. Then, the updated version number information is set as the version number information for the current parameter set data, and the status of the current parameter set entry remains valid, indicating that the new version of the parameter set is currently in use.
[0074] By tracking parameter set changes through monotonically increasing version numbers, it is possible to determine whether the parameter set has been updated without comparing it with the original data, thus reducing computational overhead.
[0075] In one embodiment, the current parameter set data contained in the NAL unit of each parameter set type can also be sent to the client separately. Therefore, while managing the version of the current parameter set data, it is also necessary to encapsulate it separately into a parameter set data packet according to the parameter set type (such as video parameter set, sequence parameter set, image parameter set) corresponding to the current parameter set data, and the parameter set data packet contains its corresponding parameter set type.
[0076] It is important to note that after performing a version update, the old version parameter set entries need to undergo a state transition, and the old version is gradually phased out through a delayed recycling mechanism.
[0077] Furthermore, when a new version is generated due to changes in the current parameter set data, the status of the old version entry is changed from valid to pending recycling, and a timed recycling mechanism is initiated. During the timed recycling window of the timed recycling mechanism, if the merge operation needs to reference the old version parameter set, the status of the old version entry is reset to valid and the timer is restarted. After the timed recycling window expires, if the old version parameter set is not referenced, the status of the old version entry is changed to recycled.
[0078] Specifically, when the current parameter set data changes and a new version is generated, the status of the old version entry is changed from valid to pending recycling, and a timed recycling mechanism is initiated. The timeout duration of the timed recycling mechanism is a preset value, such as 3 seconds.
[0079] If the old version parameter set needs to be referenced during the timed recycling window (for example, the old version parameters were used when the keyframe was encoded), the status of the old version entry will be reset from the pending recycling state to the valid state, and the timed recycling mechanism will be restarted.
[0080] After the timed recycling window expires, if the old version parameter set is not referenced by any keyframe, the old version entry status will be changed from pending recycling to recycled, releasing the memory resources it occupies.
[0081] The gradual elimination of older versions is achieved through a three-state transition between valid, pending, and reclaimed states, along with a timed reclamation mechanism. During the timed reclamation window, older versions can be referenced by in-transit keyframes; after the window expires, unreferenced older versions are automatically reclaimed, balancing memory usage and data consistency.
[0082] For example, such as Figure 3 As shown, Figure 3 This diagram illustrates the state transitions of parameter set entries in the parameter set version pool, showing the four states of a parameter set entry throughout its lifecycle and their transition relationships. When a parameter set entry is first created or its data is updated, it is in the valid state, indicating that the parameter set is currently in use and can be referenced by keyframe merging queries. When the parameter set data undergoes a version change, the old version parameter set entry transitions from the valid state to the pending recycling state, and a 3-second (timed recycling window) delayed recycling timer is started. In the pending recycling state, if a keyframe in transit needs to reference the old version parameter set, the old version parameter set entry can be reset back to the valid state. When the 3-second delayed recycling timer expires, and the old version parameter set entry is not referenced or is in the pending recycling state, it transitions to the recycled state, releasing memory resources. Furthermore, in the pending recycling state, if video streaming is paused, the old version parameter set entry can also transition to the frozen state, pausing the recycling timer, and returning to the pending recycling state to resume timing once the stream resumes.
[0083] Understandably, since the parameter set may be referenced by multiple keyframes, and there is a time difference between the encoding and transmission of keyframes, the old version cannot be deleted immediately after a version change. A timed recycling window ensures that keyframes in transit still receive the correct parameter set, avoiding decoding failures or screen artifacts. Simultaneously, delayed recycling avoids bandwidth waste caused by conflicts between old and new parameter sets.
[0084] S104. When the unified NAL type is a non-critical frame type, access unit separator, or supplementary enhancement information, the corresponding non-critical frame data packet is directly generated. Read the uniform NAL type field in the uniform descriptor to determine whether the NAL unit is a non-keyframe type, access unit separator, or supplementary enhancement information.
[0085] Specifically, when the unified NAL type value is 1, it is determined to be a non-critical frame type. This unified value is obtained by mapping between non-IDR frames in H.264 (nal_unit_type=1) and non-IDR frames in H.265 (nal_unit_type=1). Non-critical frame types include P-frames (forward prediction frames) and B-frames (bidirectional prediction frames), which need to be decoded with reference to other frames.
[0086] When the uniform NAL type value is 35, it is determined to be an Access Unit Delimiter (AUD). This uniform value is obtained by mapping the H.264 AUD (nal_unit_type=9) and the H.265 AUD (nal_unit_type=35). The Access Unit Delimiter is used to mark the start position of an access unit, assisting the decoder in frame boundary identification.
[0087] When the uniform NAL type value is 39, it is determined to be Supplemental Enhancement Information (SEI). This uniform value is obtained by mapping the SEI of H.264 (nal_unit_type=6) and the SEI of H.265 (nal_unit_type=39). Supplemental enhancement information carries auxiliary information outside the decoding process, such as timecode, user data, HDR metadata, etc.
[0088] When the value of the unified NAL type field is 1, 35 or 39, the NAL unit is determined to be a non-critical frame type, access unit separator or supplementary enhancement information, and the process of directly generating non-critical frame data packets is initiated.
[0089] For unified descriptors that are determined to be non-critical frame types, access unit separators, or supplementary enhancement information, the corresponding non-critical frame data packets are generated directly without going through parameter set version pool management and critical frame merging processes.
[0090] Specifically, a pointer to the original NAL unit's binary data is obtained from the data pointer field of the unified descriptor, and the total byte length of the NAL unit is obtained from the data length field. The binary data of the original NAL unit (including the start code, NAL header, and NAL payload) is directly encapsulated into a non-critical frame data packet. The encapsulation process of the non-critical frame data packet does not involve any data modification or format conversion, keeping the original encoder's output binary data unchanged.
[0091] Based on the different unified NAL types, the generated non-critical frame data packets are divided into the following categories: Non-key frame data packets (P-frames / B-frames): All have a unified NAL type value of 1. P-frames are predicted and coded with reference to the previous frame, while B-frames are predicted and coded bidirectionally with reference to both the previous and next frames.
[0092] Access Unit Separator Data Packet: Uniform NAL type value 35. Used to mark access unit boundaries and assist decoder synchronization.
[0093] Supplemental Enhancement Data Packet: NAL type value is 39. It carries auxiliary information and does not affect image decoding.
[0094] S105. When the unified NAL type is a keyframe type, extract identification information from the unified descriptor, and extract at least one target parameter set data corresponding to the identification information and the current version number information corresponding to each target parameter set data from the parameter set version pool.
[0095] After reading the Uniform NAL type field in the Uniform Descriptor, the Uniform NAL type may also be a keyframe type. Specifically, the Uniform NAL type value corresponding to a keyframe is 19. This Uniform NAL type value is mapped from H.264 keyframes (nal_unit_type=5) and H.265 keyframes (nal_unit_type=19 or 20). A keyframe is a frame type that can be decoded independently; the decoder can reconstruct the complete image without referencing other frames after receiving a keyframe.
[0096] When the value of the unified NAL type field is 19, the NAL unit is determined to be a keyframe type and enters the adaptive merging process.
[0097] Specifically, for a unified descriptor identified as a keyframe type, identification information is extracted from the unified descriptor and used to query the corresponding target parameter set data from the parameter set version pool. The identification information is a stream identifier (stream_id), which identifies the video stream to which the keyframe belongs; for example, the main stream identifier is 0, and the sub-stream identifier is 1.
[0098] Using the extracted stream identifier as the query condition, a read operation is initiated to the parameter set version pool. Specifically, the query retrieves all parameter set entries that are currently in an active state under the given stream identifier. The parameter set version pool returns all currently active target parameter set data for the video stream, including video parameter sets (if any), sequence parameter sets, image parameter sets, and the current version number information corresponding to each target parameter set data.
[0099] S106. Obtain the historical version number information corresponding to the historical parameter set data of the NAL unit of the key frame type that was previously merged, and determine the version change status of the target parameter set data by comparing the historical version number information and the current version number information.
[0100] After each keyframe data packet merging is completed, the version number information of each target parameter set data used in this merging is recorded as the historical version number information for the next merging. The storage structure of the historical version number information is as follows: using the stream identifier (stream_id) as the key, it stores the unified NAL type and version number of all parameter sets used in the most recent keyframe merging of this video stream.
[0101] The specific content of the historical version number information may include, but is not limited to, the unified NAL type (such as 32, 33, 34) of all target parameter sets participating in the merging under this video stream and their corresponding version numbers (such as v1, v2, etc.).
[0102] After retrieving the target parameter set data and its current version number from the parameter set version pool, the historical version number information corresponding to the video stream to which the current keyframe belongs is retrieved from the historical version number information storage area using the stream identifier of the video stream to which the keyframe belongs as the query condition. If the video stream has not undergone keyframe merging before, the historical version number information is empty, indicating that the keyframe has appeared for the first time.
[0103] Then, the current version number information is compared item by item with the historical version number information of the parameter set data recorded during the previous keyframe merging of the video stream. That is, the version number of each target parameter set data is compared with the version number of the corresponding parameter set recorded in the historical version number information, including all target parameter sets (video parameter sets (if any), sequence parameter sets, and image parameter sets) that are in a valid state under the video stream.
[0104] The comparison method can be a step-by-step comparison according to a unified NAL type. For example, compare the version number of the current sequence parameter set with the version number of the historical sequence parameter set, compare the version number of the current image parameter set with the version number of the historical image parameter set, and compare video parameter sets if they exist.
[0105] Based on the comparison results, the version change status of the target parameter set data is determined. The version change status can include: the first appearance of a keyframe, changes in core coding parameters, changes in non-core coding parameters, or the parameter set remaining stable without change.
[0106] Specifically, if the historical version number information is empty, it means that the video stream has not sent any keyframes before and there is no historical version number information for comparison, and the version change status is determined to be the first occurrence of a keyframe.
[0107] If the historical version number is not empty, and the version number of any target parameter set changes, and the version number of the sequence parameter set changes, and the core encoding parameter field in its data content is different from the historical version, then the version change status is determined to be a core encoding parameter change. Specifically, the core encoding parameter fields in the current sequence parameter set data are extracted and compared byte-by-byte with the corresponding fields in the historical sequence parameter set data; if any field differs, it is determined to be a core encoding parameter change. The core encoding parameters include, but are not limited to, resolution, frame rate, bitrate, encoding level, encoding grade, color format, and bit depth.
[0108] If the historical version number information is not empty, and the version numbers of some target parameter sets have changed, but not all target parameter sets have changed, and the core coding parameter fields in the sequence parameter set are the same as in the historical version, then the version change status is determined to be a non-core coding parameter change. Specifically, the version numbers of some target parameter sets (such as image parameter sets or video parameter sets) have changed, but the version numbers of the sequence parameter sets have not changed, or the version of the sequence parameter sets has changed but the core coding parameter fields are consistent with the historical version. Non-core coding parameters include: entropy coding mode, slice grouping mode, and deblocking filter parameters in the image parameter set; auxiliary fields in the video parameter set; and video availability information fields in the sequence parameter set.
[0109] If the historical version number information is not empty, and the version number of all current target parameter set data is completely consistent with the version number recorded in the historical version number information, then the version change status is determined to be that the parameter set is stable and unchanged.
[0110] S107. Determine the data merging mode based on the version change status.
[0111] Generally, in existing technologies, the IDR frame merging strategy is singular, packaging all parameters once regardless of the scenario. Traditionally, when processing IDR frames, a simple strategy of "buffering each parameter set as it arrives, and then concatenating the entire buffered IDR frame" is adopted. This approach repeatedly concatenates the same SPS / PPS even when the parameter set remains unchanged, adding approximately 200-500 bytes of invalid data. For a 15fps, GOP=30 bitstream, this translates to approximately 240KB-600KB of redundant data transmitted per minute.
[0112] To address this technical problem, embodiments of this application determine the data merging mode by analyzing the change status of the version number information of the target parameter set data.
[0113] In one embodiment, the data merging mode is selected based on the following two conditions: First, whether it is the first occurrence of a keyframe. If the video stream has not previously sent a keyframe, it is determined to be the first occurrence of a keyframe. Second, whether the parameter set version number has changed. The version number information of the currently acquired target parameter set data is compared with the version number information used during the last keyframe merging. If the version number of any parameter set has changed, it is determined that the parameter set version has changed; if the version numbers of all parameter sets have not changed, it is determined that the parameter sets are stable and unchanged.
[0114] Based on the above conditions, one of the following three merging modes is adaptively selected: full merging mode, incremental merging mode, and lightweight merging mode. The three modes are adaptively switched to dynamically balance bandwidth and real-time performance in different scenarios.
[0115] The full merge mode is suitable for scenarios where a keyframe appears for the first time, or when there is a major version change in the parameter set (e.g., a change in sequence parameter set content due to a resolution switch from 1080p to 720p). In this mode, all valid target parameter set data from the video stream in the parameter set version pool are stitched together, with a data volume of approximately 200 to 500 bytes. The full merge mode ensures that the client obtains complete decoding parameters upon the first transmission or when core parameters change, avoiding decoding failures.
[0116] Incremental merging mode is suitable for scenarios where some parameter set versions change (e.g., only the image parameter set is updated, while the sequence parameter set remains unchanged). In this mode, only the target parameter set data whose version number has changed is concatenated, with a data volume of approximately 50 to 200 bytes. Incremental merging mode reduces bandwidth consumption by transmitting only the changed portions when some parameters change.
[0117] Lightweight merging mode is suitable for scenarios where the parameter set is stable and unchanged (the version number of all target parameter set data is consistent with the last merge). In this mode, the cached data from the last transmission is used, eliminating the need to reassemble the parameter set, and the data volume is approximately 100 to 300 bytes. Lightweight merging mode reuses the cache when the parameters are stable, avoiding repeated transmission of unchanged parameter sets and further reducing bandwidth consumption.
[0118] Furthermore, when the version change status meets the first version change condition, the data merging mode is determined to be a full merging mode; when the version change status meets the second version change condition, the data merging mode is determined to be an incremental merging mode; when the version number information has not changed, the data merging mode is determined to be a lightweight merging mode.
[0119] The first version change condition includes at least one of the following: the version number information appears for the first time in the parameter set version pool, or the core encoding parameters of the parameter set data are changed; the second version change condition is that some parameter set data versions are changed, and the changed part of the parameter set data is non-core encoding parameters; wherein, the core encoding parameters include resolution, frame rate, bit rate, encoding level, encoding grade, color format and bit depth; the non-core encoding parameters include entropy encoding mode, slice grouping mode, deblocking filter parameters in the image parameter set, auxiliary fields in the video parameter set, and video availability information fields in the sequence parameter set.
[0120] In one embodiment, if the version change status meets the first version change condition, namely: the current version number information of the target dataset data is appearing for the first time in the parameter set version pool, meaning that the video stream has not previously sent keyframes and there is no historical version number information available for comparison; or, the core encoding parameters of the parameter set data have changed, for example, the resolution in the sequence parameter set has switched from 1080p to 720p, or the core encoding parameters such as frame rate and bit rate have changed. In this case, the data merging mode is determined to be the full merge mode.
[0121] If the change in version number information meets the second version change condition, namely: the version of some parameter set data has changed, and the changed parameter set data are non-core coding parameters, such as only the image parameter set being updated (e.g., the entropy coding mode is adjusted), while the sequence parameter set and video parameter set remain unchanged; or the version of the sequence parameter set has changed but the resolution remains unchanged. In this case, the data merging mode is determined to be the incremental merging mode.
[0122] If the version number information has not changed, that is, the version number of all target parameter set data is exactly the same as the version number used when the last keyframe was merged, then the data merging mode is determined to be the lightweight merging mode.
[0123] S108. Based on the data merging mode, the keyframe data in the unified descriptor is merged with the target parameter set data corresponding to the current version number information in the parameter set version pool to obtain the keyframe data packet.
[0124] Depending on the data merging mode, the corresponding target parameter set data is selected from the parameter set version pool. Specifically, in full merging mode, all valid target parameter set data for the video stream are selected, including video parameter sets (if any), sequence parameter sets, and image parameter sets; in incremental merging mode, only target parameter set data whose version numbers have changed are selected; in lightweight merging mode, parameter set data is not reselected, and the cached copy of the parameter set data from the last transmission is used directly.
[0125] Keyframe data is obtained from the Uniform Descriptor. Specifically, a pointer to the raw NAL unit binary data is obtained from the data pointer field of the Uniform Descriptor, and the total byte length of the keyframe data is obtained from the data length field. The keyframe data contains the start code and the NAL header, and is the raw binary data output by the encoder, without any modification.
[0126] The selected target parameter set data and keyframe data are merged according to the order specified in the encoding standard to obtain the keyframe data packet. For example, the merging can be performed in the order of video parameter set, sequence parameter set, image parameter set, and keyframe: video parameter set (if any) first, sequence parameter set next, image parameter set next, and keyframe data last.
[0127] The merging operation can directly reference the original binary data: the target parameter set data comes from the original NAL unit binary payload (including the start code) stored in the parameter set version pool, and the keyframe data comes from the original NAL unit binary data (including the start code) pointed to by the unified descriptor. The merging process does not involve parsing or format conversion of the parameter set content; it only performs sequential concatenation of the data. The structure of the merged keyframe data packet is: video parameter set NAL unit (if any) + sequence parameter set NAL unit + image parameter set NAL unit + keyframe NAL unit.
[0128] S109. The key frame data packet, the parameter set data packet, and the non-key frame data packet are respectively cached in the priority queue of the corresponding priority, and are used as cached data packets in the priority queue.
[0129] In one embodiment, the data to be sent to the client includes merged keyframe data packets and non-keyframe data packets that have completed unified descriptor mapping but have not been version-managed or merged with data.
[0130] Further, the data type of the keyframe data packet, the parameter set data packet, or the non-keyframe data packet is obtained; based on the data type, the target priority queue corresponding to the keyframe data packet, the parameter set data packet, or the non-keyframe data packet is determined; and the keyframe data packet, the parameter set data packet, or the non-keyframe data packet is cached in the target priority queue.
[0131] For keyframe data packets, the data type is "merged keyframe", which includes video parameter set (if any), sequence parameter set, image parameter set and keyframe data.
[0132] For parameter set data packets and non-critical frame data packets, read their unified NAL type field to determine the specific data type: when the unified NAL type value is 32, 33, or 34, the data type is a parameter set (not merged with the critical frame and sent separately); when the unified NAL type value is 1 and the frame is a P reference frame, the data type is a non-critical reference frame; when the unified NAL type value is 1 and the frame is a B frame, the data type is a non-reference frame; when the unified NAL type value is 39, the data type is supplementary enhancement information.
[0133] Based on the data type of the data packet (keyframe data packet, parameter set data packet, or non-keyframe data packet), the corresponding target priority queue is determined. As shown in Table 2, this application sets up four levels of priority queues and the priority scheduling strategy corresponding to each level of priority queue: Table 2 Priority Queues and Their Scheduling Strategies As shown in Table 2, the priority queue is divided into four levels, from highest to lowest: P0 Priority (Highest Priority): Stores the merged keyframe data packet (including parameter set). This priority ensures that the keyframe data packet is sent first, regardless of which video stream it comes from.
[0134] P1 Priority (High Priority): Stores separately transmitted parameter set data packets (VPS, SPS, PPS). This priority ensures that the parameter set is sent first after the key frame and must arrive at the client before non-key reference frames.
[0135] P2 Priority (Medium Priority): Stores non-critical reference frame data packets (P-reference frames). This priority ensures that reference frames are sent after the parameter set for image prediction after the keyframes.
[0136] P3 priority (low priority): Stores non-reference frame data packets (B-frames) and supplemental enhancement information (SEI) data packets. Data packets of this priority are sent when the first three priority queues are empty, and can be discarded if necessary to ensure bandwidth for critical data.
[0137] The target priority queue is determined based on the data type, and the mapping rules are shown in Table 2.
[0138] When the data type is "merged keyframe", the corresponding target priority queue is P0 (highest priority). This target priority queue corresponds to the merged keyframe data packet, which contains complete parameter set data and keyframe data, and is a necessary prerequisite for the normal operation of the decoder.
[0139] When the data type is "parameter set" (a separately sent video parameter set, sequence parameter set, or image parameter set), the corresponding target priority queue is P1 (high priority). This target priority queue ensures that the parameter set arrives at the client after the keyframe and before the non-key reference frame for decoder parameter updates.
[0140] When the data type is "non-critical reference frame" (P frame), the corresponding target priority queue is P2 (medium priority). This target priority queue ensures that the reference frame is sent after the parameter set, and is used for image prediction and continuous playback after the key frame.
[0141] When the data type is "non-reference frame" (B-frame) or "supplementary enhancement information" (SEI), the corresponding target priority queue is P3 (low priority). The packets in this target priority queue have the lowest importance and can be dropped during network congestion to ensure bandwidth for critical data.
[0142] Keyframe data packets, parameter set data packets, or non-keyframe data packets are placed into their respective target priority queues as buffered data packets. Buffered data packets within the same target priority queue are arranged in arrival order. The target priority queues are divided into four levels: the P0 priority queue stores data packets with a target priority of P0; the P1 priority queue stores data packets with a target priority of P1; the P2 priority queue stores data packets with a target priority of P2; and the P3 priority queue stores data packets with a target priority of P3. Each target priority queue uses a first-in, first-out (FIFO) data structure; newly arriving data packets are placed at the tail of the queue, and are retrieved from the head of the queue during transmission.
[0143] S110. Based on a preset priority scheduling strategy, each cached data packet in the priority queue is sent to the client in sequence according to a preset multi-level priority order, so that the client can decode and display the cached data packets.
[0144] In existing technologies, in scenarios involving simultaneous transmission from dual or multiple cameras, each video stream has its own NAL (Network Alignment) transmission thread. Traditional methods utilize the operating system's native thread priorities, which cannot achieve fine-grained control such as "keyframes from one stream being sent before non-keyframes from another." Test data shows that when two video streams need to be sent simultaneously, the P99 transmission latency of keyframes (IDRs) reaches 48ms, far exceeding the threshold perceptible to the human eye. To address this technical problem, this application implements a multi-level priority queue system combined with a priority scheduling strategy to manage the priority of multiple buffered data packets or keyframe data packets to be sent, thus clearly defining the transmission order of each buffered data packet.
[0145] For cached data packets (key frame data packets, parameter set data packets, or non-key frame data packets) in priority queues at each level, the sending operation is performed sequentially according to the preset priority scheduling strategy.
[0146] In one embodiment, the priority scheduling strategy adopts a round-robin mechanism; the round-robin mechanism includes: after each transmission is completed, checking the cached data packets in each priority queue in descending order of priority; and sending the cached data packets in the same priority queue in a first-in-first-out order.
[0147] The cached data packets are the data packets currently cached in priority queues at various levels, including keyframe data packets, parameter set data packets, or non-keyframe data packets.
[0148] Specifically, when the network transmission layer is ready (e.g., the previous data packet has been sent or there is available space in the transmission buffer), the polling mechanism of the priority scheduling strategy is triggered. The polling mechanism starts checking from the highest priority queue and proceeds downwards until a non-empty queue is found or all queues have been checked. Furthermore, after each transmission is completed, the polling mechanism re-checks the buffered data packets in each target priority queue in descending priority order.
[0149] As shown in Table 2 and Figure 4 As shown, a priority scheduling strategy can be executed through the scheduler. After each data transmission, the polling mechanism checks the following order: First, check the P0 priority queue. If a buffered data packet exists, send the first data packet in that queue. After sending, restart the check from the P0 priority queue. If the P0 priority queue is empty, check the P1 priority queue. If a buffered data packet exists, send the first data packet in that queue. After sending, restart the check from the P0 priority queue. If both the P0 and P1 priority queues are empty, check the P2 priority queue. If a buffered data packet exists, send the first data packet in that queue. After sending, restart the check from the P0 priority queue. If the P0, P1, and P2 priority queues are all empty, check the P3 priority queue. If a buffered data packet exists, send the first data packet in that queue. After sending, restart the check from the P0 priority queue. If all priority queues are empty, wait for new data packets to be enqueued, and trigger the polling mechanism again.
[0150] Buffered data packets in the same priority queue are sent in a first-in, first-out (FIFO) order. That is, the buffered data packets that enter the queue first are sent first, and the buffered data packets that enter the queue later are sent later. This ensures the timing consistency of buffered data packets of the same priority and avoids decoding anomalies caused by out-of-order delivery.
[0151] After the cached data packet is sent, it is removed from the corresponding target priority queue, the queue status information (queue length, total data volume, etc.) is updated, the memory resources occupied by the cached data packet are released (if the reference count is zero), and the next polling cycle is triggered to start checking again from the P0 priority queue.
[0152] In multi-stream video transmission scenarios, buffered data packets from different video streams (such as the main stream and sub-streams) are mixed and entered into the same priority scheduling system. Regardless of which video stream the buffered data packets originate from, they are uniformly scheduled according to the aforementioned four priority levels. For example, when the main stream (H.265, 1080p) has merged keyframe data packets in queue P0, and the sub-stream (H.264, 720p) has B-frame data packets in queue P3, the scheduler prioritizes sending the keyframes in the main stream's P0 queue, and only processes the B-frames in the sub-stream's P3 queue after the keyframes in the main stream have been sent. This polling mechanism achieves fine-grained control that prioritizes keyframes from one stream over non-keyframes from another, solving the problem of each video stream operating independently and keyframes being blocked by low-priority frames in traditional methods.
[0153] When network congestion occurs, non-reference frames (B-frames) and supplementary enhancement information (SEI) in the P3 priority queue can be selectively dropped to free up bandwidth and ensure the transmission of critical data in the P0 to P2 priority queues. Specifically, when the transmit buffer reaches a preset threshold, data packets are dropped starting from the tail of the P3 priority queue until the buffer is restored to a safe level. Data packets in the P0 to P2 priority queues are generally not dropped to ensure the reliable transmission of critical frames, parameter sets, and reference frames.
[0154] After keyframe data packets, parameter set data packets, or non-keyframe data packets are sent to the client via the RTP protocol, the client executes the following decoding and display process: Parsing the parameter set (video parameter set, sequence parameter set, image parameter set) in the keyframe data packet, extracting decoding parameters (resolution, frame rate, bit rate, etc.), and initializing the decoder. Decoding the keyframe data to obtain a complete image, and rendering and displaying it. Decoding subsequent non-key reference frames (P-frames), performing image prediction with reference to the previous frame, and displaying them continuously. If Supplemental Enhancement Information (SEI) is received, parsing the auxiliary data (such as timecode, user data, etc.) within it for synchronization or display purposes.
[0155] To facilitate a clearer understanding of the specific implementation process of the method provided in this application by those skilled in the art, this embodiment provides the following... Figure 5 The illustration shows a specific embodiment. In this specific embodiment, a real-time video transmission scenario using an IPC camera is taken as an example.
[0156] The IPC camera encoder outputs the raw NAL bitstream (H.264 / H.265 raw bitstream, containing sequence parameter sets / image parameter sets / IDR keyframes / P-frames, etc.). From the video bitstream output by the encoder, the continuous bitstream is segmented into independent NAL units by detecting the position of the start code. When a consecutive byte sequence 0x00 0x00 0x01 is detected, it is determined to be a 3-byte start code; when a consecutive byte sequence 0x00 0x00 0x00 0x01 is detected, it is determined to be a 4-byte start code. The target byte in the header information of the NAL unit is read, which is the first byte after the start code.
[0157] Extract bits 4 to 0 of the target byte. If the value range is between 1 and 12, determine the unified NAL type as H.264 format (first encoding format). Extract bits 6 to 1 of the target byte. If the value range is between 32 and 35, determine the unified NAL type as H.265 format (second encoding format). The first value range (1 to 12) and the second value range (32 to 35) do not overlap, ensuring the uniqueness of the recognition result.
[0158] The raw NAL unit is encapsulated into a unified format UnalDescriptor. The UnalDescriptor contains a unified NAL type field (codec_type, 0 for H.264, 1 for H.265), a unified NAL type field (nal_type), a raw NAL type field, a parameter set identifier field, a time domain level field, a data pointer field, a data length field, a display timestamp field, a decoding timestamp field, a stream identifier field (stream_id, 0 for main stream, 1 for sub-stream), and a reference count field.
[0159] Determine if the NAL type is a parameter set type (the unified NAL type value is 32, 33, or 34). If it is a parameter set type, proceed to the S2 parameter set version pool management process. If it is not a parameter set type, further determine if it is an IDR keyframe type (the unified NAL type value is 19). If it is not an IDR keyframe type (i.e., a non-keyframe, such as a P-frame or B-frame), send it directly to the S4 unified priority scheduler.
[0160] For a unified descriptor of a parameter set type, the stream identifier (stream_id), unified NAL type (nal_type), and parameter set identifier (param_set_id) are extracted from the unified descriptor as index data to construct the query key for the parameter set version pool. The index key structure is (stream identifier, unified NAL type, parameter set identifier). The historical parameter set entries corresponding to this index data are then queried from the parameter set version pool.
[0161] If there is no historical parameter set entry corresponding to the index data in the parameter set version pool, construct the current parameter set entry corresponding to the current parameter set data, determine the version number information of the current parameter set data as 1, and set the entry status to the valid state (ACTIVE).
[0162] If a historical parameter set entry corresponding to the indexed data exists in the parameter set version pool, compare the historical parameter set data corresponding to the historical parameter set entry with the current parameter set data. If the historical parameter set data is the same as the current parameter set data, determine the historical version number information corresponding to the historical parameter set data as the version number information of the current parameter set data, and make no changes. If the historical parameter set data is different from the current parameter set data, update the historical parameter set data to the current parameter set data, and update the version number information based on the historical version number information (e.g., from v1 to v2), and determine the updated version number information as the version number information of the current parameter set data.
[0163] When the current parameter set data changes to create a new version, a state transition is performed on the old version parameter set entries, gradually phasing out the old version through a delayed reclamation mechanism. The old version entry's state is changed from ACTIVE to STALE, and a timed reclamation mechanism is initiated with a 3-second interval. During the timed reclamation window, if keyframe merging requires referencing the old version parameter set, the old version entry's state is reset to ACTIVE and the timer is restarted. After the timed reclamation window expires, if the old version parameter set is no longer referenced, the old version entry's state is changed to EVICTED, releasing memory resources.
[0164] For the unified descriptor of the IDR keyframe type, extract the identification information (stream identifier, stream_id) from the unified descriptor, and read the target parameter set data corresponding to the identification information and the version number information of the target parameter set data from the parameter set version pool.
[0165] Based on the version number information of the target parameter set data, the data merging mode (full merging mode, incremental merging mode, or lightweight merging mode) is determined. Based on the determined data merging mode, the target parameter set data and the keyframe data in the unified descriptor are merged according to the standard order of video parameter set, sequence parameter set, image parameter set, and keyframes to obtain a keyframe data packet. The merging process directly references the original binary data and does not involve parsing the parameter set content.
[0166] The unified priority scheduler, based on a preset priority scheduling strategy, sends keyframe data packets and unmerged non-keyframe unified descriptors to the client according to a preset multi-level priority order. The priority scheduling strategy employs a round-robin mechanism: after each transmission, it re-checks the cached data packets in each priority queue in descending order of priority. First, it checks the P0 priority queue; if a cached data packet exists, it sends the packet at the head of the queue, and after sending, it restarts the check from P0. If P0 is empty, it checks P1, and so on. Cached data packets in the same priority queue are sent in a first-in, first-out (FIFO) order.
[0167] Keyframe data packets are transmitted to the client via an RTP (Real-time Transport Protocol) network. Upon receiving the keyframe data packets, the client parses the parameter sets (video parameter sets, sequence parameter sets, image parameter sets), extracts decoding parameters (resolution, frame rate, bit rate, etc.), and initializes the decoder. The keyframe data is decoded to obtain a complete image, which is then rendered and displayed. Subsequent non-key reference frames (P-frames) are decoded, and image prediction is performed with reference to the previous frame for continuous display. If supplementary enhancement information (SEI) is received, auxiliary data (such as timecode, user data, etc.) is parsed for synchronization or display purposes. Finally, the app displays the video feed in real time.
[0168] This embodiment provides a cross-encoding format data merging method. The method receives raw image data from a camera image signal processor, encodes it according to preset encoding parameters, outputs NAL units, and identifies the encoding format type and NAL type by reading the header information of the NAL units. The NAL units are then encapsulated into a unified descriptor, mapping NAL types corresponding to different encoding format types to a unified NAL type. By using a unified NAL type, NAL types with the same semantics but different original enumeration values in different encoding formats are mapped to the same enumeration value, making the upper-layer processing logic independent of the encoding format and eliminating the overhead of redundant processing caused by format differences at the encoding format identification level. When the unified NAL type is a parameter set type, the current parameter set data is extracted from the unified descriptor, input into a parameter set version pool for version management, and a version number is determined. Simultaneously, a corresponding parameter set data packet is generated. Parameter set changes are tracked through the version number, thereby avoiding the repeated transmission of unchanged parameter sets and reducing the parameter set duplication rate. When the unified NAL type is an access unit separator or supplementary enhancement information, the corresponding non-critical frame data packets are directly generated without going through parameter set version pool management and keyframe merging processes, reducing unnecessary processing overhead and improving the transmission efficiency of non-critical frames. When the unified NAL type is a keyframe type, identification information is extracted from the unified descriptor, and at least one target parameter set data corresponding to the identification information and the current version number information corresponding to each target parameter set data are extracted from the parameter set version pool. By comparing the historical version number information and the current version number information, the version change status of the target parameter set data is determined, and then the data merging mode is determined according to the version change status, so that the data merging mode is adapted to the actual needs and reduces the bandwidth waste caused by invalid data splicing. Based on the determined data merging mode, the keyframe data in the unified descriptor is merged with the target parameter set data of each target parameter set with the corresponding current version number information in the parameter set version pool to obtain the keyframe data packet. The parameter set and keyframe are sent together, and the client obtains the complete decoding parameters and keyframe data at one time, without waiting for additional parameter sets to start decoding, reducing decoding waiting latency. Keyframe data packets, parameter set data packets, and non-keyframe data packets are cached in target priority queues of corresponding priorities, serving as cached data packets in the priority queues. This achieves hierarchical caching of data packets. Based on a preset priority scheduling strategy, each cached data packet in the priority queue is sent to the client sequentially according to a preset multi-level priority order. A polling mechanism ensures that high-priority cached data packets are sent first, preventing keyframes from being blocked by non-keyframes, ensuring the real-time transmission of critical data, and improving data transmission efficiency.
[0169] Please see Figure 6 , Figure 6This is a schematic diagram of the structure of a first embodiment of a cross-encoding format data merging device provided in this application. The cross-encoding format data merging device is used to perform the aforementioned cross-encoding format data merging method.
[0170] like Figure 6 As shown, the cross-encoding format data merging device 200 includes: a data encoding module 201, an encoding type identification module 202, a version management module 203, a non-keyframe management module 204, a version number information extraction module 205, a version change status determination module 206, a merging mode determination module 207, a data merging module 208, a data caching module 209, and a data transmission module 210.
[0171] The data encoding module 201 is used to receive raw image data from the image signal processor of the camera, encode the raw image data according to preset encoding parameters, and output at least one NAL unit. The encoding type identification module 202 is used to identify the encoding format type and NAL type of the NAL unit by reading the header information of the NAL unit, and encapsulate the NAL unit into a unified descriptor of a unified format, so that the NAL types of different encoding formats are mapped to a unified NAL type. The unified NAL type includes parameter set type, keyframe type, non-keyframe type, access unit separator and supplementary enhancement information. Version management module 203 is used to extract the current parameter set data from the unified descriptor when the unified NAL type is a parameter set type, and input the current parameter set data into the parameter set version pool for version management, and determine the version number information of the current parameter set data; The non-critical frame management module 204 is used to directly generate the corresponding non-critical frame data packet when the unified NAL type is a non-critical frame type, access unit separator, or supplementary enhancement information. Version number information extraction module 205 is used to extract identification information from the unified descriptor when the unified NAL type is a key frame type, and extract at least one target parameter set data corresponding to the identification information and the current version number information corresponding to each target parameter set data from the parameter set version pool; The version change status determination module 206 is used to obtain the historical version number information corresponding to the historical parameter set data of the NAL unit of the key frame type that was merged last time, and to determine the version change status of the target parameter set data by comparing the historical version number information and the current version number information. Merge mode determination module 207 is used to determine the data merging mode based on the version change status; The data merging module 208 is used to merge the keyframe data in the unified descriptor with the target parameter set data corresponding to the current version number information in the parameter set version pool based on the data merging mode to obtain a keyframe data packet. Data caching module 209 is used to cache the key frame data packet, the parameter set data packet and the non-key frame data packet to the priority queue of the corresponding priority, as cached data packets in the priority queue; The data transmission module 210 is used to send each cached data packet in the priority queue to the client in a preset multi-level priority order based on a preset priority scheduling strategy, so that the client can decode and display the cached data packets.
[0172] It should be noted that those skilled in the art will understand that, for the sake of convenience and brevity, the specific working processes of the above-described apparatus and modules can be referred to the corresponding processes in the aforementioned cross-encoding format data merging method embodiments, and will not be repeated here.
[0173] The apparatus provided in the above embodiments can be implemented as a computer program, which can be used in, for example... Figure 7 It runs on the computer device shown.
[0174] Please see Figure 7 , Figure 7 This is a schematic block diagram of a computer device provided in an embodiment of this application. The computer device may be a server.
[0175] See Figure 7 The computer device includes a processor, memory, and network interface connected via a system bus, wherein the memory may include non-volatile storage media and internal memory.
[0176] Non-volatile storage media can store operating systems and computer programs. These computer programs include program instructions that, when executed, cause the processor to perform any cross-coding format data merging method.
[0177] The processor provides computing and control capabilities, supporting the operation of the entire computer device.
[0178] Internal memory provides an environment for the execution of computer programs on non-volatile storage media, which, when executed by a processor, enable the processor to perform any cross-encoding format data merging method.
[0179] This network interface is used for network communication, such as sending assigned tasks. Those skilled in the art will understand that... Figure 7The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.
[0180] It should be understood that the processor can be a Central Processing Unit (CPU), but it can also be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among these, a general-purpose processor can be a microprocessor or any conventional processor.
[0181] In one embodiment, the processor is configured to run a computer program stored in memory to perform the following steps: The system receives raw image data from the image signal processor of the camera, encodes the raw image data according to preset encoding parameters, and outputs at least one NAL unit. By reading the header information of the NAL unit, the encoding format type and NAL type of the NAL unit are identified, and the NAL unit is encapsulated into a unified descriptor of a unified format, so that the NAL types of different encoding formats are mapped to a unified NAL type. The unified NAL type includes parameter set type, keyframe type, non-keyframe type, access unit separator and supplementary enhancement information. When the unified NAL type is a parameter set type, the current parameter set data is extracted from the unified descriptor and input into the parameter set version pool for version management, and the version number information of the current parameter set data is determined. When the unified NAL type is a non-critical frame type, access unit separator, or supplementary enhancement information, the corresponding non-critical frame data packet is directly generated. When the unified NAL type is a keyframe type, identification information is extracted from the unified descriptor, and at least one target parameter set data corresponding to the identification information and the current version number information corresponding to each target parameter set data are extracted from the parameter set version pool. Obtain the historical version number information corresponding to the historical parameter set data of the NAL unit of the key frame type that was previously merged, and determine the version change status of the target parameter set data by comparing the historical version number information with the current version number information; Determine the data merging mode based on the version change status; Based on the data merging mode, the keyframe data in the unified descriptor is merged with the target parameter set data corresponding to the current version number information in the parameter set version pool to obtain the keyframe data packet; The keyframe data packets, the parameter set data packets, and the non-keyframe data packets are respectively cached in priority queues of corresponding priorities, and are used as cached data packets in the priority queues; Based on a preset priority scheduling strategy, each cached data packet in the priority queue is sent to the client in a preset multi-level priority order so that the client can decode and display the cached data packets.
[0182] The embodiments of this application also provide a computer-readable storage medium storing a computer program, the computer program including program instructions, and the processor executing the program instructions to implement any of the cross-encoding format data merging methods provided in the embodiments of this application.
[0183] The computer-readable storage medium may be an internal storage unit of the computer device described in the foregoing embodiments, such as the hard disk or memory of the computer device. The computer-readable storage medium may also be an external storage device of the computer device, such as a plug-in hard disk, SmartMediaCard (SMC), SecureDigital (SD) card, or FlashCard equipped on the computer device.
[0184] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in this application, and these modifications or substitutions should all be covered within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. A method for merging data across encoding formats, characterized in that, The method includes: The system receives raw image data from the image signal processor of the camera, encodes the raw image data according to preset encoding parameters, and outputs at least one NAL unit. By reading the header information of the NAL unit, the encoding format type and NAL type of the NAL unit are identified, and the NAL unit is encapsulated into a unified descriptor of a unified format, so that NAL types with different encoding formats are mapped to a unified NAL type. The unified NAL type includes parameter set type, keyframe type, non-keyframe type, access unit separator and supplementary enhancement information. When the unified NAL type is a parameter set type, the current parameter set data is extracted from the unified descriptor and input into the parameter set version pool for version management to determine the version number information of the current parameter set data; and a parameter set data package corresponding to the current parameter set data is generated. When the unified NAL type is a non-critical frame type, access unit separator, or supplementary enhancement information, the corresponding non-critical frame data packet is directly generated. When the unified NAL type is a keyframe type, identification information is extracted from the unified descriptor, and at least one target parameter set data corresponding to the identification information and the current version number information corresponding to each target parameter set data are extracted from the parameter set version pool. Obtain the historical version number information corresponding to the historical parameter set data of the NAL unit of the key frame type that was previously merged, and determine the version change status of the target parameter set data by comparing the historical version number information with the current version number information; Determine the data merging mode based on the version change status; Based on the data merging mode, the keyframe data in the unified descriptor is merged with the target parameter set data corresponding to the current version number information in the parameter set version pool to obtain the keyframe data packet; The keyframe data packets, the parameter set data packets, and the non-keyframe data packets are respectively cached in priority queues of corresponding priorities, and are used as cached data packets in the priority queues; Based on a preset priority scheduling strategy, each cached data packet in the priority queue is sent to the client in a preset multi-level priority order so that the client can decode and display the cached data packets.
2. The cross-encoding format data merging method according to claim 1, characterized in that, The step of identifying the encoding format type of the NAL unit by reading its header information includes: Read the target byte from the header information of the NAL unit; When the byte value of the target byte is within the first value range, the encoding format type of the NAL unit is determined to be the first encoding format; When the byte value of the target byte is within the second value range, the encoding format type of the NAL unit is determined to be the second encoding format, wherein the first value range and the second value range do not overlap.
3. The cross-encoding format data merging method according to claim 1, characterized in that, The step of inputting the current parameter set data into the parameter set version pool for version management and determining the version number information of the current parameter set data includes: Construct index data for the current parameter set data, the index data including stream identifier, NAL type and parameter set identifier; Query the historical parameter set entries corresponding to the index data in the parameter set version pool; If there is no historical parameter set entry corresponding to the index data in the parameter set version pool, construct the current parameter set entry corresponding to the current parameter set data and determine the version number information of the current parameter set data.
4. The cross-encoding format data merging method according to claim 3, characterized in that, After querying the historical parameter set entry corresponding to the index data in the parameter set version pool, the process further includes: When a historical parameter set entry corresponding to the index data exists in the parameter set version pool, the historical parameter set data corresponding to the historical parameter set entry is compared with the current parameter set data. When the historical parameter set data is the same as the current parameter set data, the historical version number information corresponding to the historical parameter set data is determined as the version number information of the current parameter set data; When the historical parameter set data differs from the current parameter set data, the historical parameter set data is updated to the current parameter set data, and the version number information is updated based on the historical version number information corresponding to the historical parameter set data, and the updated version number information is determined to be the version number information of the current parameter set data.
5. The cross-encoding format data merging method according to claim 4, characterized in that, After updating the historical parameter set data to the current parameter set data, and updating the version number information based on the historical version number information corresponding to the historical parameter set data, and determining that the updated version number information is the version number information of the current parameter set data, the method further includes: When the current parameter set data changes and a new version is generated, the status of the old version entries is changed from valid to pending recycling, and a timed recycling mechanism is started. During the timed recycling window of the timed recycling mechanism, if the merge operation needs to reference the old version parameter set, the status of the old version entry will be reset to a valid state and the timer will be restarted. If the old version parameter set is not referenced after the timed recycling window expires, the status of the old version entry will be changed to recycled.
6. The cross-encoding format data merging method according to claim 1, characterized in that, The step of determining the data merging mode based on the version change status includes: When the version change status meets the first version change conditions, the data merging mode is determined to be the full merging mode; When the version change status meets the conditions for the second version change, the data merging mode is determined to be the incremental merging mode; If the version number information remains unchanged, the data merging mode is determined to be the lightweight merging mode.
7. The cross-encoding format data merging method according to claim 6, characterized in that, The first version change conditions include at least one of the following: the version number information appears for the first time in the parameter set version pool, or the core encoding parameters of the parameter set data are changed; The second version change condition is that some parameter set data versions have changed, and the changed parameter set data are non-core encoded parameters; The core coding parameters include resolution, frame rate, bit rate, coding profile, coding level, chroma format, and bit depth; the non-core coding parameters include entropy coding mode, slice grouping mode, and deblocking filtering parameters in the image parameter set, auxiliary fields in the video parameter set, and video availability information fields in the sequence parameter set.
8. The cross-encoding format data merging method according to claim 1, characterized in that, The step of caching the keyframe data packet, the parameter set data packet, and the non-keyframe data packet into priority queues of corresponding priorities, as cached data packets in the priority queues, includes: Obtain the data type of the keyframe data packet, the parameter set data packet, or the non-keyframe data packet; Based on the data type, determine the target priority queue corresponding to the key frame data packet, the parameter set data packet, or the non-key frame data packet; The keyframe data packet, the parameter set data packet, or the non-keyframe data packet are cached in the target priority queue.
9. The cross-encoding format data merging method according to claim 1, characterized in that, The priority scheduling strategy adopts a round-robin mechanism; The polling mechanism includes: after each transmission is completed, checking the cached data packets in each priority queue in descending order of priority; Buffered data packets in the same priority queue are sent in a first-in, first-out (FIFO) order.
10. A data merging device across encoding formats, characterized in that, The cross-encoding format data merging device includes: The data encoding module is used to receive raw image data from the image signal processor of the camera, encode the raw image data according to preset encoding parameters, and output at least one NAL unit. The encoding type identification module is used to identify the encoding format type and NAL type of the NAL unit by reading the header information of the NAL unit, and encapsulate the NAL unit into a unified descriptor of a unified format, so that the NAL types of different encoding formats are mapped to a unified NAL type. The unified NAL type includes parameter set type, keyframe type, non-keyframe type, access unit separator and supplementary enhancement information. The version management module is used to extract the current parameter set data from the unified descriptor when the unified NAL type is a parameter set type, input the current parameter set data into the parameter set version pool for version management, determine the version number information of the current parameter set data, and generate the parameter set data package corresponding to the current parameter set data. The non-critical frame management module is used to directly generate the corresponding non-critical frame data packet when the unified NAL type is a non-critical frame type, access unit separator, or supplementary enhancement information. The version number information extraction module is used to extract identification information from the unified descriptor when the unified NAL type is a keyframe type, and to extract at least one target parameter set data corresponding to the identification information and the current version number information corresponding to each target parameter set data from the parameter set version pool. The version change status determination module is used to obtain the historical version number information corresponding to the historical parameter set data of the NAL unit of the key frame type that was merged last time, and to determine the version change status of the target parameter set data by comparing the historical version number information and the current version number information. The merge mode determination module is used to determine the data merge mode based on the version change status. The data merging module is used to merge the keyframe data in the unified descriptor with the target parameter set data corresponding to the current version number information in the parameter set version pool based on the data merging mode, so as to obtain a keyframe data packet. The data caching module is used to cache the key frame data packets, the parameter set data packets, and the non-key frame data packets into priority queues of corresponding priorities, as cached data packets in the priority queues; The data transmission module is used to send each cached data packet in the priority queue to the client in a preset multi-level priority order based on a preset priority scheduling strategy, so that the client can decode and display the cached data packets.