Video coding processing method and related products

By segmenting and recombining medical image video frames, and encoding the medical image area and the ordinary content area of ​​the screen separately, the problem of low encoding efficiency in the existing technology is solved, and efficient transmission of medical image content is achieved.

CN118741128BActive Publication Date: 2026-01-16BEIJING HONGYUN RONGTONG TECH
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202410618343.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-05-17
Publication Date
2026-01-16
Estimated Expiration
2044-05-17

AI Technical Summary

Technical Problem

Existing technologies fail to effectively distinguish the characteristics of different types of areas in the encoding and decoding of medical video, especially medical image content and ordinary screen content areas, resulting in low encoding efficiency and an inability to meet the requirements of high-fidelity encoding for medical image content.

Method used

By segmenting and recombining medical image video frames, the medical image area and the ordinary screen content area are encoded separately. Using an encoding tool that matches the image attributes, the medical image area is encoded using the medical image method, and the ordinary screen content area is encoded using the screen content encoding tool, generating a multi-stream combination.

Benefits of technology

It improves the efficiency of video encoding, ensures high-fidelity transmission of medical image content, and enhances encoding efficiency and transmission reliability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118741128B_ABST
    Figure CN118741128B_ABST
Patent Text Reader

Abstract

The application discloses a video coding processing method and related products, the method comprises the following steps: extracting a video frame subsequence from a video frame sequence to be coded, each video frame of the video frame subsequence comprising a first target image region and a second target image region; performing segmentation and reorganization on the first target image region and the second target image region comprised by each video frame to obtain a segmentation and reorganization result corresponding to each video frame, the segmentation and reorganization result at least comprising a first image block corresponding to the first target image region and a parameter mode value used for indicating whether there is other image block combination; calling an encoder corresponding to the segmentation and reorganization result to code the segmentation and reorganization result to obtain a code stream combination corresponding to the video frame subsequence. The embodiment provided by the application codes the segmentation and reorganization result by using the encoder corresponding to the segmentation and reorganization result, and effectively improves the coding efficiency of video coding.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of image processing, and in particular, to a video coding processing method and related products. BACKGROUND

[0002] Medical image videos mainly include image data obtained by large image devices such as computed tomography (CT), magnetic resonance imaging (MRI), ultrasound, or positron emission computed tomography (PET) in detecting lesions and evaluating treatment effects.

[0003] Medical image videos play a key role in multiple application scenarios. For example, through image cloud storage and PACS cloudization, doctors can realize remote reading of images and make diagnoses even at different geographical locations. Such applications not only improve diagnostic efficiency but also facilitate communication and interaction between doctors and patients. In addition, there are other application scenarios such as remote surgery. The application of medical image videos is gradually expanding from traditional diagnostic tools to more extensive fields such as intelligent analysis and interdisciplinary research, which greatly enriches the application scenarios of medical image videos and promotes the improvement of medical service quality. SUMMARY

[0004] In view of the above defects or deficiencies in the prior art, it is desirable to provide a video coding processing method and related products to improve the coding efficiency of medical image videos.

[0005] In a first aspect, an embodiment of the present application provides a video coding processing method, which comprises: extracting a video frame subsequence from a video frame sequence to be coded, each video frame of the video frame subsequence containing a first target image region and a second target image region; segmenting and recombining the first target image region and the second target image region contained in each video frame to obtain a segmentation and recombination result corresponding to each video frame, the segmentation and recombination result at least including a first image block corresponding to the first target image region and a parameter mode value indicating whether there is a combination of other image blocks; calling an encoder group corresponding to the segmentation and recombination result to code the segmentation and recombination result to obtain a code stream combination corresponding to the video frame subsequence.

[0006] In a second aspect, embodiments of the present application provide a video decoding processing method, which comprises: decoding a first code stream in a received code stream combination to obtain first image data corresponding to the first code stream and a mode parameter value; decoding a code stream corresponding to the mode parameter value in the code stream combination according to the mode parameter value to obtain image data corresponding to the mode parameter value and first coordinate values and / or second coordinate values of a marking region; reconstructing a first target image region according to the first image data, and reconstructing an image region corresponding to a first image block combination and an image region corresponding to a second image block combination according to the image data corresponding to the mode parameter value and the first coordinate values and / or the second coordinate values of the marking region; and recombining the first target image region, the image region corresponding to the first image block combination and the image region corresponding to the second image block combination according to the marking region, and repeating the above operations until a video frame sub-sequence is output.

[0007] In a third aspect, embodiments of the present application further provide a video encoding processing device, which comprises: a sub-sequence extraction unit configured to extract a video frame sub-sequence from a video frame sequence to be encoded, each video frame of the video frame sub-sequence comprising a first target image region and a second target image region; an image segmentation and recombination unit configured to segment and recombine the first target image region and the second target image region comprised in each video frame to obtain a segmentation and recombination result corresponding to each video frame, the segmentation and recombination result comprising at least a first image block corresponding to the first target image region and a parameter mode value indicating whether there is another image block combination; and an encoder group calling unit configured to call an encoder group corresponding to the segmentation and recombination result to encode the segmentation and recombination result, thereby obtaining a code stream combination corresponding to the video frame sub-sequence.

[0008] In a fourth aspect, embodiments of the present application further provide a video decoding processing device, which comprises: a first decoding unit configured to decode a first code stream in a received code stream combination to obtain first image data corresponding to the first code stream and a mode parameter value; a second decoding unit configured to decode a code stream corresponding to the mode parameter value in the code stream combination according to the mode parameter value to obtain image data corresponding to the mode parameter value and first coordinate values and / or second coordinate values of a marking region; an image reconstruction unit configured to reconstruct a first target image region according to the first image data, and reconstruct an image region corresponding to a first image block combination and an image region corresponding to a second image block combination according to the image data corresponding to the mode parameter value and the first coordinate values and / or the second coordinate values of the marking region; and a video frame recombination unit configured to recombine the first target image region, the image region corresponding to the first image block combination and the image region corresponding to the second image block combination according to the marking region, and repeat the above operations until a video frame sub-sequence is output.

[0009] In a fifth aspect, an embodiment of the present application provides a video coding processing system, which comprises the encoding apparatus as described in the third aspect and the decoding apparatus as described in the fourth aspect.

[0010] In a sixth aspect, an embodiment of the present application provides an electronic device, which comprises a memory, a processor, and a computer program stored in the memory and executable on the processor, and the processor implements the method described in the embodiments of the present application when executing the program.

[0011] In a seventh aspect, an embodiment of the present application provides a computer readable storage medium, which stores a computer program, and the computer program is executable on a processor to implement the method described in the embodiments of the present application.

[0012] The technical scheme provided by the present application has the beneficial effects that:

[0013] The present application provides a video coding processing method and related products, which first extracts a video frame sub-sequence from a sequence of video frames to be encoded, each video frame of the video frame sub-sequence comprising a first target image region and a second target image region; then performs segmentation and recombination on the first target image region and the second target image region contained in each video frame to obtain a segmentation and recombination result corresponding to each video frame, the segmentation and recombination result comprising at least a first image block corresponding to the first target image region and a parameter mode value for indicating whether there is a combination of other image blocks; and then calls an encoder group corresponding to the segmentation and recombination result to encode the segmentation and recombination result to obtain a code stream combination corresponding to the video frame sub-sequence. The video coding method provided by the present application utilizes the image characteristics of different image regions in a video frame to perform segmentation and recombination on the video frame to obtain different segmentation and recombination results, and then utilizes the encoding tools corresponding to the segmentation and recombination results to encode the segmentation and recombination results, which can effectively improve the coding efficiency of video coding. BRIEF DESCRIPTION OF DRAWINGS

[0014] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings needed in the embodiment description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative effort.

[0015] Figure 1 A schematic diagram of an application scenario of a video coding processing method provided by an embodiment of the present application is shown;

[0016] Figure 2 A flowchart of a video coding processing method 200 provided by an embodiment of the present application is shown;

[0017] Figure 3Fig. 1 shows a schematic diagram of a video frame according to an embodiment of the present application;

[0018] Figure 4 Fig. 4 shows a schematic diagram of a segmentation result of a video frame according to an embodiment of the present application;

[0019] Figure 5 Fig. 5 shows a schematic diagram of a plurality of segmentation results of a video frame according to an embodiment of the present application;

[0020] Figure 6 Fig. 6 shows a schematic diagram of a video encoding processing method 600 according to an embodiment of the present application;

[0021] Figure 7 Fig. 7 shows a schematic diagram of a video encoding processing method 700 according to an embodiment of the present application;

[0022] Figure 8 Fig. 8 shows a schematic diagram of a video encoding processing method according to another embodiment of the present application;

[0023] Figure 9 Fig. 9 shows a schematic diagram of a video decoding processing method 900 according to an embodiment of the present application;

[0024] Figure 10 Fig. 10 shows a schematic diagram of a video decoding processing method according to another embodiment of the present application;

[0025] Figure 11 Fig. 11 shows a schematic diagram of a video encoding processing apparatus 1100 according to an embodiment of the present application;

[0026] Figure 12 Fig. 12 shows a schematic diagram of a video decoding processing apparatus 1200 according to an embodiment of the present application;

[0027] Figure 13 Fig. 13 shows a schematic diagram of a video encoding and decoding processing system 1300 according to an embodiment of the present application;

[0028] Figure 14 Fig. 14 shows a schematic diagram of an electronic device 1400 according to an embodiment of the present application. DETAILED DESCRIPTION

[0029] In order to make the objectives, technical solutions and advantages of the present application clearer, the embodiments of the present application will be further described in detail below with reference to the accompanying drawings. It should be understood that the specific embodiments described herein are merely intended to explain the related application, but not to limit the application. In addition, it should be noted that only the parts related to the application are shown in the accompanying drawings for the convenience of description.

[0030] It should be noted that the embodiments in the present application and the features in the embodiments can be combined with each other in the case of no conflict. The present application will be described in detail below with reference to the drawings and in combination with the embodiments.

[0031] Medical image videos play a crucial role in online medical services, and their quality and safety are directly related to the efficiency and quality of medical services. Medical image videos are important tools for doctors to make diagnoses, especially in the fields of tumors and cardiovascular diseases. High-quality images can provide intuitive information about the location, size, and nature of lesions.

[0032] In recent years, with the development of medical image video technology, online medical services have developed rapidly, especially in the post-pandemic era, when people have paid more attention to the advantages of online medical services. In addition, doctors can perform remote surgeries by coordinating with robots through real-time video transmission. In the process of remote surgery, the transmitted content usually includes a series of screen content in addition to the medical image part, such as Figure 1 As shown in the figure, a certain frame of video in the medical image video output by the medical imaging device is shown, and the center area of the video frame is the medical image content area, and the two sides of the video frame are the screen normal content area.

[0033] Usually, in order to ensure the quality and real-time performance of the transmission of medical image videos, H.264 / AVC or H.265 / HEVC standards are used for encoding and decoding of medical image videos. However, the same encoding tools are used for the screen normal content area and the medical image content area in the medical image video, which are two different types of areas. However, these two different types of areas have completely different characteristics. The screen normal content area has less noise, sharp edges, and continuous color tones, while the medical image content area has more noise, more high-frequency information, and other characteristics.

[0034] In addition, during transmission, doctors pay more attention to medical image content and have lower sensitivity to some information around them. Obviously, the medical image content that doctors care about needs to be encoded with high fidelity, while other parts (such as screen normal content) can tolerate some distortion.

[0035] Therefore, based on the difference in screen content of medical image videos, the present application proposes a video encoding and decoding method, which can combine and encode different screen content image areas, so that better encoding results can be obtained as a whole.

[0036] In order to more clearly understand the inventive concept provided by the present application, the following will be described in combination with Figures 2-11 The video encoding and decoding method proposed by the present application is described.

[0037] Please refer to Figure 2 ,Figure 2 A flowchart of a video encoding processing method 200 provided by an embodiment of the present application is shown, which can be implemented by a video encoding processing apparatus. The method comprises the following steps:

[0038] In step S201, a video frame sub-sequence is extracted from a video frame sequence to be encoded, each video frame of the video frame sub-sequence containing a first target image region and a second target image region, wherein the first target image region comprises at least a target image.

[0039] In step S202, the first target image region and the second target image region contained in each video frame are segmented and reorganized to obtain a segmentation and reorganization result corresponding to each video frame, the segmentation and reorganization result comprising at least a first image block corresponding to the first target image region and a parameter mode value indicating whether there is a combination of other image blocks.

[0040] In step S203, a group of encoders corresponding to the segmentation and reorganization result is called to encode the segmentation and reorganization result to obtain a code stream combination corresponding to the video frame sub-sequence.

[0041] In the above steps, the video encoding apparatus processes the acquired medical image data before the medical image device transmits the medical image data to the input port of the encoder through the signal interface.

[0042] The video frame sequence to be encoded refers to the video frame sequence of the medical image video waiting to be input to the encoder. For example, continuous medical image data is acquired from a medical image device, which includes computed tomography (CT) scans, magnetic resonance imaging (MRI), ultrasound, or positron emission computed tomography (PET) scans, etc.

[0043] In the field of medical imaging, the video frame sub-sequence can be different types of image sequences obtained according to different imaging parameters and scanning protocols. These image sequences can highlight the features of different tissue types, which are crucial for disease diagnosis and treatment planning. For example, in MRI, different sequences such as T1 weighted sequences and T2 weighted sequences can be obtained by adjusting factors that affect the MR signal. Each sequence has its unique characteristics in displaying tumors or other lesion sites, so a case usually contains multiple sequences to facilitate a more comprehensive assessment of the disease.

[0044] In some embodiments, each video frame of the video frame sub-sequence can contain a first target image region and a second target image region.

[0045] The first target image region refers to the image region within each video frame of a video frame subsequence that can cover and contain the target object. The first target image region can also be called a sequence-level image region. The target object can be certain organs or target objects acquired in medical imaging.

[0046] The second target image region refers to the image region in the video frame other than the first target image region.

[0047] The image areas of video frames output by medical imaging equipment can be broadly divided into medical imaging areas and non-medical imaging areas. Medical imaging areas refer to the image areas of cells, tissues, and organs displayed in the video frames of medical imaging videos. Non-medical imaging areas refer to the image areas used to display text content or blank image areas in the video frames. The first target image area is the image area in the video frame subsequence that can cover the medical imaging area, and the second target image area is any other image area besides the first target image area. The target objects contained in the first target image area refer to cells, tissues, and organs displayed in the video frame.

[0048] In some embodiments, the first target image region and the second target image region of each video frame in the video frame subsequence are segmented and recombined. After determining the first target image region in each video frame, each video frame is divided according to the first target image region to obtain multiple image blocks.

[0049] like Figure 3 As shown, the video frame subsequence includes {video frame 1, video frame 2, ..., video frame N}. Each video frame includes a first target image region (i.e., the image region covered by the orange area) and a second target image region. The first target image region includes a medical image region (i.e., the medical region).

[0050] The second target image area can be a regular screen content area. The regular screen content area can be further divided into screen text areas and screen blank areas. Screen text areas are typically used to display text related to the medical image area.

[0051] The technical solution provided by this invention can obtain multiple image regions by dividing video frames into regions, and then combine multiple image regions according to the image attributes of the image regions. This can effectively utilize the correlation between image regions with the same image attributes to improve coding efficiency.

[0052] like Figure 4As shown, for any given video frame, in most cases, on-screen text typically appears in areas 1, 2, 3, and 4. In rare cases, due to excessive length, a small portion of text may appear in areas 5 and 8. Blank areas on the screen usually refer to images without any content, such as areas 6 and 7.

[0053] Medical imaging captures images of the body's interior using electromagnetic correlation techniques. For example, MRI uses strong magnetic fields and radio waves to create detailed images of the body's interior. Unlike X-rays and CT scans, MRI does not use radiation and is therefore considered safer. MRI is particularly good at displaying soft tissues such as muscles, ligaments, and the brain. Doctors can use MRI to examine brain diseases, musculoskeletal problems, tumors, and more.

[0054] Each type of medical imaging has its specific uses and advantages, and doctors will choose the most appropriate type of imaging based on the patient's condition and the problems that need to be diagnosed. Through these images, doctors can obtain detailed information about the patient's internal structure and function, thereby making accurate diagnoses and developing effective treatment plans.

[0055] like Figure 4 As shown, taking any video frame in a medical image video frame subsequence as an example, each video frame is divided, and the division results are then recombinated to obtain three parts. The first part combines image regions 6 and 7 (which can be called image blocks 6 and 7) to obtain image block combination A. The second part combines image regions 1, 2, 3, 4, 5, and 8 (which can be called image blocks 1-4, 5, and 8) to obtain image block combination B. The third part uses the first target image region (i.e., the sequence-level image region) as image block C.

[0056] After recombining the image blocks of each video frame, the combined result of each video frame may simultaneously contain image block combination A, image block combination B and image block C, or it may contain image block combination A and image block C, or it may contain image block combination B and image block C, or it may contain only image block C, etc.

[0057] The parameter mode value is used to indicate whether other image patch combinations exist. Indicating the existence of other image patch combinations refers to whether the segmentation and reconstruction result contains image patch A and image patch B. The parameter mode value can be represented by two digits; for example, 00 indicates that both image patch A and image patch B exist, 01 indicates that only image patch B exists, 10 indicates that only image patch A exists, and 11 indicates that neither image patch A nor image patch B exists.

[0058] In order to strengthen the result of the prediction process, each video frame in the video frame subsequence is segmented into a plurality of image regions with different image properties, and then different segmentation results are combined according to the image properties to obtain a combined result with similar or identical image properties, and then the combined result is encoded by using an encoding tool corresponding to the image properties to obtain a code stream combination. The result output by each encoder is referred to as a code stream. The results output by a plurality of encoders are combined to obtain the code stream combination.

[0059] In the segmentation and recombination result, there can be various cases, such as Figure 5 As shown, according to the position of the first target image region in the video frame, the video frame can be divided into various results, Figure 5 The first row of all the diagrams, the second row of all the diagrams, and the first diagram of the third row shown in the drawings represent that the image block combination A, the image block combination B, and the image block C exist in the segmentation and recombination result at the same time, and the difference lies in that the sizes of the regions occupied by the image block combination A and the image block combination B in the video frame are different. Figure 5 The second diagram to the fourth diagram of the third row shown in the drawings represent that the image block combination A and the image block C exist in the segmentation and recombination result, and the difference lies in that the sizes of the regions occupied by the image block combination A in the video frame are different. Figure 5 The first diagram to the third diagram of the fourth row shown in the drawings represent that the image block combination B and the image block C exist in the segmentation and recombination result, and the difference lies in that the sizes of the regions occupied by the image block combination B in the video frame are different. Figure 5 The fourth diagram of the fourth row shown in the drawings represents that only the image block C exists in the segmentation and recombination result.

[0060] After each video frame is segmented and recombined, the segmentation and recombination result has different image properties. For example, for the first image block corresponding to the first target image region, there are more medical image noises, and for such an image block, it is expected to increase the yield by filtering, and for other image blocks with other image properties, filtering processing can not be needed. For the image block C, an encoder with a medical image method is used for encoding to obtain a first code stream, which can improve the encoding yield of the image block C. For the image block combination A and the image block combination B, the noises are less, and filtering cannot bring the encoding yield, and can even reduce the original encoding yield, and then an encoder with a screen content or other specific scene tool is selected to encode the image block combination A and the image block combination B respectively to obtain a second code stream and a third code stream. For the image block combination A and the image block combination B, the possibility of similar image blocks appearing in the current video frame is very high, and it is considered to use an encoding tool matched with the image block combination A and the image block combination B to encode them respectively, and compared with using the same encoding tool as the image block C to encode them, the former can obtain better encoding yield.

[0061] The video coding method provided by the application fully utilizes the correlation between the image regions with the same image properties by combining the regions with the same image properties in the video frame, and then calls the encoder group corresponding to the segmentation and reorganization result to code the segmentation and reorganization result to obtain the code stream combination. The reorganized image region is coded by using the coding tool with the same image property, which not only matches the image property of the reorganization result, but also fully utilizes the correlation of the image, so that a better reference block can be obtained in the process of predictive coding, thereby improving the coding efficiency of the video coding as a whole. Compared with the method of coding the medical image video by using a single encoder, the novel video coding method provided by the application effectively improves the coding efficiency of the video coding.

[0062] Please refer to Figure 6 , Figure 6 A flowchart of a video coding processing method 600 provided by an embodiment of the application is shown, and the method can be implemented by a video coding processing device. The method comprises the following steps:

[0063] In step S601, a video frame sequence to be coded is obtained.

[0064] In step S602, a first target image region contained in each video frame in the video frame sequence is determined.

[0065] In step S603, a video frame subsequence is extracted from the video frame sequence according to the first target image region.

[0066] In step S604, each video frame is segmented according to the position coordinates corresponding to the first target image region, to obtain a first image block corresponding to the first target image region and other image blocks to be combined.

[0067] In step S605, the position coordinates of a marked region corresponding to each video frame are determined.

[0068] In step S606, the other image blocks to be combined are combined according to at least the position coordinates of the marked region, to obtain at least one image block combination corresponding to a second target image region.

[0069] In step S607, a parameter mode value is generated according to the at least one image block combination corresponding to the second target image region.

[0070] In step S608, an encoder group corresponding to the segmentation and reorganization result is called to code the segmentation and reorganization result, to obtain a code stream combination corresponding to the video frame subsequence.

[0071] In the above steps, a medical image video to be coded is obtained, and then the medical image video is sequenced to obtain a video frame sequence to be coded.

[0072] In the video frame sequence, a first target image region contained in each video frame can be obtained according to an image segmentation algorithm. The first target image region refers to a region that can cover a medical image region in the medical video. As shown in Figure 4 The first target image region at least includes a medical image content region, and can also include a partial image region outside the medical image content region.

[0073] Since the size of the medical image content region obtained according to different imaging parameters and scanning protocols is different, after the first target image region is determined, the corresponding imaging parameter of the first target image region can be determined, and then the first target image region that covers the medical image region as much as possible can be obtained according to the imaging parameter and the scanning protocol. The first target image region can be referred to as a sequence-level image region. The sequence-level image region can include a medical image content region and a partial screen normal content region.

[0074] According to the first target image region, a video frame sub-sequence can be extracted from the video frame sequence, and the video frame sub-sequence can also be referred to as a sequence-level image frame set. Then, for each video frame, according to the position coordinates corresponding to the first target image region, each video frame is segmented, including: determining first position coordinate values and second position coordinate values corresponding to the first target image region; performing first segmentation on the video frame according to the first position coordinate values to obtain a first segmentation result; performing second segmentation on the first segmentation result according to the second position coordinate values to obtain a first image block corresponding to the first target image region and other image blocks to be combined.

[0075] As shown in Figure 4 Each video frame is divided into a first target image region including a medical image content and a second target image region including a screen normal content. The first coordinate position widthC and the second coordinate position heightC of the first target image region are determined, and then the video frame is divided for the first time according to widthC of the first target image region, and the video frame is divided for the second time according to heightC of the first target image region to obtain the coordinate positions of the first image block C and other image blocks (i.e. image blocks 1, 2, 3, 4, 5, 6, 7, 8). The image block 1 is determined as a marker region.

[0076] According to the position coordinates of the marking area, the other image blocks to be combined are combined to obtain at least one image block combination corresponding to the second target image area, including: in the other image blocks to be combined, according to the first coordinate value of the marking area and the second coordinate value of the first target image area, the other image blocks associated with the second coordinate value are combined to obtain a first image block combination; in the other image blocks to be combined, other than the first image block and the first image block combination, according to the first coordinate value and the second coordinate value of the marking area, the other image blocks are combined to obtain a second image block combination.

[0077] As shown in Figure 4 According to the value of the horizontal coordinate x of the marking area and the value of the heightC of the first target image area, the coordinate positions of the image block 6 and the image block 7 can be spliced and combined to obtain a first image block combination A. Then, according to the values of the horizontal coordinate x and the vertical coordinate y of the marking area and the values of the widthC and the heightC of the first target image area, the coordinate positions of the image blocks 1, 2, 3, 4, 5, and 8 are spliced and combined to obtain a second image block combination B.

[0078] The first target image area can be an independent image block C, which exists in each video frame, and the size of the first target image area is different in different video frame subsequences.

[0079] For the first target image area and the second target image area existing in each video frame at the same time, a segmentation and reorganization result is obtained by segmentation and reorganization. The screen normal content area contained in the segmentation and reorganization result is variable, for example, in some cases, the first image combination block A may not exist, and in some cases, the second image combination block B may not exist.

[0080] As shown in Figure 5 After dividing the video frame, there can be multiple segmentation results. The segmentation results can be classified and arranged, and the segmentation results can have the following situations: the first image block combination A and the second image block combination B exist at the same time, only the first image block combination A exists, or only the second image block combination B exists, or neither the first image block combination nor the second image block combination exists.

[0081] In order to quickly identify whether other encoding tools need to be called, a parameter mode value can be generated according to the at least one image block combination corresponding to the second target image area, including: generating a first indication bit according to whether the first image block combination exists in the segmentation and reorganization result; generating a second indication bit according to whether the second image block combination exists in the segmentation and reorganization result; and generating a parameter mode value according to the first indication bit and the second indication bit.

[0082] The technical scheme provided by the present application can accurately identify the image block combination containing multiple image attributes in the segmentation and reorganization result, and can improve the processing speed by setting the parameter mode value corresponding to each video frame. The parameter mode value is used to indicate the case that each video frame contains the image combination block A and the image combination block B. According to the parameter mode value, it can be determined whether other encoding tools need to be called to encode the video frame.

[0083] When the parameter mode value determines that other encoding tools need to be called, the encoding tools corresponding to the first target image region and the other encoding tools that need to be called are called to encode the segmentation and reorganization result, and the combination of multiple code streams is obtained.

[0084] The video encoding method provided by the present application can segment and reorganize the video frame according to the coordinate position of the first target image region in the video frame, obtain the image combination result with similar or identical image attributes, and then encode the reorganized image region by using the encoding tool with the same image attribute. The video encoding method provided by the present application can not only match the image attribute of the reorganization result, but also fully utilize the correlation of the image, so that a better reference block can be obtained in the process of predictive encoding, thereby improving the encoding efficiency of the video encoding as a whole. Compared with the method of encoding the medical image video by using a single encoder, the novel video encoding method provided by the present application effectively improves the encoding efficiency of the video encoding.

[0085] Please refer to Figure 7 , Figure 7 A flowchart of a video encoding processing method 700 provided by an embodiment of the present application is shown, and the method can be implemented by a video encoding processing device. The method comprises the following steps:

[0086] In step S701, a video frame subsequence is extracted from a video frame sequence to be encoded, and each video frame of the video frame subsequence contains a first target image region and a second target image region. The first target image region at least includes a target image.

[0087] In step S702, each video frame is segmented according to the position coordinates corresponding to the first target image region, to obtain a first image block corresponding to the first target image region and other image blocks to be combined.

[0088] In step S703, the position coordinates of the marking region corresponding to each video frame are determined.

[0089] In step S704, the other image blocks to be combined are combined according to at least the position coordinates of the marking region, to obtain at least one image block combination corresponding to the second target image region.

[0090] Step S705: Generate parameter mode values ​​based on at least one image block combination corresponding to the second target image region.

[0091] Step S706: Determine the parameter mode value corresponding to each video frame.

[0092] Step S707: Call the first encoder corresponding to the first target image region to encode the first image block and parameter mode value to obtain the first bitstream corresponding to the first target image region.

[0093] Step S708: Call the encoder group corresponding to the parameter mode value to encode the first image block combination and the second image block combination respectively, to obtain the second bitstream corresponding to the first image block combination and the third bitstream corresponding to the second image block combination.

[0094] In some embodiments, calling an encoder group corresponding to a parameter mode value to encode a first image block combination and a second image block combination to obtain a second bitstream corresponding to the first image block combination and a third bitstream corresponding to the second image block combination includes: when it is determined that the parameter mode value indicates that the segmentation and reconstruction result includes a first image block combination and a second image block combination, calling a second encoder corresponding to the first image block combination to encode the first coordinate values ​​of the first image block combination and the marked region to obtain a second bitstream corresponding to the first image block combination; and calling a third encoder corresponding to the second image block group to encode the second coordinate values ​​of the second image block combination and the marked region to obtain a third bitstream corresponding to the second image block combination.

[0095] like Figure 8 As shown, in each video frame of the video frame subsequence, after determining the medical image region in each video frame (i.e., frame-level region determination), the sequence-level image region of each video frame is then determined (i.e., sequence-level region determination). Then, based on the coordinate positions of the sequence-level image regions, the video frames undergo frame reassembly processing to obtain the first image block corresponding to the sequence-level image region and a parameter mode value indicating whether other image block combinations exist. For example, a parameter mode value of spMode of 00 indicates that the encoding tool corresponding to the first image block combination A and the encoding tool corresponding to the second image block combination B need to be invoked.

[0096] While the encoding tool corresponding to the first image block C is called to encode the first image block C and the parameter mode value, the encoding tool corresponding to the first image block combination A is called to encode the first image block A and the horizontal coordinate value x of the marked area, and the encoding tool corresponding to the second image block combination B is called to encode the second image block combination B and the vertical coordinate value y of the marked area. Each encoding tool outputs a bitstream, thus obtaining bitstream combinations including C.bin, A.bin and B.bin.

[0097] In some embodiments, when the parameter mode value indicates that the segmentation and recombination result includes only the first image block combination, the second encoder corresponding to the first image block combination is invoked to encode the first coordinate value of the first image block combination and the marked region to obtain the second bitstream corresponding to the first image block combination.

[0098] like Figure 8 As shown, the parameter mode value spMode is 10, which means that the encoding tool corresponding to the first image block combination A needs to be called. While the encoding tool corresponding to the first image block C is called to encode the first image block C and the parameter mode value, the encoding tool corresponding to the first image block combination A is called to encode the first image block A and the x-coordinate of the marked area. Each encoding tool outputs a bitstream, that is, the bitstream combination includes C.bin and A.bin.

[0099] In some embodiments, when the parameter mode value indicates that the segmentation and reconstruction result includes only the second image block combination, the third encoder corresponding to the second image block combination is invoked to encode the second coordinate values ​​of the second image block combination and the marked region to obtain the third bitstream corresponding to the second image block combination.

[0100] like Figure 8 As shown, the parameter mode value spMode being 01 indicates that the encoding tool corresponding to the second image block combination B needs to be called. While the encoding tool corresponding to the first image block C is called to encode the first image block C and the parameter mode value, the encoding tool corresponding to the second image block combination B is called to encode the second image block B and the vertical coordinate value y of the marked area. Each encoding tool outputs a bitstream, that is, the bitstream combination includes C.bin and B.bin.

[0101] In some embodiments, the method further includes: when the determined parameter mode value indicates that the segmentation and reconstruction result does not include the first image block combination and the second image block combination, there is no need to call other encoders.

[0102] like Figure 8As shown, the parameter mode value spMode is 11, indicating that no other encoding tools need to be called, and only the encoding tool corresponding to the first image block C is called to encode the first image block C and the parameter mode value, and the encoding tool outputs a code stream, i.e., the code stream combination includes C.bin.

[0103] Optionally, the encoding tool corresponding to the first image block combination A can be a palette mode (PaletteMode, English abbreviation PLT) encoding tool. PLT is one of the newly added encoding tools of HEVC-SCC, and the core idea is to use a color palette to represent an image region. This method is particularly suitable for screen content regions with a large number of colors and a wide color range. In this mode, the encoder does not directly encode pixel values, but encodes indexes pointing to the palette, and the palette contains information of all colors in that region. In this way, the amount of data that must be encoded can be reduced, because the index is usually much less than the actual pixel data.

[0104] Optionally, the encoding tool corresponding to the second image block combination B can be an intra block copy (IntraBlock Copy, IBC) encoding tool. IBC is mainly used to efficiently compress screen content with a large amount of color information and details, such as text, graphics, and dynamic user interface elements. This technology achieves high compression rate by predicting the current coding unit (CU) as another region in the same frame that has been decoded. This prediction mode is similar to inter prediction, except that it refers to a reconstructed region in the current frame rather than a neighboring frame. IBC finds the best block vector (BV), i.e., the motion vector, for each CU when performing motion search, which indicates the displacement from the current block to the reference block.

[0105] When using the IBC technology, if the distance between the reference block and the current block is too far, it may hinder the reference process. This is because as the distance increases, the accuracy of the prediction tends to decrease, thereby affecting the compression efficiency and the quality of the reconstructed image. In actual encoding process, since IBC needs to search for the best block vector in the reconstructed region of the frame where the current CU is located, if the search region is too large or the reference block is too dispersed from the current block, it will increase the search complexity and may introduce more prediction errors.

[0106] Optionally, the encoding tool corresponding to the first image block C can be a HEVC (H.265) or H.266 (VVC) encoder, which has high compression ratio and fast processing capability. HEVC (H.265) can achieve twice the compression rate of H.264 while maintaining image quality, which means that the bit rate can be reduced by 50% under the same picture quality. Moreover, HEVC supports video encoding with a resolution of up to 4K or even 8K, which is very beneficial for the high-definition requirements of medical images.

[0107] VVC has a significant improvement in coding performance compared to HEVC, with an average of 49% improvement. VVC has better support for new video types such as 8K ultra-high-definition video, screen content, high dynamic range, and 360-degree panoramic video, which provides more flexibility for medical video transmission. VVC also supports adaptive bandwidth and resolution streaming and real-time communication applications, which are very useful for real-time and reliability requirements in medical video transmission.

[0108] The video coding processing method provided by the application fully utilizes the correlation between image data in the video frame by segmenting and recombining the video frame, and then calls the coding tool matched with the image attribute according to the image attribute of the segmentation and recombination result, encodes the image combination blocks with different image attributes in the segmentation and recombination result respectively, effectively increases the coding efficiency of video coding, and can better ensure the efficiency and reliability of medical video transmission.

[0109] Please refer to Figure 9 , Figure 9 A flowchart of a video decoding processing method 900 provided by another embodiment of the application is shown, which can be implemented by a video decoding device. The method comprises the following steps:

[0110] Step S901, decoding the first code stream in the received code stream combination to obtain the first image data corresponding to the first code stream and the mode parameter value.

[0111] Step S902, decoding the code stream corresponding to the mode parameter value in the code stream combination according to the mode parameter value to obtain the image data corresponding to the mode parameter value and the first coordinate value and / or the second coordinate value of the marked area.

[0112] Step S903, reconstructing the first target image area according to the first image data, and reconstructing the image area corresponding to the first image block combination and the image area corresponding to the second image block combination according to the image data corresponding to the mode parameter value and the first coordinate value and / or the second coordinate value of the marked area.

[0113] Step S904, recombining the first target image area, the image area corresponding to the first image block combination and the image area corresponding to the second image block combination according to the marked area, and repeating the above operation until the output video frame subsequence.

[0114] In the steps described above, after the medical image video is encoded using different encoding tools to obtain a combined bitstream, the decoder uses a decoding algorithm corresponding to the encoder to decode the bitstream. The decoder first needs to identify the bitstream format, which is usually defined by the encoder during the encoding process. Then, the decoder decompresses and decodes the bitstream according to this format to recover the original medical image video data. In this process, the decoder uses pre-set parameters and methods to decode the data in the bitstream to restore high-quality medical image video.

[0115] like Figure 10 As shown, when the decoding tool receives the compressed combined bitstream, it first performs entropy decoding on the first bitstream C.bin output by the first encoder to recover the quantized coefficients and other important information. Then, it performs inverse quantization and inverse DCT transformation on the quantized coefficients to reconstruct the residual signal of the first image data and the value corresponding to the mode parameter value spMode, thus obtaining the reconstructed region C map and the value corresponding to spMode.

[0116] Then, based on the value of the mode parameter spMode, it is determined whether the second bitstream output by the second encoder and / or the third bitstream output by the third encoder need to be decoded.

[0117] If spMode is 10, then the second bitstream A.bin needs to be decoded. The colors in the palette and the index of the escape color are used to reconstruct the pixels in the CU, and the horizontal coordinate value x of the marked area, as well as the width-A and height-A of the second image data, that is, the reconstructed map of region A and the x value.

[0118] If spMode is 01, decoding of the third bitstream B.bin is required. Block matching is used to find the optimal matching block for each coding unit (CU). After finding the optimal matching block, a block vector is calculated, indicating the relative positional relationship between the current block and the optimal matching block. Based on the calculated block vector, the corresponding pixel data is copied from the reference frame to the corresponding position in the current frame, completing the decoding of the current block. After decoding and combining all blocks according to the above process, the reconstruction result of the third image data and the vertical coordinate value y of the marked region are obtained, i.e., the reconstructed map of region B and the y-value.

[0119] If spMode is 00, then both the second bitstream A.bin and the third bitstream B.bin need to be decoded. The second bitstream and the third bitstream are decoded according to the above description to obtain the horizontal coordinate value x and the vertical coordinate value y of the marked area, as well as the second image data width-A and height-A, and the third image data width-B and height-B, that is, the reconstructed image of region A and the x value, and the reconstructed image of region B and the y value.

[0120] If spMode is 11, other code streams do not need to be decoded.

[0121] Finally, according to the value of x, the value of y, the first image data block, the position of each region is determined, and then the video frame is reorganized according to the position of each region. After all the video frames are decoded and combined according to the above process, the video frame subsequence is obtained.

[0122] The video decoding processing method provided by the application decodes the code stream according to the decoding algorithm corresponding to the encoder, can fully utilize the reference block of the prediction encoding process, fully utilize the correlation between the image data in the case of having the same image attribute, improve the decoding efficiency, and obtain high-quality medical imaging video, which not only meets the requirement of high-fidelity encoding of the medical imaging part concerned by the doctor, but also improves the video transmission efficiency in the case that some distortion exists in other parts.

[0123] Please refer to Figure 11 , Figure 11 A structure schematic diagram of a video encoding processing device 1100 provided by another embodiment of the application is shown. The device comprises:

[0124] A subsequence extraction unit 1101 is configured to extract a video frame subsequence from a video frame sequence to be encoded, each video frame of the video frame subsequence comprising a first target image region and a second target image region, the first target image region comprising at least a target image.

[0125] An image segmentation and reorganization unit 1102 is configured to segment and reorganize the first target image region and the second target image region comprised by each video frame to obtain a segmentation and reorganization result corresponding to each video frame, the segmentation and reorganization result comprising at least a first image block corresponding to the first target image region and a parameter mode value indicating whether other image blocks exist.

[0126] An encoder group calling unit 1103 is configured to call an encoder group corresponding to the segmentation and reorganization result to encode the segmentation and reorganization result, and obtain a code stream combination corresponding to the video frame subsequence.

[0127] Optionally, the subsequence extraction unit 1101 is further configured to: obtain the video frame sequence to be encoded; determine the first target image region comprised by each video frame in the video frame sequence; and extract the video frame subsequence from the video frame sequence according to the first target image region.

[0128] Optionally, the image segmentation and reorganization unit 1102 is further configured to: for each video frame, segment each video frame according to the position coordinates corresponding to the first target image region to obtain the first image block corresponding to the first target image region and other image blocks to be combined; determine the position coordinates of the marking region corresponding to each video frame; and combine the other image blocks to be combined according to at least the position coordinates of the marking region to obtain at least one image block combination corresponding to the second target image region; and generate the parameter mode value according to the at least one image block combination corresponding to the second target image region.

[0129] The image segmentation and reorganization unit 1102 is further configured to: determine the first position coordinate value and the second position coordinate value corresponding to the first target image region; perform first segmentation on the video frame according to the first position coordinate value to obtain a first segmentation result; and perform second segmentation on the first segmentation result according to the second position coordinate value to obtain the first image block corresponding to the first target image region and other image blocks to be combined.

[0130] The image segmentation and reorganization unit 1102 is further configured to: in the other image blocks to be combined, combine the other image blocks associated with the second coordinate value according to the first coordinate value of the marking region and the second coordinate value of the first target image region to obtain a first image block combination; and in the other image blocks to be combined, combine the other image blocks except the first image block and the first image block combination according to the first coordinate value and the second coordinate value of the marking region to obtain a second image block combination.

[0131] The image segmentation and reorganization unit 1102 is further configured to: generate a first indication bit according to whether the first image block combination exists in the segmentation and reorganization result; generate a second indication bit according to whether the second image block combination exists in the segmentation and reorganization result; and generate the parameter mode value according to the first indication bit and the second indication bit.

[0132] The encoder group calling unit 1103 is further configured to: determine the parameter mode value corresponding to each video frame; call the first encoder corresponding to the first target image region to encode the first image block and the parameter mode value to obtain a first code stream corresponding to the first target image region; and call the encoder group corresponding to the parameter mode value to encode the first image block combination and the second image block combination respectively to obtain a second code stream corresponding to the first image block combination and a third code stream corresponding to the second image block combination.

[0133] The encoder group calling unit 1103 is further configured to: when it is determined that the parameter mode value indicates that the segmentation and recombination result comprises the first image block combination and the second image block combination, call a second encoder corresponding to the first image block combination, encode the first image block combination and the first coordinate value of the marking area to obtain a second code stream corresponding to the first image block combination, and call a third encoder corresponding to the second image block combination to encode the second image block combination and the second coordinate value of the marking area to obtain a third code stream corresponding to the second image block combination; or when it is determined that the parameter mode value indicates that the segmentation and recombination result only comprises the first image block combination, call the second encoder corresponding to the first image block combination to encode the first image block combination and the first coordinate value of the marking area to obtain the second code stream corresponding to the first image block combination; or when it is determined that the parameter mode value indicates that the segmentation and recombination result only comprises the second image block combination, call the third encoder corresponding to the second image block combination to encode the second image block combination and the second coordinate value of the marking area to obtain the third code stream corresponding to the second image block combination.

[0134] The video encoding processing device provided in the present application fully utilizes the correlation between the image data in the video frame by performing segmentation and recombination on the video frame, and then calls an encoding tool matched with the image attribute according to the image attribute of the segmentation and recombination result to encode the image combination blocks with different image attributes in the segmentation and recombination result, effectively increases the encoding efficiency of video encoding, and can better ensure the efficiency and reliability of medical video transmission.

[0135] Please refer to Figure 12 , Figure 12 A structure diagram of a video decoding processing device 1200 provided in another embodiment of the present application is shown. The device comprises:

[0136] A first decoding unit 1201 is configured to decode a first code stream in a received code stream combination to obtain first image data corresponding to the first code stream and a mode parameter value.

[0137] A second decoding unit 1202 is configured to decode a code stream corresponding to the mode parameter value in the code stream combination according to the mode parameter value to obtain image data corresponding to the mode parameter value and first coordinate value and / or second coordinate value of a marking area.

[0138] An image reconstruction unit 1203 is configured to reconstruct a first target image area according to the first image data, and reconstruct an image area corresponding to the first image block combination and an image area corresponding to the second image block combination according to the image data corresponding to the mode parameter value and the first coordinate value and / or the second coordinate value of the marking area.

[0139] The video frame recombination unit 1204 is configured to recombine the first target image region, the image region corresponding to the first image block combination and the image region corresponding to the second image block combination according to the marking region, and repeat the above operation until a video frame subsequence is output.

[0140] Optionally, the second decoding unit 1202 further comprises a second decoder and / or a third decoder corresponding to the parameter mode value,

[0141] The second decoder is configured to decode the second code stream to obtain image data corresponding to the first image block combination and the first coordinate value of the marking region.

[0142] The third decoder is configured to decode the third code stream to obtain image data corresponding to the second image block combination and the second coordinate value of the marking region.

[0143] The video decoding processing method provided by the present application decodes the code stream according to the decoding algorithm corresponding to the encoder, which can fully utilize the reference block in the prediction encoding process, fully utilize the correlation between the image data in the case of having the same image attribute, improve the decoding efficiency, obtain high-quality medical imaging video, meet the requirement of high-fidelity encoding for the medical imaging part concerned by the doctor, and improve the video transmission efficiency in the case of allowing some distortion in other parts.

[0144] Please refer to Figure 13 , Figure 13 Fig. 13 shows a structural schematic diagram of a video coding processing system 1300 provided by another embodiment of the present application. The system comprises a video encoding processing apparatus as described above and a video decoding processing apparatus as described above. Figure 11 The video decoding processing apparatus as described above. Figure 12 The video decoding processing apparatus as described above.

[0145] The video encoding processing system provided by the present application, at the encoding end, fully utilizes the correlation between the image data in the video frame by segmenting and recombining the video frame, then calls the encoding tool matched with the image attribute according to the image attribute of the segmentation and recombination result, encodes the image combination blocks with different image attributes in the segmentation and recombination result respectively, effectively increases the encoding efficiency of the video encoding, and can better ensure the efficiency and reliability of the medical video transmission. At the decoding end, the code stream can be decoded according to the decoding algorithm corresponding to the encoder, which can fully utilize the reference block in the prediction encoding process, fully utilize the correlation between the image data in the case of having the same image attribute, improve the decoding efficiency, obtain high-quality medical imaging video, meet the requirement of high-fidelity encoding for the medical imaging part concerned by the doctor, and improve the video transmission efficiency in the case of allowing some distortion in other parts.

[0146] The computer program instructions can also be loaded onto a computer, other programmable data processing apparatus, or other processing devices to cause a series of operational steps to be performed on the computer, other programmable apparatus or other processing devices to produce a computer implemented process such that the instructions which execute on the computer or other programmable apparatus provide processes for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks.

[0147] Reference will now be made to Figure 14 , Figure 14 A structure diagram of an electronic device 1400 provided by an embodiment of the present application is shown. The electronic device can be a computer device, a desktop computer, or the like terminal. Figure 14 The structure of the electronic device is not limited. As shown in Figure 14 , the electronic device at least includes a memory 1401 and a processor 1402. For example, the electronic device can further include more or less components (such as a network interface, a display device, and the like) than those shown in Figure 14 . Embodiments of the present application also provide a computer readable storage medium having a computer program stored thereon, the computer program being executed by a processor to implement the method described above. Figure 2 , Figure 6 , Figure 7 , Figure 9

[0148] In particular, according to the embodiments provided by the present application, the processes described above with reference to the flowcharts Figure 2 , Figure 6 , Figure 7 , Figure 9 may be implemented as a computer software program. For example, the embodiments provided by the present application include a computer program product comprising a computer program carried on a machine readable medium, the computer program comprising program code for executing the method shown in the flowchart. In such embodiments, the computer program can be downloaded and installed from a network by a communication part, and / or installed from a detachable medium. When the computer program is executed by a central processing unit (CPU), the above-mentioned functions defined in the system of the present application are executed.

[0149] ​It should be noted that the computer-readable medium in the present application can be a computer-readable signal medium or a computer-readable storage medium or any combination thereof. The computer-readable storage medium may, for example, but is not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or apparatus, or any combination thereof. More specific examples of the computer-readable storage medium can include, but are not limited to, an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In the present application, the computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, device or apparatus. In the present application, the computer-readable signal medium can include a data signal carried in a baseband or as a part of a carrier wave, which carries computer-readable program code. Such a propagated data signal can take various forms, including but not limited to an electromagnetic signal, an optical signal, or any suitable combination thereof. The computer-readable signal medium can also be any computer-readable medium other than the computer-readable storage medium, which can send, transmit, propagate or transport a program for use by or in conjunction with an instruction execution system, device or apparatus. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to wireless, wireline, optical cable, RF, etc., or any suitable combination thereof.

[0150] The flowcharts and block diagrams in the drawings illustrate the possible implementation architectures, functions and operations of the methods and computer program products described according to various embodiments provided by the present application. In the flowcharts or block diagrams, each block can represent a module, a program segment or a part of code, which contains one or more executable instructions for implementing the specified logical functions. It should also be noted that in some alternative implementations, the functions noted in the blocks can occur in different orders than that noted in the drawings. For example, two blocks represented in succession can actually be executed substantially in parallel, and sometimes in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and the combination of blocks in the block diagrams and / or flowcharts, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.

[0151] The units or modules described in the embodiments of the present application can be implemented by software or hardware. The units or modules described can be arranged in a processor, for example, a processor can include a sub-sequence extraction unit, an image segmentation and reorganization unit, and an encoder group calling unit. In some cases, the names of the units or modules do not limit the units or modules themselves, for example, the sub-sequence extraction unit can also be described as "a unit for extracting a sub-sequence of video frames from a sequence of video frames to be encoded".

[0152] As another aspect, the embodiments of the present application also provide a computer readable storage medium, which can be included in the electronic device described in the above embodiments, or can exist separately and not be assembled into the electronic device. The computer readable storage medium stores one or more programs, and the programs are used by one or more processors to perform the video encoding or decoding processing method described in the present application.

[0153] The above description is merely preferred embodiments of the present application and a description of the principles of the technology used. Those skilled in the art should understand that the scope of the present application is not limited to the technical solutions formed by the specific combinations of the above technical features, and also covers other technical solutions formed by any combinations of the above technical features or equivalent features without departing from the inventive concept. For example, the above features can be replaced with the technical features disclosed in the present application (but not limited to) having similar functions to form technical solutions.

Claims

1. A method of video encoding processing, characterized by, The method comprises: extracting a video frame sub-sequence from a medical image video frame sequence to be encoded, segmenting each video frame in the video frame sub-sequence to obtain a plurality of image regions with different image attributes, each video frame in the video frame sub-sequence containing a first target image region and a second target image region, wherein the first target image region is a core region covering medical image content, and the second target image region is a normal content region of a screen without medical image content; segmenting and recombining the first target image region and the second target image region contained in each video frame to obtain a segmentation and recombination result corresponding to each video frame, the segmentation and recombination result comprising a first image block corresponding to the first target image region, at least one image block combination corresponding to the second target image region, and a parameter mode value, wherein the at least one image block combination comprises a first image block combination and / or a second image block combination, and the parameter mode value is used to indicate the case where the segmentation and recombination result includes the first image block combination and the second image block combination; determining the parameter mode value corresponding to each video frame; calling a first encoder corresponding to the first target image region to encode the first image block and the parameter mode value to obtain a first code stream corresponding to the first target image region, wherein the first encoder adopts H.265 or H.266; if the parameter mode value indicates that the segmentation and recombination result only includes the first image block combination, calling a second encoder to encode the first image block combination to obtain a second code stream corresponding to the first image block combination; if the parameter mode value indicates that the segmentation and recombination result only includes the second image block combination, calling a third encoder to encode the second image block combination to obtain a third code stream corresponding to the second image block combination; if the parameter mode value indicates that the segmentation and recombination result includes the first image block combination and the second image block combination, calling the second encoder to encode the first image block combination to obtain a second code stream corresponding to the first image block combination, and calling the third encoder to encode the second image block combination to obtain a third code stream corresponding to the second image block combination; wherein the second encoder is a palette mode encoding tool, and the third encoder is an intra block copy encoding tool.

2. The method of claim 1, wherein, The extraction of at least one video frame sub-sequence from the medical image video frame sequence to be encoded comprises: obtaining a video frame sequence to be encoded; determining a first target image region contained in each video frame in the video frame sequence; extracting the video frame sub-sequence from the video frame sequence according to the first target image region.

3. The method of claim 1, wherein, The segmentation and recombination of the first target image region and the second target image region contained in each video frame comprises: segmenting each video frame according to position coordinates corresponding to the first target image region to obtain a first image block corresponding to the first target image region and other image blocks to be combined; determining position coordinates of a marked region corresponding to each video frame; combining the other image blocks to be combined according to the position coordinates of the mark region, to obtain at least one image block combination corresponding to the second target image region; generating the parameter mode value according to the at least one image block combination corresponding to the second target image region.

4. The method of claim 3, wherein, segmenting each video frame according to the position coordinates corresponding to the first target image region, comprising: determining a first position coordinate value and a second position coordinate value corresponding to the first target image region; performing first segmentation on the video frame according to the first position coordinate value, to obtain a first segmentation result; performing second segmentation on the first segmentation result according to the second position coordinate value, to obtain a first image block corresponding to the first target image region and other image blocks to be combined.

5. The method of claim 3, wherein, combining the other image blocks to be combined according to the position coordinates of the mark region, to obtain at least one image block combination corresponding to the second target image region, comprising: combining, in the other image blocks to be combined, the other image blocks associated with the second coordinate value according to the first coordinate value of the mark region and the second coordinate value of the first target image region, to obtain a first image block combination; and / or combining, in the other image blocks to be combined, the other image blocks except the first image block and the first image block combination according to the first coordinate value and the second coordinate value of the mark region, to obtain a second image block combination.

6. The method of claim 3, wherein, generating the parameter mode value according to the at least one image block combination corresponding to the second target image region, comprising: generating a first indication bit according to whether the first image block combination exists in the segmentation reorganization result; generating a second indication bit according to whether the second image block combination exists in the segmentation reorganization result; generating the parameter mode value according to the first indication bit and the second indication bit.

7. The method of claim 3, wherein: if the parameter mode value indicates that the segmentation reorganization result only includes the first image block combination, the step of encoding the first image block combination by invoking a second encoder to obtain a second code stream corresponding to the first image block combination comprises: encoding the first image block combination and the first coordinate value of the mark region by invoking a second encoder to obtain a second code stream corresponding to the first image block combination; if the parameter mode value indicates that the segmentation reorganization result only includes the second image block combination, the step of encoding the second image block combination by invoking a third encoder to obtain a third code stream corresponding to the second image block combination comprises: encoding the second image block combination and the second coordinate value of the mark region by invoking a third encoder to obtain a third code stream corresponding to the second image block combination. If the parameter mode value indicates that the split reorganization result includes the first image block combination and the second image block combination, the step of calling a second encoder to encode the first image block combination to obtain a second code stream corresponding to the first image block combination and calling a third encoder to encode the second image block combination to obtain a third code stream corresponding to the second image block combination includes: calling the second encoder to encode the first image block combination and the first coordinate value of the marking area to obtain the second code stream corresponding to the first image block combination, and calling the third encoder to encode the second image block combination and the second coordinate value of the marking area to obtain the third code stream corresponding to the second image block combination.

8. A method of video decoding processing, comprising: The method comprises: decoding the first code stream in the received code stream combination to obtain first image data corresponding to the first code stream and a mode parameter value, wherein the first code stream is encoded using a first encoder, the first encoder is H.265 or H.266, and the decoding of the first code stream uses a decoding algorithm corresponding to the first encoder; decoding the second code stream in the code stream combination according to the mode parameter value to obtain second image data corresponding to the first image block combination and a first coordinate value of a marking area, and decoding the third code stream in the code stream combination to obtain third image data corresponding to the second image block combination and a second coordinate value of the marking area, wherein the second code stream is encoded using a second encoder, the second encoder is a palette mode encoding tool, the decoding of the second code stream uses a palette corresponding to the second encoder, the third code stream is encoded using a third encoder, and the decoding of the third code stream uses a block matching technique corresponding to the third encoder; reconstructing a first target image area according to the first image data, and reconstructing an image area corresponding to the first image block combination and an image area corresponding to the second image block combination according to the second image data, the third image data, and the first coordinate value and the second coordinate value of the marking area; reorganizing the first target image area, the image area corresponding to the first image block combination, and the image area corresponding to the second image block combination according to the marking area, and repeating the above operations until a video frame subsequence is output.

9. A video encoding processing apparatus, characterized by comprising: The device comprises: a subsequence extraction unit configured to extract a video frame subsequence from a medical image video frame sequence to be encoded, split each video frame in the video frame subsequence to obtain a plurality of image areas with different image attributes, and each video frame in the video frame subsequence comprises a first target image area and a second target image area, wherein the first target image area is a core area covering medical image content, and the second target image area is a common content area of a screen without medical image content; The image segmentation and reorganization unit is configured to segment and reorganize the first target image region and the second target image region included in each video frame to obtain a segmentation and reorganization result corresponding to each video frame, the segmentation and reorganization result including a first image block corresponding to the first target image region, at least one image block combination corresponding to the second target image region, and a parameter mode value, wherein the at least one image block combination includes a first image block combination and / or a second image block combination, and the parameter mode value is used to indicate a case where the segmentation and reorganization result includes the first image block combination and the second image block combination; The encoder group calling unit is configured to determine the parameter mode value corresponding to each video frame, call a first encoder corresponding to the first target image region to encode the first image block and the parameter mode value to obtain a first code stream corresponding to the first target image region, wherein the first encoder adopts H.265 or H.266, call a second encoder to encode the first image block combination to obtain a second code stream corresponding to the first image block combination if the parameter mode value indicates that the segmentation and reorganization result only includes the first image block combination, call a third encoder to encode the second image block combination to obtain a third code stream corresponding to the second image block combination if the parameter mode value indicates that the segmentation and reorganization result only includes the second image block combination, and call the second encoder to encode the first image block combination to obtain the second code stream corresponding to the first image block combination and call the third encoder to encode the second image block combination to obtain the third code stream corresponding to the second image block combination if the parameter mode value indicates that the segmentation and reorganization result includes the first image block combination and the second image block combination, wherein the second encoder is a palette mode encoding tool, and the third encoder is an intra block copy encoding tool.

10. A video decoding processing apparatus, characterized by comprising: The apparatus includes: The first decoding unit is configured to decode the first code stream in the received code stream combination to obtain first image data and a mode parameter value corresponding to the first code stream, wherein the first code stream is encoded by using a first encoder, the first encoder is H.265 or H.266, and a decoding algorithm corresponding to the first encoder is adopted for decoding the first code stream. a second decoding unit configured to decode a second bitstream in the bitstream combination according to the mode parameter value to obtain second image data corresponding to a first image block combination and first coordinate values of a marker region, and decode a third bitstream in the bitstream combination to obtain third image data corresponding to a second image block combination and second coordinate values of the marker region, wherein the second bitstream is encoded by a second encoder, the second encoder is a palette mode encoding tool, and a palette corresponding to the second encoder is used in decoding the second bitstream, and the third bitstream is encoded by a third encoder, the third encoder is an intra block copy encoding tool, and a block matching technique corresponding to the third encoder is used in decoding the third bitstream; an image reconstruction unit configured to reconstruct a first target image region according to the first image data, and reconstruct an image region corresponding to the first image block combination and an image region corresponding to the second image block combination according to the second image data, the third image data, and the first coordinate values and the second coordinate values of the marker region; a video frame recombination unit configured to recombine the first target image region, the image region corresponding to the first image block combination, and the image region corresponding to the second image block combination according to the marker region, and repeat the above operations until a video frame sub-sequence is output.

11. A video coding processing system, comprising the encoding apparatus of claim 9 and the decoding apparatus of claim 10.

12. An electronic device comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, The processor implements the method of any one of claims 1-8 when executing the program.

13. A computer readable storage medium, having stored thereon a computer program, the computer program being executed by a processor to implement the method of any one of claims 1-8.

Citation Information

Patent Citations

  • Object-based video transcoding method and device

    CN102630043A

  • Image encoding method, decoding method, encoding device and decoding device

    CN105704491A

  • Video coding method based on HEVC-SCC

    CN113079373A

  • Segmented video codec for high resolution and high frame rate video

    US20160094606A1