Image processing method, device, system, network equipment, terminal and storage medium

By writing synthesis indication information and feature information into the video image, the problem of multiple ROI panoramic videos being unable to be encoded is solved, enabling users to watch multiple ROIs at the same time and reducing network and hardware resource usage.

CN115883882BActive Publication Date: 2025-09-12ZTE CORP
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202310035569.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2018-09-19
Publication Date
2025-09-12
Estimated Expiration
2038-09-19

AI Technical Summary

Technical Problem

In the prior art, panoramic videos of multiple ROIs cannot be effectively encoded, resulting in users being unable to view multiple regions of interest simultaneously.

Method used

By obtaining synthesis indication information and feature information of the region of interest, writing them into the supplementary enhancement information SEI, a media stream of the video image is generated, and the synthesis display of the ROI in the video stream is controlled.

Benefits of technology

It realizes the video image encoding of multiple ROIs, meets the needs of users to watch multiple ROIs at the same time, and reduces the network and hardware resource usage.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115883882B_ABST
    Figure CN115883882B_ABST
Patent Text Reader

Abstract

Embodiments of the present invention provide an image processing method, apparatus, system, network device, terminal, and storage medium. Synthesis indication information for indicating a synthesis display mode between regions of interest in a video image is obtained; the synthesis indication information and characteristic information of the regions of interest are written into supplemental enhancement information (SEI) to generate a media stream of the video image, wherein the media stream includes the SEI; that is, the synthesis indication information and characteristic information of the regions of interest are written into a bitstream of the video image, thereby implementing a video image encoding process when multiple (at least two) regions of interest (ROIs) exist. During video playback, the synthesis display and playback of each ROI can be controlled based on the synthesis indication information, thereby satisfying a user's need to view multiple ROIs simultaneously.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present invention relate to, but are not limited to, the field of image coding and decoding technology. Specifically, they relate to, but are not limited to, an image processing method, apparatus, system, network equipment, terminal, and storage medium. Background Art

[0002] The rapid development of digital media technology, exponential increases in hardware performance, significant increases in network bandwidth and speed, and the exponential growth in the number of mobile devices have created opportunities for the development of video applications. Video applications are rapidly evolving from single-viewpoint, low-resolution, and low-bitrate formats to multi-viewpoint, high-resolution, and high-bitrate formats, offering users new video content types and presentation features, as well as a more immersive and engaging viewing experience.

[0003] 360-degree panoramic video (hereinafter referred to as panoramic video) is a new type of video content. Users can choose any viewing angle based on their subjective needs, thus achieving 360-degree viewing. Although current network performance and hardware processing capabilities are relatively high, the rapid increase in the number of users and the huge amount of panoramic video data require reducing network and hardware resource usage while ensuring the user viewing experience.

[0004] Currently, Region of Interest (ROI) technology can display a portion of a panoramic video based on user preferences, eliminating the need to process the entire video. However, related technologies typically only have a single ROI, which can only display a limited portion of the panoramic video image and cannot meet the user's need to view multiple ROIs. Therefore, when multiple ROIs are present, how to encode and indicate the composite display of each ROI is an urgent problem to be solved. Summary of the Invention

[0005] The image processing method, apparatus, system, network device, terminal, and storage medium provided by the embodiments of the present invention mainly solve the technical problem of how to implement encoding when there are multiple ROIs.

[0006] To solve the above technical problems, an embodiment of the present invention provides an image processing method, comprising:

[0007] Acquiring synthesis indication information for indicating a synthesis display mode between regions of interest in a video image;

[0008] The synthesis indication information and the characteristic information of the region of interest are written into supplemental enhancement information SEI to generate a media stream of the video image, wherein the media stream includes the SEI.

[0009] An embodiment of the present invention further provides an image processing method, comprising:

[0010] receiving a video stream and supplemental enhancement information (SEI) of a video image, wherein the SEI contains synthesis indication information indicating a synthesis display mode between regions of interest in the video image and feature information of the regions of interest;

[0011] Parsing the SEI to obtain synthesis indication information of the region of interest and feature information of the region of interest;

[0012] The synthesis playback display of the image of the region of interest in the video stream is controlled according to the synthesis instruction information and the characteristic information of the region of interest.

[0013] An embodiment of the present invention further provides an image processing method, comprising:

[0014] The network side obtains synthesis indication information for indicating a synthesis display mode of each region of interest in the video image, writes the synthesis indication information and feature information of the region of interest into supplemental enhancement information (SEI) to generate a media stream of the video image, and sends the media stream to a target node, wherein the media stream includes the SEI.

[0015] The target node receives the media stream, parses the media stream to obtain synthesis indication information of the region of interest and characteristic information of the region of interest, and controls the playback and display of the video stream in the media stream according to the synthesis indication information and the characteristic information of the region of interest.

[0016] An embodiment of the present invention further provides an image processing device, comprising:

[0017] An acquisition module, configured to acquire synthesis indication information for indicating a synthesis display mode of each region of interest in a video image;

[0018] A processing module is configured to write the synthesis indication information and the characteristic information of the region of interest into supplemental enhancement information SEI to generate a media stream of the video image, wherein the media stream includes the SEI.

[0019] An embodiment of the present invention further provides an image processing device, comprising:

[0020] a receiving module, configured to receive a video stream and supplemental enhancement information (SEI) of a video image, wherein the SEI contains synthesis indication information indicating a synthesis display mode between regions of interest in the video image and feature information of the regions of interest;

[0021] a parsing module, configured to parse the SEI to obtain synthesis indication information of a region of interest and feature information of the region of interest;

[0022] A control module is used to control the synthesis, playback and display of the image of the region of interest in the video stream according to the synthesis instruction information and the characteristic information of the region of interest.

[0023] An embodiment of the present invention further provides an image processing system, comprising the two image processing devices described above.

[0024] An embodiment of the present invention further provides a network device, comprising a first processor, a first memory, and a first communication bus;

[0025] The first communication bus is used to realize connection and communication between the first processor and the first memory;

[0026] The first processor is configured to execute one or more computer programs stored in the first memory to implement the steps of any of the above image processing methods.

[0027] An embodiment of the present invention further provides a terminal, comprising a second processor, a second memory, and a second communication bus;

[0028] The second communication bus is used to realize connection and communication between the second processor and the second memory;

[0029] The second processor is configured to execute one or more computer programs stored in the second memory to implement the steps of the image processing method described above.

[0030] An embodiment of the present invention further provides a storage medium storing one or more programs. The one or more programs can be executed by one or more processors to implement the steps of the image processing method described above.

[0031] The beneficial effects of the present invention are:

[0032] According to the image processing method, apparatus, system, network device, terminal, and storage medium provided by the embodiments of the present invention, synthesis indication information for indicating a synthesis display mode between regions of interest in a video image is obtained; the synthesis indication information and characteristic information of the regions of interest are written into supplemental enhancement information (SEI) to generate a media stream of the video image, wherein the media stream includes the SEI; that is, the synthesis indication information and characteristic information of the regions of interest are written into the bitstream of the video image, thereby implementing a video image encoding process when multiple (at least two) ROIs exist. During video playback, the synthesis display and playback of each ROI can be controlled based on the synthesis indication information, thereby meeting the user's need to view multiple ROIs simultaneously. In certain implementation processes, technical effects including but not limited to the above can be achieved.

[0033] Other features and corresponding beneficial effects of the present invention are described in the latter part of the specification, and it should be understood that at least some of the beneficial effects become obvious from the description in the specification of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0034] Figure 1 This is a flowchart of an image processing method according to a first embodiment of the present invention;

[0035] Figure 2 Schematic diagram of ROI image stitching indication in embodiment 1 of the present invention Figure 1 ;

[0036] Figure 3 Schematic diagram of ROI image stitching indication in embodiment 1 of the present invention Figure 2 ;

[0037] Figure 4 Schematic diagram of ROI image stitching indication in embodiment 1 of the present invention Figure 3 ;

[0038] Figure 5 Schematic diagram of ROI image stitching indication in embodiment 1 of the present invention Figure 4 ;

[0039] Figure 6 Schematic diagram of ROI image stitching indication in embodiment 1 of the present invention Figure 5 ;

[0040] Figure 7 Schematic diagram of ROI image fusion indication in embodiment 1 of the present invention Figure 1 ;

[0041] Figure 8 Schematic diagram of ROI image fusion indication in embodiment 1 of the present invention Figure 2 ;

[0042] Figure 9 Schematic diagram of the ROI image overlapping area according to the first embodiment of the present invention;

[0043] Figure 10 This is a schematic diagram of ROI image nesting indication according to the first embodiment of the present invention;

[0044] Figure 11 Schematic diagram of ROI image transparent channel processing according to the first embodiment of the present invention;

[0045] Figure 12 Schematic diagram of ROI image coordinate position according to the first embodiment of the present invention;

[0046] Figure 13 Schematic diagram of ROI image video stream generation according to the first embodiment of the present invention Figure 1 ;

[0047] Figure 14 Schematic diagram of ROI image video stream generation according to the first embodiment of the present invention Figure 2 ;

[0048] Figure 15 This is a flowchart of an image processing method according to a second embodiment of the present invention;

[0049] Figure 16 This is a flowchart of an image processing method according to a third embodiment of the present invention;

[0050] Figure 17 This is a structural diagram of an image processing device according to a fourth embodiment of the present invention;

[0051] Figure 18 This is a structural diagram of an image processing device according to a fifth embodiment of the present invention;

[0052] Figure 19 This is a schematic diagram of the structure of an image processing system according to a sixth embodiment of the present invention;

[0053] Figure 20 This is a schematic diagram of the network device structure of Embodiment 7 of the present invention;

[0054] Figure 21 This is a schematic diagram of the terminal structure of embodiment 8 of the present invention. DETAILED DESCRIPTION

[0055] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the following is a further detailed description of the embodiments of the present invention through specific implementation methods in conjunction with the accompanying drawings. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.

[0056] Example 1:

[0057] In order to realize how to encode when there are multiple ROIs in a video image to meet the user's demand for viewing multiple ROIs at the same time, the embodiment of the present invention provides an image processing method, which is mainly used in network side devices, encoders, etc., including but not limited to servers, base stations and other devices, see Figure 1 , including the following steps:

[0058] S101: Acquire synthesis indication information for indicating a synthesis display mode between regions of interest in a video image.

[0059] When encoding a video image, synthesis indication information is obtained to indicate the synthesis display mode between ROIs in the video image. It should be understood that if the video image does not exist or is not divided into ROIs, the process of obtaining the synthesis indication information does not occur. The corresponding synthesis indication information is only obtained when multiple ROIs exist. Even if there is only one ROI, this solution can still be used to control the display of that one ROI.

[0060] Optionally, during the encoding process, it may be determined first that an ROI exists in the video image, and then corresponding synthesis indication information may be obtained to indicate a synthesis display mode of the ROI.

[0061] ROI includes but is not limited to the following methods:

[0062] 1. Video images can be pre-analyzed using image processing, ROI recognition, and other technologies. The analysis results can then be used to identify specific content or specific spatial locations within the panoramic video, thereby forming different ROIs. For example, during a football match, a single camera could track the ball's trajectory and use that as the ROI. Alternatively, ROI recognition technology could be used to identify and track a specific target (such as a player) in the captured video image to form an ROI.

[0063] 2. According to user needs or pre-set information, the video image is manually divided into specific content or specific spatial locations to form different regions of interest.

[0064] 3. During the video image playback process, the user's area of ​​interest information is collected, and specific content or specific spatial locations in the panoramic video are automatically divided according to this information, thereby forming different areas of interest.

[0065] 4. The user selects the area of ​​interest while watching the video image.

[0066] S102: Write the synthesis indication information and the characteristic information of the region of interest into supplemental enhancement information SEI to generate a media stream of the video image, wherein the media stream includes the SEI.

[0067] A media stream for the video image is generated based on the acquired synthesis instruction information. Specifically, the synthesis instruction information is encoded and written into the code stream of the video image, thereby generating the media stream for the video image. A playback device can decode the media stream and, at a minimum, synthesize and display the ROIs in the video image.

[0068] The composite indication information includes at least one of the following indication information:

[0069] The first indication information for indicating that the regions of interest are to be spliced ​​and displayed, the second indication information for indicating that the regions of interest are to be fused and displayed, the third indication information for indicating that the regions of interest are to be nested and displayed, the fourth indication information for indicating that the regions of interest are to be scaled and displayed, the fifth indication information for indicating that the regions of interest are to be rotated and displayed, and the sixth indication information for indicating that the regions of interest are to be cropped and displayed.

[0070] The first instruction information is used to instruct the ROIs to be stitched together. The so-called stitching means that the two ROIs are adjacent and do not overlap. Figure 2 ,Regions A, B, C, and D are four regions of interest in the video image, ,which have the same size and can be stitched together according to their ,positions where they appear in the panorama.

[0071] Optional, such as Figure 3 As shown, areas A, B, C, and D can be spliced ​​together at random positions or specified positions.

[0072] Optional, such as Figure 4 As shown, the sizes of areas A, B, C, and D may be different.

[0073] Optional, such as Figure 5 As shown, the positions of areas A, B, C, and D can be arranged at will, and their sizes are also inconsistent.

[0074] Optional, such as Figure 6 As shown, areas A, B, C, and D can be joined to form any non-rectangular shape.

[0075] The second instruction information is used to instruct the fusion of the ROIs so that there is a partial overlap between the two ROIs, but not to completely superimpose one ROI on the other. Figure 7 ,A, B, C, and D regions are four regions of interest in the video ,image, which are overlapped and fused together with a specific range of ,regions.

[0076] Optional, such as Figure 8 As shown, the composite display mode of the four ROIs can be to directly cover the pixels in a fixed covering order, and the order of superposition is A→B→C→D. Therefore, the last covered D is not covered by the other three ROIs.

[0077] Optionally, for the overlapping areas generated by fusion, the pixel values ​​can be processed in the following way: Figure 9As shown, new pixel values ​​are calculated for overlapping pixels in different regions of the four ROIs. For example, the average of all pixels, different weights for pixels in different regions, or feature matching methods can be used to calculate new pixel values, achieving a natural image fusion effect. Feature matching methods are typically used on network devices with strong video processing capabilities to achieve the best possible fusion effect. While theoretically applicable to devices, they also require higher performance.

[0078] The third type of indication information is used to indicate that the ROI should be displayed in a nested manner, where one ROI is completely overlapped with another ROI. Figure 10 Regions A and B are two regions of interest in the video image. Region B is completely overlapped on region A and nested together. The nesting position can be set according to actual needs. For example, according to the image size, a relatively small ROI of the image is overlapped on a relatively large ROI, or customized by the user.

[0079] The fourth type of indication information is used to instruct the ROI to be scaled, i.e., to change the size of the image, and includes a scaling ratio value. For example, when the scaling ratio value is 2, it can indicate that the diagonal length of the ROI is enlarged to twice its original length.

[0080] The fifth type of indication information is used to instruct to rotate the ROI, including a rotation type and a rotation angle, wherein the rotation type includes but is not limited to horizontal rotation and vertical rotation.

[0081] The sixth instruction information is used to instruct to intercept and display the area of ​​interest, see Figure 11 Regions A and B are two regions of interest in the video image. The circular region in region B is intercepted, which can be achieved using an Alpha transparent channel. Optionally, the intercepted region B can be nested with region A to synthesize the image.

[0082] In practical applications, multiple types of the above six types of indication information may be combined to perform synthesis processing on corresponding ROIs, so as to better meet the user's viewing needs for multiple ROIs.

[0083] In this embodiment, the video image may be encoded using the H.264 / AVC standard or the H.265 / HEVC (High Efficiency Video Coding) standard. During the encoding process, the obtained synthesis indication information is written into the code stream of the video image.

[0084] In other examples of the present invention, feature information of a corresponding ROI in the video image may be obtained, and a media stream of the video image may be generated based on the obtained synthesis indication information and the feature information. That is, the synthesis indication information and the feature information may be simultaneously written into the bitstream of the video image.

[0085] The generated media stream includes at least two parts: description data and a video stream. In this embodiment, the acquired synthesis indication information and feature information are written into the description data. It should be noted that the description data is primarily used to instruct the decoding of the video stream and enable playback of the video images. The description data may include at least one of the following information: time synchronization information, text information, and other related information.

[0086] It should also be noted that the description data is part of the video image and optionally exists in the following two forms: first, it can be encoded together with the video stream in the form of a code stream, that is, it is part of the data in the video stream; or it can be encoded separately from the video stream and separated from the video stream.

[0087] ROI feature information includes location information and / or encoding quality indicator information; the location information includes coordinate information of a specific location in the ROI, as well as the length and width of the ROI. The specific location can be any of the four corners of the ROI, such as the upper left pixel or the lower right pixel, or the center of the ROI. The encoding quality indicator information can be the encoding quality level used during the encoding process. Different encoding quality indicators represent different encoding quality levels, and encoding at different encoding quality levels produces different image quality. For example, the encoding quality indicator information can be "1," "2," "3," "4," "5," or "6," with different values ​​representing different encoding quality levels. For example, a coding quality indicator of "1" indicates low-quality encoding; conversely, a coding quality indicator of "2" indicates medium-quality encoding, which is better than "1." The larger the value, the higher the encoding quality.

[0088] In other examples of the present invention, the ROI position information can also be represented in the following manner: Figure 12, the upper side of the ROI area 121 is located at the 300th row of the video image, the lower side is located at the 600th row of the video image, the left side is located at the 500th column of the video image, and the right side is located at the 800th column of the video image. That is, the position information of the ROI area is identified by its row and column position. For a 1920*1080 image area, the pixel position of the upper left corner is (0,0), and the pixel position of the lower right corner is (1919,1079). When it comes to two-dimensional or three-dimensional image areas, the Cartesian coordinate system can be used, or other non-Cartesian curvilinear coordinate systems, such as cylindrical, spherical, or polar coordinate systems, can also be used.

[0089] It should be understood that the length of ROI is Figure 12 The length of the upper and lower sides, that is, the distance between the left and right sides, can be used as the length value of the ROI, that is, 800-500=300 pixels, and 600-300=300 pixels can be used as the width value of the ROI. The reverse is also possible.

[0090] The synthesis indication information and feature information of ROI are shown in Table 1 below:

[0091] Table 1

[0092]

[0093] table_id: table identifier;

[0094] vers ion: version information;

[0095] length: length information;

[0096] roi_num: contains the number of regions of interest;

[0097] (roi_position_x, roi_position_y, roi_position_z): coordinate information of the region of interest in the video image;

[0098] roi_width: width of the region of interest;

[0099] roi_height: height of the region of interest;

[0100] roi_qual ity: quality information of the region of interest;

[0101] relation_type: composite indication information of the region of interest, 0 is splicing, 1 is embedding, 2 is fusion;

[0102] (roi_new_position_x, roi_new_position_y, roi_new_position_z): coordinate information of the region of interest in the new image;

[0103] scale: the scaling ratio of the region of interest;

[0104] rotation: rotation angle of the region of interest;

[0105] fl ip: flip of the region of interest, 0 for horizontal flip, 1 for vertical flip;

[0106] alpha_flag: transparent channel identifier, 0 means there is no transparent channel information, 1 means there is transparent channel information;

[0107] alpha_info(): Transparent channel information, combined with the region of interest (intercepted) to generate a new image;

[0108] filter_info(): When relation_type is fusion mode, it can indicate the filtering mode of the fusion area, such as mean, median, etc.

[0109] user_data(): user information.

[0110] The roi_info_table containing the ROI synthesis indication information and feature information is written into the description data of the video image. The description data may optionally include at least one of the following: Supplemental Enhancement Information (SEI), Video Usability Information (VUI), and a system layer media attribute description unit.

[0111] The roi_info_table is written into the supplementary enhancement information in the video stream. A specific example may be a structure as shown in Table 2 below.

[0112] Table 2

[0113]

[0114] The roi_info_table contains relevant information of the corresponding ROI (synthesis instruction information, feature information, etc.), which is written into the supplementary enhancement information. The information identified as ROI_INFO can be obtained from the SEI information, which is equivalent to using the ROI_INFO information as the identification information of the SEI information.

[0115] The roi_info_table is written into the video availability information. For a specific example, see the structure shown in Table 3 below.

[0116] Table 3

[0117]

[0118] In Table 3, when the value of roi_info_flag is 1, it indicates that there is ROI information. roi_info_table() is the roi_info_table data structure in Table 1 above, which contains ROI-related information. Region of interest information with roi_info_flag set to 1 can be obtained from the VUI information.

[0119] The roi_info_table is written into the system layer media attribute description unit, where the system layer media attribute description unit includes but is not limited to the descriptor of the transport stream, the data unit of the file format (such as in Box), and the media description information of the transport stream (such as the Media Presentation Description (MPD) and other information units).

[0120] The ROI synthesis indication information and feature information are written into the SEI, and can be further combined with its temporal motion-constrained tile sets (MCTS). Optionally, the ROI related information is combined with the temporal motion-constrained tile sets using the H.265 / HEVC standard. By closely combining the ROI synthesis indication information with the tiles, the required tile data can be flexibly extracted without adding separate encoding and decoding ROI data. This can meet the different needs of users and is more conducive to user interaction in the application. This is shown in Table 4 below.

[0121] Table 4

[0122]

[0123] Among them, roi_info_flag: 0 indicates that there is no relevant information about the region of interest, and 1 indicates that there is relevant information about the region of interest.

[0124] An example of roi_info is shown in Table 5 below.

[0125] Table 5

[0126]

[0127] length: length information;

[0128] roi_num: contains the number of regions of interest;

[0129] (roi_pos_x, roi_pos_y): coordinate information of the region of interest in the slice group (Sl ice Group) or in the tile Ti le;

[0130] roi_width: width of the region of interest;

[0131] roi_height: height of the region of interest;

[0132] roi_qual ity: quality information of the region of interest;

[0133] relat ion_type: the relationship of the region of interest, 0 is splicing, 1 is embedding, 2 is fusion;

[0134] (roi_new_pos_x, roi_new_pos_y): coordinate information of the region of interest in the new image;

[0135] scale: the scaling ratio of the region of interest;

[0136] rotation: rotation angle of the region of interest;

[0137] fl ip: flip of the region of interest, 0 for horizontal flip, 1 for vertical flip;

[0138] alpha_flag: transparent channel identifier, 0 means there is no transparent channel information, 1 means there is transparent channel information;

[0139] alpha_info(): Transparent channel information, which can be combined with the region of interest to produce a new image;

[0140] filter_info(): When relation_type is fusion mode, the filtering method of the fusion area can be indicated, such as mean, median, etc.

[0141] The video stream included in the media stream includes video image data. The process of generating the video stream includes: obtaining a region of interest of the video image, dividing the associated images of each region of interest in the same image frame into at least one slice unit and independently encoding them to generate a first video stream of the video image.

[0142] See also Figure 13As shown, a first frame of a video image is acquired, and associated images of each ROI in the first frame are determined. Assuming that there are two ROIs in the video image, namely ROI131 and ROI132, and assuming that there are associated image A1 of ROI131 and associated image B1 of ROI132 in the first frame, the associated image A1 of ROI131 is divided into at least one slice unit for independent encoding, and the associated image B1 of ROI132 is divided into at least one slice unit for independent encoding; or both the associated image A1 and the associated image B1 are divided into at least one slice unit for independent encoding; and similar steps as for the first frame are performed on all other frames of the video image in a serial or parallel manner until the encoding of all image frames of the video image is completed to generate the first video stream.

[0143] For example, the associated image A1 is divided into a slice unit a11 for independent encoding, and the associated image B1 is divided into two slice units b11 and b12 for independent encoding.

[0144] For regions 133 of the video image other than the ROI-associated image, any existing encoding method can be used for encoding, either independently or non-independently. The resulting first video stream includes at least all independently encoded slice units for each ROI. For the receiving end, when the user only needs to view the ROI image, only the slice unit corresponding to the ROI in the first video stream can be extracted (without extracting all slice units) and independently decoded without relying on other slices to complete decoding, thereby reducing the decoding performance requirements of the receiving end.

[0145] According to needs, only the ROI-related image in the video image may be encoded, and other regions except the related image may not be encoded, or the related image and other regions may be encoded separately.

[0146] The slice units include Slice of H.264 / AVC standard, Tile of H.265 / HEVC standard, etc.

[0147] For the video stream in the media stream, it can also be a second video stream, where the process of generating the second video stream is as follows: each associated image is synthesized according to the synthesis indication information as an image frame to be processed, and the image frame to be processed is divided into at least one slice unit for encoding to generate a second video stream of the area of ​​interest.

[0148] See also Figure 14 ,The difference from the first video stream is that the second video stream is,the ROI ( Figure 14The associated images (C1 and D1, respectively) of ROI 141 and ROI 142 are first synthesized according to the synthesis instruction information, assuming here that splicing synthesis is used. The synthesized image is then used as an image frame to be processed E1. The image frame to be processed E1 is then divided into at least one slice unit (for example, e11) for encoding. The encoding method here can be independent encoding, non-independent encoding, or other encoding methods. The other image frames in the video image are also processed using the above method, and each image frame can be processed in parallel or serially. In this way, the second video stream is generated.

[0149] The second video stream can be processed by the decoder using a common decoding method. After decoding, the synthesized ROI image can be directly obtained without the need to merge the ROI-related images. This encoding method helps reduce the processing load on the decoder and improve decoding efficiency. However, the synthesis process must be performed before encoding.

[0150] In other examples of the present invention, the network side or the encoding side can generate the above two video streams for the same video image.

[0151] The generated media stream can be stored or sent to the corresponding target node. For example, upon receiving a request from the target node for obtaining a video image, the media stream is triggered to be sent to the target node. Optionally, the identification information of the content to be obtained indicated by the acquisition request is parsed, and the media stream is sent to the target node based on the identification information.

[0152] Optionally, when the identification information is the first identification, the first video stream and the description data are sent to the target node; for example, the server receives a request for a video image from a terminal and sends the media stream of the video image (including the first video stream and the description data) to the terminal according to the request. The terminal can decode the media stream to fully play the video image. Of course, the terminal can also decode the media stream, extract the independently encodable slice unit data of the region of interest, and combine it with the description data to play and display the image of the region of interest.

[0153] When the identification information is the second identification information, the slice unit of the region of interest in the first video stream is extracted (without decoding) and the description data, and sent to the target node. For example, the server side may receive a request from the terminal for the region of interest, find the independently encodable slice unit data corresponding to the region of interest based on the request information, extract it, add relevant information about the region of interest (synthesis indication information and feature information, etc.) or modified region of interest information, generate a new stream and send it to the terminal. This avoids sending the entire stream to the terminal, reducing network bandwidth usage and transmission delay.

[0154] When the identification information is the third identification, the second video stream and the description data are sent to the target node. For example, the server may also select to send the second video stream and description data of the video image to the terminal based on a request sent by the terminal. After decoding the second video stream and description data, the terminal can directly obtain the synthesized ROI image without having to synthesize the ROI based on the synthesis indication information in the description data. This helps reduce terminal resource usage and improve terminal processing efficiency.

[0155] In other examples of the present invention, the video image may be a 360-degree panoramic video, a stereoscopic video, etc. When the video image is a stereoscopic video, the relevant information of the ROI (including synthesis indication information and feature information, etc.) may be applicable to both the left and right fields of view.

[0156] The image processing method provided by the embodiment of the present invention writes synthesis indication information into the video image code stream to indicate the synthesis display of the ROI image in the video image, thereby realizing the encoding process of multiple ROIs in the video image and meeting the user's viewing needs of multiple ROI images at the same time.

[0157] By independently encoding the ROI image, the decoding end can perform independent decoding without relying on other slices for decoding. In the media stream sending method, you can choose to extract the independently decodable slice unit data where the ROI is located and send it to the terminal, without having to send all the slice data to the terminal. This is beneficial to reducing network bandwidth usage and improving transmission efficiency and decoding efficiency.

[0158] Example 2:

[0159] Based on the first embodiment, the present invention provides an image processing method, which is mainly used in terminals, decoders, etc., including but not limited to mobile phones, personal computers, etc. Figure 15 , the image processing method comprises the following steps:

[0160] S151: Receive a video stream of a video image and supplemental enhancement information SEI, wherein synthesis indication information indicating a synthesis display mode between regions of interest in the video image and feature information of the regions of interest are written into the SEI.

[0161] S152: parse the SEI to obtain synthesis indication information of the region of interest and feature information of the region of interest.

[0162] Based on the different types of description data, that is, the different locations where the ROI-related information is placed, such as SEI, VUI, MPD, etc., synthetic indication information of the ROI information is extracted. A description of the synthetic indication information is provided in Example 1 and is not repeated here. Optionally, characteristic information of the ROI, including location information and encoding quality indication information, can also be obtained from the description data.

[0163] According to the relevant information of the ROI, the ROI image data, that is, the video stream data, is obtained.

[0164] S153: Control the synthesis, playback and display of the image of the region of interest in the video stream according to the synthesis instruction information and the characteristic information of the region of interest.

[0165] The ROI image is synthesized according to the synthesis instruction information and the characteristic information of the region of interest and then played and displayed.

[0166] In other examples of the present invention, before receiving the video stream and description data of the video image, it also includes sending an acquisition request to the network side (or encoding end), and the acquisition request can also be set with identification information for indicating the acquired content, so as to obtain different video streams.

[0167] For example, when the identification information is set to the first identification, it can be used to indicate the acquisition of the first video stream and description data of the corresponding video image; when the identification information is set to the second identification, it can be used to indicate the acquisition of the slice unit and description data of the area of ​​interest in the first video stream of the corresponding video image; when the identification information is set to the third identification, it can be used to indicate the acquisition of the second video stream and description data of the corresponding video image.

[0168] When the acquisition request differs, the media stream received from the network will be different, and the subsequent processing will also differ accordingly. For example, when the identification information in the acquisition request is the first identifier, the first video stream and description data of the corresponding video image will be obtained. At this time, the first video stream and description data can be decoded to obtain the complete image of the video image, and the complete image can also be played. Alternatively, the independently encodable slice unit data of the ROI image in the first video stream is extracted, and the ROI image is synthesized according to the ROI synthesis instruction information in the description data for playback and display.

[0169] When the identification information in the acquisition request is the second identification, the independently decodable slice unit and description data of the ROI of the corresponding video image will be obtained. At this time, the terminal can directly decode the ROI independently decodable slice unit, synthesize the ROI image according to the synthesis indication information in the description data, and play and display it.

[0170] When the identification information in the acquisition request is the third identification, the second video stream and description data of the corresponding video image will be obtained. At this time, the terminal can directly decode it using a conventional decoding method to obtain the synthesized ROI image, and then play and display it.

[0171] It should be understood that the acquisition request is not limited to including identification information for indicating the content to be acquired, but should also include other necessary information, such as address information of the local end and the other end, identification information of the requested video image, verification information, etc.

[0172] Example 3:

[0173] The embodiment of the present invention provides an image processing method based on the first embodiment and / or the second embodiment, which is mainly applied to a system including a network side and a terminal side. Figure 16 , the image processing method mainly includes the following steps:

[0174] S161: The network side obtains synthesis indication information for indicating a synthesis display mode of each region of interest in a video image.

[0175] S162: The network side generates a media stream of the video image based on the synthesis indication information.

[0176] S163: The network side sends the media stream to the target node.

[0177] S164: The target node receives the media stream.

[0178] S165: The target node parses the media stream to obtain synthesis indication information of the region of interest.

[0179] S166: The target node controls the playback and display of the video stream in the media stream according to the synthesis instruction information.

[0180] For details, please refer to the relevant descriptions in Example 1 and / or Example 2, which will not be repeated here.

[0181] It should be understood that the media stream generated by the network side and the media stream sent to the target node by the network side can be the same or different. As described in Example 1 and / or Example 2, the network side can flexibly select the video stream to be sent to the target node based on the acquisition request of the target node, rather than a specific video stream. Therefore, the media stream generated by the network side can be used as the first media stream, and the media stream sent to the target node can be used as the second media stream to facilitate differentiation.

[0182] Example 4:

[0183] Based on the first embodiment, the present invention provides an image processing device for implementing the steps of the image processing method described in the first embodiment. Figure 17 , the image processing device comprises:

[0184] An acquisition module 171 is configured to acquire synthesis indication information indicating a synthesis display mode for each region of interest in a video image;

[0185] The processing module 172 is configured to write the synthesis indication information and the characteristic information of the region of interest into supplemental enhancement information (SEI) to generate a media stream of the video image, wherein the media stream includes the SEI. The specific steps of the image processing method are described in Example 1 and are not repeated here.

[0186] Embodiment 5:

[0187] Based on the second embodiment, the present invention provides an image processing device for implementing the steps of the image processing method described in the second embodiment. Figure 18 , the image processing device comprises:

[0188] A receiving module 181 is configured to receive a video stream and supplemental enhancement information (SEI) of a video image, wherein the SEI contains synthesis indication information indicating a synthesis display mode between regions of interest in the video image and feature information of the regions of interest;

[0189] A parsing module 182 is configured to parse the SEI to obtain synthesis indication information of a region of interest and feature information of the region of interest;

[0190] The control module 183 is configured to control the synthesis, playback, and display of the image of the region of interest in the video stream according to the synthesis instruction information and the characteristic information of the region of interest.

[0191] The specific steps of the image processing method can be found in the description of Example 2 and will not be repeated here.

[0192] Example 6:

[0193] The embodiment of the present invention provides an image processing system based on the embodiment 3, including the image processing device 191 as described in the embodiment 4 and the image processing device 192 as described in the embodiment 5, see Figure 19 The image processing system is used to implement the image processing method described in the third embodiment.

[0194] The specific steps of the image processing method can be found in the description of Example 3 and will not be repeated here.

[0195] Embodiment seven:

[0196] The embodiment of the present invention provides a network device based on the embodiment 1, see Figure 20 , comprising a first processor 201, a first memory 202 and a first communication bus 203;

[0197] The first communication bus 203 is used to realize the connection and communication between the first processor 201 and the first memory 202;

[0198] The first processor 201 is configured to execute one or more computer programs stored in the first memory 202 to implement the steps of the image processing method described in Example 1. Please refer to the description in Example 1 for details, which will not be repeated here.

[0199] Embodiment 8:

[0200] Based on the second embodiment, the present invention provides a terminal. Figure 21 , including a second processor 211, a second memory and a second communication bus 213;

[0201] The second communication bus 213 is used to realize the connection and communication between the second processor 211 and the second memory 212;

[0202] The second processor 211 is configured to execute one or more computer programs stored in the second memory 212 to implement the steps of the image processing method described in Example 2. Please refer to the description in Example 2 for details, which will not be repeated here.

[0203] Embodiment 9:

[0204] An embodiment of the present invention provides a storage medium based on Embodiments 1 and 2. The storage medium may be a computer-readable storage medium, which stores one or more computer programs. The one or more computer programs may be executed by one or more processors to implement the steps of the image processing method described in Embodiment 1 or Embodiment 2.

[0205] Please refer to the descriptions in Examples 1 and 2 for details, which will not be repeated here.

[0206] The storage medium includes volatile or non-volatile, removable or non-removable media implemented in any method or technology for storing information (such as computer-readable instructions, data structures, computer program modules or other data). Storage media include, but are not limited to, RAM (Random Access Memory), ROM (Read-Only Memory), EEPROM (Electrically Erasable Programmable read only memory), flash memory or other memory technology, CD-ROM (Compact Disc Read-Only Memory), digital versatile disks (DVD) or other optical disk storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to store the desired information and can be accessed by a computer.

[0207] This embodiment also provides a computer program (or computer software), which can be distributed on a computer-readable medium and executed by a computing device to implement at least one step of the image processing method in the above-mentioned embodiment one and / or embodiment two; and in some cases, at least one step shown or described can be executed in an order different from that described in the above-mentioned embodiments.

[0208] This embodiment further provides a computer program product, including a computer readable device, on which the computer program as shown above is stored. In this embodiment, the computer readable device may include the computer readable storage medium as shown above.

[0209] It can be seen that those skilled in the art should understand that all or some of the steps, systems, and functional modules / units in the methods disclosed above can be implemented as software (which can be implemented using computer program code executable by a computing device), firmware, hardware, and appropriate combinations thereof. In hardware implementations, the division between the functional modules / units mentioned in the above description does not necessarily correspond to the division of physical components; for example, a physical component can have multiple functions, or a function or step can be performed by several physical components in cooperation. Some or all physical components can be implemented as software executed by a processor, such as a central processing unit, a digital signal processor, or a microprocessor, or as hardware, or as an integrated circuit, such as an application-specific integrated circuit.

[0210] In addition, it is well known to those skilled in the art that communication media generally contain computer-readable instructions, data structures, computer program modules, or other data in a modulated data signal such as a carrier wave or other transport mechanism, and may include any information delivery media. Therefore, the present invention is not limited to any specific hardware and software combination.

[0211] The above content is a further detailed description of the embodiments of the present invention in conjunction with specific implementation methods, and the specific implementation of the present invention cannot be considered to be limited to these descriptions. For those skilled in the art of the present invention, without departing from the concept of the present invention, several simple deductions or substitutions can be made, which should be considered to fall within the scope of protection of the present invention.

Claims

1. An image processing method, comprising: Acquiring synthesis indication information for indicating a synthesis display mode between regions of interest in a video image, the synthesis indication information including at least one of the following: first indication information for instructing splicing and displaying the regions of interest, second indication information for instructing fusion and displaying the regions of interest, third indication information for instructing nesting and displaying the regions of interest, fourth indication information for instructing scaling and displaying the regions of interest, fifth indication information for instructing rotating and displaying the regions of interest, and sixth indication information for instructing truncating and displaying the regions of interest; Writing the synthesis indication information and the characteristic information of the region of interest into supplemental enhancement information (SEI) to generate a media stream of the video image, wherein the media stream includes the SEI; The media stream further includes a video stream, and the image processing method further includes: The region of interest of the video image is obtained, and associated images of each region of interest in the same image frame are divided into at least one slice unit for independent encoding to generate a first video stream of the video image.

2. The image processing method according to claim 1, wherein: After obtaining synthesis indication information for indicating a synthesis display mode between the regions of interest in the video image, the image processing method further includes: Acquire characteristic information of each region of interest.

3. The image processing method according to claim 1, wherein: The characteristic information includes position information and / or encoding quality indication information; the position information includes coordinate information of a specific position of the region of interest, and a length value and a width value of the region of interest.

4. The image processing method according to claim 1, wherein: The image processing method further includes: synthesizing the associated images according to the synthesis indication information to form an image frame to be processed, dividing the image frame to be processed into at least one slice unit for encoding to generate a second video stream of the region of interest.

5. The image processing method according to claim 4, wherein: The image processing method further includes: storing or sending the media stream to a target node.

6. The image processing method according to claim 5, wherein: Before sending the media stream to the target node, the method further includes: receiving a request from the target node for obtaining the video image.

7. The image processing method according to claim 6, wherein: The sending of the media stream to the target node includes: parsing identification information of the content to be obtained as indicated by the acquisition request, and sending the media stream to the target node according to the identification information.

8. The image processing method according to claim 7, wherein: Sending the media stream to the target node according to the identification information includes: When the identification information is the first identification, sending the first video stream and the supplemental enhancement information SEI to the target node; When the identification information is the second identification, extracting the slice unit of the region of interest and the supplementary enhancement information SEI from the first video stream, and sending them to the target node; When the identification information is the third identification, the second video stream and the supplemental enhancement information SEI are sent to the target node.

9. The image processing method according to any one of claims 1 to 8, wherein: The video image is a panoramic video image.

10. An image processing method, comprising: Receiving a video stream and supplemental enhancement information (SEI) of a video image, wherein synthesis indication information indicating a synthesis display mode between regions of interest in the video image and feature information of the regions of interest are written in the SEI, wherein the synthesis indication information includes at least one of the following: first indication information for instructing to splice and display the regions of interest, second indication information for instructing to fuse and display the regions of interest, third indication information for instructing to nest and display the regions of interest, fourth indication information for instructing to scale and display the regions of interest, fifth indication information for instructing to rotate and display the regions of interest, and sixth indication information for instructing to crop and display the regions of interest; Parsing the SEI to obtain synthesis indication information of the region of interest and feature information of the region of interest; Controlling the synthesis playback display of the image of the region of interest in the video stream according to the synthesis instruction information and the characteristic information of the region of interest; The video stream includes a first video stream, which is generated by obtaining the region of interest of the video image and dividing the associated images of each region of interest in the same image frame into at least one slice unit for independent encoding.

11. An image processing method, comprising: The network side is configured to obtain synthesis indication information indicating a synthesis display mode for each region of interest in a video image, write the synthesis indication information and feature information of the region of interest into supplemental enhancement information (SEI) to generate a media stream of the video image, and send the media stream to a target node, wherein the media stream includes the SEI. The network side is further configured to, if the media stream also includes a video stream, obtain the region of interest of the video image, divide the associated images of each region of interest in the same image frame into at least one slice unit for independent encoding, to generate a first video stream of the video image, wherein the synthesis indication information includes at least one of the following: first indication information for indicating that the regions of interest are to be spliced ​​and displayed, second indication information for indicating that the regions of interest are to be fused and displayed, third indication information for indicating that the regions of interest are to be nested and displayed, fourth indication information for indicating that the regions of interest are to be scaled and displayed, fifth indication information for indicating that the regions of interest are to be rotated and displayed, and sixth indication information for indicating that the regions of interest are to be cropped and displayed. The target node receives the media stream, parses the media stream to obtain synthesis indication information of the region of interest and characteristic information of the region of interest, and controls the playback and display of the video stream in the media stream according to the synthesis indication information and the characteristic information of the region of interest.

12. An image processing device, comprising: an acquisition module, configured to acquire synthesis indication information for indicating a synthesis display mode of each region of interest in a video image, the synthesis indication information including at least one of the following: first indication information for indicating that the regions of interest are to be spliced ​​and displayed, second indication information for indicating that the regions of interest are to be fused and displayed, third indication information for indicating that the regions of interest are to be nested and displayed, fourth indication information for indicating that the regions of interest are to be scaled and displayed, fifth indication information for indicating that the regions of interest are to be rotated and displayed, and sixth indication information for indicating that the regions of interest are to be truncated and displayed; a processing module, configured to write the synthesis indication information and the characteristic information of the region of interest into supplemental enhancement information SEI to generate a media stream of the video image, wherein the media stream includes the SEI; The acquisition module is further configured to acquire the region of interest of the video image if the media stream also includes a video stream; The processing module is further configured to divide the associated images of the regions of interest in the same image frame into at least one slice unit for independent encoding, so as to generate a first video stream of the video image.

13. An image processing apparatus, comprising: A receiving module, configured to receive a video stream and supplemental enhancement information (SEI) of a video image, wherein the SEI contains synthesis indication information indicating a synthesis display mode between regions of interest in the video image and feature information of the regions of interest; the video stream includes a first video stream, the first video stream is obtained by obtaining the regions of interest of the video image, and the associated images of each region of interest in the same image frame are divided into at least one slice unit for independent encoding and generated, and the synthesis indication information includes at least one of the following: first indication information for indicating that the regions of interest are to be spliced ​​and displayed, second indication information for indicating that the regions of interest are to be fused and displayed, third indication information for indicating that the regions of interest are to be nested and displayed, fourth indication information for indicating that the regions of interest are to be scaled and displayed, fifth indication information for indicating that the regions of interest are to be rotated and displayed, and sixth indication information for indicating that the regions of interest are to be cropped and displayed; a parsing module, configured to parse the SEI to obtain synthesis indication information of a region of interest and feature information of the region of interest; A control module is used to control the synthesis, playback and display of the image of the region of interest in the video stream according to the synthesis instruction information and the characteristic information of the region of interest.

14. An image processing system comprising: The image processing device according to claim 12 and the image processing device according to claim 13.

15. A network device comprising a first processor, a first memory, and a first communication bus; The first communication bus is used to realize connection and communication between the first processor and the first memory; The first processor is configured to execute one or more computer programs stored in the first memory to implement the steps of the image processing method according to any one of claims 1 to 9.

16. A terminal comprising a second processor, a second memory, and a second communication bus; The second communication bus is used to realize connection and communication between the second processor and the second memory; The second processor is configured to execute one or more computer programs stored in the second memory to implement the steps of the image processing method as claimed in claim 11.

17. A storage medium storing one or more computer programs, wherein the one or more computer programs can be executed by one or more processors to implement the steps of the image processing method according to any one of claims 1 to 9 or claim 10.

Citation Information

Patent Citations

  • Multi-lens optical center superposing type omnibearing shooting device and panoramic shooting and retransmitting method

    CN101521745A

  • Method and apparatus for defining and reconstructing rois in scalable video coding

    US20120201306A1

  • Image processing method, image processing apparatus, and data storage media

    US6643414B1