Playing method and device for realizing image multi-level scaling by utilizing grid coding technology

The grid coding method addresses video clarity and bandwidth issues by transcoding, encoding, and decoding videos based on user gestures, ensuring clear and efficient playback of high-resolution videos on mobile devices.

CN120321453APending Publication Date: 2025-07-15中央广播电视总台
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510501235.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-21
Publication Date
2025-07-15

AI Technical Summary

Technical Problem

When existing mobile terminal devices play high-definition videos, there are problems such as unclear effects after video amplification, excessive system bandwidth usage, and excessive system resource loss.

Method used

The video is transcoding, grid encoding, splitting and encapsulated using grid coding technology, video transmission is carried out using preset protocols, and grid slice files are loaded according to the zoom gesture information to achieve multi-level scaling.

Benefits of technology

Multi-level amplification playback of 4K and 8K videos is realized on mobile terminals, maintaining clear effects while reducing system bandwidth usage and performance consumption.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120321453A_ABST
    Figure CN120321453A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a playing method and device for realizing multi-level zooming of an image by using a grid coding technology, and the method comprises the steps: carrying out the transcoding of each video issued at the upstream, and obtaining a plurality of transcoded videos corresponding to each video; for each type of transcoded video corresponding to each video, performing grid coding on the current transcoded video to obtain a coded video; splitting the coded video, and packaging the split video according to a preset protocol to obtain a packaged video; under the condition that the zoom gesture information is received, the encapsulated video is acquired and unencapsulated, and the grid slice file in the unencapsulated video is loaded according to the zoom level in the zoom gesture information, so that 4K, 8K or even higher resolution on-demand or live broadcast video can be always subjected to multi-level magnification playing at a controllable and constant lower code rate at a mobile terminal, and the user experience is improved. Therefore, the effect of magnifying and watching more clearly is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of video playback. Specifically, it relates to a playback method and device for implementing multi-level zooming of images using trellis coding technology. Background Art

[0002] Currently, on existing mobile terminal devices, video magnification operations are usually performed in the following two ways:

[0003] First, for videos of general clarity (such as 1080P videos), since the original video resolution is low and the information of each pixel point is fixed, magnifying on this basis will cause pixel stretching and result in video blurring, which will affect the user's viewing experience.

[0004] Second, for high-definition videos (such as 4K and 8K videos), since the bit rate of such videos is large (4K bit rate is 20M - 100M, 8K bit rate is 100M+), it causes a very large bandwidth occupation when users watch videos. Since decoding needs to be performed on the mobile side during viewing, it also causes huge system performance and power consumption.

[0005] Therefore, for the existing mobile terminal players, the video effect is not clear after video magnification, the system bandwidth occupation is too high, and the system resource consumption is too large. Therefore, the user experience is poor in actual use. Summary of the Invention

[0006] An embodiment of this application provides a playback method and device for implementing multi-level zooming of images using trellis coding technology.

[0007] In the first aspect of the embodiment of this application, a playback method for implementing multi-level zooming of images using trellis coding technology is provided, including:

[0008] Transcoding each video sent from upstream to obtain multiple transcoded videos corresponding to each video respectively;

[0009] For each transcoded video corresponding to each video, performing trellis coding on the current transcoded video to obtain a coded video;

[0010] Splitting the coded video and encapsulating the split video according to a preset protocol to obtain an encapsulated video;

[0011] When receiving zoom gesture information, obtaining and decrypting the encapsulated video, and loading the grid slice files in the decrypted video according to the zoom level in the zoom gesture information.

[0012] In an optional embodiment of this application, transcoding each video sent from upstream to obtain multiple transcoded videos corresponding to each video respectively includes:

[0013] Obtain the source resolution of each video sent from the upstream;

[0014] Determine the target resolution corresponding to each magnification level according to the source resolution of the current video and the preset magnification level;

[0015] Transcode the current video according to the target resolution to obtain multiple transcoded videos corresponding to the current video.

[0016] In an optional embodiment of the present application, perform grid coding on the current transcoded video to obtain a coded video, including:

[0017] Divide the entire current transcoded video image into multiple grids, encode each grid independently, and provide a corresponding syntax structure in the bitstream of the current transcoded video, where the syntax structure is used to split and combine each grid slice at the bitstream level.

[0018] In an optional embodiment of the present application, split the coded video, including:

[0019] Use grid coding to divide the coded video into multiple regions, and independently encode and encapsulate the data for each region to form a data storage structure.

[0020] In an optional embodiment of the present application, encapsulate the split video according to a preset protocol to obtain an encapsulated video, including:

[0021] Describe the grid video track of each split video according to the preset protocol, where the grid video tracks of each split video are described in sequence from left to right and from top to bottom.

[0022] In an optional embodiment of the present application, when receiving zoom gesture information, obtain and decrypt the encapsulated video, and load the grid slice files in the decrypted video according to the zoom level in the zoom gesture information, including:

[0023] Pull the encapsulated video with a resolution of 1080P;

[0024] Determine the scaled resolution according to the zoom level in the zoom gesture information;

[0025] Obtain the encapsulated video at the scaled resolution, and determine the grid slice files in the encapsulated video at the scaled resolution according to the zoom gesture information;

[0026] Switch the encapsulated video with a resolution of 1080P to the grid slice files in the encapsulated video at the scaled resolution.

[0027] In an optional embodiment of the present application, determining the grid slice file in the encapsulated video at the scaled resolution according to the zoom gesture information includes:

[0028] Analyze the SEI unit in the encapsulated video with a resolution of 1080P to obtain the absolute time code of the current display frame;

[0029] Pull the grid slice file in the encapsulated video at the scaled resolution according to the absolute time code.

[0030] In a second aspect of the embodiments of the present application, a playback device for implementing multi-level zooming of images using grid coding technology is provided, including:

[0031] A transcoding module for transcoding each video sent from upstream to obtain multiple transcoded videos corresponding to each video respectively;

[0032] An encoding module for performing grid coding on the current transcoded video for each transcoded video corresponding to each video to obtain an encoded video;

[0033] A splitting module for splitting the encoded video and encapsulating the split video according to a preset protocol to obtain an encapsulated video;

[0034] A loading module for, when receiving the zoom gesture information, obtaining and decrypting the encapsulated video, and loading the grid slice file in the decrypted video according to the zoom level in the zoom gesture information.

[0035] In a third aspect of the embodiments of the present application, a computer device is provided, including: a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the steps of any one of the above playback methods for implementing multi-level zooming of images using grid coding technology are implemented.

[0036] In a fourth aspect of the embodiments of the present application, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps of any one of the above playback methods for implementing multi-level zooming of images using grid coding technology are implemented.

[0037] The above technical solutions provided by the embodiments of the present application compared with the prior art have at least some or all of the following advantages:

[0038] The playback method for realizing multi-level image zooming by using grid coding technology according to the embodiments of the present application transcodes each video sent from the upstream to obtain multiple transcoded videos corresponding to each video respectively; for each transcoded video corresponding to each video, grid coding is performed on the current transcoded video to obtain a coded video; the coded video is split, and the split video is encapsulated according to a preset protocol to obtain an encapsulated video; in the case of receiving zoom gesture information, the encapsulated video is obtained and decapsulated, and the grid slice files in the decapsulated video are loaded according to the zoom level in the zoom gesture information, so as to realize multi-level enlarged playback of 4K, 8K or even higher-resolution on-demand or live videos on a mobile terminal always at a controllable and constantly low bit rate, so as to achieve a clearer effect when enlarged for viewing. BRIEF DESCRIPTION OF THE DRAWINGS

[0039] The drawings described herein are used to provide a further understanding of the present application, and constitute a part of the present application. The schematic embodiments and descriptions thereof of the present application are used to explain the present application, and do not constitute an improper limitation to the present application. In the drawings:

[0040] Figure 1 is a flowchart of a playback method for realizing multi-level image zooming by using grid coding technology provided by an embodiment of the present application;

[0041] Figure 2 is a flowchart of a playback method for realizing multi-level image zooming by using grid coding technology provided by another embodiment of the present application;

[0042] Figure 3 is a schematic diagram of the basic principle of grid coding provided by an embodiment of the present application;

[0043] Figure 4 is a schematic diagram of 10*10 grid segmentation provided by an embodiment of the present application;

[0044] Figure 5 is a schematic diagram of the inflection points of the maximum display selection ratio and the minimum display selection ratio provided by an embodiment of the present application;

[0045] Figure 6 is a schematic diagram of the loading area and the display area provided by an embodiment of the present application;

[0046] Figure 7 is a flowchart of the alignment of the zoomed picture provided by an embodiment of the present application;

[0047] Figure 8 is a schematic diagram of the structure of a playback device for realizing multi-level image zooming by using grid coding technology provided by an embodiment of the present application;

[0048] Figure 9Schematic diagram of a computer device provided by an embodiment of the present application. Detailed implementation manners

[0049] In order to make the technical solutions and advantages in the embodiments of the present application clearer and more understandable, the following further describes the exemplary embodiments of the present application in detail with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than an exhaustive list of all embodiments. It should be noted that, without conflict, the embodiments in the present application and the features in the embodiments can be combined with each other.

[0050] Please refer to Figure 1 and Figure 2 The method for playing an image with multi-level zooming implemented by using grid coding technology provided by an embodiment of the present application includes the following steps S100 to S400:

[0051] S100. Transcode each video sent from the upstream to obtain multiple transcoded videos corresponding to each video respectively;

[0052] S200. For each transcoded video corresponding to each video, perform grid coding on the current transcoded video to obtain a coded video;

[0053] S300. Split the coded video and encapsulate the split video according to a preset protocol to obtain an encapsulated video;

[0054] S400. When receiving the zoom gesture information, obtain and decrypt the encapsulated video, and load the grid slice files in the decrypted video according to the zoom level in the zoom gesture information.

[0055] The method for playing an image with multi-level zooming implemented by using grid coding technology of the present application uses grid coding technology, and is a method for playing live and on-demand videos with multi-level zooming on a mobile device. The high-definition videos sent from the upstream are transcoded, and then these videos are subjected to grid coding and splitting processing by using grid coding technology; the coded videos are encapsulated into a preset protocol and sent to the streaming media for the client to call and interact. By coding the grid videos and pulling the grid video stream, the effect of being clearer when zooming in is achieved.

[0056] In an alternative embodiment of the present application, the preset protocol is the HLS (HTTP Live Streaming) protocol or the dash (Dynamic Adaptive Streaming over HTTP) protocol. Among them, the dash protocol was published by ISO / IEC in 2012 and officially became an international standard, allowing the use of any coding standard and supporting CDN (Content Delivery Network).

[0057] In an alternative embodiment of the present application, in step 100, transcoding is performed on each video sent from the upstream to obtain multiple transcoded videos corresponding to each video, including:

[0058] Obtain the source resolution of each video sent from the upstream;

[0059] Determine the target resolution corresponding to each magnification level according to the source resolution of the current video and the preset magnification level;

[0060] Transcode the current video according to the target resolution to obtain multiple transcoded videos corresponding to the current video.

[0061] In an alternative embodiment of the present application, Table 1 below shows the target resolution corresponding to each magnification level for a 4K source, and Table 2 below shows the target resolution corresponding to each magnification level for an 8K source.

[0062] Table 1

[0063] Serial number Magnification level Resolution 1 1-2 1080P 2 3-4 2K 3 5-6 3K 4 7-8 4K

[0064] Table 2

[0065]

[0066]

[0067] The grid magnification of the present application adopts a step-by-step magnification mode. Different resolutions correspond to different magnification levels. Considering the magnification experience, the magnification level is set to 8 levels, and a source resolution 8-level magnification model as shown in Table 1 and Table 2 above is constructed, which can keep the video clear after magnification and ensure that the video is always in a clear display state during the magnification process.

[0068] In an alternative embodiment of the present application, in step 200, grid encoding is performed on the current transcoded video to obtain an encoded video, including:

[0069] The entire current transcoded video image is divided into multiple grids, each grid is independently encoded, and a corresponding syntax structure is provided in the bitstream of the current transcoded video, where the syntax structure is used to split and combine each grid slice at the bitstream level.

[0070] In an optional embodiment of the present application, grid coding is a tile_based coding method supported by the H.265 coding standard (a method of dividing an image or video frame into multiple small tiles for encoding). This method allows the coding units to be independently encoded in a rectangular manner, that is, it allows the entire image to be divided into multiple grids, each grid is independently encoded, and relevant syntax structures are provided in the bitstream, so that each grid slice can be split and combined at the bitstream level according to these syntax structures. The basic principle of grid coding is as Figure 3 shown. Using grid coding, a video frame in a 2K video, a 4K video, or an 8K video is divided into multiple regions, each region is independently encoded, and independent data encapsulation is performed to form a data storage Package structure of the grid model (the package structure usually refers to a way of organizing code in programming, mainly used to distinguish different namespaces, prevent naming conflicts, and facilitate code management and maintenance) for subsequent calls. When grid coding is used for regional coding in technical implementation, it has the following characteristics:

[0071] (1) Automatic division of the video grid area to achieve grid segmentation processing without resolution limitation;

[0072] (2) Support for various input audio and video encoder formats, and the video encoder formats include H.264 / H.265;

[0073] (3) Support for various output audio and video encoder formats, the video encoder format includes H.265, and the audio encoder formats include AAC / AC3 / Audio VIVID;

[0074] (4) Support for multiple-rate grid transcoding synchronization SEI (Supplemental Enhancement Information) frame insertion for one input and multiple outputs to ensure frame synchronization of multiple-rate outputs and ensure picture synchronization.

[0075] In an optional embodiment of the present application, in step 300, splitting the encoded video includes:

[0076] Using grid coding to divide the encoded video into multiple regions, and independently encoding and independently data encapsulating each region to form a data storage structure.

[0077] This application supports custom grid area division. Theoretically, the finer the grid, the smaller the blurred edge area, and the smaller the bandwidth it occupies. When designing, the way to select the sharding area is to minimize the selected display, that is: only select the middle area for coding and use the edge as a reference, which is beneficial to the overall control of the bitrate. According to calculations, the 10*10 grid division is the optimal value selected based on the measurement of the grid coding data model, considering the display area, image quality, and bandwidth comprehensively. Figure 4 Taking the horizontal resolution value of 3840 at the maximum resolution in Figure 4 as an example, according to Table 3 below and Figure 5 the corresponding relationship between it and the following expressions: In the 10*10 grid system, to ensure that the output bandwidth of each resolution of 2K / 3K / 4K is about 4Mbps, the corresponding bitrates are 2M, 8M, 16M, and 25M respectively. Then a constant output of 6Mbps for the terminal bitstream can be achieved, effectively reducing the terminal bitrate. The bitrate output results in the 10*10 grid are shown in Table 4 below.

[0078] f1(x) = (1920 * x / n + 1) 2 / x 2

[0079] where f1(x) is the maximum selection display ratio at high resolution, x is the number of 4K horizontal grid divisions, and n is the horizontal resolution value at the maximum resolution.

[0080] f2(x) = (1920 * x / n - 1) 2 / x 2

[0081] where f2(x) is the minimum selection display ratio at high resolution, x is the number of 4K horizontal grid divisions, and n is the horizontal resolution value at the maximum resolution.

[0082] Table 3

[0083]

[0084] Table 4

[0085]

[0086] The present application designs reasonable magnification levels and the number of grids, enabling the virtual frame layer to always output a predetermined and identical bit rate when selecting different-level grid encodings. A grid encoding data model is established, stipulating an 8-level magnification for 4K / 8K videos and specifying the corresponding bit rates and resolutions to ensure a clear display state during the magnification process. Through grid splitting data calculation, in a 10*10 grid system, the output bandwidth of each resolution of 2K / 3K / 4K is approximately 6Mbps, effectively reducing the system bandwidth and performance occupancy. Therefore, using grid encoding technology to perform grid encoding and splitting on videos can ensure video magnification clarity and reduce system bandwidth occupancy.

[0087] In an optional embodiment of the present application, in step 300, the split video is encapsulated according to a preset protocol to obtain an encapsulated video, including:

[0088] Describing the grid video track of each split video according to the preset protocol, where the grid video tracks of each split video are described in sequence from left to right and from top to bottom.

[0089] When the present application uses the DASH protocol to transmit the encoded video, it encapsulates the grid-encoded videos with multiple resolutions and bit rates using the DASH protocol, and uses streaming media and CDN for transmission and invocation to achieve transmission immediacy. Among them, taking advantage of the characteristics of the DASH protocol, the encoded data stream is encapsulated into the DASH protocol to interact with the streaming media. Using the DASH protocol to combine the encoding transmission structure, it docks with the streaming media to achieve transmission and distribution. The execution steps are as follows:

[0090] Describing each grid video track through the DASH protocol;

[0091] Second, an AdaptationSet (an Adaptation Set refers to a set containing different media presentation forms, such as video, audio, and subtitles, etc.) represents a grid video track;

[0092] Third, use attributes to identify whether it is in the grid video track format, where the format is the source stream id, x coordinate, y coordinate, grid width, grid height, source stream width, and source stream height;

[0093] Fourth, the grid video tracks of each grid encoding stream are described in sequence from left to right and from top to bottom in the media distribution protocol.

[0094] This application uses the Dash protocol to achieve transmission and distribution of encoding structures with different resolutions and bit rates, and realizes fast calling of streaming media to players. In the Dash protocol, multi-level grid streaming media and metadata information are embedded, and user-side access is achieved in combination with CDN transmission services. Each grid video track is described by the Dash protocol, and attributes are established to identify and distinguish whether it is a grid video track format. The data information includes: "source stream ID, X coordinate, Y coordinate, grid width, grid height, source stream width, source stream height" and the like.

[0095] In an optional embodiment of the present application, in step 400, when the zoom gesture information is received, the packaged video is obtained and unpacked, and the grid slice file in the unpacked video is loaded according to the zoom level in the zoom gesture information, including:

[0096] Pull the packaged video with a resolution of 1080P;

[0097] Determine the resolution after scaling according to the scaling level in the scaling gesture information;

[0098] Obtaining a packaged video at a scaled resolution, and determining a grid slice file in the packaged video at the scaled resolution according to the scaling gesture information;

[0099] Switch the packaged video with a resolution of 1080P to the mesh slice file in the packaged video with the scaled resolution.

[0100] In an optional embodiment of the present application, determining a grid slice file in a packaged video at a scaled resolution according to the scaling gesture information includes:

[0101] Parse the SEI unit in the packaged video with a resolution of 1080P to obtain the absolute time code of the current display frame;

[0102] Pull mesh slice files from the packed video at scaled resolution based on absolute timecode.

[0103] In an optional embodiment of the present application, the player is responsible for pulling streams, decoding and application interaction. It will obtain the video stream from the streaming media according to the user's zoom gesture and decode and play it. Since it takes time to pull the video stream, the 1080P high-definition stream is pulled by default in the design. When a high-bitrate video is obtained, it is switched to a 4K or 8K video stream. During the switching process, single-layer and multi-level grid coding inter-screen synchronization technology will be used to achieve screen synchronization of multi-bitrate video streams:

[0104] The different-resolution pictures under a picture group respectively form a video track. Since the same picture group is output by the same encoder for different-resolution videos, picture synchronization at different resolutions can be achieved through the video track. For example, the video tracks 4K-tiled-track1 and 3K-tiled-track1 are synchronized at each resolution among the track1s;

[0105] The synchronization between each picture slice under a video track, for example, between the picture slices 4K-tiled-track1-1 and 4K-tiled-track1-2, is aligned through SEI synchronization data. The SEI is a Unix timestamp.

[0106] In an optional embodiment of the present application, during the switching process, the client obtains the mesh-encoded data stream from the streaming media according to the zoom level, and pulls the high resolution / bitrate with the current interaction point as the center point through the zoom operation, so that the clarity increases with the zoom level to reach: 2K, 3K, 4K, 8K. The pulling range of the video stream will be calculated according to the zoom level. When performing the zoom operation, specific mesh slice files will be loaded from the streaming media according to the current display area, and the parts outside the display area do not need to be loaded and decoded; and in order to effectively reduce the bandwidth, the shards at the edge of the display area are not loaded, and edge blurring processing is performed by superimposing with the 1080P reference layer high-definition video, as Figure 6 shown. Among them, when pulling the video along with the zoom level, in order to ensure the synchronization of the video during zooming, the SEI information will be used to record the absolute time code for aligning the zoomed pictures, and the steps shown in Figure 7 are executed:

[0107] (1) The server will use SEI to record the absolute time code in the generated bitstream;

[0108] (2) When the client plays 1080P, it will record the absolute time code of the current display frame by parsing the SEI unit;

[0109] (3) When it is necessary to pull the mesh shards, calculate the slice positions to be processed subsequently through the absolute time code;

[0110] (4) After decoding, synchronize by comparing the absolute time codes of the two streams.

[0111] In this application, the client obtains grid-encoded data from the streaming media according to the zoom level. After zooming, the high resolution / bitrate will be pulled with the current interaction point as the center point, and the clarity will increase with the zoom level to reach: 2K, 3K, 4K, 8K. Dragging will calculate the sliding interaction area, and the video stream pulling range will be calculated based on the calculation results; after zooming or dragging, the slice file of a specific grid will be loaded from the server according to the current display area, and the part beyond the display area does not need to be loaded and decoded; the slices at the edge of the non-display area will be blurred.

[0112] The present application discloses a method for playing back images by using grid coding technology to realize multi-level zooming, which realizes multi-level zooming operation on images and has the advantages of clear zooming effect, smooth picture zooming, low bandwidth occupancy, and low system performance consumption.

[0113] It should be understood that, although the various steps in the flow chart are displayed in sequence according to the indication of the arrows, these steps are not necessarily executed in sequence according to the order indicated by the arrows. Unless there is a clear description in this article, the execution of these steps is not strictly limited in order, and these steps can be executed in other orders. Moreover, at least a portion of the steps in the figure may include multiple sub-steps or multiple stages, and these sub-steps or stages are not necessarily executed at the same time, but can be executed at different times, and the execution order of these sub-steps or stages is not necessarily to be carried out in sequence, but can be executed in turn or alternately with other steps or at least a portion of the sub-steps or stages of other steps.

[0114] See also Figure 8 An embodiment of the present application provides a playback device 800 for implementing multi-level image scaling using a grid coding technology, comprising:

[0115] The transcoding module 810 is used to transcode each video sent from the upstream to obtain a plurality of transcoded videos corresponding to each video;

[0116] The encoding module 820 is used for performing grid encoding on the current transcoded video for each type of transcoded video corresponding to each video to obtain an encoded video;

[0117] The splitting module 830 is used to split the encoded video and encapsulate the split video according to a preset protocol to obtain an encapsulated video;

[0118] The loading module 840 is used to obtain and unpack the packaged video when receiving the zoom gesture information, and load the grid slice file in the unpacked video according to the zoom level in the zoom gesture information.

[0119] For the specific limitations of the above-mentioned device 800, reference can be made to the limitations of the method for implementing multi-level zooming of images using grid coding technology in the above text, which will not be elaborated here. Each module in the above-mentioned device 800 can be implemented in whole or in part by software, hardware, and their combination. Each of the above modules can be embedded in the processor of the computer device in hardware form or be independent of it, or can be stored in the memory of the computer device in software form, so as to facilitate the processor to call and execute the operations corresponding to each of the above modules.

[0120] In one embodiment, a computer device is provided. The internal structure diagram of the computer device can be as Figure 9 shown. The computer device includes a processor, a memory, a network interface, and a database connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store data. The network interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, it implements a method for playing an image with multi-level zooming using grid coding technology as described above. It includes: including a memory and a processor, the memory stores a computer program, and when the processor executes the computer program, it implements any step in the method for playing an image with multi-level zooming using grid coding technology as described above.

[0121] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by the processor, it can implement any step in the method for playing an image with multi-level zooming using grid coding technology as described above.

[0122] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0123] This application is described with reference to the flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram, as well as the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the functions specified in one or more flows Figure 1 or more flows and / or blocks Figure 1 or means for implementing the functions specified in one or more blocks.

[0124] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufacture including instruction means for implementing the functions specified in one or more flows Figure 1 or more flows and / or blocks Figure 1 or means for implementing the functions specified in one or more blocks.

[0125] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, so that the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one or more flows Figure 1 or more flows and / or blocks Figure 1 or means for implementing the functions specified in one or more blocks.

[0126] Although the preferred embodiments of the present application have been described, those skilled in the art can make additional changes and modifications once they know the basic creative concepts. Therefore, the appended claims are intended to be construed to include the preferred embodiments and all changes and modifications falling within the scope of the present application.

[0127] Obviously, those skilled in the art can make various changes and modifications to the present application without departing from the spirit and scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the claims of the present application and their equivalent technologies, the present application is also intended to include these changes and modifications.

Claims

1. A playback method for implementing multi-level zooming of an image using trellis coding technology, characterized in that, It includes: Transcode each video sent from upstream respectively to obtain multiple transcoded videos corresponding to each video respectively; For each transcoded video corresponding to each video, perform grid encoding on the current transcoded video to obtain an encoded video; Split the encoded video and encapsulate the split video according to a preset protocol to obtain an encapsulated video; When receiving pinch gesture information, obtain and decrypt the encapsulated video, and load the grid slice files in the decrypted video according to the zoom level in the pinch gesture information.

2. The method according to claim 1, characterized in that, Transcode each video sent from upstream respectively to obtain multiple transcoded videos corresponding to each video respectively, including: Obtain the source resolution of each video sent from upstream; Determine the target resolution corresponding to each zoom level according to the source resolution of the current video and the preset zoom level; Transcode the current video according to the target resolution to obtain multiple transcoded videos corresponding to the current video.

3. The method according to claim 1, wherein Perform grid encoding on the current transcoded video to obtain an encoded video, including: Divide the entire current transcoded video image into multiple grids, independently encode each grid, and provide a corresponding syntax structure in the bitstream of the current transcoded video, where the syntax structure is used to split and combine each grid slice at the bitstream level.

4. The method according to claim 1, wherein Split the encoded video, including: Use grid encoding to divide the encoded video into multiple regions, and independently encode and independently encapsulate the data for each region to form a data storage structure.

5. The method according to claim 1, wherein Encapsulate the split video according to a preset protocol to obtain an encapsulated video, including: Describe the grid video track of each split video according to a preset protocol, where the grid video tracks of each split video are described in sequence from left to right and from top to bottom.

6. The method according to claim 1, characterized in that, When receiving pinch gesture information, obtain and decrypt the encapsulated video, and load the grid slice files in the decrypted video according to the zoom level in the pinch gesture information, including: Pull the encapsulated video with a resolution of 1080P; Determine the scaled resolution according to the zoom level in the pinch gesture information; Obtain the encapsulated video at the scaled resolution, and determine the grid slice files in the encapsulated video at the scaled resolution according to the pinch gesture information; Switch the encapsulated video with a resolution of 1080P to the grid slice files in the encapsulated video at the scaled resolution.

7. The method according to claim 6, wherein Determine the grid slice files in the encapsulated video at the scaled resolution according to the pinch gesture information, including: Parse the SEI unit in the encapsulated video with a resolution of 1080P to obtain the absolute time code of the current display frame; Pull the grid slice files in the encapsulated video at the scaled resolution according to the absolute time code.

8. A playback device for realizing multi-level zooming of an image by using grid coding technology, characterized in that, It includes: A transcoding module for transcoding each video sent from upstream respectively to obtain multiple transcoded videos corresponding to each video respectively; An encoding module for performing grid encoding on the current transcoded video for each transcoded video corresponding to each video to obtain an encoded video; A splitting module for splitting the encoded video and encapsulating the split video according to a preset protocol to obtain an encapsulated video; A loading module, configured to, when receiving zoom gesture information, obtain and decrypt the encapsulated video, and load grid slice files in the decrypted video according to the zoom level in the zoom gesture information.

9. A computer device, comprising: It includes a memory and a processor, and the memory stores a computer program. It is characterized in that when the processor executes the computer program, the steps of the method for playing images with multi-level zoom using grid coding technology according to any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, the steps of the method for playing images with multi-level zoom using grid coding technology according to any one of claims 1 to 7 are implemented.