Image Encoding Method, Apparatus, Storage Medium, and Device

By using a method combining global and local motion vectors in video encoding technology, the target reference image unit is determined, which solves the problem of insufficient accuracy of motion vectors in the prior art and improves the image encoding quality.

CN114697678BActive Publication Date: 2025-05-27CAMBRICON TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202011618683.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-12-30
Publication Date
2025-05-27
Estimated Expiration
2040-12-30

AI Technical Summary

Technical Problem

In the existing video encoding technology, the accuracy of motion vectors is too low, which affects the quality of video encoding, and is especially poor in videos in different scenarios.

Method used

By acquiring the global motion vector between the image to be encoded and the reference image and the object attribute characteristics in the image to be encoded, the target reference image unit is determined, and encoding processing is performed based on the local and global motion vectors to generate the target coded data.

Benefits of technology

It improves the accuracy of motion vectors during inter-frame encoding, thereby improving the quality of image encoding and is suitable for videos in different scenes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114697678B_ABST
    Figure CN114697678B_ABST
Patent Text Reader

Abstract

An embodiment of the present application discloses an image encoding method, apparatus, storage medium, and device. The method includes: obtaining an image to be encoded, obtaining a reference image corresponding to the image to be encoded, where the image to be encoded includes image units to be encoded. Obtaining a global motion vector between the image to be encoded and the reference image, obtaining object attribute features in the image unit to be encoded, and determining a target reference image unit corresponding to the image unit to be encoded in the reference image according to the object attribute features and the global motion vector. Obtaining a local motion vector corresponding to the image unit to be encoded according to the target reference image unit. Encoding the image unit to be encoded according to the local motion vector and the global motion vector to generate target encoded data corresponding to the image unit to be encoded. Through the present application, the accuracy of the motion vector in the inter-frame encoding process can be improved, and thus the quality of image encoding can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer technology, and in particular, to an image encoding method, apparatus, storage medium, and device. Background Art

[0002] Video utilizes the principle of the human eye's visual persistence. By playing a series of images, it gives people the feeling of motion. Since the data volume of video is relatively large, if only the video image frames are transmitted, it requires huge transmission resources and storage space. Therefore, video can be encoded. The main function of video encoding is to encode video pixel data into a video bitstream, thereby reducing the data volume of the video and achieving the purpose of reducing the network bandwidth during transmission and reducing the storage space.

[0003] In existing video encoding technologies, a video can be divided into a series of video frames, and by calculating the motion vectors between the reference frame and the video frames, the encoding process of the video frames can be achieved. However, since the video contains rich pixel data information, such as videos in different scenarios can contain different information, the motion vectors calculated in the existing technologies lack specific information for videos in different scenarios, resulting in too low accuracy of the motion vectors, and thus affecting the video encoding quality. Summary of the Invention

[0004] The technical problem to be solved by the embodiments of the present application is to provide an image encoding method, apparatus, storage medium, and device, which can improve the accuracy of the motion vectors in the inter-frame encoding process, and thus can improve the quality of image encoding.

[0005] One aspect of the embodiments of the present application provides an image encoding method, including:

[0006] Obtain an image to be encoded, and obtain a reference image corresponding to the image to be encoded, where the image to be encoded includes image units to be encoded;

[0007] Obtain the global motion vector between the image to be encoded and the reference image, obtain the object attribute features in the image unit to be encoded, and determine the target reference image unit corresponding to the image unit to be encoded in the reference image according to the object attribute features and the global motion vector;

[0008] Obtain the local motion vector corresponding to the image unit to be encoded according to the target reference image unit;

[0009] Perform encoding processing on the image unit to be encoded according to the local motion vector and the global motion vector to generate the target encoding data corresponding to the image unit to be encoded.

[0010] Wherein, the global motion vector includes a first motion vector;

[0011] Obtaining the global motion vector between the image to be encoded and the reference image includes:

[0012] Dividing the image to be encoded into N coding image units, and obtaining coding image unit U among the N coding image units i ; N is a positive integer, and i is a positive integer less than or equal to N;

[0013] Traversing the reference image, determining K reference image units associated with the coding image unit Ui, and obtaining the first matching degrees between the coding image unit Ui and the K reference image units respectively; K is a positive integer;

[0014] Obtaining the associated coding image unit and the associated reference image unit corresponding to the maximum first matching degree, and obtaining the motion vector between the associated coding image unit and the associated reference image unit; the associated coding image unit belongs to the N coding image units, and the associated reference image unit belongs to the K reference image units;

[0015] Determining the motion vector between the associated coding image unit and the associated reference image unit as the first motion vector between the image to be encoded and the reference image.

[0016] Wherein, the image to be encoded includes the identification information corresponding to the image unit to be encoded;

[0017] Obtaining the object attribute feature in the image unit to be encoded, and determining the target reference image unit corresponding to the image unit to be encoded in the reference image according to the object attribute feature and the global motion vector includes:

[0018] Inputting the image to be encoded into a region detection model, and determining the image unit to be encoded in the image to be encoded according to the identification information;

[0019] Obtaining the object attribute feature corresponding to the image unit to be encoded according to the region detection model;

[0020] Determining the search area range and the motion offset direction corresponding to the image unit to be encoded according to the object attribute feature;

[0021] Determining the target reference image unit corresponding to the image unit to be encoded in the reference image according to the search area range, the motion offset direction and the global motion vector.

[0022] Among them, determining the target reference picture element corresponding to the picture element to be coded in the reference picture according to the search area range, the motion offset direction, and the global motion vector includes:

[0023] Determining the search starting point corresponding to the picture element to be coded in the reference picture according to the global motion vector;

[0024] Determining the target reference picture element corresponding to the picture element to be coded in the reference picture according to the search starting point, the search area range, and the motion offset direction.

[0025] Among them, determining the target reference picture element corresponding to the picture element to be coded in the reference picture according to the search starting point, the search area range, and the motion offset direction includes:

[0026] Determining the search area corresponding to the picture element to be coded in the reference picture according to the search starting point, the search area range, and the motion offset direction;

[0027] Obtaining M candidate reference picture elements covered by the search area, and obtaining the second matching degrees between the picture element to be coded and the M candidate reference picture elements respectively;

[0028] Determining the candidate reference picture element corresponding to the maximum second matching degree as the target reference picture element corresponding to the picture element to be coded.

[0029] Among them, the global motion vector further includes a second motion vector;

[0030] Obtaining the global motion vector between the picture to be coded and the reference picture includes:

[0031] Inputting the picture to be coded into a target detection model, and obtaining the target coding object in the picture to be coded in the target detection model;

[0032] Determining the target reference object associated with the target coding object in the reference picture;

[0033] Determining the motion vector between the target coding object and the target reference object as the second motion vector between the picture to be coded and the reference picture.

[0034] Among them, coding and processing the picture element to be coded according to the local motion vector and the global motion vector to generate the target coding data corresponding to the picture element to be coded includes:

[0035] Obtain the rate-distortion cost values corresponding to the local motion vector and the global motion vector respectively, and determine the motion vector corresponding to the minimum rate-distortion cost value as the target motion vector;

[0036] Encode the image unit to be encoded according to the target motion vector to generate target encoded data corresponding to the image unit to be encoded.

[0037] An embodiment of the present application provides an image encoding device on the one hand, including:

[0038] A first acquisition module, configured to acquire an image to be encoded, and acquire a reference image corresponding to the image to be encoded, where the image to be encoded includes an image unit to be encoded;

[0039] A determination module, configured to acquire a global motion vector between the image to be encoded and the reference image, acquire an object attribute feature in the image unit to be encoded, and determine a target reference image unit corresponding to the image unit to be encoded in the reference image according to the object attribute feature and the global motion vector;

[0040] A second acquisition module, configured to acquire a local motion vector corresponding to the image unit to be encoded according to the target reference image unit;

[0041] A generation module, configured to encode the image unit to be encoded according to the local motion vector and the global motion vector to generate target encoded data corresponding to the image unit to be encoded.

[0042] Wherein, the global motion vector includes a first motion vector, and the determination module includes:

[0043] A first division unit, configured to divide the image to be encoded into N encoded image units, and acquire an encoded image unit U in the N encoded image units i ; N is a positive integer, and i is a positive integer less than or equal to N;

[0044] A first acquisition unit, configured to traverse the reference image to determine K reference image units associated with the encoded image unit U i , and acquire the first matching degrees between the encoded image unit U i and the K reference image units respectively; K is a positive integer;

[0045] A second acquisition unit, configured to acquire the associated encoded image unit and the associated reference image unit corresponding to the maximum first matching degree, and acquire the motion vector between the associated encoded image unit and the associated reference image unit; the associated encoded image unit belongs to the N encoded image units, and the associated reference image unit belongs to the K reference image units;

[0046] A first determination unit, configured to determine the motion vector between the associated coded image unit and the associated reference image unit as the first motion vector between the to-be-coded image and the reference image.

[0047] Wherein, the to-be-coded image includes identification information corresponding to the to-be-coded image unit;

[0048] The determination module further includes:

[0049] A second determination unit, configured to input the to-be-coded image into a region detection model, and determine the to-be-coded image unit in the to-be-coded image according to the identification information;

[0050] A third acquisition unit, configured to acquire the object attribute features corresponding to the to-be-coded image unit according to the region detection model;

[0051] A third determination unit, configured to determine the search area range and the motion offset direction corresponding to the to-be-coded image unit according to the object attribute features;

[0052] A fourth determination unit, configured to determine the target reference image unit corresponding to the to-be-coded image unit in the reference image according to the search area range, the motion offset direction, and the global motion vector.

[0053] Wherein, the fourth determination unit is specifically configured to:

[0054] Determine the search start point corresponding to the to-be-coded image unit in the reference image according to the global motion vector;

[0055] Determine the target reference image unit corresponding to the to-be-coded image unit in the reference image according to the search start point, the search area range, and the motion offset direction.

[0056] Wherein, the fourth determination unit is specifically configured to:

[0057] Determine the search area corresponding to the to-be-coded image unit in the reference image according to the search start point, the search area range, and the motion offset direction;

[0058] Obtain M candidate reference image units covered by the search area, and obtain the second matching degrees between the to-be-coded image unit and the M candidate reference image units respectively;

[0059] Determine the candidate reference image unit corresponding to the maximum second matching degree as the target reference image unit corresponding to the to-be-coded image unit.

[0060] Among them, the global motion vector further includes a second motion vector;

[0061] The determination module further includes:

[0062] A fourth acquisition unit, configured to input the image to be encoded into a target detection model, and acquire a target encoded object in the image to be encoded in the target detection model;

[0063] A fifth determination unit, configured to determine a target reference object associated with the target encoded object in the reference image;

[0064] A sixth determination unit, configured to determine the motion vector between the target encoded object and the target reference object as the second motion vector between the image to be encoded and the reference image.

[0065] Among them, the generation module includes:

[0066] A fifth acquisition unit, configured to acquire the rate-distortion cost values corresponding to the local motion vector and the global motion vector respectively, and determine the motion vector corresponding to the minimum rate-distortion cost value as the target motion vector;

[0067] A generation unit, configured to perform encoding processing on the image unit to be encoded according to the target motion vector, and generate target encoded data corresponding to the image unit to be encoded.

[0068] On the one hand, the present application provides a computer device, including: a processor and a memory;

[0069] Among them, the memory is used to store a computer program, and the processor is used to call the above computer program to execute the following steps:

[0070] Acquire an image to be encoded, and acquire a reference image corresponding to the image to be encoded, where the image to be encoded includes an image unit to be encoded;

[0071] Acquire the global motion vector between the image to be encoded and the reference image, acquire the object attribute features in the image unit to be encoded, and determine the target reference image unit corresponding to the image unit to be encoded in the reference image according to the object attribute features and the global motion vector;

[0072] Acquire the local motion vector corresponding to the image unit to be encoded according to the target reference image unit;

[0073] Perform encoding processing on the image unit to be encoded according to the local motion vector and the global motion vector, and generate target encoded data corresponding to the image unit to be encoded.

[0074] One aspect of the embodiments of the present application provides a computer-readable storage medium storing a computer program, which is adapted to be loaded and executed by a processor to perform the following steps:

[0075] Obtain an image to be encoded, and obtain a reference image corresponding to the image to be encoded, where the image to be encoded includes image units to be encoded;

[0076] Obtain a global motion vector between the image to be encoded and the reference image, obtain object attribute features in the image unit to be encoded, and determine a target reference image unit corresponding to the image unit to be encoded in the reference image according to the object attribute features and the global motion vector;

[0077] Obtain a local motion vector corresponding to the image unit to be encoded according to the target reference image unit;

[0078] Encode the image unit to be encoded according to the local motion vector and the global motion vector to generate target encoded data corresponding to the image unit to be encoded.

[0079] One aspect of the present application provides a computer program product or a computer program, which includes computer instructions stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions to cause the computer device to execute the method in the above aspect.

[0080] In the embodiments of the present application, by obtaining an image to be encoded and obtaining a reference image corresponding to the data of the image to be encoded, the data of the image to be encoded includes image units to be encoded. Obtain a global motion vector between the image to be encoded and the reference image, obtain object attribute features in the image unit to be encoded, and determine a target reference image unit corresponding to the image unit to be encoded in the reference image according to the object attribute features and the global motion vector. Through the object attribute features and the global motion vector in the image unit to be encoded, the search area in the reference image can be reduced, the complexity of determining the target reference image unit in the reference image can be reduced, and the accuracy can be improved. Obtain a local motion vector corresponding to the image unit to be encoded according to the target reference image unit. Encode the image unit to be encoded according to the local motion vector and the global motion vector to generate target encoded data corresponding to the image unit to be encoded. By obtaining the local motion vector and the global motion vector corresponding to the image unit to be encoded, the present application combines the motion factors corresponding to the local and the global respectively, which can improve the accuracy of the motion vector in the inter-frame encoding process, and further improve the quality of image encoding. BRIEF DESCRIPTION OF THE DRAWINGS

[0081] To more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the accompanying drawings required for the description of the embodiments or the prior art. Obviously, the accompanying drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can be obtained based on these drawings.

[0082] Figure 1 It is a schematic structural diagram of an image coding system provided by an embodiment of the present application;

[0083] Figure 2 It is a schematic flow diagram of an image coding method provided by an embodiment of the present application;

[0084] Figure 3 It is a schematic diagram of a method for obtaining a reference image corresponding to an image to be coded provided by an embodiment of the present application;

[0085] Figure 4 It is a schematic diagram of obtaining a first matching degree provided by an embodiment of the present application;

[0086] Figure 5 It is a schematic diagram of a method for determining a target reference image unit corresponding to an image unit to be coded provided by an embodiment of the present application;

[0087] Figure 6 It is a schematic flow diagram of an image coding method provided by an embodiment of the present application;

[0088] Figure 7 It is a schematic structural diagram of an image coding device provided by an embodiment of the present application;

[0089] Figure 8 It is a schematic structural diagram of a computer device provided by an embodiment of the present application. Detailed implementation manners

[0090] The following will clearly and completely describe the technical solutions in the embodiments of the present application with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only some embodiments of the present application, rather than all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present application.

[0091] See Figure 1 , Figure 1 It is a schematic structural diagram of an image coding system provided by an embodiment of the present application. As Figure 1 shown, the image coding system may include a server 10 and a user terminal cluster. The user terminal cluster may include one or more user terminals, and the number of user terminals will not be limited here. AsFigure 1 As shown, it may specifically include user terminals 100a, 100b, 100c, …, 100n. As Figure 1 shown, the user terminals 100a, 100b, 100c, …, 100n can be respectively connected to the above-mentioned server 10 through a network, so that each user terminal can perform data interaction with the server 10 through this network connection.

[0092] Among them, each user terminal in the user terminal cluster can include: intelligent terminals with image encoding such as smart phones, tablet computers, laptop computers, desktop computers, wearable devices, smart homes, and head-mounted devices. It should be understood that, as Figure 1 shown, each user terminal in the user terminal cluster can be installed with a target application (i.e., an application client). When the application client runs on each user terminal, it can perform data interaction with the above-mentioned Figure 1 shown server 10 respectively.

[0093] Among them, as Figure 1 shown, the server 10 can be used to obtain the global motion vector between the image to be encoded and the reference image and the local motion vector corresponding to the image unit to be encoded, and perform encoding processing on the image unit to be encoded according to the local motion vector and the global motion vector to generate the target encoding data corresponding to the image unit to be encoded. The server 10 can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms.

[0094] For ease of understanding, an embodiment of the present application can select one user terminal from the Figure 1 shown multiple user terminals as the target user terminal. For example, an embodiment of the present application can use the Figure 1 shown user terminal 100a as the target user terminal, and a target application (i.e., an application client) with image encoding can be integrated in the target user terminal. At this time, the target user terminal can achieve data interaction with the server 10 through the service data platform corresponding to the application client. For example, the target user terminal can send the image to be encoded and the reference image corresponding to the image to be encoded to the server 10. The server 10 can obtain the global motion vector between the image to be encoded and the reference image and the local motion vector corresponding to the image unit to be encoded, perform encoding processing on the image unit to be encoded according to the local motion vector and the global motion vector to generate the target encoding data corresponding to the image unit to be encoded, and send the target encoding data to the target user terminal.

[0095] Please refer to Figure 2 , Figure 2 which is a schematic flowchart of an image encoding method provided by an embodiment of the present application. This image encoding method can be executed by a computer device, which can be a server (such as server 11 in the above Figure 1 ), or a user terminal (such as any user terminal in the user terminal cluster in the above Figure 1 ), or a system composed of a server and a user terminal. The present application does not make any limitations in this regard. As Figure 2 shown, this image encoding method can include steps S101 - S104.

[0096] S101, obtain an image to be encoded, and obtain a reference image corresponding to the image to be encoded, where the image to be encoded includes image units to be encoded.

[0097] Specifically, after the video shooting end obtains video data through a photographic device, the computer device can encode the video data to obtain image encoding data corresponding to the video data, send the image encoding data to the decoding end, and the decoding end restores the video data according to the image encoding data. By encoding the video data, the number of bits for video transmission can be greatly reduced, improving the transmission efficiency. When the computer device receives the video data, it can determine the image currently required to be encoded from the video data as the image to be encoded. After determining the image to be encoded, determine the reference image corresponding to the image to be encoded from the video data, and divide the image to be encoded into multiple image units to be encoded, that is, the image to be encoded can be divided into n×n image units to be encoded. For example, the image to be encoded can be divided into 16×16 image units to be encoded. After the image to be encoded is divided into n×n image units to be encoded, encode the n×n image units to be encoded to obtain image encoding data corresponding to the image to be encoded.

[0098] Specifically, the computer device can divide the image to be encoded into one I-frame and multiple B or P frames. The I-frame is an intra-coded frame (also known as a key frame), which is an independent frame with all information. It can be decoded independently without referring to other images. The first frame in a video sequence is always an I-frame. That is, when encoding an I-frame image, the pixel data in the I-frame image is directly encoded. The P-frame is a forward-predicted frame (also known as a forward-reference frame), which needs to refer to the previous I-frame for encoding. When decoding, it needs to rely on the pixels of the I-frame image it references and its corresponding motion vectors to restore the picture. The image to be encoded in the embodiments of the present application can refer to the image corresponding to the P-frame. If the image to be encoded can be the image corresponding to the P-frame, the reference image corresponding to the image to be encoded can refer to the image corresponding to the I-frame forward of the image to be encoded. The B-frame is a bi-directional predicted coding frame (also known as a bi-directional reference frame). What the B-frame records is the difference between this frame and the previous and next frames. That is, when decoding the image corresponding to the B-frame, it is necessary to combine the encoded data of the previous frame image of the B-frame and the encoded data of the next frame image of the B-frame. That is, the picture of the image is restored by superimposing the encoded data corresponding to the previous and next two frames and the encoded data corresponding to this B-frame image. The image to be encoded in the embodiments of the present application can refer to the image corresponding to the B-frame. If the image to be encoded can be the image corresponding to the B-frame, the reference image corresponding to the image to be encoded can refer to the image corresponding to the I-frame or P-frame forward of the image to be encoded, and the image corresponding to the I-frame or P-frame backward of the image to be encoded. Different types of image frames correspond to different reference relationships.

[0099] As Figure 3 shown, Figure 3 is a schematic diagram of a method for obtaining a reference image corresponding to an image to be encoded provided by an embodiment of the present application. As Figure 3 shown, if the image to be encoded in the present application is the second frame image, and the image frame type of this second frame image is of B-frame type, the reference image corresponding to the image to be encoded can be the first frame image and the third frame image; if the image to be encoded in the present application is the fifth frame image, and the image frame type of this second frame image is of P-frame type, the reference image corresponding to the image to be encoded can be the first frame image. Therefore, the reference image corresponding to the image to be encoded can be determined according to the image frame type corresponding to the image to be encoded.

[0100] S102. Obtain the global motion vector between the image to be encoded and the reference image, obtain the object attribute features in the image unit to be encoded, and determine the target reference image unit corresponding to the image unit to be encoded in the reference image according to the object attribute features and the global motion vector.

[0101] Specifically, after receiving the image to be encoded and the reference image, the computer device can perform motion estimation on the image to be encoded and the reference image to obtain the global motion vector between the image to be encoded and the reference image. The global motion vector refers to the offset distance of the image to be encoded relative to the reference image, that is, the relative displacement. Obtain the object attribute features in the image unit to be encoded. The object attribute features can refer to the mobility of the object in the image unit to be encoded. For example, if the image unit to be encoded includes a house and a vehicle, the object attribute feature of the house can refer to relatively small relative movement, and the object attribute feature of the vehicle can be relatively large relative movement. According to the object number attribute features and the global motion vector in the image unit to be encoded, determine the target reference image unit corresponding to the image unit to be encoded in the reference image. The target reference image unit refers to the reference image unit with the highest matching degree with the image unit to be encoded in the reference image, that is, the best matching reference image unit. Determining the target reference image unit corresponding to the image unit to be encoded according to the object attribute features and the global motion vector in the image unit to be encoded can increase the accuracy of determining the target reference unit in the reference image and reduce the complexity of determining the target reference image unit.

[0102] Optionally, the global motion vector includes a first motion vector. The specific process for the computer device to obtain the global motion vector between the image to be encoded and the reference image may include: The computer device divides the image to be encoded into N coding image units, and obtains the coding image unit U in the N coding image units i , where N is a positive integer and i is a positive integer less than or equal to N. Traverse the reference image to obtain K reference image units associated with the coding image unit Ui, and obtain the first matching degrees between the coding image unit U i and the K reference image units respectively. K is a positive integer. Obtain the associated coding image unit and the associated reference image unit corresponding to the maximum first matching degree, perform motion estimation on the associated coding image unit and the associated reference image unit, and obtain the motion vector between the associated coding image unit and the associated reference image unit. The associated coding image unit belongs to the N coding image units, and the associated reference image unit belongs to the K reference image units. Determine the motion vector between the associated coding image unit and the associated reference image unit as the first motion vector between the image to be encoded and the reference image.

[0103] Specifically, the computer device can divide the image to be encoded into N coding image units, and obtain the coding image unit U in the N coding image units i. Among them, N is a positive integer, i is a positive integer less than or equal to N, the size of the coded image unit is less than or equal to the size of the image to be coded, and is much larger than the image unit to be coded. It should be noted that the method of dividing the image to be coded into N coded image units can be the same as or different from the method of dividing the image to be coded to obtain multiple image units to be coded above. The computer device can, according to the coded image unit U i , traverse the reference image to obtain K reference image units associated with the coded image unit U i , and obtain the first matching degree between the coded image unit Ui and each of the K reference image units. A coded image unit and a reference image unit jointly correspond to a first matching degree.

[0104] As Figure 4 shown, Figure 4 is a schematic diagram of obtaining the first matching degree provided by an embodiment of the present application. As Figure 4 shown, when the computer device needs to obtain the first matching degree corresponding to the coded image unit U 1 , it can, according to the coded image unit U 1 , traverse the reference image to determine K reference image units associated with the coded image unit U 1 , and U 1 belongs to N coded image units. For example, the computer device can determine the detection frame for detection in the reference image according to the unit size corresponding to the coded image unit U 1 , and traverse the reference image according to this detection frame to obtain K reference image units associated with the coded image unit U 1 . Calculate the first matching degrees between the coded image unit U 1 and the K reference image units respectively, that is, the first matching degree between the coded image unit U 1 and the reference image unit 1, the first matching degree between the coded image unit U 1 and the reference image unit 2, until the first matching degree between the coded image unit U 1 and the reference image unit K is calculated. In this way, the K first matching degrees corresponding to the coded image unit can be obtained. By analogy, obtain the K first matching degrees corresponding to each coded image unit among the N coded image units.

[0105] Among them, the computer device can calculate the feature maps corresponding to the coded image unit and the reference image unit respectively through a convolutional neural network to obtain two corresponding groups of feature map vectors. For example, when calculating the first matching degree between the coded image unit U 1 and the reference image unit 1, the computer device can calculate the coded image unit U 1Feature maps corresponding to the reference image units 1 respectively to obtain the encoded image unit U 1 Feature map vectors corresponding to the reference image units 1 respectively. According to the feature vector maps corresponding to each encoded image unit and the corresponding reference image unit respectively, by calculating the Euclidean distance, calculate the Euclidean distance between each encoded image unit and the corresponding reference image unit. Determine the first matching degree according to the Euclidean distance between each encoded image unit and the corresponding reference image unit, that is, the shorter the Euclidean distance, the higher the corresponding first matching degree; the greater the Euclidean distance, the lower the corresponding first matching degree. And so on, obtain the K first matching degrees corresponding to each of the N encoded image units, that is, N*K first matching degrees. From the N*K first matching degrees corresponding to the N encoded image units, obtain the associated encoded image unit and the associated reference image unit corresponding to the maximum first matching degree (i.e., the shortest Euclidean distance). The associated encoded image unit belongs to the N encoded image units, and the associated reference image unit belongs to the K reference image units. If the encoded image unit U 1 has the largest matching degree with the reference image unit 2, then the encoded image unit U 1 is the associated encoded image unit corresponding to the maximum first matching degree, and the reference image unit 2 is the associated reference image unit corresponding to the maximum first matching degree. Perform motion estimation on the associated encoded image unit and the associated reference image unit, obtain the motion vector between the associated encoded image unit and the associated reference image unit, and determine the motion vector between the associated encoded image unit and the associated reference image unit as the first motion vector between the to-be-encoded image and the reference image.

[0106] Optionally, the global motion vector further includes a second motion vector. The specific manner for the computer device to obtain the global motion vector between the to-be-encoded image and the reference image may further include: input the to-be-encoded image into the target detection model, and obtain the target encoding object in the to-be-encoded image in the target detection model. Determine the target reference object associated with the target encoding object in the reference image, and determine the motion vector between the target encoding object and the target reference object as the second motion vector between the to-be-encoded image and the reference image.

[0107] Specifically, the computer device can also input the image to be encoded into the target detection model to obtain the target encoding object in the image to be encoded, such as a person, a house, or a vehicle, etc. The target encoding object is any one of one or more objects in the image to be encoded. After determining the target encoding object in the image to be encoded, the target reference object associated with the target encoding object can be determined in the reference image according to the feature information of the target encoding object, that is, the pixel information after the movement of the target encoding object is found in the reference image. Motion estimation is performed on the target encoding object and the target reference object to obtain the motion vector between the target encoding object and the target reference object, and the motion vector between the target encoding object and the target reference object is determined as the second motion vector between the image to be encoded and the reference image. Since there is one or more objects in the image to be encoded, the number of second motion vectors is one or more, and one object corresponds to one second motion vector. When the motion speed of the object in the coding image unit is very fast, the motion trajectory of the target encoding object in the coding image unit can be detected by the target detection model. Combining the category information of the objects in the image to be encoded, different objects correspond to different second motion vectors, and the second motion vector corresponding to the coding image unit can be obtained, which can improve the efficiency of the motion vector in the encoding process.

[0108] Optionally, the image to be encoded includes the identification information corresponding to the image unit to be encoded. The specific process of the computer device determining the target reference image unit corresponding to the image unit to be encoded in the reference image according to the object attribute feature and the global motion vector may include: inputting the image to be encoded into the region detection model, and determining the image unit to be encoded in the image to be encoded according to the identification information. According to the region detection model, the object attribute feature corresponding to the image unit to be encoded is obtained, and according to the object attribute feature, the search area range and the motion offset direction corresponding to the image unit to be encoded are determined. The target reference image unit corresponding to the image unit to be encoded is determined in the reference image according to the search area range, the motion offset direction, and the global motion vector.

[0109] Specifically, the computer device can input the image to be encoded into the region detection model, and determine the image unit to be encoded in the image to be encoded according to the identification information corresponding to the image unit to be encoded. The identification information corresponding to the image unit to be encoded can be used to indicate the position information of the image unit to be encoded in the image to be encoded, so that the region detection model can find the region corresponding to the image unit to be encoded in the image to be encoded. Feature extraction is performed on the image unit to be encoded according to the region detection model to determine the objects in the image unit to be encoded, such as objects like houses, vehicles, people, etc., and obtain the object attribute features corresponding to the objects in the image unit to be encoded. The object attribute feature can refer to the mobility of the object in the image unit to be encoded. For example, the relative movement of a house is small, while the relative movement of a vehicle is large. According to the object attribute features of the objects in the image unit to be encoded, determine the search area range and the motion offset direction corresponding to the image unit to be encoded. For example, since the relative movement of a house is small, the corresponding search area range will be small, and since the relative movement of a vehicle is large, the corresponding search area range will be large. After obtaining the search area range and the motion offset direction corresponding to the image unit to be encoded, determine the target reference image unit corresponding to the image unit to be encoded in the reference image according to the search area range, the motion offset direction, and the global motion vector. Through the region detection model, objects in different scenarios can be detected, thereby obtaining the search area in the reference image, searching for the target reference image unit in this search area, and then performing motion estimation on the target reference image unit and the image unit to be encoded to obtain the local motion vector corresponding to the image unit to be encoded. It can be seen that searching for the target reference image unit in the search area can improve the accuracy of motion estimation and reduce the complexity of motion estimation, thereby improving the efficiency of image encoding.

[0110] Optionally, the specific manner in which the computer device determines the target reference image unit corresponding to the image unit to be encoded in the reference image according to the search area range, the motion offset direction, and the global motion vector may include: determining the search start point corresponding to the image unit to be encoded in the reference image according to the global motion vector. Determining the target reference image unit corresponding to the image unit to be encoded in the reference image according to the search start point, the search area range, and the motion offset direction.

[0111] Specifically, the computer device can determine the corresponding coordinate information in the reference image according to the global motion vector. Since the global motion vector is a vector, such as (x, y), the corresponding coordinate information can be determined in the reference image. According to the coordinate information corresponding to the global motion vector, determine the search start point corresponding to the image unit to be encoded in the reference image. Since there are multiple pieces of data for the global motion vector, the determined search start points are also multiple. Determine the target reference image unit corresponding to the image unit to be encoded in the reference image according to the search start point, the search area range, and the motion offset direction.

[0112] Optionally, the specific process for the computer device to determine the target reference picture element corresponding to the to-be-encoded picture element in the reference picture according to the search starting point, the search area range, and the motion offset direction may include: determining the search area corresponding to the to-be-encoded picture element in the reference picture according to the search starting point, the search area range, and the motion offset direction. Obtaining M candidate reference picture elements covered by the search area, obtaining the second matching degrees between the to-be-encoded picture element and the M candidate reference picture elements respectively, and determining the candidate reference picture element corresponding to the maximum second matching degree as the target reference picture element corresponding to the to-be-encoded picture element.

[0113] Specifically, the computer device may determine the search area corresponding to the to-be-encoded picture element in the reference picture according to the search starting point, the search area range, and the motion offset direction. As Figure 5 shown, Figure 5 FIG. is a schematic diagram of a method for determining a target reference picture element corresponding to a to-be-encoded picture element provided in an embodiment of the present application. As Figure 5 shown, the position in the reference picture having the same coordinates as the to-be-encoded picture element is determined as the coordinate origin, that is, (0, 0). The search area range may be adjusted according to the object attribute characteristics corresponding to the to-be-encoded picture element. If the object attribute characteristics corresponding to the to-be-encoded picture element have less mobility, the corresponding search area range may be smaller. If the object attribute characteristics corresponding to the to-be-encoded picture element have greater mobility, the corresponding search area range may be larger. As Figure 5 shown, when the object attribute characteristics corresponding to the to-be-encoded picture element have very large mobility, the default search area range may not cover the position of the actual to-be-encoded picture element in the reference picture, resulting in a low accuracy rate of the finally obtained target reference picture element. Therefore, the size of the default search area range may be adjusted to increase the search area range. Since the data of the obtained global motion vectors is multiple, the corresponding search starting points are also multiple. The search starting points within the adjusted search area range may be determined as valid starting points, and the search starting points outside the adjusted search area range are invalid starting points. As Figure 5As shown, search start point 1, search start point 2, and search start point 3 are valid start points, while search start point 4 and search start point 5 are invalid start points. A search area is determined based on the valid search start points and a preset search size. Each valid search start point has a corresponding search area, and a full search is performed within the search area corresponding to each valid search start point to determine the target reference image unit corresponding to the image unit to be encoded. Among them, each image unit to be encoded in the image to be encoded has a corresponding second motion vector, that is, the target coding object included in the image unit to be encoded is determined, and the second motion vector corresponding to the target coding object is used as the second motion vector corresponding to the image unit to be encoded. Among them, the search order of the search areas corresponding to multiple search start points can be determined according to the motion offset direction corresponding to the image unit to be encoded. For example, the search area corresponding to the search start point in the motion offset direction is preferentially searched. After determining the search area in the reference image, M candidate reference image units covered by the search area can be obtained. The second matching degrees between the image unit to be encoded and the M reference image units are obtained, and the candidate reference image unit corresponding to the largest second matching degree is determined as the target reference image unit (i.e., the best matching reference unit) corresponding to the image unit to be encoded. In this way, the search efficiency for searching the target reference image unit in the reference image can be improved.

[0114] S103. Obtain the local motion vector corresponding to the image unit to be encoded according to the target reference image unit.

[0115] Specifically, after the computer device determines the target reference image unit in the reference image, it can obtain the local motion vector corresponding to the image unit to be encoded according to the target reference image unit. Among them, the relative displacement between the target reference image unit and the image unit to be encoded can be calculated, and the local motion vector corresponding to the image unit to be encoded is determined according to the relative displacement. Among them, reference can be made to Figure 3 to determine the reference image corresponding to the image to be encoded, and to determine the reference image unit corresponding to the image unit to be encoded in the reference image, perform motion estimation on the image unit to be encoded and the reference image unit, and obtain the local motion vector corresponding to the image unit to be encoded. The local motion vector corresponding to the image unit to be encoded may refer to the optimal motion vector selected from the temporal motion vector and the spatial motion vector corresponding to the image unit to be encoded.

[0116] S104. Perform encoding processing on the image unit to be encoded according to the local motion vector and the global motion vector to generate the target encoded data corresponding to the image unit to be encoded.

[0117] Specifically, the computer device can determine a target motion vector based on a local motion vector and a global motion vector, and perform encoding processing on the image unit to be encoded according to the target motion vector to generate target encoding data corresponding to the image unit to be encoded. The local motion vector and the global motion vector can be used as candidate motion vectors corresponding to the image unit to be encoded, and the number of global motion vectors is also multiple. The target motion vector can be determined from the local motion vector and the global motion vector.

[0118] In the embodiments of the present application, by obtaining an image to be encoded, obtaining a reference image corresponding to the data of the image to be encoded, the data of the image to be encoded includes an image unit to be encoded. Obtaining a global motion vector between the image to be encoded and the reference image, obtaining an object attribute feature in the image unit to be encoded, and determining a target reference image unit corresponding to the image unit to be encoded in the reference image according to the object attribute feature and the global motion vector. By the object attribute feature corresponding to the image unit to be encoded, determining a search direction and a search area range in the reference image, and determining a search starting point in the reference image according to the global motion vector, combining the search direction, the search area range and the search starting point, the target reference image unit corresponding to the image unit to be encoded can be found more quickly and accurately in the reference image. It can be seen that by the object attribute feature and the global motion vector in the image unit to be encoded, the search area in the reference image can be reduced, the complexity of determining the target reference image unit in the reference image can be reduced, and the accuracy can be improved. According to the target reference image unit, obtaining a local motion vector corresponding to the image unit to be encoded. Performing encoding processing on the image unit to be encoded according to the local motion vector and the global motion vector to generate target encoding data corresponding to the image unit to be encoded. In the present application, by obtaining the local motion vector and the global motion vector corresponding to the image unit to be encoded, combining the motion factors corresponding to the local and the global respectively, the accuracy of the motion vector in the inter-frame encoding process can be improved, and further the quality of the image encoding can be improved. At the same time, in this solution, the target motion vector corresponding to the image unit to be encoded is determined by machine learning (such as a region detection module and a graph matching technology, etc.), which can improve the efficiency of the image encoding.

[0119] Such as Figure 6 shown, Figure 6 is a schematic diagram of an image encoding method provided by an embodiment of the present application. This method can be executed by a computer device. This method can be executed by a computer device. The computer device can be a server (such as the server 11 above Figure 1 ), or a user terminal (such as any user terminal in the user terminal cluster above Figure 1 ), or a system composed of a server and a user terminal. The present application does not make any limitation on this. Such as Figure 6 described, the steps of the image encoding method include S201-205.

[0120] S201, obtain the image to be encoded, obtain the reference image corresponding to the image to be encoded, and the image to be encoded includes image units to be encoded.

[0121] S202, obtain the global motion vector between the image to be encoded and the reference image, obtain the object attribute features in the image unit to be encoded, and determine the target reference image unit corresponding to the image unit to be encoded in the reference image according to the object attribute features and the global motion vector.

[0122] S203, obtain the local motion vector corresponding to the image unit to be encoded according to the target reference image unit.

[0123] In the embodiments of the present application, the specific implementation manners of steps S201-203 can be referred to Figure 2 the descriptions in the corresponding embodiments, and the embodiments of the present application will not be repeated here.

[0124] S204, obtain the rate-distortion cost values corresponding to the local motion vector and the global motion vector respectively, and determine the motion vector corresponding to the minimum rate-distortion cost value as the target motion vector.

[0125] Specifically, the computer device can obtain the rate-distortion cost values corresponding to the local motion vector and the global motion vector respectively, and use the motion vector corresponding to the minimum rate-distortion cost value as the target motion vector corresponding to the image unit to be encoded. Among them, the rate-distortion cost values corresponding to the local motion vector and the global motion vector can be calculated respectively through the rate-distortion calculation function. Among them, the rate-distortion cost can be used as an evaluation criterion for the coding performance and is used for the selection of multiple options. Therefore, in the process of predicting and processing the image unit to be encoded, the computer device can determine the rate-distortion cost corresponding to the image unit to be encoded based on the predicted image unit (which refers to the predicted image unit generated according to the reference image unit and the corresponding motion vector), the image unit to be encoded, and the coding auxiliary parameters corresponding to the non-motion estimation mode (for example, the coding bit rate parameter, the coding distortion parameter), and selecting the coding mode corresponding to the minimum rate-distortion cost can obtain the optimal coding performance. Among them, the coding bit rate also represents the degree of data compression. The lower the coding bit rate, the more severely the video data is compressed. Specifically, the calculation formula of the rate-distortion cost can be referred to the following formula (1):

[0126] rdcost = dist + bit × λ (1)

[0127] Among them, rdcost refers to the rate-distortion cost, dist represents the coding distortion parameter, bit refers to the coding bit rate parameter associated with the coding mode (for example, the non-motion estimation mode), and λ is the Lagrange factor. S205, perform coding processing on the image unit to be encoded according to the target motion vector to generate the target coding data corresponding to the image unit to be encoded.

[0128] Specifically, after determining the target motion vector, the computer device can generate a predicted picture element corresponding to the picture element to be coded according to the target motion vector corresponding to the picture element to be coded (the relative displacement between the picture element to be coded and the target reference picture element) and the pixel corresponding to the target reference picture element in the reference picture. After obtaining the predicted picture element, the computer device can obtain the residual (pixel difference) between the predicted picture element and the picture element to be coded, and generate differential picture data according to the residual. Further, the computer device can transform the differential picture data, perform quantization processing on the transformed differential picture data, and generate target coded data corresponding to the picture element to be coded in the encoder. Of course, the picture to be coded may include multiple picture elements to be coded, and each picture element to be coded can obtain the target coded data corresponding to each picture to be coded based on the above operations, so as to obtain the coded data corresponding to the entire picture to be coded. Among them, transformation refers to performing an orthogonal transformation on the video frame picture to remove the correlation between spatial pixels. The orthogonal transformation makes the energy originally distributed on each pixel concentrated on a few low-frequency coefficients in the frequency domain, which represents most of the information of the picture. This characteristic of the frequency coefficient is beneficial to the quantization method based on the HVS (Human Visual System) characteristic of humans. Among them, the transformation method may include, but is not limited to: K-L transform (Karhunen-Loeve Transform), discrete cosine transform (Discrete Cosine Transform, DCT), discrete wavelet transform (Discrete Wavelet Transformation, DWT). Quantization refers to the process of reducing the data representation precision. By quantization, the amount of data to be coded can be reduced. Quantization is a lossy compression technology. Quantization can include vector quantization and scalar quantization. Vector quantization jointly quantizes a group of data, and scalar quantization independently quantizes each input data.

[0129] After the computer device generates the encoded data corresponding to the image to be encoded, it can send the encoded data of the image to be encoded to the decoding end, so that the decoding end decodes the encoded data and restores the picture corresponding to the image to be encoded. For example, when user A and user B connect for a video call, after user A's device obtains the video data, it can encode the video data through the solution of the application embodiment, that is, divide the video data into a series of video frames, and obtain the encoded data of the video data by encoding each video frame separately. Then, the encoded data of the video data can be sent to the decoding device of user B's device, so that the decoding device of user B's device decodes the encoded data corresponding to the video data to obtain the restored video data, and the restored video data is displayed on the display device (such as a mobile phone) corresponding to user B's device. In this way, user A and user B can share things with each other. Through this application, the network bandwidth during transmission can be reduced and the storage space can be reduced. At the same time, the accuracy and efficiency of restoring the image can also be improved.

[0130] In the embodiment of the present application, by obtaining the image to be encoded and the reference image corresponding to the image data to be encoded, the image data to be encoded includes image units to be encoded. Obtain the global motion vector between the image to be encoded and the reference image, obtain the object attribute features in the image unit to be encoded, and determine the target reference image unit corresponding to the image unit to be encoded in the reference image according to the object attribute features and the global motion vector. By the object attribute features corresponding to the image unit to be encoded, determine the search direction and the search area range in the reference image, and determine the search starting point in the reference image according to the global motion vector. Combining the search direction, the search area range and the search starting point, the target reference image unit corresponding to the image unit to be encoded can be found more quickly and accurately in the reference image. It can be seen that through the object attribute features and the global motion vector in the image unit to be encoded, the search area in the reference image can be reduced, the complexity of determining the target reference image unit in the reference image can be reduced, and the accuracy can be improved. According to the target reference image unit, obtain the local motion vector corresponding to the image unit to be encoded. Encode the image unit to be encoded according to the local motion vector and the global motion vector to generate the target encoded data corresponding to the image unit to be encoded. In this application, by obtaining the local motion vector and the global motion vector corresponding to the image unit to be encoded, the motion factors corresponding to the local and the global are combined. From the multiple candidate motion vectors corresponding to the local motion vector and the global motion vector, the motion vector corresponding to the maximum rate-distortion cost value is determined as the target motion vector, and encoding the image unit to be encoded according to the target motion vector can improve the accuracy of the motion vector in the inter-frame encoding process, and thus can improve the quality of image encoding. At the same time, this solution can improve the efficiency of image encoding by determining the target motion vector corresponding to the image unit to be encoded through machine learning (such as region detection modules and graph matching techniques, etc.).

[0131] Please refer to Figure 7 , Figure 7 which is a schematic structural diagram of an image encoding device provided by an embodiment of the present application. The above image encoding device may be a computer program (including program code) running on a computer device. For example, the image encoding device is an application software; the device may be used to execute corresponding steps in the image encoding method provided by the embodiment of the present application. As Figure 7 shown, the image encoding device may include: a first acquisition module 11, a determination module 12, a second acquisition module 13, and a generation module 14.

[0132] The first acquisition module 11 is configured to acquire an image to be encoded, and acquire a reference image corresponding to the image to be encoded, where the image to be encoded includes an image unit to be encoded;

[0133] The determination module 12 is configured to acquire a global motion vector between the image to be encoded and the reference image, acquire an object attribute feature in the image unit to be encoded, and determine a target reference image unit corresponding to the image unit to be encoded in the reference image according to the object attribute feature and the global motion vector;

[0134] The second acquisition module 13 is configured to acquire a local motion vector corresponding to the image unit to be encoded according to the target reference image unit;

[0135] The generation module 14 is configured to perform encoding processing on the image unit to be encoded according to the local motion vector and the global motion vector, and generate target encoded data corresponding to the image unit to be encoded.

[0136] Among them, the global motion vector includes a first motion vector, and the determination module 12 includes:

[0137] The first partitioning unit 1201 is configured to partition the image to be encoded into N encoded image units, and acquire an encoded image unit U i in the N encoded image units; N is a positive integer, and i is a positive integer less than or equal to N;

[0138] The first acquisition unit 1202 is configured to traverse the reference image, determine K reference image units associated with the encoded image unit U i , and acquire a first matching degree between the encoded image unit U i and each of the K reference image units; K is a positive integer;

[0139] The first acquisition unit 1203 is configured to acquire the associated coded image unit and the associated reference image unit corresponding to the maximum first matching degree, and acquire the motion vector between the associated coded image unit and the associated reference image unit; the associated coded image unit belongs to the N coded image units, and the associated reference image unit belongs to the K reference image units;

[0140] The first determination unit 1204 is configured to determine the motion vector between the associated coded image unit and the associated reference image unit as the first motion vector between the to-be-coded image and the reference image.

[0141] Wherein, the to-be-coded image includes identification information corresponding to the to-be-coded image unit;

[0142] The determination module 12 further includes:

[0143] The second determination unit 1205 is configured to input the to-be-coded image into a region detection model, and determine the to-be-coded image unit in the to-be-coded image according to the identification information;

[0144] The third acquisition unit 1206 is configured to acquire the object attribute feature corresponding to the to-be-coded image unit according to the region detection model;

[0145] The third determination unit 1207 is configured to determine the search area range and the motion offset direction corresponding to the to-be-coded image unit according to the object attribute feature;

[0146] The fourth determination unit 1208 is configured to determine the target reference image unit corresponding to the to-be-coded image unit in the reference image according to the search area range, the motion offset direction, and the global motion vector.

[0147] Wherein, the fourth determination unit 1208 is specifically configured to:

[0148] Determine the search start point corresponding to the to-be-coded image unit in the reference image according to the global motion vector;

[0149] Determine the target reference image unit corresponding to the to-be-coded image unit in the reference image according to the search start point, the search area range, and the motion offset direction.

[0150] Wherein, the fourth determination unit 1208 is specifically configured to:

[0151] Determine the search area corresponding to the to-be-coded image unit in the reference image according to the search start point, the search area range, and the motion offset direction;

[0152] Obtain M candidate reference image units covered by the search area, and obtain the second matching degrees between the to-be-encoded image unit and the M candidate reference image units respectively;

[0153] Determine the candidate reference image unit corresponding to the maximum second matching degree as the target reference image unit corresponding to the to-be-encoded image unit.

[0154] Wherein, the global motion vector further includes a second motion vector;

[0155] The determining module 12 further includes:

[0156] The fourth obtaining unit 1209 is configured to input the to-be-encoded image into a target detection model, and obtain a target encoding object in the to-be-encoded image in the target detection model;

[0157] The fifth determining unit 1210 is configured to determine a target reference object associated with the target encoding object in the reference image;

[0158] The sixth determining unit 1211 is configured to determine the motion vector between the target encoding object and the target reference object as the second motion vector between the to-be-encoded image and the reference image.

[0159] Wherein, the generating module 14 includes:

[0160] The fifth obtaining unit 1401 is configured to obtain the rate-distortion cost values corresponding to the local motion vector and the global motion vector respectively, and determine the motion vector corresponding to the minimum rate-distortion cost value as the target motion vector;

[0161] The generating unit 1402 is configured to perform encoding processing on the to-be-encoded image unit according to the target motion vector, and generate target encoding data corresponding to the to-be-encoded image unit.

[0162] According to an embodiment of the present application, Figure 2 Or Figure 6 The steps involved in the image encoding method shown can be performed by Figure 7 Each module in the image encoding device shown. For example, Figure 2 The step S101 shown in can be performed by Figure 1 The first obtaining module 11 in; Figure 2 The step S102 shown in can be performed by Figure 7 The determining module 12 in; Figure 2 The step S103 shown in can be performed by Figure 7 The second obtaining module 13 in; Figure 2 The step S104 shown in can be performed by Figure 7Execute by the generation module 14 in etc.

[0163] In the embodiments of the present application, by obtaining the image to be encoded, obtaining the reference image corresponding to the image data to be encoded, the image data to be encoded includes the image unit to be encoded. Obtain the global motion vector between the image to be encoded and the reference image, obtain the object attribute features in the image unit to be encoded, and determine the target reference image unit corresponding to the image unit to be encoded in the reference image according to the object attribute features and the global motion vector. Through the object attribute features corresponding to the image unit to be encoded, determine the search direction and the search area range in the reference image, and determine the search start point in the reference image according to the global motion vector. Combining the search direction, the search area range and the search start point, the target reference image unit corresponding to the image unit to be encoded can be found more quickly and accurately in the reference image, improving the efficiency and accuracy of motion estimation. According to the target reference image unit, obtain the local motion vector corresponding to the image unit to be encoded. Perform encoding processing on the image unit to be encoded according to the local motion vector and the global motion vector to generate the target encoded data corresponding to the image unit to be encoded. The present application combines the motion factors corresponding to the local and the global by obtaining the local motion vector and the global motion vector corresponding to the image unit to be encoded. From the multiple candidate motion vectors corresponding to the local motion vector and the global motion vector, determine the target motion vector as the motion vector corresponding to the maximum rate-distortion cost value, and encode the image unit to be encoded according to the target motion vector, which can improve the accuracy of the motion vector in the inter-frame encoding process, and thus can improve the quality of image encoding. At the same time, the present solution can improve the efficiency of image encoding by determining the target motion vector corresponding to the image unit to be encoded through machine learning (such as region detection module and graph matching technology, etc.).

[0164] According to an embodiment of the present application, Figure 7 Each module in the image encoding device shown can be separately or all combined into one or several units to form, or some of them can be further split into multiple smaller sub-units with smaller functions, and the same operations can be achieved without affecting the realization of the technical effects of the embodiments of the present application. The above modules are divided based on logical functions. In practical applications, the function of one module can also be realized by multiple units, or the functions of multiple modules can be realized by one unit. In other embodiments of the present application, the image encoding device may also include other units. In practical applications, these functions can also be assisted by other units and can be realized by the cooperation of multiple units.

[0165] According to an embodiment of the present application, it can be achieved by running on a general computer device such as a computer including processing elements and storage elements such as a central processing unit (CPU), a random access storage medium (RAM), and a read-only storage medium (ROM) that can execute such asFigure 2 or Figure 6 a computer program (including program code) for each step involved in the corresponding method shown in Figure 7 to construct an image encoding device as shown in

[0166] Please refer to Figure 8 , Figure 8 which is a schematic structural diagram of a computer device provided by an embodiment of the present application. As shown in Figure 8 , the above computer device 1000 may include: a processor 1001, a network interface 1004, and a memory 1005. In addition, the above computer device 1000 may further include: a user interface 1003, and at least one communication bus 1002. Among them, the communication bus 1002 is used to realize the connection and communication between these components. Among them, the user interface 1003 may include a display screen (Display) and a keyboard (Keyboard). Optionally, the user interface 1003 may further include a standard wired interface and a wireless interface. The network interface 1004 may optionally include a standard wired interface and a wireless interface (such as a WI-FI interface). The memory 1005 may be a high-speed RAM memory or a non-volatile memory, such as at least one disk memory. Optionally, the memory 1005 may further be at least one storage device located far from the aforementioned processor 1001. As shown in Figure 8 , the memory 1005, as a computer-readable storage medium, may include an operating system, a network communication module, a user interface module, and a device control application program.

[0167] In Figure 8 , in the computer device 1000 shown, the network interface 1004 can provide network communication functions; while the user interface 1003 is mainly used to provide an input interface for users; and the processor 1001 can be used to call the device control application program stored in the memory 1005 to implement:

[0168] Optionally, the processor 1001 can be used to call the device control application program stored in the memory 1005 to implement:

[0169] obtain an image to be encoded, obtain a reference image corresponding to the image to be encoded, and the image to be encoded includes image units to be encoded;

[0170] Obtain the global motion vector between the image to be encoded and the reference image, obtain the object attribute features in the image unit to be encoded, and determine the target reference image unit corresponding to the image unit to be encoded in the reference image according to the object attribute features and the global motion vector;

[0171] Obtain the local motion vector corresponding to the image unit to be encoded according to the target reference image unit;

[0172] Encode the image unit to be encoded according to the local motion vector and the global motion vector to generate the target encoded data corresponding to the image unit to be encoded.

[0173] Optionally, the processor 1001 can be used to call the device control application program stored in the memory 1005 to implement:

[0174] Divide the image to be encoded into N encoded image units, and obtain the encoded image unit Ui in the N encoded image units i ; N is a positive integer, and i is a positive integer less than or equal to N;

[0175] Traverse the reference image to determine K reference image units associated with the encoded image unit Ui i , and obtain the first matching degrees between the encoded image unit Ui i and the K reference image units respectively; K is a positive integer;

[0176] Obtain the associated encoded image unit and the associated reference image unit corresponding to the maximum first matching degree, and obtain the motion vector between the associated encoded image unit and the associated reference image unit; the associated encoded image unit belongs to the N encoded image units, and the associated reference image unit belongs to the K reference image units;

[0177] Determine the motion vector between the associated encoded image unit and the associated reference image unit as the first motion vector between the image to be encoded and the reference image.

[0178] Optionally, the processor 1001 can be used to call the device control application program stored in the memory 1005 to implement:

[0179] Input the image to be encoded into the region detection model, and determine the image unit to be encoded in the image to be encoded according to the identification information;

[0180] Obtain the object attribute features corresponding to the image unit to be encoded according to the region detection model;

[0181] Determine the search area range and the motion offset direction corresponding to the image unit to be encoded according to the object attribute characteristics;

[0182] Determine the target reference image unit corresponding to the image unit to be encoded in the reference image according to the search area range, the motion offset direction, and the global motion vector.

[0183] Optionally, the processor 1001 may be used to call the device control application program stored in the memory 1005 to implement:

[0184] Determine the search starting point corresponding to the image unit to be encoded in the reference image according to the global motion vector;

[0185] Determine the target reference image unit corresponding to the image unit to be encoded in the reference image according to the search starting point, the search area range, and the motion offset direction.

[0186] Optionally, the processor 1001 may be used to call the device control application program stored in the memory 1005 to implement:

[0187] Determine the search area corresponding to the image unit to be encoded in the reference image according to the search starting point, the search area range, and the motion offset direction;

[0188] Obtain M candidate reference image units covered by the search area, and obtain the second matching degrees between the image unit to be encoded and the M candidate reference image units respectively;

[0189] Determine the candidate reference image unit corresponding to the maximum second matching degree as the target reference image unit corresponding to the image unit to be encoded.

[0190] Optionally, the processor 1001 may be used to call the device control application program stored in the memory 1005 to implement:

[0191] Input the image to be encoded into the target detection model, and obtain the target encoding object in the image to be encoded in the target detection model;

[0192] Determine the target reference object associated with the target encoding object in the reference image;

[0193] Determine the motion vector between the target encoding object and the target reference object as the second motion vector between the image to be encoded and the reference image.

[0194] Optionally, the processor 1001 may be used to call the device control application program stored in the memory 1005 to implement:

[0195] Obtain the rate-distortion cost values corresponding to the local motion vector and the global motion vector respectively, and determine the motion vector corresponding to the minimum rate-distortion cost value as the target motion vector;

[0196] Encode the image unit to be encoded according to the target motion vector to generate target encoded data corresponding to the image unit to be encoded.

[0197] It should be understood that the computer device 1000 described in the embodiments of the present application can execute the description of the above image encoding method in the corresponding embodiments described above, or can also execute the description of the above image encoding device in the corresponding embodiments described above, which will not be elaborated here. In addition, the description of the beneficial effects of adopting the same method will not be elaborated either. Figure 2 Or Figure 6 The description of the above image encoding method in the corresponding embodiments described above, which will not be elaborated here. In addition, the description of the beneficial effects of adopting the same method will not be elaborated either. Figure 7 It should be understood that the computer device 1000 described in the embodiments of the present application can execute the description of the above image encoding method in the corresponding embodiments described above, or can also execute the description of the above image encoding device in the corresponding embodiments described above, which will not be elaborated here. In addition, the description of the beneficial effects of adopting the same method will not be elaborated either.

[0198] According to one aspect of the present application, there is provided a computer program product or a computer program, the computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device can execute the description of the image encoding method in the corresponding embodiments described above, which will not be elaborated here. In addition, the description of the beneficial effects of adopting the same method will not be elaborated either. Figure 2 Or Figure 6 The description of the image encoding method in the corresponding embodiments described above, which will not be elaborated here. In addition, the description of the beneficial effects of adopting the same method will not be elaborated either.

[0199] As an example, the above program instructions can be deployed to be executed on a computer device, or can be deployed to be executed on multiple computer devices located at one place. Or, they can be executed on multiple computer devices distributed at multiple places and interconnected through a communication network. The multiple computer devices distributed at multiple places and interconnected through a communication network can form a blockchain network.

[0200] Those of ordinary skill in the art can understand that all or part of the processes of implementing the methods in the above embodiments can be completed by instructing relevant hardware through a computer program. The above program can be stored in a computer-readable storage medium, and when the program is executed, it can include the processes of the embodiments of the above methods. Among them, the above storage medium can be a magnetic disk, an optical disk, a read-only memory (ROM), or a random access memory (RAM), etc.

[0201] The above-disclosed are only the preferred embodiments of the present application. Of course, the scope of the rights of the present application cannot be limited thereby. Therefore, equivalent changes made according to the claims of the present application still fall within the scope covered by the present application.

Claims

1. An image encoding method, characterized in that, comprising: obtaining an image to be encoded, obtaining a reference image corresponding to the image to be encoded, the image to be encoded including image units to be encoded; obtaining a global motion vector between the image to be encoded and the reference image, obtaining object attribute features in the image unit to be encoded, and determining a target reference image unit corresponding to the image unit to be encoded in the reference image according to the object attribute features and the global motion vector; obtaining a local motion vector corresponding to the image unit to be encoded according to the target reference image unit; encoding and processing the image unit to be encoded according to the local motion vector and the global motion vector to generate target encoded data corresponding to the image unit to be encoded; wherein, the image to be encoded includes identification information corresponding to the image unit to be encoded; the obtaining object attribute features in the image unit to be encoded, and determining a target reference image unit corresponding to the image unit to be encoded in the reference image according to the object attribute features and the global motion vector, includes: inputting the image to be encoded into a region detection model, and determining the image unit to be encoded in the image to be encoded according to the identification information; obtaining the object attribute features corresponding to the image unit to be encoded according to the region detection model; determining the size of the search region range and the motion offset direction corresponding to the image unit to be encoded according to the object attribute features; wherein, the search region range takes the position with the same coordinates as the image unit to be encoded in the reference image as the coordinate origin; determining a target reference image unit corresponding to the image unit to be encoded in the reference image according to the search region range, the motion offset direction, and the global motion vector; the determining a target reference image unit corresponding to the image unit to be encoded in the reference image according to the search region range, the motion offset direction, and the global motion vector, includes: determining a search start point corresponding to the image unit to be encoded in the reference image according to the global motion vector; there are multiple global motion vectors, there are multiple search start points, and the multiple search start points correspond to the multiple global motion vectors; determining a target reference image unit corresponding to the image unit to be encoded in the reference image according to the search start point, the search region range, and the motion offset direction; the determining a target reference image unit corresponding to the image unit to be encoded in the reference image according to the search start point, the search region range, and the motion offset direction, includes: Determine the search starting points within the search area from among the multiple search starting points as valid starting points. When there are multiple valid starting points, determine multiple search areas corresponding to the multiple valid starting points based on the multiple valid starting points and a preset search size, determine the search order of the multiple search areas corresponding to the multiple valid starting points according to the motion offset direction, and perform a search in accordance with the search order to obtain M candidate reference image units covered by the multiple search areas; Obtain a second matching degree between the to-be-encoded image unit and each of the M candidate reference image units; Determine the candidate reference image unit corresponding to the maximum second matching degree as the target reference image unit corresponding to the to-be-encoded image unit.

2. The method according to claim 1, wherein, the global motion vector includes a first motion vector; the obtaining of the global motion vector between the to-be-encoded image and the reference image includes: Divide the image to be encoded into N encoded image units, and obtain an encoded image unit U from the N encoded image units i ; N is a positive integer, and i is a positive integer less than or equal to N; Traverse the reference image to determine the K reference image units associated with the coded image unit U i and obtain the first matching degrees between the coded image unit U i and the K reference image units respectively; K is a positive integer; obtaining an associated encoded image unit and an associated reference image unit corresponding to the maximum first matching degree, and obtaining a motion vector between the associated encoded image unit and the associated reference image unit; the associated encoded image unit belongs to the N encoded image units, and the associated reference image unit belongs to the K reference image units; determine the motion vector between the associated encoded image unit and the associated reference image unit as the first motion vector between the to-be-encoded image and the reference image.

3. The method according to claim 2, wherein, the global motion vector further includes a second motion vector; the obtaining of the global motion vector between the to-be-encoded image and the reference image includes: input the to-be-encoded image into a target detection model, and obtain a target encoded object in the to-be-encoded image in the target detection model; determine a target reference object associated with the target encoded object in the reference image; determine the motion vector between the target encoded object and the target reference object as the second motion vector between the to-be-encoded image and the reference image.

4. The method according to claim 1, wherein, the encoding the to-be-encoded image unit according to the local motion vector and the global motion vector to generate target encoded data corresponding to the to-be-encoded image unit includes: obtaining rate-distortion cost values corresponding to the local motion vector and the global motion vector respectively, and determining the motion vector corresponding to the minimum rate-distortion cost value as the target motion vector; encoding the to-be-encoded image unit according to the target motion vector to generate target encoded data corresponding to the to-be-encoded image unit.

5. An image encoding apparatus, wherein, comprises: a first obtaining module, configured to obtain a to-be-encoded image, and obtain a reference image corresponding to the to-be-encoded image, where the to-be-encoded image includes a to-be-encoded image unit; A determination module, configured to obtain a global motion vector between the image to be encoded and the reference image, obtain object attribute features in the image unit to be encoded, and determine a target reference image unit corresponding to the image unit to be encoded in the reference image according to the object attribute features and the global motion vector; A second acquisition module, configured to obtain a local motion vector corresponding to the image unit to be encoded according to the target reference image unit; A generation module, configured to perform encoding processing on the image unit to be encoded according to the local motion vector and the global motion vector, and generate target encoded data corresponding to the image unit to be encoded; Wherein, the image to be encoded includes identification information corresponding to the image unit to be encoded; the determination module further includes: A second determination unit, configured to input the image to be encoded into a region detection model, and determine the image unit to be encoded in the image to be encoded according to the identification information; A third acquisition unit, configured to obtain the object attribute features corresponding to the image unit to be encoded according to the region detection model; A third determination unit, configured to determine the size of the search area range and the motion offset direction corresponding to the image unit to be encoded according to the object attribute features; wherein, the search area range takes the position with the same coordinates as the image unit to be encoded in the reference image as the coordinate origin; A fourth determination unit, configured to determine a target reference image unit corresponding to the image unit to be encoded in the reference image according to the search area range, the motion offset direction, and the global motion vector; The fourth determination unit is specifically configured to: Determine a search starting point corresponding to the image unit to be encoded in the reference image according to the global motion vector; there are multiple global motion vectors, there are multiple search starting points, and the multiple search starting points correspond to the multiple global motion vectors; Determine the search starting points located within the search area range as valid starting points from the multiple search starting points. When there are multiple valid starting points, determine multiple search areas corresponding to the multiple valid starting points based on the multiple valid starting points and a preset search size, determine the search order of the multiple search areas corresponding to the multiple valid starting points according to the motion offset direction, and perform a search in accordance with the search order to obtain M candidate reference image units covered by the multiple search areas; Obtain second matching degrees between the image unit to be encoded and the M candidate reference image units respectively; Determine the candidate reference image unit corresponding to the maximum second matching degree as the target reference image unit corresponding to the image unit to be encoded.

6. A computer device, Characterized in that, It includes: A processor and a memory; The memory stores a computer program, and when the computer program is executed by the processor, the processor executes the method according to any one of claims 1 to 4.

7. A computer-readable storage medium, Characterized in that, The computer-readable storage medium stores a computer program, and this computer program is suitable for being loaded and executed by a processor to execute the method according to any one of claims 1 - 4.

Citation Information

Patent Citations

  • Adaptive motion search range

    CN101366279A

  • Motion estimating method and device thereof

    CN108702512A

  • Coding method

    JP2008011455A