Image recognition method and device, electronic equipment and readable storage medium

By generating and recognizing the contour image of the target image, the problem of computational resource consumption caused by decoding processing in the image recognition process is solved, and more efficient image recognition is achieved.

CN116563771BActive Publication Date: 2026-07-21TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
TENCENT TECHNOLOGY (SHENZHEN) CO LTD
Filing Date
2022-01-28
Publication Date
2026-07-21

AI Technical Summary

Technical Problem

In existing technologies, the decoding process in image recognition requires a large amount of computing resources, resulting in low recognition efficiency.

Method used

By acquiring the encoded data of the target image, and utilizing the relationship between the texture complexity of macroblocks and the encoded data, a contour image of the target image is generated, and the contour image is recognized, thus avoiding the decoding process of the target image.

Benefits of technology

It saves computing resources and improves image recognition efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116563771B_ABST
    Figure CN116563771B_ABST
Patent Text Reader

Abstract

The application discloses an image recognition method and device, electronic equipment and readable storage medium, and belongs to the technical field of image processing. The method comprises the following steps: acquiring encoding data of a target image, wherein the encoding data of the target image comprises encoding data of a plurality of macroblocks in the target image, the consumed information amount of any macroblock in the plurality of macroblocks during encoding processing is proportional to the texture complexity of the any macroblock, and the texture complexity of the any macroblock is related to the pixel value of each pixel point in the any macroblock; acquiring a contour image corresponding to the target image based on the encoding data of each macroblock, wherein the contour image is used for reflecting the contour of an object in the target image; and performing image recognition processing on the contour image to obtain an image recognition result. Since the encoding data of the target image does not need to be decoded to obtain the target image, a large amount of computing resources can be saved, and the image recognition efficiency is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image processing technology, and in particular to an image recognition method, apparatus, electronic device, and readable storage medium. Background Technology

[0002] In the field of image processing technology, image recognition technology, used to identify target objects in images, is of great importance, and its applications are becoming increasingly widespread. For example, image recognition of landscape images can identify animals within them.

[0003] In related technologies, the image data acquired by electronic devices may be coded data obtained after encoding an image. In this case, the electronic device needs to first decode the coded image data to obtain the final image, and then perform image recognition. Since decoding the coded image data consumes significant computational resources, it reduces the efficiency of image recognition. Summary of the Invention

[0004] This application provides an image recognition method, apparatus, electronic device, and readable storage medium, which can be used to solve the problem of low image recognition efficiency caused by the large amount of computing resources required for decoding processing. The technical solution includes the following contents.

[0005] On one hand, embodiments of this application provide an image recognition method, the method comprising:

[0006] The encoded data of the target image is obtained. The encoded data of the target image includes the encoded data of multiple macroblocks in the target image. The amount of information consumed when encoding any macroblock in the multiple macroblocks is proportional to the texture complexity of the macroblock. The texture complexity of the macroblock is related to the pixel value of each pixel in the macroblock.

[0007] The contour image corresponding to the target image is obtained based on the encoded data of each macroblock, and the contour image is used to reflect the contour of the object in the target image;

[0008] The contour image is subjected to image recognition processing to obtain the image recognition result.

[0009] On the other hand, embodiments of this application provide an image recognition device, the device comprising:

[0010] The acquisition module is used to acquire the encoded data of the target image. The encoded data of the target image includes the encoded data of multiple macroblocks in the target image. The amount of information consumed when any macroblock in the multiple macroblocks is encoded is proportional to the texture complexity of the macroblock. The texture complexity of the macroblock is related to the pixel value of each pixel in the macroblock.

[0011] The acquisition module is further configured to acquire a contour image corresponding to the target image based on the encoded data of each macroblock, wherein the contour image is used to reflect the contour of the object in the target image;

[0012] The image recognition module is used to perform image recognition processing on the contour image to obtain the image recognition result.

[0013] In one possible implementation, the acquisition module is configured to perform statistical processing on the encoded data of each macroblock to obtain the amount of information consumed by each macroblock; determine the largest first amount of information from the information consumed by each macroblock; and determine the contour image corresponding to the target image based on the information consumed by each macroblock and the first amount of information.

[0014] In one possible implementation, the acquisition module is configured to, for any given macroblock, perform statistical processing on at least one coded component contained in the coded data of the given macroblock to obtain the amount of information consumed by each coded component corresponding to the given macroblock, wherein each coded component is any one of macroblock type, macroblock prediction, coded block mode, quantization parameter offset, and residual; and determine the sum of the amounts of information consumed by each coded component corresponding to the given macroblock as the amount of information consumed by the given macroblock.

[0015] In one possible implementation, the acquisition module is configured to, for any given macroblock, determine the ratio between the amount of information consumed by the macroblock and the first amount of information; map the ratios corresponding to each macroblock to a grayscale range to obtain the contour image corresponding to the target image.

[0016] In one possible implementation, the acquisition module is configured to: divide any given macroblock into multiple sub-macroblocks; perform statistical processing on the encoded data of the given macroblock to obtain the amount of information consumed by each sub-macroblock of the given macroblock; determine the largest second amount of information from the amount of information consumed by each sub-macroblock of the multiple macroblocks; and determine the contour image corresponding to the target image based on the amount of information consumed by each sub-macroblock of the multiple macroblocks and the second amount of information.

[0017] In one possible implementation, the acquisition module is configured to, for any macroblock, perform statistical processing on at least one coded component contained in the coded data of the macroblock to obtain the amount of information consumed by each coded component corresponding to each sub-macroblock of the macroblock, wherein each coded component is any one of macroblock type, macroblock prediction, coded block mode, quantization parameter offset, and residual; and for any sub-macroblock, determine the sum of the amount of information consumed by each coded component corresponding to the sub-macroblock as the amount of information consumed by the sub-macroblock.

[0018] In one possible implementation, the acquisition module is configured to, for any coded component contained in the coded data of any macroblock, in response to any of the following: macroblock type, macroblock prediction, coded block mode, or quantization parameter offset, calculate the amount of information consumed by the coded component to obtain the amount of information consumed by the coded component corresponding to each sub-macroblock of the macroblock.

[0019] In one possible implementation, the acquisition module is configured to, for any coded component contained in the coded data of any macroblock, in response to the fact that the any coded component is the residual, and the number of residual coefficient matrices included in the residual is not less than the number of sub-macroblocks of the macroblock, determine the residual coefficient matrix corresponding to any sub-macroblock of the macroblock from the residual coefficient matrices included in the residual, and calculate the amount of information consumed by the residual coefficient matrix corresponding to the sub-macroblock; in response to the fact that the any coded component is the residual, and the number of residual coefficient matrices included in the residual is less than the number of sub-macroblocks of the macroblock, determine the amount of information consumed by each sub-macroblock corresponding to the residual coefficient matrix based on the number of sub-macroblocks corresponding to the residual coefficient matrix and the amount of information consumed by the residual coefficient matrix.

[0020] In one possible implementation, the acquisition module is configured to, for any given sub-macroblock, determine the ratio between the amount of information consumed by the given sub-macroblock and the second amount of information; map the ratios corresponding to each sub-macroblock of the plurality of macroblocks to a grayscale range to obtain the contour image corresponding to the target image.

[0021] In one possible implementation, the device further includes:

[0022] The processing module is configured to, in response to the image recognition result indicating that the contour image contains sensitive content, filter the encoded data of the target image, or decode the encoded data of the target image to obtain the target image, and perform masking processing on the sensitive content of the target image.

[0023] On the other hand, embodiments of this application provide an electronic device, which includes a processor and a memory. The memory stores at least one computer program, which is loaded and executed by the processor to enable the electronic device to implement any of the image recognition methods described above.

[0024] On the other hand, a computer-readable storage medium is also provided, wherein at least one computer program is stored in the computer-readable storage medium, the at least one computer program being loaded and executed by a processor to enable a computer to implement any of the image recognition methods described above.

[0025] On the other hand, a computer program or computer program product is also provided, wherein at least one computer program is stored in the computer program or computer program product, and the at least one computer program is loaded and executed by a processor to enable the computer to implement any of the above-described image recognition methods.

[0026] The technical solution provided in this application has at least the following beneficial effects:

[0027] The technical solution provided in this application is based on the encoded data of each macroblock included in the encoded data of the target image to obtain the contour image corresponding to the target image, and then performing image recognition on the contour image. Since it is not necessary to decode the encoded data of the target image to obtain the target image, a large amount of computing resources can be saved and the image recognition efficiency can be improved. Attached Figure Description

[0028] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0029] Figure 1 This is a schematic diagram of the implementation environment of an image recognition method provided in an embodiment of this application;

[0030] Figure 2 This is a flowchart of an image recognition method provided in an embodiment of this application;

[0031] Figure 3 This is a schematic diagram of a contour image provided in an embodiment of this application;

[0032] Figure 4 This is a schematic diagram of the encoded data of a macroblock provided in an embodiment of this application;

[0033] Figure 5This is a schematic diagram of video processing provided in an embodiment of this application;

[0034] Figure 6 This is a schematic diagram of an image recognition process provided in an embodiment of this application;

[0035] Figure 7 This is a schematic diagram of the structure of an image recognition device provided in an embodiment of this application;

[0036] Figure 8 This is a schematic diagram of the structure of a terminal device provided in an embodiment of this application;

[0037] Figure 9 This is a schematic diagram of the structure of a server provided in an embodiment of this application. Detailed Implementation

[0038] To make the objectives, technical solutions, and advantages of this application clearer, the embodiments of this application will be described in further detail below with reference to the accompanying drawings. First, the terms involved in the embodiments of this application will be explained and described.

[0039] Pixel field: The state of video pixel data.

[0040] Compression domain: The state of compressed video data obtained after video pixel data has been encoded (such as H.264 / AVC encoding). The encoding process is the process of converting video from the pixel domain to the compression domain, and conversely, the decoding process is the process of restoring video from the compression domain back to the pixel domain.

[0041] Contour map: also known as edge map, is a grayscale or binarized image formed by outlining the edges of objects in an image.

[0042] Figure 1 This is a schematic diagram of the implementation environment of an image recognition method provided in an embodiment of this application, such as... Figure 1 As shown, the implementation environment includes a terminal device 101 and a server 102. The image recognition method in this embodiment can be executed by the terminal device 101, by the server 102, or by both the terminal device 101 and the server 102.

[0043] Terminal device 101 can be a smartphone, game console, desktop computer, tablet computer, laptop computer, smart TV, smart in-vehicle device, smart voice interaction device, smart home appliance, etc. Server 102 can be a single server, a server cluster consisting of multiple servers, or any of the following: cloud computing platform and virtualization center. This application embodiment does not limit this. Server 102 can communicate with terminal device 101 via a wired network or wireless network. Server 102 can have functions such as data processing, data storage, and data transmission and reception. This application embodiment does not limit this. The number of terminal devices 101 and servers 102 is not limited and can be one or more.

[0044] The image recognition method in this application embodiment can be implemented based on cloud technology. Cloud technology refers to a hosting technology that unifies a series of resources such as hardware, software, and networks within a wide area network or local area network to realize data computation, storage, processing, and sharing.

[0045] Cloud technology is a collective term for network technologies, information technologies, integration technologies, management platform technologies, and application technologies applied to the cloud computing business model. It can form resource pools, providing flexible and convenient on-demand access. Cloud computing technology will become a crucial support. Backend services of technical network systems require substantial computing and storage resources, such as video websites, image websites, and many portal websites. With the rapid development and application of the internet industry, every item may have its own identification mark in the future, requiring transmission to backend systems for logical processing. Data at different levels will be processed separately, and various industry data will all require robust system support, which can only be achieved through cloud computing.

[0046] Based on the above implementation environment, this application provides an image recognition method to... Figure 2 The flowchart shown in this embodiment of the application illustrates an image recognition method. This method can be implemented by... Figure 1 The method can be executed by either terminal device 101 or server 102, or jointly by both. For ease of description, the terminal device 101 or server 102 executing the image recognition method in this embodiment is referred to as an electronic device, and the method can be executed by an electronic device. Figure 2 As shown, the method includes steps 201 to 203.

[0047] Step 201: Obtain the encoded data of the target image. The encoded data of the target image includes the encoded data of multiple macroblocks in the target image. The amount of information consumed when encoding any macroblock is processed is proportional to the texture complexity of any macroblock. The texture complexity of any macroblock is related to the pixel value of each pixel in any macroblock.

[0048] This application does not limit the method of obtaining the encoded data of the target image. For example, the target image is encoded using H.264 / Advanced Video Coding (AVC) to obtain the encoded data of the target image, where H.264 is a digital video compression format. By encoding the target image, it is possible to convert the target image from the pixel domain to the compressed domain, reducing data redundancy in the target image.

[0049] Optionally, the target image can be any frame from a video, which can be a live stream or a pre-recorded video. For live streams, the camera of the first terminal device (such as the broadcaster's terminal device) captures video in real time, encodes the captured video to obtain a video stream, and sends the video stream to an electronic device. The video stream can also be called a bitstream or bitrate. The electronic device extracts the bitstream of any frame from the video stream; the bitstream of that frame is the encoded data of the target image.

[0050] The target image consists of multiple macroblocks, each macroblock comprising one luma pixel block and two chroma pixel blocks. Therefore, the encoded data of the target image includes the encoded data of multiple macroblocks. For live video, the encoded data of any macroblock can also be referred to as the bitstream or bitstream of that macroblock.

[0051] In this embodiment, the target image is encoded, such as by encoding the target image using H.264 / AVC, to obtain the encoded data of the target image. Generally, when encoding a target image, if the macroblock texture is relatively complex, the amount of information consumed during macroblock encoding is relatively large; conversely, if the macroblock texture is relatively simple, the amount of information consumed during macroblock encoding is relatively small. That is, the amount of information consumed during macroblock encoding is directly proportional to the texture complexity of the macroblock, and the texture complexity of the macroblock is related to the pixel values ​​of each pixel in the macroblock. In other words, the amount of information consumed during macroblock encoding can be obtained based on the macroblock encoded data.

[0052] Step 202: Obtain the contour image corresponding to the target image based on the encoded data of each macroblock. The contour image is used to reflect the contour of the object in the target image.

[0053] As mentioned above, the amount of information consumed during macroblock encoding can be obtained based on the macroblock's encoding data. The texture complexity of a macroblock is related to the pixel values ​​of each pixel within it. Therefore, based on the encoding data of each macroblock, the pixel information of each macroblock can be obtained, thereby generating the contour image corresponding to the target image. The contour image reflects the shape of the object in the target image. In this embodiment, the shape of the object refers to its outline; the two terms have the same meaning.

[0054] Please see Figure 3 , Figure 3 This is a schematic diagram of a contour image provided in an embodiment of this application. (a) is the contour image corresponding to an I-frame. An I-frame, also known as an intra-picture, typically refers to the first frame, which, after appropriate compression, serves as a reference point for random access and can be considered as an image. (b) is the contour image corresponding to a P-frame, which is predicted from the preceding B-frame or I-frame. (c) is the contour image corresponding to a B-frame, which is predicted from the adjacent preceding frame, the current frame, and the following frame. The main content of the contour image can be seen from (a), (b), and (c); therefore, the contour image can reflect the shape of objects in the image.

[0055] This application embodiment can acquire the contour image corresponding to the target image based on the coded data of each macroblock using two methods. See implementation method A1 and implementation method A2 below for details.

[0056] Implementation method A1 involves obtaining the contour image corresponding to the target image based on the encoded data of each macroblock, including: performing statistical processing on the encoded data of each macroblock to obtain the amount of information consumed by each macroblock; determining the largest first amount of information from the information consumed by each macroblock; and determining the contour image corresponding to the target image based on the information consumed by each macroblock and the first amount of information.

[0057] In this embodiment, statistical processing is performed on the encoded data of any macroblock to obtain the amount of information consumed by the macroblock. The amount of information consumed by the macroblock includes, but is not limited to, the number of bits consumed by the macroblock. The number of bits consumed by the macroblock can also be referred to as the bit overhead of the macroblock. The amount of information consumed by the macroblock refers to the amount of information consumed during the encoding process of the macroblock, and this amount of information consumed can be used to represent the information entropy of the macroblock. It should be noted that the encoded data of any macroblock includes at least one encoded component, through which the encoded information of the macroblock is recorded.

[0058] In one possible implementation, statistical processing is performed on the encoded data of each macroblock to obtain the amount of information consumed by each macroblock. This includes: for any macroblock, statistical processing is performed on at least one encoded component contained in the encoded data of any macroblock to obtain the amount of information consumed by each encoded component corresponding to any macroblock. Each encoded component is any one of macroblock type, macroblock prediction, encoded block mode, quantization parameter offset, and residual. The sum of the amounts of information consumed by each encoded component corresponding to any macroblock is determined as the amount of information consumed by any macroblock.

[0059] In this embodiment, each macroblock corresponds to a macroblock type, which can be a skip type, a direct type, a pulse code modulation (PCM) type, or other types. These other types can be referred to as non-skip types, non-direct types, and non-PCM types. The encoded components included in the encoded data of a macroblock may differ depending on its macroblock type.

[0060] Please see Figure 4 , Figure 4 This is a schematic diagram of the encoded data of a macroblock provided in an embodiment of this application. Wherein, Figure 4 When the macroblock type is not Skip, not Direct, or not PCM, the encoded data of the macroblock includes all the encoded components. The encoded data of a macroblock includes, but is not limited to, the macroblock type (Mb_Type), macroblock prediction (Mb_Pred), coded block pattern (CBP), quantization parameter offset (QP_Off), and residuals.

[0061] The macroblock type records the encoding type of the macroblock, which includes prediction method, segmentation size, and inter-frame reference direction. Prediction methods include intra-frame prediction and inter-frame prediction. In intra-frame prediction, both the predicted and actual values ​​are located in the current frame image. Intra-frame prediction primarily eliminates spatial redundancy in the image. Intra-frame prediction has a relatively low compression ratio, can be decoded independently, and does not depend on data from other frames besides the current frame image. Keyframe images in a video can use intra-frame prediction. In inter-frame prediction, the actual value is located in the current frame image, and the predicted value is located in the reference frame image. Inter-frame prediction primarily eliminates temporal redundancy in the image. Inter-frame prediction has a higher compression ratio than intra-frame prediction, but cannot be decoded independently; the current frame image can only be reconstructed after the data from the reference frame image is obtained. The segmentation size describes the size information of the macroblock. Macroblock segmentation sizes include 16×16, 16×8, 8×16, 8×8, 4×4, etc. The inter-frame reference direction is the direction of the reference frame image relative to the current frame image. Inter-frame reference directions include forward, backward, and bidirectional. Forward refers to the playback order of the video, with the reference frame image preceding the current frame image. Backward refers to the playback order of the video, with the reference frame image following the current frame image. Bidirectional refers to the playback order of the video, with the reference frame image appearing both before and after the current frame image.

[0062] Macroblock prediction includes information related to macroblock prediction. Macroblocks are divided into I-macroblocks, P-macroblocks, and B-macroblocks. I-macroblocks include Intra Prediction Mode (IPM), meaning they perform intra-frame prediction, while P-macroblocks and B-macroblocks perform inter-frame prediction. The segmentation sizes for both P-macroblocks and B-macroblocks include 16×16, 16×8, 8×16, and 8×8. For P-macroblocks or B-macroblocks with segmentation sizes of 16×16, 16×8, and 8×16, they include a Reference Index (Ref_Idx) and Motion Vector Difference (MVD). For P-macroblocks or B-macroblocks with a segmentation size of 8×8, they include a Sub-Macroblock Type (Sub_Mb_Type), a Reference Index, and Motion Vector Difference.

[0063] The block mode of the encoding represents the residual encoding scheme of the macroblock. Four bits are used to record whether the luminance residual coefficient matrix contains non-zero values, and two bits are used to record whether the chrominance residual coefficient matrix contains non-zero values.

[0064] The quantization parameter offset represents the offset of the macroblock's quantization parameter value relative to the quantization parameter value of the current frame image. Each macroblock can have one quantization parameter value, which is used to quantize the macroblock's residual coefficients to obtain quantized residual coefficients. The macroblock's residual coefficients reflect the accuracy of the macroblock's prediction values. The current frame image also has one quantization parameter value, which is used to quantize the current frame image's residual coefficients. The current frame image's residual coefficients reflect the accuracy of the current frame image's prediction values.

[0065] The residuals include the quantization residual coefficients for macroblock luminance and macroblock chrominance, where macroblock luminance includes one luminance component Y, and macroblock chrominance includes two chrominance components U and V.

[0066] When the macroblock type is Skip, Direct, or PCM, the encoded data of the macroblock includes at least one encoded component. This application embodiment does not limit the encoded components included in the encoded data of the macroblock.

[0067] In this embodiment, statistical processing of the macroblock types contained in the encoded data of any macroblock yields the amount of information consumed by the corresponding macroblock type. Statistical processing of the macroblock predictions contained in the encoded data of any macroblock yields the amount of information consumed by the corresponding macroblock prediction. Statistical processing of the encoded block modes contained in the encoded data of any macroblock yields the amount of information consumed by the corresponding encoded block mode. Statistical processing of the quantization parameter offsets contained in the encoded data of any macroblock yields the amount of information consumed by the corresponding quantization parameter offset. Statistical processing of the residuals contained in the encoded data of any macroblock yields the amount of information consumed by the corresponding residual.

[0068] Next, the information consumed by any one of the following is calculated: the information consumed by the macroblock type corresponding to any macroblock, the information consumed by macroblock prediction corresponding to that macroblock, the information consumed by the block mode of encoding corresponding to that macroblock, the information consumed by the quantization parameter offset corresponding to that macroblock, and the information consumed by the residual corresponding to that macroblock. This summation yields the information consumed by any macroblock. In this way, the information consumed by each macroblock in the target image can be determined.

[0069] Next, the largest amount of information consumed by each macroblock in the target image is determined and recorded as the first amount of information. Based on the information consumed by each macroblock and the first amount of information, the contour image corresponding to the target image is determined.

[0070] In one possible implementation, the contour image corresponding to the target image is determined based on the amount of information consumed by each macroblock and the first amount of information, including: for any macroblock, determining the ratio between the amount of information consumed by any macroblock and the first amount of information; mapping the ratio corresponding to each macroblock to a grayscale range to obtain the contour image corresponding to the target image.

[0071] In this embodiment of the application, the amount of information consumed by any macroblock and the first amount of information are used to normalize the amount of information consumed by the macroblock. The normalization process can be carried out by calculating the ratio between the amount of information consumed by any macroblock and the first amount of information. The ratio is the ratio corresponding to any macroblock, and the value range of the ratio is [0, 1].

[0072] Next, the ratio corresponding to any macroblock is multiplied by 255 to map the ratio to a grayscale range. The mapped ratio for any macroblock is recorded as its mapped value, which ranges from [0, 255]. After mapping the ratios of all macroblocks in the target image to grayscale ranges, the contour image corresponding to the target image is obtained, and this contour image is a grayscale image.

[0073] By implementing method A1, the amount of information consumed by each macroblock is statistically obtained, and based on the amount of information consumed by each macroblock, the contour image corresponding to the target image is determined. This contour image can accurately reflect the main content of the target image. To better reflect the main content of the target image and improve the clarity of the contour image, implementation method A2 can be used to determine the contour image corresponding to the target image.

[0074] Implementation method A2 involves obtaining the contour image corresponding to the target image based on the encoded data of each macroblock, including: for any macroblock, dividing the macroblock into multiple sub-macroblocks; performing statistical processing on the encoded data of any macroblock to obtain the amount of information consumed by each sub-macroblock of the macroblock; determining the largest second amount of information from the amount of information consumed by each sub-macroblock of the multiple macroblocks; and determining the contour image corresponding to the target image based on the amount of information consumed by each sub-macroblock of the multiple macroblocks and the second amount of information.

[0075] For any macroblock in the target image, if the macroblock is located in the m-th row and n-th column of the target image, then the macroblock can be denoted as {MB}. (m,n) |0≤m≤M-1,0≤n≤N-1}, where M is the number of macroblocks contained in the width direction of the target image and N is the number of macroblocks contained in the length direction of the target image. That is, the target image includes M rows and N columns of macroblocks. The M rows are denoted as row 0 to row (M-1) and the N columns are denoted as column 0 to column (N-1).

[0076] Any macroblock can be divided into multiple sub-macroblocks, and the embodiments of this application do not limit the number of sub-macroblocks. For example, any macroblock can be divided into 16 sub-macroblocks with a partition size of 4×4. In this case, any macroblock includes 4 rows and 4 columns of sub-macroblocks, denoted as row 0 to row 3 and column 0 to column 3 respectively. Then, the macroblock MB... (m,n) The sub-macroblock in row a and column b is

[0077] Next, statistical processing is performed on the encoded data of any macroblock to obtain the amount of information consumed by each sub-macroblock of that macroblock. The amount of information consumed by any sub-macroblock includes, but is not limited to, the number of bits consumed by that sub-macroblock, which can also be called the bit overhead of that sub-macroblock. The amount of information consumed by a sub-macroblock refers to the amount of information consumed during the encoding process of the sub-macroblock, and this amount of information consumed can be used to represent the information entropy of the sub-macroblock.

[0078] In one possible implementation, statistical processing is performed on the encoded data of any macroblock to obtain the amount of information consumed by each sub-macroblock of any macroblock. This includes: for any macroblock, statistical processing is performed on at least one encoded component contained in the encoded data of any macroblock to obtain the amount of information consumed by each encoded component corresponding to each sub-macroblock of any macroblock. Each encoded component is any one of macroblock type, macroblock prediction, encoded block mode, quantization parameter offset, and residual; for any sub-macroblock, the sum of the amount of information consumed by each encoded component corresponding to any sub-macroblock is determined as the amount of information consumed by any sub-macroblock.

[0079] It should be noted that the coding components contained in the coding data of any macroblock have been described above, therefore, the coding components will not be described again in the embodiments of this application.

[0080] In this embodiment of the application, for any one of the coding components in the coding data of any macroblock, including macroblock type, macroblock prediction, coding block mode, quantization parameter offset, and residual, statistical processing is performed on the coding component to obtain the amount of information consumed by the coding component corresponding to each sub-macroblock of the macroblock. The statistical processing method will be described in detail below.

[0081] In one possible implementation, statistical processing is performed on at least one coded component contained in the coded data of any macroblock to obtain the amount of information consumed by each coded component corresponding to each sub-macroblock of any macroblock. This includes: for any coded component contained in the coded data of any macroblock, in response to any coded component being a macroblock type, macroblock prediction, coded block mode, or quantization parameter offset, the amount of information consumed by any coded component is statistically analyzed to obtain the amount of information consumed by any coded component corresponding to each sub-macroblock of any macroblock.

[0082] The four coded components—macroblock type, macroblock prediction, coded block mode, and quantization parameter offset—belong to the macroblock itself, not to any specific sub-macroblock. In other words, these four coded components are shared information across all the macroblock's sub-macroblocks. Therefore, any one of these four coded components contained in the coded data of any macroblock is equivalent to any one of these coded components contained in the coded data of all its sub-macroblocks.

[0083] By statistically processing the macroblock types contained in the encoded data of any macroblock, we can obtain the amount of information consumed by the corresponding macroblock type (bit{mb_type}). (m,n)}. For any sub-macroblock within this macroblock, the amount of information consumed by the macroblock type corresponding to this macroblock is bit{mb_type}. (m,n)} represents the amount of information consumed by the macroblock type corresponding to this sub-macroblock.

[0084] By performing statistical processing on the macroblock predictions contained in the encoded data of any macroblock, we can obtain the amount of information consumed by the macroblock predictions corresponding to that macroblock, in bits{mb_pred}. (m,n)}. For any sub-macroblock within this macroblock, the amount of information consumed by the macroblock prediction corresponding to this macroblock is bit{mb_pred}. (m,n)} represents the amount of information consumed in macroblock prediction corresponding to that sub-macroblock.

[0085] By performing statistical processing on the coded block patterns contained in the coded data of any macroblock, the amount of information consumed by the coded block patterns corresponding to that macroblock in bits{CBP} can be obtained. (m,n)}. For any sub-macroblock within this macroblock, the amount of information consumed by the block mode corresponding to the macroblock's encoding is bit{CBP}. (m,n)} represents the amount of information consumed by the block mode of the encoding corresponding to the sub-macroblock.

[0086] By statistically processing the quantization parameter offsets contained in the encoded data of any macroblock, we can obtain the amount of information consumed by the quantization parameter offsets of that macroblock, in bits{QP_off}. (m,n)}. For any sub-macroblock within this macroblock, the amount of information consumed by the quantization parameter offset corresponding to this macroblock is bit{QP_off}. (m,n)}, which is the amount of information consumed by the quantization parameter offset corresponding to the sub-macroblock.

[0087] In one possible implementation, statistical processing is performed on at least one coded component contained in the coded data of any macroblock to obtain the amount of information consumed by each coded component corresponding to each sub-macroblock of any macroblock. This includes: for any coded component contained in the coded data of any macroblock, in response to any coded component being a residual, and the number of residual coefficient matrices included in the residual being not less than the number of sub-macroblocks of any macroblock, then the residual coefficient matrix corresponding to any sub-macroblock of any macroblock is determined from the residual coefficient matrix included in the residual, and the amount of information consumed by the residual coefficient matrix corresponding to any sub-macroblock is statistically analyzed; in response to any coded component being a residual, and the number of residual coefficient matrices included in the residual being less than the number of sub-macroblocks of any macroblock, then the amount of information consumed by each sub-macroblock corresponding to any residual coefficient matrix is ​​determined based on the number of sub-macroblocks corresponding to any residual coefficient matrix and the amount of information consumed by any residual coefficient matrix.

[0088] The residuals include the quantization residual coefficients for macroblock luma and quantization residual coefficients for macroblock chroma. Since the quantization parameter coefficients for luma of different sub-macroblocks of a macroblock may differ, and the quantization parameter coefficients for chroma of different sub-macroblocks of a macroblock may also differ, it is necessary to determine the residuals contained in the coded data of any sub-macroblock of a macroblock based on the residuals contained in the coded data of any macroblock. For ease of description, the residuals contained in the coded data of any sub-macroblock of a macroblock are described below from the perspective of a macroblock.

[0089] The coded data of a macroblock contains residuals including multiple residual coefficient matrices. At the same time, the macroblock is divided into multiple sub-macroblocks. Therefore, the residuals contained in the coded data of any sub-macroblock can be determined based on the relationship between the number of residual coefficient matrices and the number of sub-macroblocks.

[0090] When the number of residual coefficient matrices is not less than the number of sub-macroblocks (meaning each sub-macroblock corresponds to at least one residual coefficient matrix), the residual coefficient matrix corresponding to any sub-macroblock is determined from the multiple residual coefficient matrices. This residual coefficient matrix represents the residuals contained in the encoded data of that sub-macroblock. Then, statistical processing is performed on the residual coefficient matrix corresponding to that sub-macroblock to obtain the amount of information consumed by it.

[0091] For example, for macroblocks (MB) (m,n)A 4×4 discrete cosine transform yields 16 residual coefficient matrices. Each residual coefficient matrix is ​​4×4 in size and can be denoted as... macroblock MB (m,n) The macroblock is divided into 16 sub-macroblocks of size 4×4. At this point, each sub-macroblock is aligned with a 4×4 residual coefficient matrix; that is, one sub-macroblock corresponds to one residual coefficient matrix. Therefore, any residual coefficient matrix represents the residual coefficient matrix corresponding to one sub-macroblock. Statistical processing of any residual coefficient matrix yields the amount of information consumed by the residual coefficient matrix corresponding to a sub-macroblock.

[0092] When the number of residual coefficient matrices is less than the number of sub-macroblocks, meaning one residual coefficient matrix corresponds to at least one sub-macroblock, then one residual coefficient matrix represents the residuals contained in the encoded data of at least one sub-macroblock. For any sub-macroblock corresponding to a residual coefficient matrix, statistical processing can be performed on the residual coefficient matrix to obtain the amount of information consumed by that residual coefficient matrix. Then, combined with the number of sub-macroblocks corresponding to that residual coefficient matrix, the amount of information consumed by the residual coefficient matrix corresponding to that sub-macroblock can be determined.

[0093] For example, for macroblocks (MB) (m,n) Using an 8×8 discrete cosine transform, four residual coefficient matrices are obtained. At this point, any residual coefficient matrix is ​​8×8 in size, and can be denoted as... macroblock MB (m,n) The macroblock is divided into 16 sub-macroblocks of size 4×4. At this point, each sub-macroblock is aligned with a 4×4 residual coefficient matrix, meaning one residual coefficient matrix corresponds to four sub-macroblocks. Therefore, any residual coefficient matrix is ​​the residual coefficient matrix corresponding to four sub-macroblocks. Statistical processing of any residual coefficient matrix yields the information consumed by the residual coefficient matrices corresponding to the four sub-macroblocks. The information consumed by the residual coefficient matrix corresponding to any sub-macroblock can be denoted as...

[0094] Next, the information consumed by any one of the following is calculated: the amount of information consumed by the macroblock type corresponding to any sub-macroblock, the amount of information consumed by the macroblock prediction corresponding to the sub-macroblock, the amount of information consumed by the block mode of the encoding corresponding to the sub-macroblock, the amount of information consumed by the quantization parameter offset corresponding to the sub-macroblock, and the amount of information consumed by the residual coefficient matrix corresponding to the sub-macroblock. The sum of these information values ​​is then used to obtain the amount of information consumed by the sub-macroblock.

[0095] Optionally, the amount of information consumed by any sub-macroblock is determined according to formulas (1) and (2) as shown below, or the amount of information consumed by any sub-macroblock is determined according to formulas (1) and (3) as shown below.

[0096]

[0097]

[0098]

[0099] in, This indicates the amount of information consumed by any sub-macroblock, with + and = representing the accumulation sign, and bit{mb_type (m,n)} represents the amount of information consumed by the macroblock type corresponding to this sub-macroblock, bit{mb_pred (m,n)} represents the amount of information consumed in macroblock prediction corresponding to this sub-macroblock, bit{CBP (m,n)} represents the amount of information consumed by the block mode of the encoding corresponding to this sub-macroblock, bit{QP_off (m,n)} indicates the amount of information consumed by the quantization parameter offset corresponding to this sub-macroblock. and Both represent the amount of information consumed by the residual coefficient matrix corresponding to the sub-macroblock.

[0100] Using the above method, the amount of information consumed by each sub-macroblock of each macroblock in the target image can be determined; that is, the amount of information consumed by each sub-macroblock in the target image can be determined. Next, the largest amount of information consumed by each sub-macroblock in the target image is determined and denoted as the second amount of information. Based on the amount of information consumed by each sub-macroblock in the target image and the second amount of information, the contour image corresponding to the target image is determined.

[0101] In one possible implementation, the contour image corresponding to the target image is determined based on the amount of information consumed and the amount of information second of each sub-macroblock of multiple macroblocks, including: for any sub-macroblock, determining the ratio between the amount of information consumed and the amount of information second of any sub-macroblock; mapping the ratio corresponding to each sub-macroblock of multiple macroblocks to a grayscale value range to obtain the contour image corresponding to the target image.

[0102] In this embodiment, for any sub-macroblock in the target image, the information consumed by the sub-macroblock is normalized based on the amount of information consumed and the amount of information consumed by the sub-macroblock. The normalization process can be performed by calculating the ratio between the amount of information consumed by the macroblock and the amount of information consumed by the macroblock. This ratio is the ratio corresponding to the sub-macroblock, and the value range of this ratio is [0, 1]. The normalization process is shown in the following formula (4).

[0103]

[0104] in, This represents the ratio corresponding to any sub-macroblock in the target image. This indicates the amount of information consumed by the sub-macroblock. This indicates the second piece of information.

[0105] Next, for any sub-macroblock in the target image, the ratio corresponding to that macroblock is multiplied by 255, i.e., the result is calculated. The ratio corresponding to the sub-macroblock is mapped to a grayscale value range. The mapped ratio corresponding to any sub-macroblock is recorded as the mapped value of that sub-macroblock, and the value range of this mapped value is [0, 255]. After mapping the ratios corresponding to all sub-macroblocks in the target image to the grayscale value range, the contour image corresponding to the target image can be obtained, and this contour image is a grayscale image.

[0106] By implementing method A2, each macroblock in the target image is divided into multiple sub-macroblocks, and the amount of information consumed by each sub-macroblock is statistically analyzed, resulting in finer statistical granularity. This improves the sharpness and accuracy of the obtained contour image based on the amount of information consumed by all sub-macroblocks in the target image, thereby enhancing the accuracy of image recognition results in subsequent image recognition processing.

[0107] Step 203: Perform image recognition processing on the contour image to obtain the image recognition result.

[0108] In this embodiment, any pixel-domain image recognition algorithm or model can be used to perform image recognition on the contour image to obtain the image recognition result. The image recognition result indicates that the target image contains sensitive content or does not contain sensitive content. This embodiment does not limit the sensitive content; for example, sensitive content refers to vulgar content.

[0109] If the target image does not contain sensitive content, the electronic device can decode the encoded data of the target image to obtain the target image, and then render the target image to display it on the electronic device. Alternatively, for live video, the electronic device can send the encoded data of the target image to a second terminal device (such as a viewer's terminal device), which can then decode the encoded data of the target image to obtain the target image, render it, and display it on the second terminal device.

[0110] It should be noted that for live video, the electronic device sends the video stream to the second terminal device. This video stream includes the target image stream; that is, the video stream includes the encoded data of the target image.

[0111] In one possible implementation, after performing image recognition processing on the contour image to obtain the image recognition result, the method further includes: in response to the image recognition result indicating that the contour image contains sensitive content, filtering the encoded data of the target image, or decoding the encoded data of the target image to obtain the target image, and then performing masking processing on the sensitive content of the target image.

[0112] Optionally, if the target image contains sensitive content, the electronic device filters out the encoded data of the target image to prevent its display and dissemination. For live video, the electronic device can stop sending video streams to the second terminal device to prevent the display and dissemination of the target image.

[0113] Optionally, if the target image contains sensitive content, the electronic device decodes the encoded data of the target image to obtain the target image, and then de-masks the sensitive content of the target image to mask the sensitive content, resulting in a de-masked target image. The electronic device can then render the de-masked target image for display. For live video, the electronic device can re-encode the rendered de-masked target image to obtain encoded data of the de-masked target image, send the encoded data of the de-masked target image to a second terminal device, and the second terminal device decodes the encoded data of the de-masked target image to obtain the de-masked target image, and then renders the de-masked target image for display.

[0114] It should be noted that for live video, the video stream received by the electronic device from the first terminal device (which can be referred to as the first video stream) and the video stream sent by the electronic device to the second terminal device (which can be referred to as the second video stream) can be the same or different. After receiving the first video stream containing sensitive content, the electronic device can filter the encoded data of the target image containing sensitive content or de-mask the sensitive content in the target image, and send the second video stream without sensitive content to the second terminal device, thereby preventing the display and dissemination of sensitive content and improving the quality of the live video.

[0115] The above method obtains the contour image corresponding to the target image based on the encoded data of each macroblock included in the encoded data of the target image, and then performs image recognition on the contour image. Since it does not require decoding the encoded data of the target image to obtain the target image, it can save a lot of computing resources and improve the efficiency of image recognition.

[0116] The above describes the image recognition method from the perspective of its steps. The following section will provide a detailed explanation and illustration using a live streaming scenario. In the live streaming scenario, the electronic device executing steps 201 to 203 is a server.

[0117] Please see Figure 5 , Figure 5 This is a schematic diagram illustrating video processing according to an embodiment of this application. In this embodiment, the broadcaster's terminal device captures video (live video) in real time and performs preprocessing such as beautification and adding special effects. The video is related to the broadcaster, such as a video of the broadcaster playing a game or dancing. In this embodiment, the real-time captured video can be called a video pixel stream. The broadcaster's terminal device encodes the video pixel stream to obtain a video bit stream, which is encoded data in the form of a bit stream (also called a bitstream) corresponding to the video. The broadcaster's terminal device sends the video bit stream to the server. The video bit stream includes encoded data of multiple frames, and the encoded data of any frame can be the encoded data of the target image mentioned above.

[0118] The server receives the video bitstream. It extracts the frame image bitstream from the video bitstream. Then, using the frame image bitstream, it constructs the corresponding contour image. Next, it performs image recognition processing on the contour image to obtain the image recognition result. If the image recognition result indicates that the frame image does not contain sensitive content, the server sends the video bitstream to the viewer's terminal device. If the image recognition result indicates that the frame image contains sensitive content, the server stops receiving the video bitstream sent by the broadcaster's terminal device and stops sending the video bitstream to the viewer's terminal device. Optionally, the server can perform image recognition processing on each frame image in the video bitstream. If the image recognition result for a frame image indicates that the frame image contains sensitive content, the server stops receiving the video bitstream sent by the broadcaster's terminal device and stops sending the video bitstream to the viewer's terminal device.

[0119] After receiving the video bitstream from the server, the viewer's terminal device decodes the video bitstream to obtain the video pixel stream. Then, the video pixel stream is rendered so that the viewer's terminal device can display the video for easy viewing.

[0120] Next, please see Figure 6 , Figure 6 This is a schematic diagram of an image recognition process provided in an embodiment of this application. Figure 6 The image recognition process shown is Figure 5 The content states that "the server extracts the bitstream of the frame image from the video bitstream, uses the bitstream of the frame image to construct the contour image corresponding to the frame image, and performs image recognition processing on the contour image."

[0121] In this embodiment, the server can receive a video bitstream, extract the frame image bitstream from the video bitstream, and extract the macroblock bitstream from the frame image bitstream. The frame image bitstream corresponds to the encoded data of the target image mentioned above, and the macroblock bitstream corresponds to the encoded data of the macroblock mentioned above. The macroblock is divided into sub-macroblocks. Since the macroblock bitstream includes at least one encoded component, statistical processing is performed on each encoded component to obtain the number of bits consumed by each sub-macroblock. The number of bits consumed by each sub-macroblock corresponds to the amount of information consumed by the sub-macroblock mentioned above. The number of bits consumed by each sub-macroblock is normalized to obtain the ratio corresponding to each sub-macroblock. The ratio corresponding to each sub-macroblock is mapped to a grayscale value range to obtain the contour image corresponding to the frame image. Then, image recognition processing is performed on the contour image.

[0122] It's important to note that while the broadcaster's terminal device uses H.264 / AVC to encode the live video into a video bitstream, the viewer's terminal device also needs to use H.264 / AVC to decode the video bitstream to obtain the live video. H.264 / AVC is a video codec protocol with wide and unified applications. Therefore, for the broadcaster's terminal device, it can be seamlessly applied to all live streaming architectures without additional adaptation work. Simultaneously, for the cloud server, a contour image corresponding to each frame is constructed based on the video bitstream. This contour image is a pixel-domain image, and all pixel-domain image recognition algorithms and models can be used to perform image recognition processing on the contour image.

[0123] When decoding a video bitstream using H.264 / AVC, the decoding process can be divided into three operations: prediction and compensation, inverse transform and inverse quantization, and entropy decoding. In this embodiment, for frame images of different frame types, the time required for each operation is statistically analyzed, and the time ratio of each operation is calculated based on the time required for each operation, resulting in Table 1 below.

[0124] Table 1

[0125]

[0126] As shown in Table 1, when using H.264 / AVC to decode video bitstreams, the time allocated to prediction and compensation is significantly greater than that allocated to inverse transform and inverse quantization, and the time allocated to prediction and compensation is also significantly greater than that allocated to entropy decoding. Entropy decoding corresponds to the server constructing a contour image corresponding to a frame image using the bitstream of the frame image in this embodiment of the application.

[0127] In other words, in this embodiment, the server only needs to perform entropy decoding to obtain the contour image, saving the time spent on prediction and compensation, inverse transformation and inverse quantization. For I-frame images, the time overhead can be reduced by 72.07%; for P-frame images, by 80.29%; and for B-frame images, by 85.81%. Therefore, the image recognition method of this embodiment can save significant computing resources and improve image recognition efficiency.

[0128] Figure 7 The diagram shown is a structural schematic of an image recognition device provided in an embodiment of this application. Figure 7 As shown, the device includes:

[0129] The acquisition module 701 is used to acquire the encoded data of the target image. The encoded data of the target image includes the encoded data of multiple macroblocks in the target image. The amount of information consumed when any macroblock in the multiple macroblocks is encoded is proportional to the texture complexity of any macroblock. The texture complexity of any macroblock is related to the pixel value of each pixel in any macroblock.

[0130] The acquisition module 701 is also used to acquire the contour image corresponding to the target image based on the encoded data of each macroblock. The contour image is used to reflect the contour of the object in the target image.

[0131] The image recognition module 702 is used to perform image recognition processing on the contour image to obtain the image recognition result.

[0132] In one possible implementation, the acquisition module 701 is used to perform statistical processing on the encoded data of each macroblock to obtain the amount of information consumed by each macroblock; determine the largest first amount of information from the information consumed by each macroblock; and determine the contour image corresponding to the target image based on the information consumed by each macroblock and the first amount of information.

[0133] In one possible implementation, the acquisition module 701 is used to perform statistical processing on at least one coded component contained in the coded data of any macroblock for any macroblock, to obtain the amount of information consumed by each coded component corresponding to the macroblock, wherein any coded component is any one of macroblock type, macroblock prediction, coded block mode, quantization parameter offset and residual; and the sum of the amount of information consumed by each coded component corresponding to the macroblock is determined as the amount of information consumed by the macroblock.

[0134] In one possible implementation, the acquisition module 701 is used to determine, for any given macroblock, the ratio between the amount of information consumed by the macroblock and the first amount of information; and to map the ratios corresponding to each macroblock to a grayscale range to obtain the contour image corresponding to the target image.

[0135] In one possible implementation, the acquisition module 701 is used to divide any macroblock into multiple sub-macroblocks; perform statistical processing on the encoded data of any macroblock to obtain the amount of information consumed by each sub-macroblock of any macroblock; determine the largest second amount of information from the amount of information consumed by each sub-macroblock of multiple macroblocks; and determine the contour image corresponding to the target image based on the amount of information consumed by each sub-macroblock of multiple macroblocks and the second amount of information.

[0136] In one possible implementation, the acquisition module 701 is used to perform statistical processing on at least one coded component contained in the coded data of any macroblock for any macroblock, to obtain the amount of information consumed by each coded component corresponding to each sub-macroblock of any macroblock, wherein any coded component is any one of macroblock type, macroblock prediction, coded block mode, quantization parameter offset, and residual; for any sub-macroblock, the sum of the amount of information consumed by each coded component corresponding to any sub-macroblock is determined as the amount of information consumed by any sub-macroblock.

[0137] In one possible implementation, the acquisition module 701 is used to, for any coded component contained in the coded data of any macroblock, in response to any coded component being any of the macroblock type, macroblock prediction, coded block mode, or quantization parameter offset, calculate the amount of information consumed by any coded component, and obtain the amount of information consumed by any coded component corresponding to each sub-macroblock of any macroblock.

[0138] In one possible implementation, the acquisition module 701 is configured to, for any coded component contained in the coded data of any macroblock, in response to any coded component being a residual and the number of residual coefficient matrices included in the residual being not less than the number of sub-macroblocks of any macroblock, determine the residual coefficient matrix corresponding to any sub-macroblock of any macroblock from the residual coefficient matrices included in the residual, and calculate the amount of information consumed by the residual coefficient matrix corresponding to any sub-macroblock; in response to any coded component being a residual and the number of residual coefficient matrices included in the residual being less than the number of sub-macroblocks of any macroblock, determine the amount of information consumed by each sub-macroblock corresponding to any residual coefficient matrix based on the number of sub-macroblocks corresponding to any residual coefficient matrix and the amount of information consumed by any residual coefficient matrix.

[0139] In one possible implementation, the acquisition module 701 is used to determine, for any given sub-macroblock, the ratio between the amount of information consumed by the sub-macroblock and the amount of information consumed by the sub-macroblock; and to map the ratios corresponding to each sub-macroblock of the multiple macroblocks to a grayscale range to obtain the contour image corresponding to the target image.

[0140] In one possible implementation, the device further includes:

[0141] The processing module is used to filter the encoded data of the target image or decode the encoded data of the target image to obtain the target image in response to the image recognition result indicating that the contour image contains sensitive content, and to perform masking processing on the sensitive content of the target image.

[0142] The aforementioned device obtains the contour image corresponding to the target image based on the encoded data of each macroblock included in the encoded data of the target image, and then performs image recognition on the contour image. Since it does not require decoding the encoded data of the target image to obtain the target image, it can save a lot of computing resources and improve image recognition efficiency.

[0143] It should be understood that the above Figure 7 The provided device, in implementing its functions, is only illustrated by the division of the above-described functional modules. In practical applications, the functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. Furthermore, the device and method embodiments provided in the above embodiments belong to the same concept, and their specific implementation processes are detailed in the method embodiments, and will not be repeated here.

[0144] Figure 8 This diagram illustrates a structural block diagram of a terminal device 800 provided in an exemplary embodiment of this application. The terminal device 800 includes a processor 801 and a memory 802.

[0145] Processor 801 may include one or more processing cores, such as a quad-core processor, an octa-core processor, etc. Processor 801 may be implemented using at least one hardware form selected from DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), and PLA (Programmable Logic Array). Processor 801 may also include a main processor and a coprocessor. The main processor, also known as a CPU (Central Processing Unit), is used to process data in the wake-up state; the coprocessor is a low-power processor used to process data in the standby state. In some embodiments, processor 801 may integrate a GPU (Graphics Processing Unit), which is responsible for rendering and drawing the content required to be displayed on the screen. In some embodiments, processor 801 may also include an AI (Artificial Intelligence) processor, which is used to handle computational operations related to machine learning.

[0146] The memory 802 may include one or more computer-readable storage media, which may be non-transitory. The memory 802 may also include high-speed random access memory and non-volatile memory, such as one or more disk storage devices or flash memory devices. In some embodiments, the non-transitory computer-readable storage media in the memory 802 are used to store at least one computer program, which is executed by the processor 801 to implement the image recognition method provided in the method embodiments of this application.

[0147] In some embodiments, the terminal device 800 may also optionally include a peripheral device interface 803 and at least one peripheral device. The processor 801, memory 802, and peripheral device interface 803 can be connected via a bus or signal line. Each peripheral device can be connected to the peripheral device interface 803 via a bus, signal line, or circuit board. Specifically, the peripheral device includes at least one of the following: a radio frequency circuit 804, a display screen 805, a camera assembly 806, an audio circuit 807, and a power supply 808.

[0148] Peripheral device interface 803 can be used to connect at least one I / O (Input / Output) related peripheral device to processor 801 and memory 802. In some embodiments, processor 801, memory 802 and peripheral device interface 803 are integrated on the same chip or circuit board; in some other embodiments, any one or two of processor 801, memory 802 and peripheral device interface 803 can be implemented on separate chips or circuit boards, which is not limited in this embodiment.

[0149] The radio frequency (RF) circuit 804 is used to receive and transmit RF (Radio Frequency) signals, also known as electromagnetic signals. The RF circuit 804 communicates with communication networks and other communication devices via electromagnetic signals. The RF circuit 804 converts electrical signals into electromagnetic signals for transmission, or converts received electromagnetic signals back into electrical signals. Optionally, the RF circuit 804 includes: an antenna system, an RF transceiver, one or more amplifiers, a tuner, an oscillator, a digital signal processor, a codec chipset, a user identity module card, etc. The RF circuit 804 can communicate with other terminals through at least one wireless communication protocol. This wireless communication protocol includes, but is not limited to: the World Wide Web, metropolitan area networks, intranets, various generations of mobile communication networks (2G, 3G, 4G, and 5G), wireless local area networks, and / or WiFi (Wireless Fidelity) networks. In some embodiments, the RF circuit 804 may also include circuitry related to NFC (Near Field Communication), which is not limited in this application.

[0150] Display screen 805 is used to display a UI (User Interface). This UI may include graphics, text, icons, videos, and any combination thereof. When display screen 805 is a touch display screen, it also has the ability to collect touch signals on or above its surface. These touch signals can be input as control signals to processor 801 for processing. In this case, display screen 805 can also be used to provide virtual buttons and / or a virtual keyboard, also known as soft buttons and / or a soft keyboard. In some embodiments, there may be one display screen 805, disposed on the front panel of terminal device 800; in other embodiments, there may be at least two display screens, disposed on different surfaces of terminal device 800 or in a folded design; in still other embodiments, display screen 805 may be a flexible display screen, disposed on a curved or folded surface of terminal device 800. Furthermore, display screen 805 may be configured as a non-rectangular irregular shape, i.e., a non-rectangular screen. Display screen 805 may be made of materials such as LCD (Liquid Crystal Display) or OLED (Organic Light-Emitting Diode).

[0151] The camera assembly 806 is used to acquire images or videos. Optionally, the camera assembly 806 includes a front-facing camera and a rear-facing camera. Typically, the front-facing camera is located on the front panel of the terminal, and the rear-facing camera is located on the back of the terminal. In some embodiments, there are at least two rear-facing cameras, which are any one of a main camera, a depth-sensing camera, a wide-angle camera, and a telephoto camera, to achieve background blurring by fusion of the main camera and the depth-sensing camera, panoramic shooting by fusion of the main camera and the wide-angle camera, VR (Virtual Reality) shooting, or other fusion shooting functions. In some embodiments, the camera assembly 806 may also include a flash. The flash can be a single-color temperature flash or a dual-color temperature flash. A dual-color temperature flash refers to a combination of a warm-light flash and a cool-light flash, which can be used for light compensation at different color temperatures.

[0152] The audio circuit 807 may include a microphone and a speaker. The microphone is used to collect sound waves from the user and the environment, converting the sound waves into electrical signals that are input to the processor 801 for processing, or input to the radio frequency circuit 804 for voice communication. For stereo sound acquisition or noise reduction purposes, multiple microphones may be used, each located at a different part of the terminal device 800. The microphone may also be an array microphone or an omnidirectional microphone. The speaker is used to convert the electrical signals from the processor 801 or the radio frequency circuit 804 into sound waves. The speaker may be a conventional diaphragm speaker or a piezoelectric ceramic speaker. When the speaker is a piezoelectric ceramic speaker, it can convert electrical signals not only into audible sound waves but also into inaudible sound waves for purposes such as distance measurement. In some embodiments, the audio circuit 807 may also include a headphone jack.

[0153] Power supply 808 is used to supply power to the various components in terminal device 800. Power supply 808 can be AC ​​power, DC power, a disposable battery, or a rechargeable battery. When power supply 808 includes a rechargeable battery, the rechargeable battery can be a wired rechargeable battery or a wireless rechargeable battery. A wired rechargeable battery is a battery that is charged via a wired line, while a wireless rechargeable battery is a battery that is charged via a wireless coil. The rechargeable battery can also be used to support fast charging technology.

[0154] In some embodiments, the terminal device 800 further includes one or more sensors 809. The one or more sensors 809 include, but are not limited to, an accelerometer 811, a gyroscope 812, a pressure sensor 813, an optical sensor 814, and a proximity sensor 815.

[0155] Accelerometer 811 can detect the magnitude of acceleration along the three coordinate axes of a coordinate system established by terminal device 800. For example, accelerometer 811 can be used to detect the components of gravitational acceleration along the three coordinate axes. Processor 801 can control display screen 805 to display the user interface in either a landscape or portrait view based on the gravitational acceleration signal acquired by accelerometer 811. Accelerometer 811 can also be used for games or for acquiring user motion data.

[0156] The gyroscope sensor 812 can detect the orientation and rotation angle of the terminal device 800. The gyroscope sensor 812, in conjunction with the accelerometer sensor 811, can collect 3D motion data from the user on the terminal device 800. Based on the data collected by the gyroscope sensor 812, the processor 801 can perform the following functions: motion sensing (e.g., changing the UI based on the user's tilt), image stabilization during shooting, game control, and inertial navigation.

[0157] The pressure sensor 813 can be disposed on the side bezel of the terminal device 800 and / or on the lower layer of the display screen 805. When the pressure sensor 813 is disposed on the side bezel of the terminal device 800, it can detect the user's grip signal on the terminal device 800, and the processor 801 can perform left / right hand recognition or quick operation based on the grip signal collected by the pressure sensor 813. When the pressure sensor 813 is disposed on the lower layer of the display screen 805, the processor 801 can control the operable controls on the UI interface based on the user's pressure operation on the display screen 805. The operable controls include at least one of button controls, scroll bar controls, icon controls, and menu controls.

[0158] An optical sensor 814 is used to collect ambient light intensity. In one embodiment, the processor 801 can control the display brightness of the display screen 805 based on the ambient light intensity collected by the optical sensor 814. Specifically, when the ambient light intensity is high, the display brightness of the display screen 805 is increased; when the ambient light intensity is low, the display brightness of the display screen 805 is decreased. In another embodiment, the processor 801 can also dynamically adjust the shooting parameters of the camera assembly 806 based on the ambient light intensity collected by the optical sensor 814.

[0159] The proximity sensor 815, also known as a distance sensor, is typically located on the front panel of the terminal device 800. The proximity sensor 815 is used to detect the distance between the user and the front of the terminal device 800. In one embodiment, when the proximity sensor 815 detects that the distance between the user and the front of the terminal device 800 is gradually decreasing, the processor 801 controls the display screen 805 to switch from a screen-on state to a screen-off state; when the proximity sensor 815 detects that the distance between the user and the front of the terminal device 800 is gradually increasing, the processor 801 controls the display screen 805 to switch from a screen-off state to a screen-on state.

[0160] Those skilled in the art will understand that Figure 8 The structure shown does not constitute a limitation on the terminal device 800, and may include more or fewer components than shown, or combine certain components, or use different component arrangements.

[0161] Figure 9This is a schematic diagram of the server structure provided in the embodiments of this application. The server 900 can vary considerably due to different configurations or performance. It may include one or more processors 901 and one or more memories 902. The one or more memories 902 store at least one computer program, which is loaded and executed by the one or more processors 901 to implement the image recognition methods provided in the above-described method embodiments. For example, the processor 901 is a CPU. Of course, the server 900 may also have wired or wireless network interfaces, a keyboard, and input / output interfaces for input and output. The server 900 may also include other components for implementing device functions, which will not be elaborated here.

[0162] In an exemplary embodiment, a computer-readable storage medium is also provided, which stores at least one computer program that is loaded and executed by a processor to enable an electronic device to implement any of the above-described image recognition methods.

[0163] Optionally, the aforementioned computer-readable storage medium may be a read-only memory (ROM), a random access memory (RAM), a compact disc read-only memory (CD-ROM), magnetic tape, floppy disk, and optical data storage device, etc.

[0164] In an exemplary embodiment, a computer program or computer program product is also provided, which stores at least one computer program that is loaded and executed by a processor to enable the computer to implement any of the above-described image recognition methods.

[0165] It should be understood that "multiple" as used in this article refers to two or more. "And / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, or B alone. The character " / " generally indicates that the preceding and following related objects have an "or" relationship.

[0166] The sequence numbers of the embodiments in this application are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.

[0167] The above description is merely an exemplary embodiment of this application and is not intended to limit this application. Any modifications, equivalent substitutions, improvements, etc., made within the principles of this application should be included within the protection scope of this application.

Claims

1. An image recognition method, characterized in that, The method includes: The encoded data of the target image is obtained. The encoded data of the target image includes the encoded data of multiple macroblocks in the target image. The amount of information consumed when encoding any macroblock in the multiple macroblocks is proportional to the texture complexity of the macroblock. The texture complexity of the macroblock is related to the pixel value of each pixel in the macroblock. For any given macroblock, divide the macroblock into multiple sub-macroblocks; Statistical processing is performed on the encoded data of any one macroblock to obtain the amount of information consumed by each sub-macroblock of the given macroblock. From the amount of information consumed by each sub-macroblock of the plurality of macroblocks, determine the largest second amount of information; For any sub-macroblock, determine the ratio between the amount of information consumed by the sub-macroblock and the second amount of information; map the ratio corresponding to the sub-macroblock to the grayscale value range to obtain the mapping value corresponding to the sub-macroblock; Based on the mapping values ​​corresponding to each sub-macroblock of the plurality of macroblocks, a grayscale image corresponding to the target image is determined, and the grayscale image is used to reflect the objects in the target image; The grayscale image is subjected to image recognition processing to obtain the image recognition result.

2. The method according to claim 1, characterized in that, The method further includes: Statistical processing is performed on the encoded data of each macroblock to obtain the amount of information consumed by each macroblock; Determine the largest first information quantity from the information consumed by each macroblock; Based on the amount of information consumed by each macroblock and the first amount of information, the grayscale image corresponding to the target image is determined.

3. The method according to claim 2, characterized in that, The statistical processing of the encoded data of each macroblock to obtain the amount of information consumed by each macroblock includes: For any macroblock, statistical processing is performed on at least one coded component contained in the coded data of the macroblock to obtain the amount of information consumed by each coded component corresponding to the macroblock. Any coded component is any one of macroblock type, macroblock prediction, coded block mode, quantization parameter offset and residual. The sum of the information consumed by each coding component corresponding to any given macroblock is determined as the information consumed by any given macroblock.

4. The method according to claim 2, characterized in that, Determining the grayscale image corresponding to the target image based on the amount of information consumed by each macroblock and the first amount of information includes: For any macroblock, determine the ratio between the amount of information consumed by the macroblock and the first amount of information; The ratios corresponding to each macroblock are mapped to grayscale value ranges to obtain the grayscale image corresponding to the target image.

5. The method according to claim 1, characterized in that, The statistical processing of the encoded data of any macroblock to obtain the amount of information consumed by each sub-macroblock of any macroblock includes: For any macroblock, statistical processing is performed on at least one coded component contained in the coded data of the macroblock to obtain the amount of information consumed by each coded component corresponding to each sub-macroblock of the macroblock. Any coded component is any one of macroblock type, macroblock prediction, coded block mode, quantization parameter offset and residual. For any given sub-macroblock, the sum of the information consumed by each coding component corresponding to that sub-macroblock is determined as the information consumed by that sub-macroblock.

6. The method according to claim 5, characterized in that, The step of statistically processing at least one coded component contained in the coded data of any macroblock to obtain the amount of information consumed by each coded component corresponding to each sub-macroblock of the macroblock includes: For any coded component contained in the coded data of any macroblock, in response to any coded component being any of the macroblock type, the macroblock prediction, the coded block mode, or the quantization parameter offset, the amount of information consumed by the coded component is calculated to obtain the amount of information consumed by the coded component corresponding to each sub-macroblock of the macroblock.

7. The method according to claim 5, characterized in that, The step of statistically processing at least one coded component contained in the coded data of any macroblock to obtain the amount of information consumed by each coded component corresponding to each sub-macroblock of the macroblock includes: For any coded component contained in the coded data of any macroblock, in response to the fact that any coded component is the residual, and the number of residual coefficient matrices included in the residual is not less than the number of sub-macroblocks of any macroblock, the residual coefficient matrix corresponding to any sub-macroblock of any macroblock is determined from the residual coefficient matrix included in the residual, and the amount of information consumed by the residual coefficient matrix corresponding to any sub-macroblock is calculated. In response to the fact that any one of the coding components is the residual, and the number of residual coefficient matrices included in the residual is less than the number of sub-macroblocks of any one macroblock, the amount of information consumed by each sub-macroblock corresponding to any one residual coefficient matrix is ​​determined based on the number of sub-macroblocks corresponding to any one residual coefficient matrix and the amount of information consumed by any one residual coefficient matrix.

8. The method according to any one of claims 1 to 7, characterized in that, After performing image recognition processing on the grayscale image to obtain the image recognition result, the method further includes: In response to the image recognition result indicating that the grayscale image contains sensitive content, the encoded data of the target image is filtered, or the encoded data of the target image is decoded to obtain the target image, and the sensitive content of the target image is masked.

9. An image recognition device, characterized in that, The device includes: The acquisition module is used to acquire the encoded data of the target image. The encoded data of the target image includes the encoded data of multiple macroblocks in the target image. The amount of information consumed when any macroblock in the multiple macroblocks is encoded is proportional to the texture complexity of the macroblock. The texture complexity of the macroblock is related to the pixel value of each pixel in the macroblock. The acquisition module is further configured to: divide any macroblock into multiple sub-macroblocks; perform statistical processing on the encoded data of any macroblock to obtain the amount of information consumed by each sub-macroblock of the macroblock; determine the largest second amount of information from the information consumed by each sub-macroblock of the multiple macroblocks; for any sub-macroblock, determine the ratio between the amount of information consumed by the sub-macroblock and the second amount of information; map the ratio corresponding to the sub-macroblock to a grayscale value range to obtain the mapping value corresponding to the sub-macroblock; and based on the mapping values ​​corresponding to each sub-macroblock of the multiple macroblocks, determine the grayscale image corresponding to the target image, wherein the grayscale image is used to reflect the objects in the target image. The image recognition module is used to perform image recognition processing on the grayscale image to obtain the image recognition result.

10. The apparatus according to claim 9, characterized in that, The acquisition module is further configured to perform statistical processing on the encoded data of each macroblock to obtain the amount of information consumed by each macroblock; and to determine the largest first amount of information from the amount of information consumed by each macroblock. Based on the amount of information consumed by each macroblock and the first amount of information, the grayscale image corresponding to the target image is determined.

11. The apparatus according to claim 10, characterized in that, The acquisition module is used to perform statistical processing on at least one coded component contained in the coded data of any macroblock for any given macroblock, to obtain the amount of information consumed by each coded component corresponding to the given macroblock, wherein any coded component is any one of macroblock type, macroblock prediction, coded block mode, quantization parameter offset, and residual; and to determine the sum of the amount of information consumed by each coded component corresponding to the given macroblock as the amount of information consumed by the given macroblock.

12. The apparatus according to claim 10, characterized in that, The acquisition module is used to determine, for any given macroblock, the ratio between the amount of information consumed by the macroblock and the first amount of information; and to map the ratios corresponding to each macroblock to a grayscale value range to obtain a grayscale image corresponding to the target image.

13. The apparatus according to claim 9, characterized in that, The acquisition module is used to perform statistical processing on at least one coded component contained in the coded data of any macroblock for any given macroblock, to obtain the amount of information consumed by each coded component corresponding to each sub-macroblock of the given macroblock, wherein any coded component is any one of macroblock type, macroblock prediction, coded block mode, quantization parameter offset, and residual; for any sub-macroblock, the sum of the amount of information consumed by each coded component corresponding to the given sub-macroblock is determined as the amount of information consumed by the given sub-macroblock.

14. The apparatus according to claim 13, characterized in that, The acquisition module is configured to, for any coded component contained in the coded data of any macroblock, in response to any of the following: macroblock type, macroblock prediction, coded block mode, quantization parameter offset, calculate the amount of information consumed by the coded component, and obtain the amount of information consumed by the coded component corresponding to each sub-macroblock of the macroblock.

15. The apparatus according to claim 13, characterized in that, The acquisition module is configured to, for any coded component contained in the coded data of any macroblock, in response to the fact that any coded component is the residual, and the number of residual coefficient matrices included in the residual is not less than the number of sub-macroblocks of any macroblock, determine the residual coefficient matrix corresponding to any sub-macroblock of any macroblock from the residual coefficient matrix included in the residual, and count the amount of information consumed by the residual coefficient matrix corresponding to any sub-macroblock; In response to the fact that any one of the coding components is the residual, and the number of residual coefficient matrices included in the residual is less than the number of sub-macroblocks of any one macroblock, the amount of information consumed by each sub-macroblock corresponding to any one residual coefficient matrix is ​​determined based on the number of sub-macroblocks corresponding to any one residual coefficient matrix and the amount of information consumed by any one residual coefficient matrix.

16. The apparatus according to any one of claims 9 to 15, characterized in that, The device further includes: The processing module is configured to, in response to the image recognition result indicating that the grayscale image contains sensitive content, filter the encoded data of the target image, or decode the encoded data of the target image to obtain the target image, and then perform masking processing on the sensitive content of the target image.

17. An electronic device, characterized in that, The electronic device includes a processor and a memory, the memory storing at least one computer program, which is loaded and executed by the processor to enable the electronic device to implement the image recognition method as described in any one of claims 1 to 8.

18. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores at least one computer program, which is loaded and executed by a processor to enable the computer to implement the image recognition method as described in any one of claims 1 to 8.

19. A computer program product, characterized in that, The computer program product stores at least one computer program, which is loaded and executed by a processor to enable the computer to implement the image recognition method as described in any one of claims 1 to 8.