Image processing method and apparatus

By decoding and desensitizing the luminance components of RAW images, a new encoded bitstream is generated, solving the privacy protection problem of RAW images in end-to-end autonomous driving and deep learning models, and achieving both privacy protection and hardware optimization.

WO2026060596A1PCT designated stage Publication Date: 2026-03-26YINWANG INTELLIGENT TECHNOLOGIES CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-09-19
Publication Date
2026-03-26

AI Technical Summary

Technical Problem

Existing image processing technologies lack the ability to detect and desensitize RAW images, especially in end-to-end autonomous driving and deep learning models, which fails to effectively protect privacy information.

Method used

By decoding, desensitizing, detecting, and processing the luminance component of a RAW image, a new luminance component image is generated. This image is then combined with color and difference components to generate an encoded bitstream, thus achieving desensitization of the RAW image while reducing hardware and transmission costs without altering the existing architecture.

Benefits of technology

It achieves privacy protection for RAW images, reduces the chip area and transmission cost of hardware decoders, and minimizes the impact of image quality degradation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024119760_26032026_PF_FP_ABST
    Figure CN2024119760_26032026_PF_FP_ABST
Patent Text Reader

Abstract

The present application provides an image processing method and apparatus, used for implementing desensitization detection and desensitization processing of RAW images. The method comprises: acquiring a first encoded bitstream, wherein the first encoded bitstream comprises encoded bitstreams of a plurality of components of a first RAW image, and the plurality of components include a first luminance component, a first color component, and a first difference component; decoding the encoded bitstream of the first luminance component to obtain an image of the first luminance component of the first RAW image; performing desensitization detection and desensitization processing on the basis of the image of the first luminance component to obtain an image of a second luminance component; encoding the image of the second luminance component to obtain an encoded bitstream of the second luminance component; obtaining a second encoded bitstream on the basis of the encoded bitstream of the second luminance component, the encoded bitstream of the first color component, and the encoded bitstream of the first difference component; and storing the second encoded bitstream, or sending the second encoded bitstream to a cloud device.
Need to check novelty before this filing date? Find Prior Art

Description

Image processing method and device TECHNICAL FIELD

[0001] The present application relates to the technical field of image processing, and in particular to an image processing method and device. BACKGROUND

[0002] With the rapid development of intelligent driving technology and automatic driving technology, in some scenarios, intelligent driving vehicles need to collect external environment information of the road where the vehicle is located and various data of the vehicle itself, for analyzing the surrounding road conditions of the intelligent driving vehicle and the running state of the intelligent driving vehicle, or for training of artificial intelligence models and map customization, etc. Among them, the intelligent driving vehicle usually collects images or videos through a camera device, and obtains the required data through analysis of the images or videos. If the images or video frames contain sensitive information such as faces, license plates, and other private information, the images or video frames need to be desensitized.

[0003] Current desensitization methods are all for the input of desensitization models of images encoded in YUV format or RGB format. With the development of end-to-end automatic driving and deep image signal processor (DeepISP) technology based on deep learning models, the collection and model training based on original image files (such as RAW images) are becoming increasingly important, but there is a lack of desensitization detection and desensitization processing technology for RAW images.

[0004] SUMMARY

[0005] The present application provides an image processing method and device for realizing desensitization detection and desensitization processing of original RAW images.

[0006] In a first aspect, the present application provides an image processing method, which can be implemented by a terminal device. The method can include: obtaining a first coded stream, the first coded stream including coded streams of a plurality of components of a first original RAW image, the plurality of components including a first luminance component, a first color component, and a first difference component; decoding the coded stream of the first luminance component to obtain an image of the first luminance component of the first RAW image; performing desensitization detection and desensitization processing based on the image of the first luminance component to obtain an image of a second luminance component; encoding the image of the second luminance component to obtain a coded stream of the second luminance component; obtaining a second coded stream based on the coded stream of the second luminance component, the coded stream of the first color component, and the coded stream of the first difference component; and storing the second coded stream or sending the second coded stream to a cloud device.

[0007] By the above method, the desensitization detection and desensitization processing of the RAW image can be realized without changing the existing desensitization architecture and algorithm, the user privacy is ensured, the chip area and transmission cost of the vehicle-side decoder are reduced on the hardware, and the influence of the secondary encoding on the image quality caused by the image desensitization is reduced.

[0008] In a possible implementation, the terminal device comprises a desensitization model, and the desensitization detection and desensitization processing based on the image of the first luminance component to obtain the image of the second luminance component comprises: inputting the image of the first luminance component into three input channels of the desensitization model respectively, and obtaining the image of the second luminance component after the desensitization detection and desensitization processing in the desensitization model, wherein the desensitization model is trained for images in YUV format or RGB format, Y is a luminance component, U and V are color components, and R, G and B represent red, green and blue respectively.

[0009] In a possible implementation, the bit width depth of the image supported by the desensitization model is a target bit, and the bit width depth of the first RAW image is greater than or equal to the target bit.

[0010] In a possible implementation, if the bit width depth of the first RAW image is greater than the target bit, the inputting the image of the first luminance component into the three input channels of the desensitization model respectively comprises: intercepting an image of the target bit from the image of the first luminance component; and inputting the image of the target bit into the three input channels of the desensitization model respectively.

[0011] In a possible implementation, the target bit is 8 bits.

[0012] In a possible implementation, the bit width depth of the first RAW image is any one of the following: 8 bits, 10 bits, 12 bits or 16 bits.

[0013] In a possible implementation, the method further comprises: obtaining a second RAW image; pre-processing the second RAW image to obtain the first RAW image, the first RAW image comprising an image in YCoCgDg format, Y being a luminance component, Co and Cg being color components, and Dg being a difference component; encoding the images of multiple components of the first RAW image respectively to obtain the first encoded code stream; and storing the first encoded code stream.

[0014] In a possible implementation, the obtaining the first encoded code stream comprises: obtaining the first encoded code stream from a storage medium.

[0015] In a possible implementation, the pre-processing of the second RAW image to obtain the first RAW image comprises: performing non-linear transformation on the second RAW image to obtain a transformed second RAW image; and performing color gamut transformation on the transformed second RAW image to obtain the first RAW image.

[0016] In a second aspect, the present application provides an image processing apparatus, comprising: an acquisition module configured to acquire a first coded stream, wherein the first coded stream comprises coded streams of a plurality of components of a first RAW image, and the plurality of components comprises a first luminance component, a first color component and a first difference component; a decoding module configured to decode the coded stream of the first luminance component to obtain an image of the first luminance component of the first RAW image; a desensitization module configured to perform desensitization detection and desensitization processing based on the image of the first luminance component to obtain an image of a second luminance component; an encoding module configured to encode the image of the second luminance component to obtain a coded stream of the second luminance component; and a storage module configured to store the second coded stream, or a sending module configured to send the second coded stream to a cloud device.

[0017] In a possible implementation, the desensitization module comprises a desensitization model, and the desensitization module is specifically configured to: input the image of the first luminance component into three input channels of the desensitization model respectively, and obtain the image of the second luminance component after desensitization detection and desensitization processing in the desensitization model, wherein the desensitization model is trained based on images in YUV format or RGB format, Y represents a luminance component, U and V represent color components, and R, G and B represent red, green and blue respectively.

[0018] In a possible implementation, a bit width depth of an image supported by the desensitization model is a target bit, and a bit width depth of the first RAW image is greater than or equal to the target bit.

[0019] In a possible implementation, if the bit width depth of the first RAW image is greater than the target bit, the desensitization module is specifically configured to: cut the image of the target bit from the image of the first luminance component; and input the image of the target bit into the three input channels of the desensitization model respectively.

[0020] In a possible implementation, the target bit is 8 bits.

[0021] In a possible implementation, a bit width depth of the first RAW image is any one of 8 bits, 10 bits, 12 bits, or 16 bits.

[0022] In a possible implementation, the obtaining unit is further configured to obtain a second RAW image; the apparatus further includes a preprocessing module configured to preprocess the second RAW image to obtain the first RAW image, the first RAW image including an image in a YCoCgDg format, Y being a luminance component, Co and Cg being color components, and Dg being a difference component; the encoding module is further configured to encode images of multiple components of the first RAW image respectively to obtain the first encoded code stream; and the storage module is configured to store the first encoded code stream.

[0023] In a possible implementation, the obtaining unit is specifically configured to obtain the first encoded code stream from a storage medium.

[0024] In a possible implementation, the preprocessing module is specifically configured to perform nonlinear transformation on the second RAW image to obtain a second RAW image after transformation, and perform color gamut transformation on the second RAW image after transformation to obtain the first RAW image.

[0025] In a third aspect, the present application provides a computing device, including a processor coupled with a memory: the processor is configured to execute computer programs or instructions stored in the memory, so that the apparatus executes the method in the first aspect and any possible implementation of the first aspect.

[0026] In a fourth aspect, the present application provides an image processing system, including a transceiver, a memory and a processor; the transceiver is configured to receive and send data; the memory is configured to store computer program instructions and data; and the processor is configured to execute the computer program instructions and data in the memory, so that the image processing system executes the method in the first aspect and any possible implementation of the first aspect.

[0027] In a fifth aspect, the present application provides a vehicle, including units or modules for implementing the method in the first aspect and any possible implementation of the first aspect.

[0028] In a sixth aspect, the present application provides a computer readable storage medium, the computer readable storage medium storing computer programs or instructions, when the computer programs or instructions are executed by a computer, the computer executes the method in the first aspect and any possible implementation of the first aspect.

[0029] In a seventh aspect, the present application provides a computer program product, which comprises computer programs or instructions, and when the computer programs or instructions are run on a computer, the computer is caused to execute the method according to the first aspect and any possible implementation manner of the first aspect.

[0030] In an eighth aspect, the present application provides a chip, which comprises a processor coupled with a memory, and the processor is configured to execute computer programs or instructions stored in the memory, and when the computer programs or instructions are executed, the method according to the first aspect and any possible implementation manner of the first aspect is implemented.

[0031] On the basis of the implementation of the above aspects, the present application can be further combined to provide more implementations.

[0032] The beneficial effects of the above-mentioned second aspect to eighth aspect can refer to the description in the first aspect, and will not be repeated here. BRIEF DESCRIPTION OF DRAWINGS

[0033] FIG. 1 is a structural schematic diagram of an end-cloud system according to an embodiment of the present application;

[0034] FIG. 2 is a hardware structural schematic diagram of a terminal device 110 according to an embodiment of the present application;

[0035] FIG. 3 is a flow schematic diagram of an image processing method according to an embodiment of the present application;

[0036] FIGS. 4a-4e are schematic diagrams of RAW images in different formats according to an embodiment of the present application;

[0037] FIG. 5 is a schematic diagram of a principle of desensitization detection and desensitization processing based on a RAW image according to an embodiment of the present application;

[0038] FIG. 6 is another schematic diagram of a principle of desensitization detection and desensitization processing based on a RAW image according to an embodiment of the present application;

[0039] FIG. 7a is a modular structural schematic diagram of a terminal device according to an embodiment of the present application;

[0040] FIGS. 7b and 7c are flow schematic diagrams of an image processing method according to an embodiment of the present application;

[0041] FIG. 8 is a structural schematic diagram of an image processing apparatus according to an embodiment of the present application;

[0042] FIG. 9 is a structural schematic diagram of another image processing apparatus according to an embodiment of the present application. DETAILED DESCRIPTION

[0043] Before introducing the technical solutions provided in the present application, first, some of the terms involved in the present application are explained and described in order to facilitate understanding by those skilled in the art.

[0044] (1) Information / data desensitization: refers to the transformation of certain sensitive information through desensitization rules to achieve reliable protection of sensitive private data.

[0045] Data desensitization can be divided into two types: static data desensitization and dynamic data desensitization. Among them, static data desensitization is generally applied to data export scenarios, for example, data needs to be exported to developers, testers, analysts, etc. Static data desensitization will save the changed data, and then provide it for the data user. Dynamic data desensitization is generally applied to scenarios that directly connect to production data, such as operation and maintenance personnel directly connecting to production databases for operation and maintenance, and customer service personnel directly accessing personal information in production through applications. Dynamic data desensitization will change the data during data acquisition, without modifying the original data.

[0046] In the vehicle field, the vehicle end needs to detect (called desensitization detection) and implement desensitization processing on the privacy data (such as face, license plate, etc.) in the collected data before uploading the collected data to the cloud device, in order to protect the privacy safety of others. Among them, a desensitization model can be set in the computing platform of the vehicle, and the desensitization model can be used to realize desensitization detection and desensitization processing. The desensitization processing method used can include but is not limited to any one of the following: rule-based desensitization, encryption desensitization, disguise desensitization, data perturbation desensitization, and data shielding desensitization, which is not limited in the present application.

[0047] (2) Image encoding and image decoding: image encoding can also be called image compression, which refers to a technology that represents an image or information contained in an image with fewer bits under the condition of meeting certain quality requirements (such as signal-to-noise ratio requirements or subjective evaluation scores). While image decoding is the inverse process of image encoding.

[0048] (3) Bayer raw image, also known as Bayer image or raw image. Raw means "unprocessed" in its original sense, which can be understood as: the Bayer raw image refers to the original data converted by a charge coupled device (CCD) image sensor and a complementary metal oxide semiconductor (CMOS) image sensor and the like of a camera from a light source signal captured to a digital signal, that is, the original image inside the camera. Therefore, the Bayer raw image can also be conceptualized as "original image encoding data" or more figuratively as "digital negative".

[0049] In addition, in the embodiment of the present application, the Bayer raw image includes three color components, and each pixel point in the Bayer raw image has only one color component, and the value of the color component can be equivalent to the pixel value of the pixel point. In one example, the three color components are red (R) component, blue (B) component, and green (G) component. In another example, the three color components are R component, B component, and yellow (Y') component. Which color components are included in the Bayer raw image is related to the color filter in the camera.

[0050] (4) RGB color space and YUV color space:

[0051] Generally, an image is composed of the smallest unit of pixels, and each pixel information is also composed of different brightness RGB information, that is, the original image can include information of three components of red R, green G, and blue B. If the RGB signal is directly used for image signal transmission, it cannot be compatible with black and white televisions, and the cost of occupied bandwidth is high, so the traditional image signal processing (ISP) device converts the image from the RGB color space to the YUV color space for transmission.

[0052] In the YUV color space, an image information is divided into one luminance information and two chrominance information, the luminance information is represented by Y, and the chrominance information is composed of hue and saturation, and the hue and saturation are represented by UV. When performing image signal transmission, the analog component video or YUV signal needs to be digitally sampled, that is, the luminance information and the chrominance information need to be sampled. Common YUV sampling methods include YUV444, YUV422, YUV420, YUV411, etc.

[0053] To solve the above problems, the embodiment of the present application provides an image processing method and device for realizing desensitization detection and desensitization processing of an original RAW image. The method and the device are based on the same technical concept. Since the principles of the method and the device for solving problems are similar, the implementation of the device and the method can be mutually referred to, and the repeated parts will not be described again. In addition, in each embodiment of the present application, if there is no special description and logical conflict, the terms and / or descriptions between the embodiments are consistent and can be mutually referred to. The technical features in different embodiments can be combined to form new embodiments according to their inherent logical relationship.

[0054] It should be noted that "at least one" in the embodiments of the present application means one or more, and "multiple" means two or more. The "and / or" describes the association relationship of the associated objects, which means that there can be three kinds of relationships, for example, A and / or B can represent the following cases: A exists alone, A and B exist together, and B exists alone, where A and B can be singular or plural. The character " / " generally represents an "or" relationship between the associated objects before and after it. "At least one of the following" or similar expressions means any combination of these items, including any combination of single item or multiple items. For example, at least one of a, b, or c can represent: a, b, c, a and b, a and c, b and c, or a and b and c, where a, b, and c can be single or multiple.

[0055] In addition, unless otherwise specified, the ordinal numbers mentioned in the embodiments of the present application are used to distinguish a plurality of objects, and are not used to limit the priority or importance of the plurality of objects. For example, the first RAW image and the second RAW image are only used to distinguish different RAW images, and do not represent the difference in priority or importance of the two RAW images.

[0056] The embodiments of the present application can be applied to the scene that the terminal device collects images or videos and uploads them to the cloud device, that is, the end-to-cloud collaborative scene. Referring to FIG. 1, it is a structural schematic diagram of an end-to-cloud system provided by the embodiments of the present application. The "end" of the end-to-cloud collaboration refers to the terminal device, and the "cloud" refers to the cloud device. The cloud device can also be a cloud server or a cloud platform. The cloud device can have the function of massive computing power. The end-to-cloud system can include a terminal device 110 and a cloud device 120. The terminal device 110 can be connected to the cloud device 120 through a wireless network.

[0057] In an embodiment, the cloud device 120 can be a computer server or a server cluster composed of multiple servers, and the implementation architecture of the cloud device 120 is not limited in the present application. The terminal device 110 can be a device with a photographing function and a network function. The terminal device 110 can also have a computing processing function. The terminal device 110 can be a mobile terminal such as a mobile phone, a tablet computer, a wearable device, a vehicle-mounted device, an augmented reality (AR) / virtual reality (VR) device, a notebook computer, an ultra-mobile personal computer (UMPC), a netbook, a personal digital assistant (PDA), or the like, or can also be a professional photographing device such as a digital camera, a single-lens reflex camera / micro-single camera, a sports video camera, a gimbal camera, a drone, and the like, and the specific type of the terminal device is not limited in the embodiments of the present application. The number of terminal devices 110 included in the terminal-cloud system can be one or multiple. The types of the multiple terminal devices 110 can be the same or different.

[0058] In an example, taking a vehicle implemented by the terminal device 110 as an example, the hardware structure of the vehicle is introduced. Referring to FIG. 2, a schematic diagram of a hardware structure of a vehicle is shown.

[0059] The vehicle can include a processor 210, an external memory interface 220, an internal memory 221, an automotive bus interface 230, a communication module 240, a sensing system 250, a display screen 260, etc. Among them, the automotive bus interface 230 can include but is not limited to a controller area network (CAN) bus interface, a FlexRay bus interface, a LIN bus interface, etc. Through various bus interfaces, the interconnection between various components inside the vehicle can be realized, and the interconnection between the vehicle and peripheral devices can also be realized. The vehicle can also communicate with servers or other vehicles through the communication module 240 and the network. The sensing system 250 can include at least one sensor, and the processor can identify the environment or scene in which the vehicle is located through the sensing information provided by the at least one sensor, to assist the vehicle to realize automatic driving or intelligent auxiliary driving function. Illustratively, the sensing system 250 can include but is not limited to at least one of an image sensor (such as a camera), a light sensor, a distance sensor, a light detection and ranging (LIDAR), a millimeter-wave radar (RADAR), etc. The number of different types of sensors can be one or more, and the deployment position can be inside the vehicle or outside the vehicle.

[0060] It can be understood that the structure illustrated in the embodiments of the present application does not constitute a specific limitation on the terminal device 110. In other embodiments of the present application, the terminal device can include more or fewer components than the illustration, or combine certain components, or split certain components, or different component arrangements. The illustrated components can be implemented in hardware, software, or a combination of software and hardware.

[0061] In specific implementation, the processor 210 can be deployed on a related vehicle-mounted device of the vehicle, for example, deployed in a mobile data center (MDC) or a cockpit domain controller (CDC) of the vehicle, or a vehicle control unit (VCU), or a vehicle domain controller (VDC), or an advanced driver assistance system (ADAS) domain controller, or a control unit of other components of the vehicle. The product form and deployment manner of the processor are not limited in the embodiments of the present application.

[0062] The processor 210 can include one or more processing units, for example: the processor 210 can include an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU), etc. Among them, different processing units can be independent devices, or can be integrated in one or more processors.

[0063] The controller can generate operation control signals according to instruction operation codes and timing signals, and complete the control of fetching and executing instructions.

[0064] The processor 210 can also be provided with a memory for storing instructions and data. In some embodiments, the memory in the processor 210 is a cache memory. The memory can save instructions or data that the processor 210 has just used or repeatedly uses. If the processor 210 needs to use the instructions or data again, it can directly call from the memory. Avoiding repeated access, reducing the waiting time of the processor 210, thus improving the efficiency of the system.

[0065] In one example, the terminal device can realize image shooting function and processing of image through at least one of camera, ISP, video codec, GPU, display screen 260 and application processor, etc.

[0066] For example, the camera can be used to capture still images or videos. An object projects an optical image through a lens to a photosensitive element. The photosensitive element can be a CCD or CMOS phototransistor. The photosensitive element converts the optical signal to an electrical signal, which is then passed to an ISP in the processor 210 to convert to a digital image signal. The ISP outputs the digital image signal to the DSP for processing. The DSP converts the digital image signal to a standard RGB, YUV, etc. format image signal. In some embodiments, the processor 210 can trigger the camera to capture at least one image according to a program or instructions in the internal memory 221, and process the at least one image according to the program or instructions, such as image post-processing (e.g., skin beautification, super resolution processing to enhance clarity, etc.). The processed image can be displayed by the display 260. In some embodiments, the vehicle can include one or N2 cameras, where N2 is a positive integer greater than 1. For example, the vehicle can include a front-facing camera, a rear-facing camera, or a 360-degree surround view camera. The vehicle can also include a cabin camera, for example.

[0067] The ISP in the processor 210 can be used to process data fed back by the image sensor. For example, when taking a picture, the shutter is opened, light passes through the lens to the camera photosensitive element, the optical signal is converted to an electrical signal, and the camera photosensitive element then performs analogue-to-digital (A / D) conversion on the electrical signal to output a corresponding digital signal. The digital signal is passed to the ISP for processing to convert to a visible image. The digital signal output by the camera sensor to the ISP can be understood as the original image captured by the camera, i.e. a RAW image. The ISP can perform ISP processing on the RAW image to ultimately generate a YUV image.

[0068] Exemplarily, the ISP processing can include: black frame correction, bad pixel correction (DPC), RAW domain noise reduction, black level correction (BLC), lens shading correction (LSC), auto white balance (AWB) gain, green balance correction, demosaic color interpolation, color correction matrix (CCM), dynamic range compression (DRC), gamma, 3D look up table (LUT), YUV domain noise reduction, sharpening, detail enhancement, etc. The ISP can also optimize the exposure, color temperature, etc. of the shooting scene.

[0069] In some embodiments, part of the functions of the ISP can be arranged in the image sensor, and other ISP processing can be retained in the ISP device of the processor. For example, the black frame correction, bad pixel correction, RAW domain noise reduction, lens shading correction, black level correction, auto white balance gain, green balance correction, DRC, etc. in the above-mentioned ISP processing can be arranged in the camera or other sensor with camera function. Optionally, the function of the ISP for optimizing the exposure, color temperature, etc. of the shooting scene can also be integrated in the camera end. The demosaic color interpolation, color correction, gamma, 3D look up table, YUV domain noise reduction, sharpening, detail enhancement, etc. in the above-mentioned ISP processing can be arranged in the processor.

[0070] The digital signal processor is used to process digital signals, which can process not only digital image signals but also other digital signals. For example, when the terminal device selects a frequency point, the digital signal processor is used to perform Fourier transform on the frequency point energy, etc.

[0071] The video codec is used to compress or decompress digital video. The vehicle can support one or more video codecs. In this way, the vehicle can play or record videos in multiple encoding formats, such as: moving picture experts group (MPEG) 1, MPEG2, MPEG3, MPEG4, etc.

[0072] The NPU is a neural-network (NN) computing processor that quickly processes input information by drawing on the structure of a biological neural network, such as the transmission mode between human brain neurons, and can also continuously self-learn. Through the NPU, intelligent cognition and other applications of a terminal device can be implemented, such as image recognition, face recognition, speech recognition, text understanding, and the like.

[0073] The external memory interface 220 can be used to connect an external memory card, such as a Micro SD card, to implement the expansion of the storage capability of the terminal device. The external memory card communicates with the processor 210 through the external memory interface 220 to implement a data storage function. For example, music, video, and the like are saved in the external memory card.

[0074] The internal memory 221 can be used to store computer executable program codes, which include instructions. The internal memory 221 can include a program storage area and a data storage area. The program storage area can store an operating system, at least one application program (such as a camera application) required by a function, and the like. The data storage area can store data (such as images captured by a camera) created during the use of the terminal device, and the like. In addition, the internal memory 221 can include a high-speed random access memory, and can also include a non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, a universal flash storage (UFS), and the like. The processor 210 executes various function applications and data processing of the terminal device by running instructions stored in the internal memory 221 and / or instructions stored in a memory disposed in the processor.

[0075] It can be understood that the structure shown in FIG. 2 does not constitute a specific limitation on the terminal device. In some embodiments, the terminal device can also include more or fewer components than those shown in FIG. 2, or combine certain components, or split certain components, or different component arrangements, and the like. Alternatively, some components shown in FIG. 2 can be implemented in hardware, software, or a combination of software and hardware.

[0076] In addition, when the terminal device is another tablet computer, wearable device, AR / VR device, notebook computer, UMPC, netbook, PDA, and the like mobile terminal, or a digital camera, single-lens reflex camera / micro-single camera, action camera, gimbal camera, unmanned aerial vehicle, and the like professional shooting device, the specific structure of these other terminal devices can also be referred to that shown in FIG. 2. Exemplarily, the other terminal device can add or reduce components on the basis of the structure given in FIG. 2, which will not be repeated here.

[0077] It should also be understood that one or more photographing applications can be run in the vehicle, so that the photographing function is realized by running the photographing application. For example, the photographing application can include a system-level application "camera" application. For another example, the photographing application can also include other applications capable of photographing installed in the terminal device.

[0078] After the terminal device obtains the RAW image through the camera, the terminal device can detect and desensitize the privacy data in the RAW image to protect privacy security.

[0079] Taking the terminal device implemented as a vehicle as an example, on the one hand, the vehicle can perform a series of ISP processing on the RAW image, and then transmit the relevant image data to the automatic driving system or the intelligent auxiliary driving system of the vehicle, so as to realize the automatic driving function or the intelligent driving auxiliary function of the vehicle. On the other hand, the vehicle can perform desensitization detection and desensitization processing on the RAW image, and then encode the RAW image to obtain a corresponding encoded code stream. The vehicle can store the encoded code stream in the local storage medium of the vehicle, or can send the encoded code stream to the cloud device. On the other hand, the vehicle can directly process the RAW image to obtain a corresponding encoded code stream, which can be stored in the local storage medium of the vehicle. In the case of uploading the encoded code stream to the cloud device, the vehicle can obtain the encoded code stream from the storage medium, and perform desensitization detection and desensitization processing on the decoded RAW image, and then perform secondary encoding to obtain a corresponding encoded code stream, and then send the secondary encoded code stream to the cloud device. Correspondingly, the cloud device can receive the encoded code stream from the terminal device, and decode the encoded code stream, perform post-processing to restore the desensitized RAW image, and then perform a related visual task. For example, in a cloud photographing scenario, the desensitized RAW image is encoded into a corresponding encoded code stream and sent to the cloud device. By using the powerful computing power of the cloud device for image processing, the problem of limited image processing effect caused by limited processing performance and memory resources of the terminal device can be solved, and better image processing effect can be obtained. Or for example, in other scenarios, the desensitized RAW image is encoded into a corresponding encoded code stream and sent to the cloud device. By using the powerful computing power of the cloud device to build a machine vision model, the limitation of the sensitive frequency range of the human eye is overcome, so as to adapt to various machine vision models and improve the inference accuracy of the related model. For ease of understanding, the method flowchart will be described in detail below.

[0080] As shown in FIG. 3, the method can include the following steps:

[0081] S310: The terminal device obtains a first encoded code stream.

[0082] In the embodiments of the present application, the first encoded code stream can include encoded code streams of a plurality of components of the first original RAW image, and the plurality of components can include a first luminance component, a first color component and a first difference component. For example, the first RAW image includes an image in YCoCgDg format, Y is a luminance component, Co and Cg are color components, and Dg is a difference component.

[0083] In actual implementation, before S310, the terminal device can acquire a second RAW image, and pre-process the second RAW image to obtain a first RAW image. The terminal device can encode images of a plurality of components of the first RAW image respectively to obtain a first encoded code stream, and store the first encoded code stream.

[0084] For example, the terminal device can include an image acquisition module, a pre-processing module, an encoding module and a storage module. The image acquisition module can be, for example, a camera in FIG. 2, the pre-processing module can be, for example, a module integrated in the processor 210 in FIG. 2, and the storage module can be, for example, an internal storage or an external storage in FIG. 2. The encoding module can be, for example, a codec module integrated in the processor 210 in FIG. 2, or can be a separate deployed codec. When implementing S310, the pre-processing module of the terminal device can acquire a RAW image corresponding to a current target scene in real time through the image acquisition module to obtain a second RAW image. Alternatively, the pre-processing module can acquire the second RAW image provided by the image acquisition module from a storage medium. The pre-processing module can pre-process the second RAW image to obtain a first RAW image. The pre-processing module can send the first RAW image to the encoding module, and the encoding module can encode images of a plurality of components of the first RAW image respectively to obtain a first encoded code stream. The encoding module can send the first encoded code stream to the storage module for storing the first encoded code stream in the storage module.

[0085] In one example, when the image acquisition module does not have an ISP function, the second RAW image is an original RAW image captured by the image acquisition module on the vehicle, for example, a RAW image captured based on a Bayer mode. For example, taking vehicle A as an example, when the image acquisition module located on vehicle A captures an original RAW image based on certain environmental information, the image acquisition module can send the original RAW image to the pre-processing module. Alternatively, the image acquisition module can also cache the original RAW image. When the image acquisition module caches the original RAW image, the image acquisition module can send the cached original RAW image to the pre-processing module in response to an image request of the pre-processing module.

[0086] In another example, the image acquisition module can include partial ISP functions, and the second RAW image is a RAW image obtained after an original RAW image captured by the image acquisition module is processed by the internal ISP.

[0087] In another example, if the image acquisition module does not have ISP functions, the second RAW image can be a RAW image obtained after an original RAW image captured by the image acquisition module on the vehicle is processed by an external ISP. For example, still taking vehicle A as an example, during the driving process of vehicle A, the image acquisition module on vehicle A can send the original RAW image to the ISP module outside the image acquisition module after capturing the original RAW image of the environmental information. The external ISP module can process and optimize the original RAW image to obtain the second RAW image after receiving the original RAW image.

[0088] The internal ISP module or the external ISP module can perform at least one of the following processing on the original RAW image to obtain the second RAW image: black frame correction, bad pixel correction, RAW domain noise reduction, lens brightness correction, black level correction, automatic white balance gain, green balance correction, DRC, etc. The internal ISP module or the external ISP module can send the second RAW image to the preprocessing module, or can cache the second RAW image. When the internal ISP module or the external ISP module caches the second RAW image, the cached second RAW image can be sent to the preprocessing module in response to an image request of the preprocessing module. In order to facilitate the distinction, the above-mentioned ISP processing on the original RAW image can also be referred to as first ISP. Then, the terminal device can further perform other ISP processing on the second RAW image. For example, the external ISP module can perform demosaicing color interpolation, color correction, gamma, 3D lookup table, YUV domain noise reduction, sharpening, and detail enhancement on the second RAW image to obtain an image in RGB format or YUV format adapted to the human eye.

[0089] In the embodiments of the present application, the Bayer raw format of the second RAW image can include, but is not limited to, an RGGB format or a GRGB format, etc. The image signals of the four channels are different from the image signals of the traditional three-channel RGB format or YUV format, and are not suitable for direct encoding using a traditional encoder, and need to be further processed before encoding. The further processing is referred to as preprocessing of the second RAW image in the embodiments of the present application, and the RAW image obtained after preprocessing is referred to as a first RAW image.

[0090] Exemplarily, the first RAW image can include an image in a YCoCgDg format, Y is a luminance component, Co and Cg are chrominance components, and Dg is an interpolation component. The image in the YCoCgDg format can be obtained by directly performing a color gamut transformation on the image in the Bayer raw format, without converting the second RAW image into an RGB format by performing demosaicing color interpolation, etc.

[0091] In one example, when implementing the above preprocessing step, the preprocessing module of the terminal device can directly perform a color gamut transformation on the second RAW image to obtain the images of the four components in the YCoCgDg format of the first RAW image. In another example, when implementing the above preprocessing step, the preprocessing module of the terminal device can also perform a nonlinear transformation on the second RAW image to obtain a transformed second RAW image, and then perform a color gamut transformation on the transformed second RAW image to obtain the images of the four components in the YCoCgDg format of the first RAW image.

[0092] The nonlinear transformation on the second RAW image can make the RAW image data distribution more suitable for subsequent encoding processing, thereby improving the image encoding performance and reducing the image encoding delay. The color gamut transformation on the second RAW image (including the nonlinearly transformed second RAW image) can remove the correlation between different channel data, remove the redundant information between color signals, improve the image compression performance, and reduce the computational complexity (or image processing complexity), for example, the computational complexity can be saved by more than 60%, thereby achieving the improvement of the image compression performance. When performing the nonlinear transformation on the second RAW image, exemplarily, the bit depth of the second RAW image can be 24 bits, and the basic principle of the nonlinear transformation on the second RAW image can be to map the 24-bit RAW data to the human eye sensitivity domain of a selectable bit depth (referred to as bit depth). The basic operations of the above nonlinear transformation can include the following two kinds:

[0093] (1) 24bit RAW data is mapped to 16 / 12 / 10 / 8bit bit depth; (2) gamma-like transformation, the original RAW data is transformed to the human eye sensitive domain through a non-linear gamma transformation, and the transformation form can be a function, a lookup table, key position points, etc.

[0094] Taking a function-based non-linear transformation as an example, the pre-processing module can perform a non-linear transformation on the second RAW image through the following expression (1).

[0095] wherein X is used to represent a first normalized pixel value of a certain pixel point in the second RAW image, Y is used to represent a second normalized pixel value (i.e. the first normalized pixel value after non-linear transformation) of the pixel point, and the value of gamma can be any suitable value. Optionally, the pre-processing module can also use other non-linear transformation methods to perform non-linear transformation on the plurality of pixel points included in the second RAW image, which will not be described here.

[0096] It should be understood that in the embodiments of the present application, the bit depth of the original RAW image obtained by the image acquisition module can be 24bit, and the RAW image data of 24bit is too large, and the transmission pressure is also large. Before being transmitted to the pre-processing module via a wired cable, the related device (including the image acquisition module integrated with the first ISP function or an external ISP module) can also compress the original RAW image in a pixel-wise linear (PWL) compression manner, such as compressing the RAW data from 24bit to 12bit, reducing the transmission bandwidth requirement by 1 times. After the data is transmitted to the pre-processing module, the 12bit compressed data is restored to 24bit through the PWL decompression of the pre-processing module, and the second RAW image is obtained.

[0097] When performing color gamut transformation on the second RAW image (including the second RAW image after non-linear transformation), the pre-processing module can extract the RGGB four components of the RAW image in Bayer format into a four-channel image in RGGB format or GRGB format, and then transform the RAW image in RGGB format or GRGB format into a RAW image in YCoCgDg format by using a color gamut transformation matrix.

[0098] As an example of the Bayer raw format of the second RAW image being RGGB format, as shown in FIG. 4a, the color component of the pixel point in the first row and the first column of the second RAW image is R, the color component of the pixel point in the first row and the second column is G, the color component of the pixel point in the second row and the first column is G, and the color component of the pixel point in the second row and the second column is B, and the pixel points of the four color components appear in a cycle. As an example of the Bayer raw format of the second RAW image being GRGB format, as shown in FIG. 4b, the color component of the pixel point in the first row and the first column of the second RAW image is G, the color component of the pixel point in the first row and the second column is R, the color component of the pixel point in the second row and the first column is G, and the color component of the pixel point in the second row and the second column is B, and the pixel points of the four color components appear in a cycle.

[0099] As an example of the first value being a Y value, the second value being a Dg value, the third value being a Co value, and the fourth value being a Cg value, taking a certain image unit (such as image unit A) included in the second RAW image as an example, the image unit A includes four pixel points (such as pixel point 1, pixel point 2, pixel point 3, and pixel point 4). Wherein, it is assumed that the pixel value (or color component) of the pixel point 1 included in the image unit A is G1, the pixel value of the pixel point 2 is G2, the pixel value of the pixel point 3 is R, and the pixel value of the pixel point 4 is B. The pre-processing module can calculate the Y value, the Dg value, the Co value, and the Cg value corresponding to the image unit A through the following expression (2):

[0100] As shown in FIG. 4c, the color component of the pixel point in the first row and the first column of the first RAW image is Y, the color component of the pixel point in the first row and the second column is Dg, the color component of the pixel point in the second row and the first column is Co, and the color component of the pixel point in the second row and the second column is Cg, and the pixel points of the four color components appear in a cycle.

[0101] It should be understood that the pixel value G1 of the pixel point 1, the pixel value G2 of the pixel point 2, the pixel value R of the pixel point 3 and the pixel value B of the pixel point 4 are all pixel values after nonlinear transformation. The color gamut conversion process based on the expression (2) is relatively simple, generally only involving addition and shift. Alternatively, the pre-processing module can also use the existing color gamut conversion method or other color gamut conversion method to realize the color gamut conversion of the second RAW image after nonlinear transformation. In another example, when the second RAW image (including the second RAW image after nonlinear transformation) is subjected to color gamut conversion, the Star Tetrix transform method can also be used to obtain a four-channel image by referring to the pixels of the first four rows, and the four-channel image can be represented as a YCbCr△ (Delta) format image, and the second RAW image can be an image including a YCbCr△ format image. The YbBCr△ format is a different name of the YCoCgDg format, and the essence is still a YCoCgDg format image.

[0102] In specific implementation, the pre-processing module can generate chrominance components Cb and Cr by predicting R and B from four green components Gx (x={l, r, t, b}) around the target pixel point in the second RAW image, as shown in FIG. 4d. The prediction process can satisfy the following expression (3):

[0103] Wherein, l, r, t and b respectively represent the sample positions of the left, right, top and bottom of the current pixel.

[0104] Then, the pre-processing module can generate luminance components Y1 and Y2 by updating G from four surrounding and (x={l, r, t, b}). The process can satisfy the following expression (4):

[0105] Then, the pre-processing module can generate a luminance difference component△ by predicting Y1 from four surrounding Y2, satisfying the following expression (5):

[0106] Then, the pre-processing module can generate the final luminance component Y by updating Y2 from four surrounding, satisfying the following expression (6):

[0107] After the above Star Tetrix transform, a four-channel image can be obtained, as shown in FIG. 4e, including an image of Y component, an image of Cr component, an image of Cb component and an image of Delta component. Y component is a luminance component, Cr and Cb are color components, and Delta is a difference component.

[0108] S320: The terminal device decodes the encoded code stream of the first luminance component to obtain an image of the first luminance component of the first RAW image.

[0109] S330: The terminal device performs desensitization detection and desensitization processing based on the image of the first luminance component to obtain an image of the second luminance component.

[0110] In the embodiments of the present application, the terminal device can include a desensitization model. When S330 is implemented, the desensitization model can be used to perform desensitization detection and desensitization processing on the image of the first luminance component of the first RAW image to obtain the image of the second luminance component. In one example, the desensitization model can be arranged in a separate desensitization module. In another example, the desensitization model can be built into a related processing module of the terminal device, for example, built into the preprocessing module mentioned above. The embodiments of the present application do not make specific limitations in this regard.

[0111] In a specific implementation, in one example, the desensitization model can be trained for images in YUV format or images in RGB format. The terminal device can input the image of the first luminance component of the first RAW image into three input channels corresponding to Y, U, and V or three input channels corresponding to R, G, and B of the desensitization model, respectively, to obtain the image of the desensitized luminance component after desensitization detection and desensitization processing by the desensitization model, which is represented as the image of the second luminance component. This method can be considered as copying the image of the Y channel three times to input into the desensitization model.

[0112] As shown in FIG. 5, taking the desensitization model trained for images in RGB format based on a bit width depth of 8 bits as an example, the image of the Y component of 8 bits is corresponded to the three input channels of R, G, and B, respectively, represented as Y(R), Y(G), and Y(B). These component images constitute an 8-bit RGB image in the desensitization model, and the image of the desensitized luminance component can be output after desensitization detection and desensitization processing in the desensitization model.

[0113] It should be understood that in the embodiments of the present application, the bit width depth of the image supported by the desensitization model is a target bit, which can be, for example, 8 bits. The bit width depth of the second RAW image can be greater than or equal to the target bit, for example, the bit width depth of the second RAW image is any one of the following: 8 bits, 10 bits, 12 bits, or 16 bits. If the bit width depth of the second RAW image is greater than the target bit, when S330 is implemented, the terminal device can input the image of the first luminance component into the three input channels of the desensitization model, and can cut the image of the target bit from the image of the first luminance component and input the image of the target bit into the three input channels of the desensitization model, respectively.

[0114] Taking the target bit as 8 bits and the bit width depth of the second RAW image as 12 bits as an example, as shown in FIG. 6, the image of the Y component of the first RAW image obtained by preprocessing the second RAW image is also 12 bits, which can be based on target bit truncation, for example, truncating the upper 8 bits or the lower 8 bits of the image of the Y component of 12 bits as the first luminance component image to be input, the terminal device can copy the truncated 8-bit Y component image three times (indicated as YYY) and input into the three input channels of the desensitization model R, G, and B, to form a pure luminance RGB image (8 bits), after the desensitization model detects and desensitizes the RGB image, the image of the desensitized Y component is obtained. The terminal device can restore the desensitized Y component image to 12 bits.

[0115] It should be understood that FIG. 6 is only an example of desensitization detection and desensitization processing of a RAW image with a high bit width depth, taking 12 bits as an example. The method is also applicable to desensitization detection and desensitization processing of RAW images with 10 bits, 16 bits, or other higher bit width depths, which will not be described here.

[0116] S340: The terminal device encodes the image of the second luminance component to obtain an encoded code stream of the second luminance component.

[0117] S350: The terminal device obtains a second encoded code stream based on the encoded code stream of the second luminance component, the encoded code stream of the first color component, and the encoded code stream of the first difference component.

[0118] In the embodiment of the application, the terminal device can include an encoding module (or encoder), which can perform the encoding operations of S340 and S350. The bit width depth of the image supported by the encoding module can be the target bit. The encoding standard adopted by the encoding module is not limited in the embodiment of the application.

[0119] S360: The terminal device stores the second encoded code stream or sends the second encoded code stream to a cloud device.

[0120] In the embodiment of the application, the terminal device can store the second encoded code stream in a local non-volatile storage medium. The type of the non-volatile storage medium is not limited in the embodiment of the application.

[0121] The cloud device can be a cloud server or a cloud platform as described above. The cloud device can decode the second encoded code stream according to a set video encoding standard or by multiplexing an existing encoding and decoding standard, to obtain a desensitized first RAW image from the second encoded code stream. Then, the cloud device can perform post-processing on the desensitized first RAW image to restore it to a desensitized second RAW image.

[0122] Exemplarily, the first RAW image comprises an image in YCoCgDg format, Y is a luminance component, Co and Cg are color components, and Dg is a difference component. The post-processing can comprise at least one of color gamut inverse transformation processing, linear transformation, etc. The cloud device can directly perform color gamut inverse transformation on the desensitized first RAW image to obtain an image in RGGB format or an image in GRGB format, thereby obtaining a desensitized second RAW image. Alternatively, the cloud device can perform linear transformation on the desensitized first RAW image to obtain a transformed first RAW image, and then perform color gamut inverse transformation on the transformed first RAW image to obtain an image in RGGB format or an image in GRGB format, thereby obtaining a desensitized second RAW image.

[0123] The cloud device can transmit the desensitized second RAW image to a downstream task module. The downstream task module can be a module comprising a neural network model based on a RAW domain, or can be a module comprising a machine vision model or a human eye vision model based on an RGB format. The cloud device can directly transmit the desensitized second RAW image to the downstream task module comprising the neural network model based on the RAW domain. The cloud device can perform relatively complex ISP processing (for example, the second ISP described above) on the desensitized second RAW image, and then transmit an RGB format image obtained by the ISP processing to the module comprising the machine vision model or the human eye vision model based on the RGB format.

[0124] In an optional embodiment, the terminal device described above can also support an image encoding and compression capability for an RGB format or a YUV format. In this scheme, the terminal device can comprise an ISP module, which can have the capability of performing relatively complex ISP processing on the second RAW image to obtain an RGB format or a YUV format image suitable for human eyes. For example, the terminal device can perform the above-mentioned processing on the second RAW image through the ISP module to obtain a third image, and the third image comprises a YUV format image, Y is a luminance component, and U and V are chrominance components. The encoding module of the terminal device can also encode the third image to obtain a third encoding bitstream. The terminal device can store the third encoding bitstream or send the third encoding bitstream to the cloud device. The encoding module used for encoding the third image in S340 can be the same encoding module as the encoding module, or can be a different encoding module, that is, the encoding module of the embodiment of the present application can support both encoding processing of RAW images and encoding processing of RGB format or YUV format images, thereby realizing multiplexing of existing encoding and decoding standards.

[0125] Thus, by the above method, the desensitization detection and desensitization processing of the RAW image can be realized without changing the existing desensitization architecture and algorithm, the user privacy is protected, the chip area of the vehicle-side decoder and the transmission cost are reduced on the hardware, and the influence of the secondary encoding on the image quality caused by the image desensitization is reduced.

[0126] For ease of understanding, the implementation details of the above image processing method are exemplarily described below in combination with the modular structure of the terminal device.

[0127] As shown in FIG. 7a, in an example, the vehicle can include a camera, an MDC, an encoding module 1, a storage module 1, a decoding module 1, a desensitization module 1, an encoding module 2 and a communication module 1. The camera can be based on the image acquisition module introduced in FIG. 3. The preprocessing module based on FIG. 3 can be integrated in the MDC (or CDC and other vehicle components with a preprocessing module). The desensitization module 1 can be integrated in the MDC or deployed independently of the MDC. FIG. 7a is only an example and does not constitute any limitation. The desensitization module 1 can cooperate with other modules to realize the RAW domain image desensitization process involved in the image processing method of the embodiments of the present application. As shown in FIG. 7b, the method can include the following steps, for example: S701: After triggering the camera shooting or video shooting function on the vehicle side, the camera can acquire an image or a video frame and send the acquired image or video frame to the MDC. Optionally, part of the ISP function (for example, the first ISP processing introduced above) can be integrated in the camera to realize simple processing of the original RAW image, for example, the first ISP processing introduced above.

[0128] S702: The deserializer in the MDC can receive the image or video frame from the camera. Wherein, the camera can send a serial data stream to the MDC in the form of a media stream. The deserializer can convert the serial data stream into a parallel data stream to facilitate the subsequent preprocessing module or ISP module 1 or other modules to process the RAW domain image.

[0129] S703: The preprocessing module can obtain the RAW image through the deserializer, take the obtained RAW image as a second RAW image, and preprocess the second RAW image to obtain a first RAW image. The preprocessing module can send the first RAW image to the encoding module 1.

[0130] Exemplarily, the preprocessing step can include at least one of nonlinear transformation processing and color domain transformation processing on the second RAW image. The first RAW image can be, for example, an image in YCoCgDg format, Y is a luminance component, Co and Cg are color components, and Dg is a difference component. For details, refer to the related description in the foregoing description of FIGS. 3-6, which will not be repeated here.

[0131] S704: The encoding module 1 can encode the images of the plurality of components of the first RAW image respectively to obtain the first encoded code stream. The encoding module 1 can store the first encoded code stream into the storage module 1.

[0132] S705: If necessary, for example, when receiving a cloud uploading instruction, the decoding module 1 can obtain the first encoded code stream from the storage module 1, and decode the encoded code stream of the first luminance component in the first encoded code stream to obtain the image of the first luminance component of the first RAW image. The decoding module 1 can send the image of the first luminance component to the desensitization module 1, and send the images of the first color component and the first difference component of the first RAW image to the encoding module 2.

[0133] S706: The desensitization module 1 can perform desensitization detection and desensitization processing based on the image of the first luminance component to obtain the image of the second luminance component. The desensitization module 1 can send the image of the second luminance component to the encoding module 2.

[0134] S707: The encoding module 2 can encode the image of the second luminance component to obtain the encoded code stream of the second luminance component, and obtain the second encoded code stream based on the encoded code stream of the second luminance component, the encoded code stream of the first color component and the encoded code stream of the first difference component.

[0135] S708: The encoding module 2 can store the second encoded code stream through a storage module (for example, the storage module 1). Alternatively, the encoding module 2 can send the second encoded code stream to a cloud device through the communication module 1.

[0136] It should be understood that the desensitization module 1 independently deployed from the MDC in FIG. 7a is only an example and does not constitute any limitation. In an alternative embodiment, a desensitization model can be built in a preprocessing module, and the image of the first luminance component of the first RAW image can be subjected to desensitization detection and desensitization processing through the desensitization model to obtain the image of the second luminance component after desensitization. The preprocessing module can send the image of the second luminance component, the image of the first color component and the image of the first difference component of the first RAW image to the encoding module 1, and the encoding module 1 can encode the images of the second luminance component, the first color component and the first difference component of the first RAW image respectively to obtain the first encoded code stream. The encoding module 1 can store the first encoded code stream or send the first encoded code stream to a cloud device.

[0137] In another example, the MDC of the vehicle can further include an ISP module 1, a security automation (SA) module, a vision preprocessing core (VPC) related distortion correction module, a stitching module, an image preprocessing module, etc. The ISP module 1 can cooperate with other modules to implement the process of image processing in the RGB domain or YUV domain involved in the image processing method of the embodiments of the present application.

[0138] As shown in FIG. 7c, the method may, for example, include the following steps:

[0139] S711: After triggering the camera shooting or video shooting function on the vehicle side, the camera can collect image or video frames and send the collected image or video frames to the MDC. Optionally, the camera can be integrated with part of the ISP function (for example, the first ISP processing introduced above) to realize simple processing of the original RAW image, for example, the first ISP processing introduced above.

[0140] S712: The deserializer in the MDC can receive the image or video frames from the camera. Wherein, the camera may, for example, send a serial data stream to the MDC in the form of a media stream. The deserializer can convert the serial data stream into a parallel data stream to facilitate subsequent processing by the preprocessing module or the ISP module 1 or other modules.

[0141] S713: The ISP module 1 can obtain the RAW image through the deserializer and perform ISP processing on the obtained RAW image to convert the RAW image into an image in the RGB format suitable for the human eye. Optionally, the ISP module 1 can also convert the image in the RGB format into the YUV format and sample it in the YUV420 manner (or other sampling manners).

[0142] S714: The ISP module 1 can provide the YUV format image data obtained after sampling to the autonomous driving system / intelligent driving assistance system of the vehicle, so that the autonomous driving system / intelligent driving assistance system realizes the autonomous driving function or the assisted driving function of the vehicle based on the YUV format image.

[0143] For example, the ISP module 1 can realize the vehicle parking function after the YUV format image data obtained after sampling is subjected to SA processing and VPC stitching processing.

[0144] Or for example, the ISP module 1 can perform fusion perception processing after the YUV format image data obtained after sampling is subjected to SA processing, VPC distortion correction, and VPC image preprocessing, to realize the autonomous driving function or the assisted driving function of the vehicle.

[0145] S715: The ISP module 1 can send the YUV format image data obtained after sampling to the encoding module 3 after SA processing and VPC distortion correction processing. The encoding module 3 can encode the YUV format data according to an existing encoding standard to obtain a third encoding code stream.

[0146] S716: The encoding module 3 can store the third encoding code stream through the storage module 2 or send the third encoding code stream to a cloud device.

[0147] S717: If necessary, the third encoding code stream can be obtained by the decoding module 2 from the storage module 2 and decoded. The decoding module 2 can send the decoded data to the desensitization module 2.

[0148] S718: The desensitization module 2 sends relevant data to the encoding module 4 after desensitization detection and desensitization processing on the received image data.

[0149] S719: The encoding module 4 performs secondary encoding on the received relevant data to obtain a fourth encoding code stream. The encoding module 4 can send the fourth encoding code stream to a cloud device through the communication module 2.

[0150] It should be understood that the modules in FIG. 7a are only examples and are not any limitation. In a specific implementation, the encoding module 1, the encoding module 2 and the encoding module 3 in FIG. 7a can be the same encoding module, and the decoding module 1 and the decoding module 2 can be the same decoding module. The implementation of each module is not limited in the embodiments of the present application.

[0151] Correspondingly, the cloud device can include a receiving module, a decoding module 3, a decoding module 4, a post-processing module, an ISP module 2 and a downstream task module (not shown in the figure). Optionally, the cloud device can further include a storage module (not shown in the figure). The receiving module can receive the first encoding code stream and / or the second encoding code stream from the terminal device and send the first encoding code stream and / or the second encoding code stream to the decoding module 3. The decoding module 3 can decode the first encoding code stream and / or the second encoding code stream to obtain a desensitized YCoCgDg format image. The decoding module 3 can send the desensitized YCoCgDg format image to the post-processing module. The post-processing module can post-process the desensitized YCoCgDg format image to obtain a desensitized second RAW image. For example, the post-processing step can include at least one of linear transformation processing, color domain inverse transformation processing and the like on the desensitized YCoCgDg format image. The post-processing module can send the desensitized second RAW image to the downstream task module for machine vision or related task processing based on a RAW domain neural network model in the downstream task module.

[0152] Alternatively, the receiving module can also receive the third encoded bitstream from the terminal device and send the third encoded bitstream to the decoding module 3. The decoding module 3 can multiplex an existing encoding and decoding standard to decode the desensitized second RAW image from the third encoded bitstream, including the image in the YCoCgDg format. The decoding module 3 can send the desensitized second RAW image to the post-processing module. The post-processing module can post-process the desensitized second RAW image to obtain the desensitized first RAW image. The post-processing module can send the desensitized first RAW image to the downstream task module so as to implement machine vision or related task processing based on the RAW domain neural network model in the downstream task module. Optionally, the decoding module 3 can send the desensitized second RAW image obtained by decoding to the ISP module 2, and the ISP module 2 can perform other ISP processing on the desensitized second RAW image to obtain an image adapted to the human eye, facilitating the downstream task module to implement human eye vision related tasks.

[0153] Alternatively, the receiving module can also receive the third encoded bitstream (or the fourth encoded bitstream) from the terminal device and send the third encoded bitstream (or the fourth encoded bitstream) to the decoding module 4, and the decoding module 4 can decode the third encoded bitstream (or the fourth encoded bitstream) to obtain a third image. The third image can include an image in the YUV format, Y being a luminance component and U and V being chrominance components. The decoding module 3 can send the third image to the downstream task module so as to implement human eye vision related tasks in the downstream task module.

[0154] It should be understood that, on the terminal device side, the encoding module 1, the encoding module 2 and the encoding module 3 can be the same encoder of the terminal device, and the decoding module 1 and the decoding module 2 can be the same decoder. The encoder and the decoder can also be the same codec. On the cloud device side, the decoding module 3 and the decoding module 4 can be the same decoder of the cloud device, and the implementation mode of each module is not limited in the embodiments of the present application. In specific implementation, the pre-processing module can also call the ISP module 1 to implement the pre-processing process of the RAW domain image. For example, the pre-processing module can call the ISP module 1 to implement the nonlinear transformation processing of the RAW domain image.

[0155] Based on the same idea, the embodiments of the present application further provide an image processing device suitable for the system architecture shown in FIG. 1. Exemplarily, the image processing device can be a vehicle as shown in FIG. 1, or can also be a functional element (such as a plug-in, a component or a chip, etc.) provided in the vehicle, which has the function of implementing the image processing method. In one example, the image processing device can be any device (such as a computing device or a control unit, etc.) in the vehicle with image processing capability, such as a camera, an MDC or a CDC, etc. in the vehicle. In another example, the image processing device can be other devices (such as a server or a cloud, etc.) outside the vehicle, or can also be a functional element provided in the other devices with the image processing function, which has the function of implementing the image processing method. Optionally, the image processing device can be used to implement the image processing method provided in the above embodiments, or the modules (such as chips) of the image processing device can be used to implement the image processing method provided in the above embodiments, thus also achieving the beneficial effects possessed by the above embodiments.

[0156] As shown in FIG. 8, when the image processing apparatus 800 is used to implement the image processing method shown in FIG. 3, FIG. 7b or FIG. 7c, the image processing apparatus 800 can include: an acquisition module 801 configured to acquire a first coded stream, the first coded stream including coded streams of a plurality of components of a first original RAW image, the plurality of components including a first luminance component, a first color component and a first difference component; a decoding module 802 configured to decode the coded stream of the first luminance component to obtain an image of the first luminance component of the first RAW image; a desensitization module 803 configured to perform desensitization detection and desensitization processing based on the image of the first luminance component to obtain an image of a second luminance component; an encoding module 804 configured to encode the image of the second luminance component to obtain a coded stream of the second luminance component; and obtain a second coded stream based on the coded stream of the second luminance component, the coded stream of the first color component and the coded stream of the first difference component; a storage module 805 configured to store the second coded stream, or a sending module 806 configured to send the second coded stream to a cloud device. For specific implementation, refer to the method steps implemented by the above method embodiments in combination with FIG. 3, FIG. 7b or FIG. 7c, which will not be described here. It should be noted that the division of modules in the embodiments of the present application is illustrative, and is only a logical function division. In actual implementation, another division mode can be used. In addition, the function units in each embodiment of the present application can be integrated in one processing unit, or can be physically separated, or two or more units can be integrated in one unit. For example, taking the pre-processing module and the encoding module as an example, the pre-processing module and the encoding module can be integrated in one module, or the pre-processing module and the encoding module are the same module. The integrated unit can be realized in the form of hardware or software function unit.

[0157] When the integrated unit is realized in the form of a software function unit and sold or used as an independent product, it can be stored in a computer readable storage medium. Based on this understanding, the technical solutions of the present application essentially or the part that contributes to the prior art or the whole or part of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium, including a plurality of instructions to make a computer device (which can be a personal computer, or a server, etc.) or a processor execute all or part of the steps of the embodiments of the present application. The foregoing storage medium includes: a U disk, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, and various media that can store program codes.

[0158] In a simple embodiment, those skilled in the art can conceive that the image processing apparatus in the above embodiments can all adopt the form shown in FIG. 9. The image processing apparatus can be used to implement the technical solutions related to the image processing apparatus in the above method embodiments, and thus can also achieve the beneficial effects possessed by the image processing apparatus in the above method embodiments.

[0159] The image processing apparatus 900 shown in FIG. 9 includes a transceiver 910, a processor 920, and optionally a memory 930. The transceiver 910, the processor 920, and the memory 930 are connected to each other. When the image processing apparatus 900 is used to implement the image processing method provided in the above embodiments, the transceiver 910 can be used to implement the data transceiving function of the terminal device or the cloud device; the processor 920 can be used to implement the data processing function of the RAW domain image, or can also be used to implement the data processing function of the RGB domain image or the YUV domain image.

[0160] Optionally, the transceiver 910, the processor 920, and the memory 930 are connected to each other through a bus 940. The bus 940 can be a peripheral component interconnect (PCI) bus or an extended industry standard architecture (EISA) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For the convenience of representation, only one thick line is used in FIG. 9, but it does not mean that there is only one bus or only one type of bus.

[0161] The transceiver 910 is configured to receive and send data. For example, when the transceiver 910 is deployed in a terminal device, the transceiver 910 can be used to implement the communication with an image acquisition module or a cloud device. In one example, the transceiver can be a transceiver device integrated with a data transceiving function. In another example, the transceiver can also be composed of a transmitter and a receiver, wherein the transmitter is configured to send data, and the receiver is configured to receive data.

[0162] Optionally, the transceiver 910 can include a transmitter and / or a receiver. The transmitter is configured to send signals, messages, information, or data, etc. The receiver is configured to receive signals, messages, information, or data, etc. For example, the transmitter sends signals, messages, information, or data, etc. under the control of the processor 920. The receiver receives signals, messages, information, or data, etc. under the control of the processor 920.

[0163] The functions of the processor 920 can refer to the method embodiments shown in FIGS. 2-6 described above, and will not be repeated here. The processor 920 can be a central processing unit (CPU), a network processor (NP), or a combination of CPU and NP, and the like. The processor 920 can further include a hardware chip. The hardware chip can be an application-specific integrated circuit (ASIC), a programmable logic device (PLD), or a combination thereof. The PLD can be a complex programmable logic device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL), or any combination thereof. The processor 920 can be implemented by hardware, and of course can execute corresponding software to implement the functions.

[0164] The memory 930 can include a volatile memory, such as a random access memory (RAM), and can also include a non-volatile memory, such as at least one disk memory.

[0165] The memory 930 stores executable program codes, and the processor 920 executes the executable program codes to respectively implement the functions of the foregoing image processing system (such as the desensitization module and the preprocessing module), thereby implementing the image processing method provided in the embodiments of the present application. That is, the memory 930 stores computer program instructions for executing the image processing method.

[0166] Alternatively, the memory 930 stores executable codes, and the processor 920 executes the executable codes to implement the functions of the foregoing image processing device (such as the preprocessing module or the ISP module), thereby implementing the image processing method provided in the embodiments of the present application. That is, the memory 930 stores computer program instructions for executing the image processing method provided in the embodiments of the present application.

[0167] Based on the same idea, the embodiments of the present application further provide a possible image processing system, which can include one or more of the MDC, the CDC, the industrial computer or the server (or cloud). Optionally, the image processing system can also include a display module for displaying the RAW image. For example, the image processing system includes the MDC and the CDC. In one example, the preprocessing module is deployed on the MDC, and the desensitization module is deployed on the CDC. In another example, the preprocessing module is deployed on the image acquisition device, and the desensitization module is respectively deployed on the MDC and the CDC. For example, the number of MDCs or CDCs included in the image processing system can be one or more, and the number of preprocessing modules and desensitization modules can be one or more, which are not limited by the embodiments of the present application. The image processing system can be deployed on a vehicle. Accordingly, the embodiments of the present application further provide a vehicle including an image acquisition device and the above-mentioned image processing system. The image acquisition device can be used to acquire an image (such as a RAW image) corresponding to the surrounding environment information of the vehicle.

[0168] Based on the same idea, the embodiments of the present application further provide a computer program product, which includes computer programs or instructions, when the computer programs or instructions are run on a computer, so that the computer executes the image processing method provided by the above embodiments.

[0169] Based on the same idea, the embodiments of the present application further provide a computer readable storage medium, which stores computer programs or instructions, when the computer programs or instructions are executed by a computer, so that the computer executes the image processing method provided by the above embodiments.

[0170] For example, but not limited to, the computer readable medium can include RAM, ROM, EEPROM, CD-ROM or other optical disk storage, magnetic disk storage medium or other magnetic storage device, or any other medium that can be used to carry or store desired program code in the form of instructions or data structures and can be accessed by a computer.

[0171] Based on the same idea, the embodiments of the present application further provide a chip, which is coupled with a memory, and is used to read the computer program stored in the memory to implement the image processing method provided by the above embodiments.

[0172] Based on the same idea, the embodiments of the present application also provide a chip system, which comprises a processor for supporting a computer device to implement the functions related to the image processing system in the above embodiments. In a possible design, the chip system further comprises a memory for storing the necessary programs and data of the computer device. The chip system can be composed of a chip, or can include a chip and other discrete devices.

[0173] The method provided by the embodiments of the present application can be implemented by software, hardware, firmware or any combination thereof. When implemented by software, the method can be implemented in the form of a computer program product entirely or partially. The computer program product includes one or more computer instructions. When the computer instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present application are entirely or partially generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network or other programmable apparatus. The computer instructions can be stored in a computer readable storage medium or transferred from one computer readable storage medium to another computer readable storage medium, for example, the computer instructions can be transferred from one website, computer, server or data center to another website, computer, server or data center through wired (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (such as infrared, wireless, microwave, etc.) mode. The computer readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server, data center, etc. integrated with one or more available media. The available media can be magnetic media (such as floppy disk, hard disk, magnetic tape), optical media (such as high-density digital video disc (digital video disc, DVD)) or semiconductor media (such as solid state drive (solid state drive, SSD)) and the like.

[0174] The steps of the method described in the embodiments of the present application can be directly embedded in hardware, software units executed by a processor, or a combination of the two. The software units can be stored in a RAM, a ROM, an EEPROM, a register, a hard disk, a removable disk, a CD-ROM or any other form of storage medium in the art. The storage medium can be connected to the processor so that the processor can read information from the storage medium and can write information to the storage medium. Alternatively, the storage medium can also be integrated into the processor. The processor and the storage medium can be arranged in an ASIC.

[0175] The computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process such that the instructions which execute on the computer or other programmable apparatus provide steps for implementing the functions specified in the flowchart block or blocks or in conjunction with the flowchart block or blocks.

[0176] These computer program instructions can also be stored in a computer- readable memory that can direct a computer or other programmable data processing apparatus to function in a particular manner, such that the instructions stored in the computer-readable memory produce an article of manufacture including instructions which implement the function specified in the flowchart block or blocks or in conjunction with the flowchart block or blocks.

[0177] It will be apparent to those skilled in the art that various modifications and variations can be made to the present application without departing from the spirit or scope of the application. Thus, it is intended that the present application cover the modifications and variations of this application provided they come within the scope of the appended claims and their equivalents.

Claims

1. An image processing method, characterized by, Applied to a terminal device, comprising: obtaining a first encoding code stream, the first encoding code stream comprising an encoding code stream of a plurality of components of a first original RAW image, the plurality of components comprising a first luminance component, a first color component and a first difference component; decoding the encoding code stream of the first luminance component to obtain an image of the first luminance component of the first RAW image; based on the image of the first luminance component, desensitization detection and desensitization processing are performed to obtain an image of a second luminance component; encoding the image of the second luminance component to obtain an encoding code stream of the second luminance component; based on the encoding code stream of the second luminance component, the encoding code stream of the first color component and the encoding code stream of the first difference component, a second encoding code stream is obtained; storing the second encoding code stream or sending the second encoding code stream to a cloud device.

2. The method of claim 1, wherein, The terminal device comprises a desensitization model, and the desensitization detection and desensitization processing based on the image of the first luminance component to obtain an image of a second luminance component comprises: inputting the image of the first luminance component into three input channels of the desensitization model respectively, and obtaining the image of the second luminance component after desensitization detection and desensitization processing in the desensitization model, wherein the desensitization model is trained for YUV format images or RGB format images, Y is a luminance component, U and V are color components, and R, G and B represent red, green and blue respectively.

3. The method of claim 2, wherein, The bit width depth of the image supported by the desensitization model is a target bit, and the bit width depth of the first RAW image is greater than or equal to the target bit.

4. The method of claim 3, wherein, If the bit width depth of the first RAW image is greater than the target bit, the inputting of the image of the first luminance component into the three input channels of the desensitization model respectively comprises: cutting the image of the target bit from the image of the first luminance component; inputting the image of the target bit into the three input channels of the desensitization model respectively.

5. The method according to claim 3 or 4, characterized in that, The target bit is 8 bits.

6. The method according to any one of claims 3-5, characterized in that, The bit width depth of the first RAW image is any one of the following: 8 bits, 10 bits, 12 bits or 16 bits.

7. The method according to any one of claims 1 to 6, characterized in that, The method further comprises: obtaining a second RAW image; preprocessing the second RAW image to obtain the first RAW image, the first RAW image comprising a YCoCgDg format image, Y being a luminance component, Co and Cg being color components, and Dg being a difference component; encoding the images of the plurality of components of the first RAW image respectively to obtain the first encoding code stream; storing the first encoding code stream.

8. The method of claim 7, wherein, The obtaining of the first encoding code stream comprises: obtaining the first encoding code stream from a storage medium.

9. The method according to claim 7 or 8, characterized in that, The preprocessing of the second RAW image to obtain the first RAW image comprises: performing nonlinear transformation on the second RAW image to obtain a transformed second RAW image; performing color gamut transformation on the transformed second RAW image to obtain the first RAW image.

10. An image processing apparatus characterized by comprising: Comprise: The acquisition module is configured to acquire a first encoded code stream, the first encoded code stream comprising encoded code streams of a plurality of components of a first original RAW image, the plurality of components comprising a first luminance component, a first color component, and a first difference component; The decoding module is configured to decode the encoded code stream of the first luminance component to obtain an image of the first luminance component of the first RAW image; The desensitization module is configured to perform desensitization detection and desensitization processing based on the image of the first luminance component to obtain an image of a second luminance component; The encoding module is configured to encode the image of the second luminance component to obtain an encoded code stream of the second luminance component; The second encoded code stream is obtained based on the encoded code stream of the second luminance component, the encoded code stream of the first color component, and the encoded code stream of the first difference component; The storage module is configured to store the second encoded code stream, or the sending module is configured to send the second encoded code stream to a cloud device. The desensitization module comprises a desensitization model, and the desensitization module is specifically configured to:

11. The apparatus of claim 10, wherein, input the image of the first luminance component into three input channels of the desensitization model respectively, and obtain the image of the second luminance component after desensitization detection and desensitization processing in the desensitization model, wherein the desensitization model is trained based on images in YUV format or RGB format, Y represents a luminance component, U and V represent color components, and R, G, and B represent red, green, and blue respectively. A bit width depth of an image supported by the desensitization model is a target bit, and a bit width depth of the first RAW image is greater than or equal to the target bit.

12. The apparatus of claim 11, wherein, If the bit width depth of the first RAW image is greater than the target bit, the desensitization module is specifically configured to:

13. The apparatus of claim 12, wherein, cut the image of the target bit from the image of the first luminance component; input the image of the target bit into the three input channels of the desensitization model respectively. The target bit is 8 bits.

14. The apparatus of claim 12 or 13, wherein, The bit width depth of the first RAW image is any one of the following: 8 bits, 10 bits, 12 bits, or 16 bits.

15. The apparatus of any one of claims 12-14, wherein, The acquisition unit is further configured to:

16. The apparatus of any one of claims 10-15, wherein, acquire a second RAW image; The apparatus further comprises a preprocessing module configured to pre-process the second RAW image to obtain the first RAW image, the first RAW image comprising an image in YCoCgDg format, Y representing a luminance component, Co and Cg representing color components, and Dg representing a difference component; The encoding module is further configured to encode the images of the plurality of components of the first RAW image respectively to obtain the first encoded code stream; The storage module is configured to store the first encoded code stream. The acquisition unit is specifically configured to:

17. The apparatus of claim 16, wherein, acquire the first encoded code stream from a storage medium. The preprocessing module is specifically configured to:

18. The apparatus of claim 16 or 17, wherein, perform nonlinear transformation on the second RAW image to obtain a transformed second RAW image; perform color gamut transformation on the transformed second RAW image to obtain the first RAW image. The apparatus comprises a processor coupled with a memory:

19. A computing device, comprising: ​ The processor is configured to execute computer programs or instructions stored in the memory to cause the apparatus to perform the method of any one of claims 1-9.

20. An image processing system, characterized by comprising a transceiver, a memory, and a processor; The transceiver is configured to receive and transmit data. The memory is configured to store computer program instructions and data. The processor is configured to execute computer programs or instructions stored in the memory to cause the apparatus to perform the method of any one of claims 1-9.

21. A vehicle characterized by comprising a transceiver, a memory, and a processor; 22. A computer-readable storage medium, characterized in that, The computer readable storage medium stores computer programs or instructions, which, when executed by a computer, cause the computer to perform the method of any one of claims 1-9.

23. A computer program product, characterised in that, The computer program product comprises computer programs or instructions, which, when executed on a computer, cause the computer to perform the method of any one of claims 1-9.

24. A chip, characterized by The chip comprises a processor coupled with a memory, and the processor is configured to execute computer programs or instructions stored in the memory, which, when executed, implement the method of any one of claims 1-9.

Citation Information

Patent Citations

  • Image processing method and device, equipment and storage medium

    CN114760480A

  • Video desensitization method, access system, equipment and medium

    CN115270156A

  • Image data processing method and device, equipment, storage medium and vehicle

    CN117951739A

  • Image decoding method

    KR1020030085336A