Image optimization method and apparatus, device, medium, product

By generating image masks and calling corresponding models to optimize image frames, the problem of inconsistent display effects when converting SDR images to HDR was solved, improving image display quality and user immersion.

CN116188296BActive Publication Date: 2026-05-22GUANGZHOU FANGGUI INFORMATION TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
GUANGZHOU FANGGUI INFORMATION TECHNOLOGY CO LTD
Filing Date
2022-12-22
Publication Date
2026-05-22

Smart Images

  • Figure CN116188296B_ABST
    Figure CN116188296B_ABST
Patent Text Reader

Abstract

The application relates to an image optimization method and device, equipment, medium and product, the method comprising: acquiring an original image frame; generating an image mask of the original image frame, so that the image mask corresponds to the contour boundaries and distinguishing features of each content object in the original image frame; and converting the original image frame into an optimized image frame by applying the image mask, wherein the original image in the contour boundary range is optimized according to the distinguishing features corresponding to each contour boundary in the image mask and the image optimization model corresponding to the distinguishing features. The application realizes the conversion of the original image frame into the optimized image frame by applying the corresponding image optimization model to different content objects, and the optimized image frame has a better display effect through corresponding optimization for different content objects.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the fields of image processing and live streaming technology, and in particular to an image optimization method and its apparatus, device, medium, and product. Background Technology

[0002] High Dynamic Range Imaging (HDRI or HDR) is a set of techniques used in computer graphics and cinematography to achieve a greater dynamic range of exposure (i.e., a greater difference between light and dark areas) than ordinary digital imaging techniques.

[0003] Standard video imaging (SDR, Standard Dynamic Range) can be converted to HDR to enhance image quality and obtain high-quality images.

[0004] In traditional technology, SDR to HDR conversion often involves converting the entire image at once. However, the quality of mainstream monitors on the market varies greatly, and there is no unified standard for SDR to HDR conversion parameters. Different scenarios and different targets will produce very different results when using the same parameters. Furthermore, display issues such as brightness or color differences will occur on different devices. For example, when watching live online broadcasts on a monitor, the live video stream rendered with HDR specifications has a better video viewing experience. For live streaming platforms, there is also the challenge of how to make live video streams display in HDR on monitors with varying display quality to enhance the immersive experience for users watching live broadcasts. Summary of the Invention

[0005] The purpose of this application is to solve the above-mentioned problems by providing an image optimization method and corresponding apparatus, devices, computer-readable storage media, and computer program products.

[0006] According to one aspect of this application, an image optimization method is provided, comprising the following steps:

[0007] Obtain the original image frame;

[0008] Generate an image mask for the original image frame, so that the image mask corresponds to the outline boundaries and distinguishing features of each content object in the original image frame;

[0009] The original image frame is converted into an optimized image frame by applying the image mask, wherein the original image within the contour boundary range is optimized by calling the image optimization model corresponding to the distinguishing features of each contour boundary in the image mask.

[0010] Optionally, generating the image mask for the original image frame includes:

[0011] Target recognition is performed based on the original image frame to determine the content region of the content object and the corresponding object type of the content object;

[0012] An image mask is created for the original image frame, wherein the regular contour of the content region of the content object is used to determine its contour boundary, and a distinguishing feature corresponding to the object type of the content object is preset within each contour boundary range of the image mask.

[0013] Optionally, generating the image mask for the original image frame includes:

[0014] Target recognition is performed based on the original image frame to determine the content region of the content object and the corresponding object type of the content object;

[0015] The outline boundary of the content object is determined from the original image of the content object corresponding to the content area;

[0016] Create an image mask for the original image frame, and pre-set distinguishing features corresponding to the object type of its content object within each contour boundary of the image mask.

[0017] Optionally, generating the image mask for the original image frame includes:

[0018] The distinguishing features are represented as pixel values. Each object type has a specific pixel value corresponding to its distinguishing feature. Different distinguishing features use different pixel values. The pixel values ​​used by the distinguishing features are pre-mapped with the image optimization model corresponding to the distinguishing feature.

[0019] Optionally, applying the image mask to convert the original image frame into an optimized image frame includes:

[0020] Based on the distinguishing features of each contour boundary in the image mask, multiple preset image optimization models corresponding to the distinguishing features are invoked;

[0021] The corresponding correction values ​​of pixel values ​​within the contour boundary range are calculated by the various image optimization models called, and then fused into the original image frame to obtain the optimized image frame.

[0022] Optionally, in the step of calling multiple preset image optimization models corresponding to the distinguishing features of each contour boundary in the image mask, the multiple image optimization models include any one or more of the following:

[0023] The first image optimization model is used to determine the correction values ​​corresponding to the conversion of the original image of a specific content object into the target image obtained by high dynamic range imaging.

[0024] The second image optimization model is used to determine the correction value corresponding to the target brightness effect obtained by outputting the original image of a specific content object to a standard display device;

[0025] The third image optimization model is used to determine the correction values ​​corresponding to the target color effect of the color components obtained by outputting the original image of a specific content object to a standard display device.

[0026] Optionally, after generating the image mask of the original image frame, the original image frame and the image mask are transmitted to the GPU to apply the image mask to convert the original image frame into an optimized image frame.

[0027] According to another aspect of this application, an image optimization apparatus is provided, comprising:

[0028] The image reading module is configured to acquire raw image frames;

[0029] The mask generation module is configured to generate an image mask for the original image frame, such that the image mask corresponds to the outline boundaries and distinguishing features of each content object in the original image frame.

[0030] The image conversion module is configured to apply the image mask to convert the original image frame into an optimized image frame, wherein the original image within the contour boundary range is optimized by calling the image optimization model corresponding to the distinguishing features of each contour boundary in the image mask.

[0031] According to another aspect of this application, an electronic device is provided, including a central processing unit and a memory, wherein the central processing unit is configured to invoke and run a computer program stored in the memory to perform the steps of the image optimization method described in this application.

[0032] According to another aspect of this application, a computer-readable storage medium is provided that stores, in the form of computer-readable instructions, a computer program implemented according to the image optimization method, wherein the computer program, when invoked by a computer, performs the steps included in the method.

[0033] According to another aspect of this application, a computer program product is provided, including a computer program / instructions that, when executed by a processor, implement the steps of the method described in any embodiment of this application.

[0034] Compared to existing technologies, this application generates an image mask after the original image frame, so that the image mask corresponds to the outline boundaries and distinguishing features of each content object in the original image frame. Based on the distinguishing features corresponding to each outline boundary in the image mask, the image optimization model corresponding to the distinguishing features is called to optimize the original image within the outline boundary range. This realizes the application of the corresponding image optimization model to distinguish different content objects and convert the original image frame into an optimized image frame. By performing corresponding optimization for different content objects, the content objects in the original image can be converted from the original SDR image specification to the HDR image specification for output display, so that the optimized image frame can obtain a higher quality and more delicate display effect. Attached Figure Description

[0035] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0036] Figure 1 A schematic diagram illustrating the model architecture of an image optimization implementation network exemplified in this application;

[0037] Figure 2 This is a flowchart illustrating one embodiment of the image optimization method of this application;

[0038] Figure 3 A schematic diagram of the network architecture for an exemplary live streaming application scenario of this application;

[0039] Figure 4 This is a flowchart illustrating the process of creating an image mask in one embodiment of this application;

[0040] Figure 5 This is a flowchart illustrating the process of creating an image mask in another embodiment of this application;

[0041] Figure 6 This is a schematic block diagram of the image optimization device of this application;

[0042] Figure 7 This is a schematic diagram of the structure of an electronic device used in this application. Detailed Implementation

[0043] The models cited or potentially cited in this application, including traditional machine learning models or deep learning models, can be deployed on a remote server and invoked remotely on the client, or deployed on a client with the capability to invoke directly, unless explicitly specified in the text. In some embodiments, when running on the client, the corresponding intelligence can be obtained through transfer learning in order to reduce the requirements on the client's hardware operating resources and avoid excessive occupation of the client's hardware operating resources.

[0044] Please see Figure 1 This application provides an exemplary model architecture for an end-to-end image optimization implementation network to facilitate the reader's understanding of the implementation of this application. For example... Figure 1 As shown, the image optimization implementation network includes a target recognition model, an image segmentation model, and an image transformation module.

[0045] The target recognition model is used to identify different types of content objects from a given overall original image and to label the corresponding object types of each content object. The target recognition model can be any mature model capable of image target recognition, and can be a corresponding model implemented based on machine learning, deep learning, or other feasible algorithms. For example, it can be implemented using the Yo-Lo model based on deep learning.

[0046] The image segmentation model is used to extract the contour boundaries of the original image of a given content object, and can represent the image region within the contour boundaries in the form of a mask. The image segmentation model can be any mature model that can perform image segmentation, and can be a corresponding model based on machine learning, deep learning, or other feasible algorithms such as edge detection algorithms. For example, the U-net model based on deep learning can be used for implementation.

[0047] The image transformation module is used to comprehensively determine the image mask corresponding to the overall original image where the content objects indicated by the contour boundaries of each content object are located based on the contour boundaries of each content object. Then, based on the object type of the content object indicated by each contour boundary in the image mask, one or more image optimization models of the corresponding type are called to determine the correction values ​​of the original image within the corresponding contour boundary range. These correction values ​​are then fused into the overall original image to transform the overall original image into an optimized image.

[0048] Figure 1 The image optimization implementation network described herein is for illustrative purposes only. In some embodiments, the image segmentation model can be omitted from the model architecture of this example, and the content region of the content object obtained by the target recognition model can be directly used as the outline boundary of the content object.

[0049] Based on the above explanation of principles, please refer to Figure 2 According to an image optimization method provided in this application, in one embodiment, the method includes the following steps:

[0050] Step S1100: Obtain the original image frame;

[0051] The original image frame can be image data suitable for computer processing retrieved from any source. The original image frame can originate from a single image file, a video image frame obtained by decompressing a video file, or a preview image frame generated after the camera unit of the terminal device has started recording.

[0052] In one exemplary application scenario, after the camera unit of the terminal device is activated, the camera unit begins to capture environmental images to generate a preview video stream, triggering the execution of the image optimization method of this application. This involves reading preview image frames from the preview video stream as raw image frames for optimization. After obtaining optimized image frames, these frames are output to the display unit of the terminal device for display. Simultaneously, each preview image frame can be optimized as a corresponding raw image frame to generate a corresponding optimized image frame. The video stream composed of these optimized image frames is the optimized video stream obtained after image optimization based on the preview video stream, resulting in a higher quality display.

[0053] In another exemplary application scenario, it can be done in such as Figure 3 In the media server for the live streaming, the video image frames obtained after decoding the live video stream submitted by the broadcaster trigger the image optimization method of this application. This method decodes and optimizes each video image frame in the live video stream to obtain its corresponding optimized image frame. Based on the optimized image frame, it is then encoded and output to the terminal devices of the viewers in the live streaming room, maintained by the application server, so that the viewers can obtain the corresponding optimized video stream and achieve a better display effect. Of course, as an alternative equivalent method, the image optimization method of this application can also be applied to the video image frames in the live stream of the broadcaster in the live streaming room on the viewers' terminal devices, which is essentially equivalent to optimization on the media server.

[0054] In another example application scenario, a video file in any storage space can be called, decoded, and then the image optimization method of this application can be applied to each video image frame to obtain the corresponding optimized image frame. Then, it can be re-stored as a new video file or the original video file can be replaced to obtain the optimized video file.

[0055] It is easy to understand that the original image frame can be image data from any channel and in any form. The image optimization method of this application can be triggered in any business scenario to obtain the corresponding original image frame through this step, so as to perform optimization processing.

[0056] Step S1200: Generate an image mask for the original image frame, so that the image mask corresponds to the outline boundaries and distinguishing features of each content object in the original image frame;

[0057] The original image frame may contain one or more content objects. These content objects are image content perceptible to the human eye, referring to objective objects in the physical world, which can be any concrete or abstract object. Concrete objects can be colors, such as yellow, green, red, etc.; abstract objects can be a type of natural scenery, such as faces, bodies, food, plants, sky, animals, buildings, etc. Classification criteria can be pre-defined for various content objects to determine different object types, allowing for differentiated processing based on the corresponding object type.

[0058] For the original image frame, any feasible image target recognition method can be applied to identify various content objects within it, thereby determining these content objects and their corresponding object types. Based on the area occupied by these content objects in the original image frame, their contour boundaries can be determined. These contour boundaries can be regular forms such as rectangles or arbitrary shapes matching the outer contours of the content objects. Generally, the area defined by the contour boundary can cover all the pixels occupied by the content object. Therefore, each contour boundary actually marks a corresponding image region of a content object.

[0059] As mentioned earlier, each content object has its corresponding object type. To facilitate indexing, each content object can be converted into a corresponding distinguishing feature, so that the distinguishing feature can be embedded in the image mask corresponding to the original image frame. Furthermore, the corresponding object type can be identified at the computer level through the distinguishing feature.

[0060] After determining the outline boundaries and distinguishing features of each content object in the original image frame, an image mask corresponding to the original image frame can be constructed so that the specific area of ​​the original image frame occupied by each content object and the individual pixels it covers, as well as the object type corresponding to these specific areas, can be indicated through the image mask.

[0061] Based on the above principles, a step can be set here: the distinguishing features are represented as pixel values, each object type has a distinguishing feature corresponding to a specific pixel value, different distinguishing features use different pixel values, and the pixel values ​​used by the distinguishing features are pre-mapped with the image optimization model corresponding to the distinguishing feature.

[0062] In one embodiment, in the image mask, each pixel within the range corresponding to the contour boundary of each content object is used to represent the distinguishing feature corresponding to the object type to which the content object pointed to by the contour boundary belongs. For example, for the first object type, the RGB values ​​of the pixels within its corresponding contour boundary range are represented as (1,0,0); for the second object type, the RGB values ​​of the pixels within its corresponding contour boundary range are represented as (2,0,0), and so on. This achieves the representation of different object types with different pixel values, making these pixel values ​​themselves carriers of the distinguishing features of different content objects. The mapping relationship data between different object types and their corresponding distinguishing feature values ​​can be preset for direct retrieval.

[0063] As can be seen, the image mask generated in the above manner is the same size as the original image frame. Moreover, for each content object identified in the original image frame, each pixel point covered by the contour boundary of the content object is determined by the contour boundary of the content object. Each pixel point covered by each contour boundary is represented as a pixel value corresponding to the object type of the content object pointed to by the contour boundary, which serves as a distinguishing feature that facilitates direct recognition by the computer.

[0064] In one embodiment, the image mask may be a single mask that represents the outline boundaries and distinguishing features of all content objects. In an alternative embodiment, the image mask may also comprise multiple specific image masks, each maintaining the same size as the original image frame, but each specific image mask only represents the outline boundaries and distinguishing features of a single content object within the total content objects. Those skilled in the art can implement this as needed.

[0065] Step S1300: Apply the image mask to convert the original image frame into an optimized image frame, wherein the image optimization model corresponding to the distinguishing features of each contour boundary in the image mask is called to optimize the original image within the contour boundary range.

[0066] After obtaining the image mask, the original image frame can be optimized based on the image mask to convert it into an optimized image frame, thereby improving its image quality and making the display effect better.

[0067] When performing image conversion on the original image frame, the corresponding content object is identified through the distinguishing features in the image mask. Then, an image optimization model with a pre-established one-to-one mapping relationship with the distinguishing features is obtained to adjust the parameters of each pixel containing the distinguishing feature. The corresponding pixel values ​​after parameter adjustment are used as the pixel values ​​of the corresponding optimized image frame, thereby optimizing the original image of the content object corresponding to the distinguishing feature.

[0068] In one embodiment, since each distinguishing feature represents an object type, the image optimization model corresponding to the corresponding object type can be determined through the distinguishing features. That is, each content object can determine its corresponding image optimization model based on the distinguishing features it represents. The image optimization model is provided according to the object type to which the content object belongs. Therefore, for content objects of different object types, their respective image optimization models can be applied for corresponding optimization.

[0069] To this end, we can pre-define the image optimization model corresponding to each object type, and then establish a mapping relationship between the image optimization type and its corresponding distinguishing features. Subsequently, we can directly call its corresponding image optimization model based on the distinguishing features.

[0070] In one embodiment, different image optimization models corresponding to different optimization parameter types can be preset for each object type according to the different optimization parameter types to be implemented, so that each parameter type has a corresponding image optimization model. When the image optimization model needs to be called, one or more of them can be called according to actual needs, so as to realize the parameter tuning control of the corresponding parameter type.

[0071] The image optimization model can be pre-modeled and implemented using corresponding preset formulas and algorithms. The image optimization model can be modeled and implemented as a machine learning model or a deep learning model, trained until convergence, and then put into use. The image optimization model can be used to determine the corresponding correction value of the pixel value based on the pixel value of the pixel point in the original image of the corresponding specific content object, and use the correction value to adjust the original pixel value to obtain the pixel value of the corresponding pixel point in the optimized image frame, thereby realizing pixel-level point-by-point optimization of the corresponding content object.

[0072] In one embodiment, several image optimization models are provided in advance to correspond to different parameter types, allowing for convenient on-demand use of one or more of them to optimize pixels. The provided image optimization models include: a first image optimization model for determining the correction value corresponding to the target image obtained by converting the original image of a specific content object into a high dynamic range imaging image; a second image optimization model for determining the correction value corresponding to the target brightness effect obtained by outputting the original image of the specific content object to a standard display device; and a third image optimization model for determining the correction value corresponding to the target color effect of the color components obtained by outputting the original image of the specific content object to a standard display device.

[0073] To obtain the first image optimization model, in one embodiment, various control parameters corresponding to high dynamic range imaging generated during the conversion of the image of each type of content object from SDR to HDR can be tested in advance, as well as the effect of the converted image based on these control parameters. Then, the control parameters with excellent conversion effect are recorded, and a first conversion function H(x,y) is constructed based on these control parameters, the original image, and the result image to model and obtain the first image optimization model, which can convert the pixel value x of the original image into the corresponding corrected value y' and thus obtain the corrected pixel value y.

[0074] In order to obtain the second image optimization model, in one embodiment, the original image is output to different display devices to test its display effect, mainly the brightness display effect. The brightness difference between different display devices and a preset standard display device is calculated. Based on this brightness difference, the brightness correction function is set as L(x,y). Then, modeling is performed to obtain the second image optimization model, which can convert the pixel value x of the original image into the corresponding correction value y' and thus obtain the corrected pixel value y.

[0075] To obtain the third image optimization model, in one embodiment, the original image is output to different display devices to test its display effect, including the target color display effect corresponding to each color component. The differences between the different display devices and the preset standard display device for each color component are calculated, i.e., color difference. Based on this color difference, a color difference correction function is set as C(x,y). For each specific color component, such as the red, green, and blue channels, it can be represented as R(x,y), G(x,y), and B(x,y), respectively. Then, a model is built for each color component to obtain a third image optimization model corresponding to each color component. This model can convert the pixel value x of the original image into the corresponding corrected value y', thereby obtaining the corrected pixel value y. Therefore, the third image optimization model can include multiple basic models set for multiple color components.

[0076] After preparing the image optimization model based on the above process, the optimization of the original image frame can be achieved according to the following process:

[0077] Step S1310: Based on the distinguishing features of each contour boundary in the image mask, call up multiple preset image optimization models corresponding to the distinguishing features;

[0078] Step S1320: The corresponding correction values ​​of the pixel values ​​within the contour boundary range are calculated by the respective image optimization models and fused into the original image frame to obtain the optimized image frame.

[0079] In one embodiment, when applying the image mask to optimize and transform the original image frame, it can be implemented directly according to the following formula:

[0080] New(x,y)=F(x,y)*H(x,y)*L(x,y)+(R(x,y)+G(x,y)+B(x,y))

[0081] in:

[0082] F(x,y) is the SDR source image, i.e., the original image frame;

[0083] H(x,y) is the SDR to HDR conversion function for the corresponding object type. The image optimization model corresponding to the H(x,y) function is different for different object types.

[0084] L(x,y) is the brightness correction function for the corresponding object type on the current device. It can correct the brightness to normal brightness or reduce the brightness difference. The image optimization model corresponding to the L(x,y) function is different for different object types.

[0085] R(x,y), G(x,y), and B(x,y) are the correction functions corresponding to the RGB color components of the corresponding object type on the current device. They can correct color differences to normal or reduce color differences. The image optimization models corresponding to the R(x,y), G(x,y), and B(x,y) functions are different for different object types.

[0086] New(x,y) is the final HDR image.

[0087] As can be seen, by applying the above formula, based on the contour boundaries of each content object provided by the image mask, and applying the image optimization models corresponding to the distinguishing features represented in each contour boundary, the correction values ​​corresponding to different image optimization models can be determined. Then, by fusing with the original image frame, an optimized image frame can be obtained, which can achieve a better display effect.

[0088] In the above embodiments, a third image optimization model corresponding to each color component is constructed for each object type based on each color component. By performing more refined optimization on the original image frame according to the color component, the image quality and display effect of the obtained optimized image frame can be further improved.

[0089] In one embodiment, step S1300 can be deployed to run on the GPU of the terminal device. After the original image frame is determined in step S1100 and the image mask of the original image frame is determined in step S1200, the original image frame and the image mask are passed to the GPU, and the GPU runs the corresponding instruction set to optimize the original image frame according to the information provided by the image mask and generate the corresponding optimized image frame.

[0090] As can be seen from the above embodiments, this application generates an image mask after the original image frame, so that the image mask corresponds to the outline boundary and distinguishing features of each content object in the original image frame. According to the distinguishing features corresponding to each outline boundary in the image mask, the image optimization model corresponding to the distinguishing features is called to optimize the original image within the outline boundary range, thereby realizing the application of the corresponding image optimization model to distinguish different content objects and convert the original image frame into an optimized image frame. By performing corresponding optimization for different content objects, the optimized image frame obtains a higher quality and more delicate display effect.

[0091] Secondly, this method can be applied to online live streaming services. By using image masks and different image optimization models, the video frame images in the live video stream can be optimized accordingly for different content objects in the video frame images. For example, key live content such as the anchor, users, or products in the live video frame images can be identified as content objects for targeted image optimization. This allows the key live content to be converted from the original SDR image specifications to HDR image specifications for output display, thereby improving the image quality of the key live content output to the viewer's monitor. This makes the key live content stand out in the live video stream and enhances the viewer's immersive experience of watching the live video.

[0092] Based on any embodiment of this application, please refer to Figure 4 Generating the image mask for the original image frame includes:

[0093] Step S1211: Based on the original image frame, perform target recognition to determine the content area of ​​the content object and the corresponding object type of the content object;

[0094] To generate the image mask for the original image frame, target recognition can be performed on the original image frame first, for example, using methods such as... Figure 1The target recognition model in the model architecture shown inputs the original image frame into the target recognition model, identifies the content objects of various object types in it, and obtains the candidate boxes of these content objects and their corresponding object types. The candidate box can be represented as the coordinate information of each corner point of the candidate box in the original image frame, thus actually defining a rectangular area, that is, the content area of ​​the corresponding content object, which covers all the image content of the corresponding content object.

[0095] Step S1212: Create an image mask for the original image frame, wherein the regular contour of the content region of the content object is used to determine its contour boundary, and a distinguishing feature corresponding to the object type of the content object is preset within each contour boundary range of the image mask.

[0096] In this embodiment, after obtaining the content regions of each content object in the original image frame, the line connecting the coordinate information of the four corner points of each content region can be used as the outline boundary of the corresponding content object. Based on this, the corresponding distinguishing features are determined according to the representation method corresponding to the object type of each content object. In an image mask of the same size created in advance corresponding to the original image frame, the pixel value of each pixel point within the outline boundary range is preset as the distinguishing feature, so that the distinguishing feature is consistent with the object type of the content object pointed to by the outline boundary.

[0097] It is easy to understand that since each pixel within the outline boundary of each content object is represented by a distinguishing feature, the subsequent image conversion optimization can be performed by applying the image optimization model corresponding to the distinguishing feature to each pixel, resulting in a more refined optimization effect.

[0098] In the above embodiments, the outline boundary of the content object is defined directly using the candidate box obtained by the target recognition model, which eliminates the need to identify the edge of the content object. This can improve the processing efficiency during image optimization, save system overhead, and quickly obtain optimized image frames.

[0099] Based on any embodiment of this application, please refer to Figure 5 Generating the image mask for the original image frame includes:

[0100] Step S1221: Based on the original image frame, perform target recognition to determine the content area of ​​the content object and the corresponding object type of the content object;

[0101] Similarly, in order to generate the image mask of the original image frame, target recognition can be performed on the original image frame first, for example, using methods such as... Figure 1The target recognition model in the model architecture shown inputs the original image frame into the target recognition model, identifies the content objects of various object types in it, and obtains the candidate boxes of these content objects and their corresponding object types. The candidate box can be represented as the coordinate information of each corner point of the candidate box in the original image frame, thus actually defining a rectangular area, that is, the content area of ​​the corresponding content object, which covers all the image content of the corresponding content object.

[0102] Step S1222: Determine the outline boundary of the content object from the original image of the content object corresponding to the content area;

[0103] In one embodiment, in order to utilize Figure 1 The image segmentation model shown determines the corresponding contour boundaries of each content object. It can crop the original image frame using the coordinate information of its corresponding candidate boxes for each content region, obtaining the original image of each content object. Taking a mature model like U-Net as an example, the image segmentation model is pre-trained to convergence, enabling it to identify the edge contours of content objects from a given original image and output a mask representation. Accordingly, the original images of each content object obtained from the original image frame are input into the image segmentation model to obtain the corresponding edge contours of each content object, serving as the corresponding contour boundaries. Using a deep learning model to determine the contour boundaries of the content objects is relatively accurate.

[0104] In another embodiment, the image segmentation model is implemented using an edge detection algorithm such as the Canny edge detection algorithm. Therefore, by applying the corresponding edge detection algorithm to perform edge detection on the original image of each content object, the corresponding contour boundaries can also be obtained. Using an edge detection algorithm to determine the contour boundaries of the content objects is more efficient and faster.

[0105] Step S1223: Create an image mask for the original image frame, and pre-set distinguishing features corresponding to the object type of its content object within each contour boundary of the image mask.

[0106] Once the outline boundaries of each content object in the original image frame are determined, an image mask for the original image frame can be created to perform image optimization on the original image frame.

[0107] In one embodiment, when the image segmentation model can output image masks corresponding to each content object, the pixel values ​​of the pixels within the contour boundaries can be modified to reflect the object type of the content object indicated by the contour boundary, based on the image mask directly output by the image segmentation model. Then, these image masks with pre-set distinguishing features can be used to optimize each content object in the original image frame separately, or the image masks of these multiple content objects can be merged into a single image mask, and then all content objects in the original image frame can be optimized at once based on this image mask.

[0108] In another embodiment, regardless of whether the image segmentation model obtains the contour boundaries of each content object separately or detects the contour boundaries of all multiple content objects at once, a single image mask can be created, in which corresponding distinguishing features are preset for the contour boundaries of each content object, and the entire image mask containing the distinguishing features of all content objects is obtained at once. Then, based on these image masks, all content objects in the original image frame are fully optimized.

[0109] When multiple content objects in the same original image frame need to have their edge contours determined, in one embodiment, multiple threads can be created for each content object. Each thread calls the image segmentation model to obtain the edge contours of each content object in parallel, thereby shortening the time required to determine the contour boundaries when there are multiple content objects. In this way, when optimizing each video image frame of a video stream, the optimization time can be significantly reduced. In applications such as live streaming, this provides a time efficiency advantage, allowing end-users to perceive no time consumption in image optimization while still obtaining high-quality image results.

[0110] According to the above embodiments, after the original image frame obtains the content area and corresponding object type of each content object through target recognition, further image segmentation methods can be used to accurately identify the edge contour of each content object, determine its contour boundary based on its real edge contour, and then accurately optimize the original image based on the contour boundary. The resulting optimized image frame has a more realistic and delicate image quality and display effect.

[0111] Please see Figure 6An image optimization apparatus according to one aspect of this application includes an image reading module 1100, a mask generation module 1200, and an image conversion module 1300, wherein: the image reading module 1100 is configured to acquire an original image frame; the mask generation module 1200 is configured to generate an image mask of the original image frame, such that the image mask corresponds to the contour boundaries and distinguishing features of each content object in the original image frame; and the image conversion module 1300 is configured to apply the image mask to convert the original image frame into an optimized image frame, wherein the original image within the contour boundary range is optimized by calling an image optimization model corresponding to the distinguishing features of each contour boundary in the image mask.

[0112] Based on any embodiment of this application, the mask generation module 1200 includes: a first target recognition unit, configured to perform target recognition based on the original image frame, and determine the content area of ​​the content object and the corresponding object type of the content object; and a first mask creation unit, configured to create an image mask of the original image frame, wherein the regular contour of the content area of ​​the content object is used to determine its contour boundary, and distinguishing features corresponding to the object type of the content object are preset within each contour boundary range of the image mask.

[0113] Based on any embodiment of this application, generating an image mask for the original image frame includes: a second target recognition unit, configured to perform target recognition based on the original image frame, and determine the content region of the content object and the corresponding object type of the content object; a second image segmentation unit, configured to determine the contour boundary of the content object from the original image of the content object corresponding to the content region; and a second mask creation unit, configured to create an image mask for the original image frame, and preset distinguishing features corresponding to the object type of the content object within each contour boundary range of the image mask.

[0114] Based on any embodiment of this application, the mask generation module 1200 includes: a feature identification unit, configured to represent the distinguishing features as pixel values, wherein each object type of distinguishing feature corresponds to a certain pixel value, and different distinguishing features use different pixel values, and the pixel values ​​used by the distinguishing features are pre-mapped with the image optimization model corresponding to the distinguishing feature.

[0115] Based on any embodiment of this application, the image conversion module 1300 includes: a model matching unit, configured to call multiple preset image optimization models corresponding to the distinguishing features of each contour boundary in the image mask; and an image fusion unit, configured to calculate the corresponding correction values ​​of pixel values ​​within the corresponding contour boundary range by each called image optimization model, and fuse them into the original image frame to obtain an optimized image frame.

[0116] Based on any embodiment of this application, the plurality of image optimization models include any one or more of the following: a first image optimization model, used to determine the correction value corresponding to the target image obtained by converting the original image of a specific content object into a high dynamic range imaging target image; a second image optimization model, used to determine the correction value corresponding to the target brightness effect obtained by outputting the original image of the specific content object to a standard display device; and a third image optimization model, used to determine the correction value corresponding to the target color effect of the color components obtained by outputting the original image of the specific content object to a standard display device.

[0117] Based on any embodiment of this application, the image conversion module 1300 runs in a GPU, which receives the original image frame and the image mask to apply the image mask to convert the original image frame into an optimized image frame.

[0118] Another embodiment of this application also provides an electronic device. For example... Figure 7 The diagram shows the internal structure of an electronic device. This electronic device includes a processor, a computer-readable storage medium, a memory, and a network interface connected via a system bus. The computer-readable storage medium stores an operating system, a database, and computer-readable instructions. The database can store information sequences, and when executed by the processor, the computer-readable instructions enable the processor to implement an image optimization method.

[0119] The processor of this electronic device provides computing and control capabilities to support the operation of the entire device. The memory of this electronic device can store computer-readable instructions, which, when executed by the processor, cause the processor to perform the image optimization method of this application. The network interface of this electronic device is used for communication with a terminal.

[0120] Those skilled in the art will understand that Figure 7 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the electronic device to which the present application is applied. The specific electronic device may include more or fewer components than shown in the figure, or combine certain components, or have different component arrangements.

[0121] In this embodiment, the processor is used to execute... Figure 6 The specific functions of each module are described, and the memory stores the program code and various data required to execute the aforementioned modules or sub-modules. The network interface is used to enable data transmission between user terminals or the server. In this embodiment, the computer-readable storage medium stores the program code and data required to execute all modules in the image optimization apparatus of this application, and the server can call the server's program code and data to execute the functions of all modules.

[0122] This application also provides a computer-readable storage medium storing computer-readable instructions, which, when executed by one or more processors, cause the one or more processors to perform the steps of the image optimization method of any embodiment of this application.

[0123] This application also provides a computer program product, including a computer program / instructions that, when executed by one or more processors, implement the steps of the method described in any embodiment of this application.

[0124] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments of this application can be implemented by a computer program instructing related hardware. This computer program can be stored in a computer-readable storage medium, and when executed, it can include the processes of the embodiments of the methods described above. The aforementioned storage medium can be a magnetic disk, optical disk, read-only memory (ROM), or random access memory (RAM), etc.

[0125] In summary, this application achieves the conversion of original image frames into optimized image frames by applying corresponding image optimization models to different content objects. By performing corresponding optimizations for different content objects, the optimized image frames achieve a higher quality and more delicate display effect.

Claims

1. An image optimization method, characterized in that, include: Obtain the original image frame; Generating an image mask for the original image frame, such that the image mask corresponds to the contour boundaries and distinguishing features of each content object in the original image frame, includes: performing target recognition based on the original image frame to determine the content region of the content object and the corresponding object type of the content object; determining the contour boundary of the content object from the original image of the content object corresponding to the content region; creating an image mask for the original image frame, and pre-setting distinguishing features corresponding to the object type of the content object within each contour boundary range of the image mask; representing the distinguishing features as pixel values, with each object type's distinguishing feature corresponding to a specific pixel value, and different distinguishing features using different pixel values, wherein the pixel values ​​used by the distinguishing features are pre-mapped with the image optimization model corresponding to the distinguishing feature; The original image frame is converted into an optimized image frame by applying the image mask, wherein the original image within the contour boundary range is optimized by calling the image optimization model corresponding to the distinguishing features of each contour boundary in the image mask.

2. The image optimization method according to claim 1, characterized in that, The process of converting the original image frame into an optimized image frame using the image mask includes: Based on the distinguishing features of each contour boundary in the image mask, multiple preset image optimization models corresponding to the distinguishing features are invoked; The corresponding correction values ​​of pixel values ​​within the contour boundary range are calculated by the various image optimization models called, and then fused into the original image frame to obtain the optimized image frame.

3. The image optimization method according to claim 2, characterized in that, In the step of calling multiple preset image optimization models corresponding to the distinguishing features of each contour boundary in the image mask, the multiple image optimization models include any one or any combination of the following: The first image optimization model is used to determine the correction values ​​corresponding to the conversion of the original image of a specific content object into the target image obtained by high dynamic range imaging. The second image optimization model is used to determine the correction value corresponding to the target brightness effect obtained by outputting the original image of a specific content object to a standard display device; The third image optimization model is used to determine the correction values ​​corresponding to the target color effect of the color components obtained by outputting the original image of a specific content object to a standard display device.

4. The image optimization method according to any one of claims 1 to 3, characterized in that, After generating the image mask for the original image frame, the original image frame and the image mask are transmitted to the GPU to apply the image mask to convert the original image frame into an optimized image frame.

5. An image optimization device, characterized in that, include: The image reading module is configured to acquire raw image frames; The mask generation module is configured to generate an image mask for the original image frame, such that the image mask corresponds to the contour boundaries and distinguishing features of each content object in the original image frame. This includes: performing target recognition based on the original image frame to determine the content region of the content object and the corresponding object type; determining the contour boundary of the content object from the original image of the content object corresponding to the content region; creating an image mask for the original image frame, and pre-setting distinguishing features corresponding to the object type of the content object within each contour boundary range of the image mask; representing the distinguishing features as pixel values, with each object type's distinguishing feature corresponding to a specific pixel value, and different pixel values ​​used for different distinguishing features. The pixel values ​​used for the distinguishing features are pre-mapped with an image optimization model corresponding to that distinguishing feature. The image conversion module is configured to apply the image mask to convert the original image frame into an optimized image frame, wherein the original image within the contour boundary range is optimized by calling the image optimization model corresponding to the distinguishing features of each contour boundary in the image mask.

6. The image optimization apparatus according to claim 5, characterized in that, The image conversion module includes: The model matching unit is configured to call multiple preset image optimization models corresponding to the distinguishing features of each contour boundary in the image mask. The image fusion unit is configured to calculate the corresponding correction values ​​of pixel values ​​within the contour boundary range of each called image optimization model, and fuse them into the original image frame to obtain an optimized image frame.

7. The image optimization apparatus according to claim 6, characterized in that, The plurality of image optimization models include any one or more of the following: The first image optimization model is used to determine the correction values ​​corresponding to the conversion of the original image of a specific content object into the target image obtained by high dynamic range imaging. The second image optimization model is used to determine the correction value corresponding to the target brightness effect obtained by outputting the original image of a specific content object to a standard display device; The third image optimization model is used to determine the correction values ​​corresponding to the target color effect of the color components obtained by outputting the original image of a specific content object to a standard display device.

8. The image optimization apparatus according to any one of claims 5 to 7, characterized in that, The image conversion module runs on the GPU and receives the original image frame and the image mask to apply the image mask to convert the original image frame into an optimized image frame.

9. An electronic device comprising a central processing unit and a memory, characterized in that, The central processing unit is used to invoke and run a computer program stored in the memory to perform the steps included in the method as described in any one of claims 1 to 4.

10. A computer-readable storage medium, characterized in that, It stores a computer program in the form of computer-readable instructions, which, when invoked by a computer, performs the steps included in the method as described in any one of claims 1 to 4.