Texture enhancement method and electronic device
Patent Information
- Application Number
- CN202311295587.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-09-28
- Publication Date
- 2026-09-11
- Estimated Expiration
- 2043-09-28
AI Technical Summary
然而,不同视频图像均采用相同的纹理调整参数,不具有自适应性,可能会出现调整过度或调整过小的问题,导致视频显示效果较差
Smart Images

Figure CN119762385B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of image processing technology, and in particular to a texture enhancement method and electronic device. Background Technology
[0002] Video is one of the ways to record and store information, allowing users to access information more intuitively and conveniently. With the development of technology, video is widely used in various fields such as life, work, and education. At the same time, people's requirements for video quality are getting higher and higher.
[0003] Currently, to improve video quality, electronic devices, after receiving images from a video stream, need to adjust the pixel values of each pixel in the image using specified texture adjustment parameters to enhance the image's texture clarity. However, using the same texture adjustment parameters for different video images lacks adaptability and may result in over-adjustment or under-adjustment, leading to poor video display quality. Summary of the Invention
[0004] In view of this, this application provides a texture enhancement method and an electronic device for improving video quality and thus video display effect.
[0005] Firstly, this application provides a texture enhancement method. After acquiring a video stream, the electronic device, for each frame of the first image in the video stream, obtains a scene detection result corresponding to the first image based on the difference between the pixel value of the target channel of the first image and the pixel value of the target channel of the previous frame of the first image; wherein, the scene detection result corresponding to the first image indicates whether the first image and the previous frame belong to the same scene, that is, it indicates whether the scene corresponding to the first image has changed.
[0006] If the scene detection result indicates that the first image and the previous frame image belong to the same scene, it means that the scene corresponding to the first image has not changed. The electronic device can use the texture adjustment parameters corresponding to the previous frame image as the texture adjustment parameters corresponding to the first image. The texture adjustment parameters corresponding to the previous frame image are texture adjustment parameters that match the scene of the previous frame image.
[0007] If the scene detection result indicates that the first image and the previous frame do not belong to the same scene, it means that the scene corresponding to the first image has changed. The electronic device can re-determine the texture adjustment parameters corresponding to the first image based on the first image, that is, determine that the texture adjustment parameters corresponding to the first image are texture adjustment parameters that match the scene of the first image.
[0008] Based on the texture adjustment parameters corresponding to the first image, the clarity of the high-frequency information of the first image is adjusted, that is, the texture clarity, to obtain the target first image. Then, the electronic device can display the target first image.
[0009] The fact that the first image and the previous frame belong to the same scene can be understood as the content (or objects) in these two frames being relatively similar, or in other words, the display effects of these two frames being relatively similar.
[0010] In this application, upon receiving the current frame image (i.e., the first image), it is determined whether the first image and the previous frame image belong to the same scene. If they belong to the same scene, the electronic device can use the texture adjustment parameters corresponding to the previous frame image as the texture adjustment parameters corresponding to the first image, without needing to re-predict the texture adjustment parameters, thus achieving rapid determination of the texture adjustment parameters and improving the processing efficiency of the video image. If they do not belong to the same scene, the electronic device can re-predict the texture adjustment parameters corresponding to the first image to obtain texture adjustment parameters that match the texture scene of the first image. The electronic device can then use these texture adjustment parameters to adjust the clarity of the high-frequency information of the first image, achieving adaptive texture adjustment, improving the quality of the video image, and thus improving the texture display effect of the video. Furthermore, since only the clarity of the high-frequency information of the first image is enhanced, the amount of data processing can be reduced while ensuring the display effect of the first image, thus guaranteeing the real-time processing of the video image. Also, since there is an error in the prediction of texture adjustment parameters, meaning that for the same scene, the predicted texture adjustment parameters may have errors. Therefore, when the current frame and the previous frame belong to the same scene, directly using the texture adjustment parameters corresponding to the previous frame to adjust the current frame can avoid the difference in texture display effects between two adjacent frames belonging to the same scene due to errors, thus ensuring the continuity of video frame display effects and improving the user's visual experience.
[0011] In one possible design approach, the process of determining the difference between the pixel values of the target channel of the first image and the pixel values of the target channel of the previous frame image may include:
[0012] For each pixel in the first image, the electronic device can calculate the difference between the pixel value of the target channel of that pixel and the pixel value of the target channel of the corresponding pixel in the previous frame image, thus obtaining the difference of the target channel corresponding to that pixel in the first image. In other words, for each first pixel in the first image, the electronic device can calculate the difference between the pixel value of the target channel of that first pixel and the pixel value of the target channel of the corresponding second pixel in the previous frame image, thus obtaining the difference of the target channel corresponding to that first pixel. The position of the first pixel in the first image is the same as the position of the corresponding second pixel in the previous frame image.
[0013] The electronic device can obtain the difference between the pixel value of the target channel of the first image and the pixel value of the target channel of the previous frame image based on the average of the differences of the target channels corresponding to each first pixel in the first image.
[0014] In this application, the difference between the pixel values of the target channel of the first image and the pixel values of the target channel of the previous image is determined based on the difference between the pixel values of the target channel of the current frame and the previous frame image. This determines the scene detection result, thereby ensuring the accuracy of the scene detection result by determining the overall difference between two adjacent frames.
[0015] In one possible design approach, the number of target channels is one or more. Accordingly, the method of obtaining the difference between the pixel values of the target channels in the first image and the pixel values of the target channels in the previous frame image based on the average of the differences between the target channels corresponding to each pixel in the first image includes:
[0016] For each target channel, the electronic device can calculate the average of the differences between the target channels corresponding to each pixel, thus obtaining the difference between the target channels of the first image. Then, the electronic device can calculate the average of the differences between the target channels of the first image, thus obtaining the difference between the pixel value of the target channel in the first image and the pixel value of the target channel in the previous frame image.
[0017] In this application, the electronic device can determine the difference between the pixel value of the target channel of the first image and the pixel value of the target channel of the previous frame image based on the difference between different target channels, that is, determine the scene detection result, realize the degree of difference between two frames of images from different dimensions, and ensure the accuracy of the scene detection result.
[0018] In one possible design approach, the process of obtaining the scene detection result corresponding to the first image based on the difference between the pixel value of the target channel of the first image and the pixel value of the target channel of the previous frame of the first image may include:
[0019] If the difference between the pixel value of the target channel of the first image and the pixel value of the target channel of the previous frame image is less than or equal to the first threshold, it indicates that the difference in pixel values between the two frames is small, that is, the difference in the displayed object is small. The electronic device can determine that the scene detection result corresponding to the first image indicates that the first image and the previous frame image belong to the same scene, that is, the scene of the first image has not changed.
[0020] If the difference between the pixel value of the target channel of the first image and the pixel value of the target channel of the previous frame image is greater than the first threshold, it indicates that the pixel value difference between the two frames is large, that is, the display object is large. The scene detection result corresponding to the first image indicates that the first image and the previous frame image do not belong to the same scene, that is, the scene of the first image has changed, thereby achieving accurate scene detection of the first image.
[0021] In one possible design approach, the target channels mentioned above include at least one of the hue (H) channel, saturation (S) channel, and brightness (V) channel.
[0022] In one possible design approach, the process of adjusting the sharpness of the high-frequency information of the first image may include: adjusting the pixel values of each pixel (or high-frequency pixel) in the high-frequency information of the first image based on the texture adjustment parameters to obtain the texture adjustment value of each high-frequency pixel in the high-frequency information.
[0023] For each pixel in the first image (i.e., the first pixel), the texture-enhanced pixel value of the first pixel is obtained based on the pixel value of the first pixel and the texture adjustment value corresponding to the high-frequency pixel in the high-frequency information. The position of the first pixel in the first image is the same as the position of the high-frequency pixel corresponding to the first pixel in the high-frequency information of the first image, thereby achieving texture enhancement of the high-frequency information in the first image.
[0024] In one possible design approach, since the texture is mainly related to the pixel value of the luminance Y channel, the texture adjustment value of the aforementioned pixel (that is, the texture adjustment value corresponding to the high-frequency pixel) includes the texture adjustment value of the luminance Y channel of the high-frequency pixel. Correspondingly, the pixel value after texture enhancement includes the pixel value of the Y channel after texture enhancement.
[0025] The above-described texture enhancement process for pixels in the first image may include:
[0026] The texture adjustment value of the Y channel of each high-frequency pixel in the high-frequency information is obtained by multiplying the pixel value of the Y channel of each high-frequency pixel in the high-frequency information by the texture adjustment parameter.
[0027] For each first pixel in the first image, the pixel value of the Y channel of the first pixel and the texture adjustment value of the Y channel of the corresponding high-frequency pixel are calculated to obtain the texture-enhanced Y channel pixel value of the first pixel, thereby achieving accurate enhancement of the high-frequency information of the first image and ensuring the texture enhancement effect of the high-frequency information.
[0028] In one possible design approach, the process of determining the high-frequency information of the first image may include:
[0029] An electronic device can perform mean filtering on a first image to obtain a first filtered image. Then, for each pixel in the first image, the pixel value of the Y channel of that pixel is subtracted from the pixel value of the Y channel of the corresponding pixel in the first filtered image to obtain the high-frequency information of the first image. In this application, using mean filtering to determine the high-frequency information of the first image can improve the determination speed of high-frequency information while ensuring the accuracy of the determined high-frequency information.
[0030] In one possible design approach, the process of determining the texture color adjustment parameters corresponding to the first image may include:
[0031] The electronic device can use the first image as the input parameter of the image processing model, run the image processing model, and obtain the texture adjustment parameters corresponding to the first image. The image processing model is a lightweight model that can accurately predict the texture adjustment parameters that match the scene of the first image. Furthermore, because the image processing model is a lightweight model, it has high performance, fast image processing speed, and can reduce power consumption.
[0032] In one possible design approach, to improve image processing speed, the electronic device can downsample the first image to obtain a downsampled first image. Then, based on the difference between the pixel values of the target channel in the downsampled first image and the pixel values of the target channel in the previous downsampled frame of the first image, the scene detection result corresponding to the first image is obtained, enabling rapid determination of the scene detection result.
[0033] In a second aspect, this application provides an electronic device, the electronic device including a display screen, a memory and one or more processors; the display screen, the memory and the processor are coupled; the display screen is used to display an image generated by the processor, the memory is used to store computer program code, the computer program code including computer instructions; when the processor executes the computer instructions, the electronic device performs the method described above.
[0034] Thirdly, this application provides a computer storage medium including computer instructions that, when executed on an electronic device, cause the electronic device to perform the method described above.
[0035] Fourthly, this application provides a computer program product that, when run on an electronic device, causes the electronic device to perform the method described above.
[0036] It is understood that the beneficial effects achieved by the electronic device described in the second aspect, the computer storage medium described in the third aspect, and the computer program product described in the fourth aspect can be referred to the beneficial effects in the first aspect and any of its possible design embodiments, which will not be repeated here. Attached Figure Description
[0037] Figure 1A A schematic diagram of a video scene provided for an embodiment of this application;
[0038] Figure 1B A video scene illustration provided for an embodiment of this application. Figure 2 ;
[0039] Figure 1C A video scene illustration three is provided for an embodiment of this application;
[0040] Figure 1D A video scene illustration provided for an embodiment of this application. Figure 4 ;
[0041] Figure 1E A video scene illustration provided for an embodiment of this application. Figure 5 ;
[0042] Figure 1F A video scene illustration provided for an embodiment of this application. Figure 6 ;
[0043] Figure 1G A schematic diagram of a texture enhancement process provided in this application embodiment;
[0044] Figure 1H A schematic diagram of a texture enhancement process provided in this application embodiment. Figure 2 ;
[0045] Figure 2 A structural block diagram of an electronic device provided in an embodiment of this application;
[0046] Figure 3A A schematic diagram three illustrating a texture enhancement process provided in this application embodiment;
[0047] Figure 3B This application provides a schematic diagram of a model training process.
[0048] Figure 4 A schematic diagram of a mapping curve provided for an embodiment of this application;
[0049] Figure 5 A flowchart illustrating a texture enhancement method provided in an embodiment of this application;
[0050] Figure 6This is a schematic diagram of a scene detection process provided in an embodiment of this application;
[0051] Figure 7 A schematic diagram of a texture enhancement process provided in this application embodiment. Figure 4 ;
[0052] Figure 8 A schematic diagram of a texture enhancement process provided in this application embodiment. Figure 5 ;
[0053] Figure 9 A schematic diagram of a segmentation network model structure provided in an embodiment of this application;
[0054] Figure 10 A schematic diagram of a texture enhancement process provided in this application embodiment. Figure 6 . Detailed Implementation
[0055] Hereinafter, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature. In the description of this embodiment, unless otherwise stated, "a plurality of" means two or more.
[0056] Video is composed of multiple frames of images. With the development of technology, video is widely used in various fields, such as office work, movies, meetings, games, collaboration, and filming. Electronic devices can play videos, that is, display images from a video. For example, in an office setting using electronic devices (such as user A's device), an electronic device might share the desktop of another device (such as user B's device). User B's device sends its desktop display data to user A's device, and user A's device displays user B's desktop (such as user A's device). Figure 1A As shown, the desktop display data can be sent to user A's device in the form of streaming media, that is, in the form of a video stream. This desktop display data is equivalent to an image. Simply put, user B's device continuously sends images including the desktop display content to user A's device for user A's device to display the image. Therefore, when user B's device displays the content of a PPT file, the desktop display data includes the content of the displayed PPT file; that is, the image includes the content of the PPT file, which can be considered an image. It should be understood that a PPT file can also be considered an image in other situations, such as recording the content displayed on an electronic device, where the displayed content can include the content of the PPT file.
[0057] For example, in such Figure 1BIn the collaborative scenario shown, the mobile phone sends screen projection data to the PC in real time. This screen projection data is sent in streaming media format and includes the screen data displayed on the mobile phone. The PC displays the mobile phone's screen based on the received screen projection data. Here, "screen" can refer to images in a video.
[0058] For example, in such Figure 1C As shown, the phone displays the game screen. Since the game screen is displayed continuously, one frame of the game screen can be considered as one frame of the video.
[0059] For example, such as Figure 1D As shown, the phone is playing the movie XX and displaying scenes from the movie XX. Here, XX movie is equivalent to a video, and the scenes in the movie are equivalent to images.
[0060] For example, such as Figure 1E As shown, the phone is in video recording mode, generating a corresponding video stream, which may include recorded video images. When the user wants to stop recording, the user can click the stop recording button 10 as shown in 1D. The phone responds to the click operation of the stop recording button 10, stops recording, and generates a corresponding video file, which may include multiple frames of images.
[0061] For example, such as Figure 1F As shown, the mobile phone plays a short video and displays images from the short video, which can also be referred to as a video.
[0062] The above Figures 1A to 1F The type of video shown is only one example. Videos can also be used in other aspects, that is, videos can be of other types, such as TV series, etc. This application does not limit them.
[0063] To improve video quality, electronic devices can enhance the texture (or texture sharpness) of video images to improve their display quality. In one implementation, for each frame of a video, the electronic device can directly adjust the pixel values of each pixel in the image using specified adjustment parameters to enhance the texture sharpness. For example, pixel values include H-channel pigment values; multiplying the H-channel pigment values of each pixel by the specified texture adjustment parameters enhances the texture sharpness. However, different scenarios require different amounts of texture adjustment parameters. For instance, when an electronic device plays a movie, some scenes have high texture sharpness, thus requiring smaller texture adjustment parameters. If the texture adjustment parameters are too large, over-sharpening may occur, leading to image distortion (e.g., if clothing in an image has "XX" written on it, over-sharpening will make the "XX" appear pasted on rather than as if it were printed on the clothing itself). In some scenes, the texture clarity is low, so a larger texture adjustment parameter is required. If the texture adjustment parameter is too small, the adjustment may be too small, that is, the sharpening degree is too small, the texture clarity of the image is not effectively improved, and the details are still blurry.
[0064] In another implementation, such as Figure 1G As shown, electronic devices utilize conventional sharpening algorithms, such as unsharp masking (USM), to sharpen video images and enhance their clarity. Electronic devices can use sharpening operators (such as Roberts, Prewitt, Sobel, Laplacian, and Kirsch operators) to extract high-frequency information (or edge information) from images. Sharpening operators can be 2x2 or 3x3 masks; by convolving or performing related operations with the mask on the image, edge information can be obtained. The principle behind determining image edge information is to calculate gradient values in different directions using first and / or second derivatives to achieve edge detection. After obtaining the edge information, adding or subtracting it from the original image (i.e., the edge information) can sharpen the image and enhance its texture clarity. However, this implementation lacks adaptability and cannot adjust to the scene depicted in the video.
[0065] In another implementation, the edge information extraction of the above image can be equivalent to high-pass filtering in the frequency domain. Electronic devices can use a two-dimensional discrete Fourier transform to transform the image from the spatial domain (or time domain) to the frequency domain. In the frequency domain, high-frequency regions represent areas with relatively complex textures, while low-frequency regions represent areas with relatively smooth textures. Therefore, to improve the texture clarity of the image, electronic devices can perform high-pass filtering on the frequency domain image to extract the high-frequency regions. Then, the electronic devices can use an inverse two-dimensional discrete Fourier transform to transform the high-frequency regions of the image from the frequency domain to the time domain. Finally, the electronic devices can add the high-frequency regions in the time domain to the original image to obtain the texture-enhanced image.
[0066] However, the above implementation method lacks adaptability and cannot adaptively adjust according to the scene of the images in the video.
[0067] Therefore, to address the aforementioned problems, this application proposes a texture enhancement method. After receiving a video stream, the electronic device can perform scene detection for each first image in the video stream. This involves comparing the pixel values of the target channel of the first image with the pixel values of the target channel of the previous frame image to obtain the scene detection result corresponding to the first image. This determines whether the scene of the first image has changed, i.e., whether the difference in texture information between the first image and the previous frame image is significant. If the scene detection result corresponding to the first image indicates a scene change, it means that the scene of the first image differs significantly from the scene of the previous frame image. In other words, the difference in texture information between the scene of the first image and the previous frame image is significant, resulting in a significant difference in texture display effect. Therefore, if... Figure 1HAs shown, the electronic device needs to reuse the trained first MobileNetV3 model to predict the texture adjustment parameters corresponding to the first image, obtaining texture adjustment parameters (e.g., 0.23) that are adapted to the scene of the first image. If the scene detection result for the first image indicates that the scene has not changed, it means that the scene of the first image is relatively similar to that of the previous frame. In other words, the texture information of the first image differs significantly from that of the previous frame, resulting in a smaller difference in texture display effect. Therefore, the electronic device can directly use the texture adjustment parameters corresponding to the previous frame as the texture adjustment parameters for the first image, avoiding unnecessary prediction of texture adjustment parameters. Subsequently, the electronic device can use the texture adjustment parameters corresponding to the first image to adjust the texture information of the first image. For example, based on the texture adjustment parameters, combined with a sharpening algorithm, the texture information of the first image can be adjusted to obtain a texture-enhanced first image, improving the texture clarity of the first image and achieving real-time enhancement of the video image's clarity. Furthermore, scene detection determines whether it is necessary to re-predict the texture adjustment parameters, avoiding unnecessary predictions, reducing data processing volume, and thus lowering the power consumption of the electronic device. Afterwards, the electronic device can display the texture-enhanced first image, ensuring the texture display effect of the first image, improving the clarity of the first image, thereby improving video quality and ensuring the user's visual experience.
[0068] For example, the electronic device in this application embodiment may be a mobile phone, tablet computer, wearable device, computer (such as PC), personal digital assistant (PDA), or other device capable of displaying images in a video. This application embodiment does not impose any special restrictions on the specific form of the electronic device.
[0069] For example, Figure 2 A schematic diagram of the structure of electronic device 200 is shown. For example... Figure 2 As shown, the electronic device 200 may include a processor 210, an external memory interface 220, an internal memory 221, a universal serial bus (USB) interface 230, a charging management module 211, a power management module 212, a battery 213, an antenna 1, an antenna 2, a mobile communication module 240, a wireless communication module 250, an audio module 270, a speaker 270A, a receiver 270B, a microphone 270C, a headphone jack 270D, a sensor module 280, buttons 290, a motor 291, an indicator 292, a camera 293, a display screen 294, and a subscriber identification module (SIM) card interface 295, etc.
[0070] It is understood that the structures illustrated in the embodiments of the present invention do not constitute a specific limitation on the electronic device 200. In other embodiments of this application, the electronic device 200 may include more or fewer components than illustrated, or combine some components, or split some components, or have different component arrangements. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.
[0071] Processor 210 may include one or more processing units, such as application processor (AP), modem processor, graphics processing unit (GPU), image signal processor (ISP), controller, memory, video codec, digital signal processor (DSP), baseband processor, and / or neural network processing unit (NPU). Different processing units may be independent devices or integrated into one or more processors.
[0072] The controller can be the nerve center and command center of the electronic device 200. The controller can generate operation control signals based on the instruction opcode and timing signals to control the fetching and execution of instructions.
[0073] The processor 210 may also include a memory for storing instructions and data. In some embodiments, the memory in the processor 210 is a cache memory. This memory can store instructions or data that the processor 210 has just used or that are used repeatedly. If the processor 210 needs to use the instruction or data again, it can directly retrieve it from the memory. This avoids repeated accesses, reduces the waiting time of the processor 210, and thus improves the efficiency of the system.
[0074] In some embodiments, the processor 210 may include one or more interfaces. Interfaces may include an inter-integrated circuit (I2C) interface, an inter-integrated circuit sound (I2S) interface, a pulse code modulation (PCM) interface, a universal asynchronous receiver / transmitter (UART) interface, a mobile industry processor interface (MIPI), a general-purpose input / output (GPIO) interface, a subscriber identity module (SIM) interface, and / or a universal serial bus (USB) interface, etc.
[0075] It is understood that the interface connection relationships between the modules illustrated in the embodiments of the present invention are merely illustrative and do not constitute a structural limitation on the electronic device 200. In other embodiments of this application, the electronic device 200 may also employ different interface connection methods or combinations of multiple interface connection methods as described in the above embodiments.
[0076] The charging management module 211 receives charging input from a charger. The charger can be a wireless charger or a wired charger. In some wired charging embodiments, the charging management module 211 receives charging input from the wired charger via a USB interface 230. In some wireless charging embodiments, the charging management module 211 receives wireless charging input via the wireless charging coil of the electronic device 200. While charging the battery 213, the charging management module 211 can also supply power to the electronic device via the power management module 212.
[0077] The wireless communication function of electronic device 200 can be implemented through antenna 1, antenna 2, mobile communication module 240, wireless communication module 250, modem processor, and baseband processor.
[0078] Antenna 1 and antenna 2 are used to transmit and receive electromagnetic wave signals. Each antenna in electronic device 200 can be used to cover one or more communication frequency bands. Different antennas can also be multiplexed to improve antenna utilization. For example, antenna 1 can be multiplexed as a diversity antenna for a wireless local area network. In some other embodiments, the antennas can be used in conjunction with a tuning switch.
[0079] The mobile communication module 240 can provide solutions for wireless communication, including 2G / 3G / 4G / 5G, applied to the electronic device 200. The mobile communication module 240 may include at least one filter, switch, power amplifier, low noise amplifier (LNA), etc. The mobile communication module 240 can receive electromagnetic waves via antenna 1, and perform filtering, amplification, and other processing on the received electromagnetic waves before transmitting them to a modem processor for demodulation. The mobile communication module 240 can also amplify the signal modulated by the modem processor and convert it into electromagnetic waves for radiation via antenna 1. In some embodiments, at least some functional modules of the mobile communication module 240 may be housed in the processor 210. In some embodiments, at least some functional modules of the mobile communication module 240 and at least some modules of the processor 210 may be housed in the same device.
[0080] The modem processor may include a modulator and a demodulator. The modulator modulates the low-frequency baseband signal to be transmitted into a mid-to-high frequency signal. The demodulator demodulates the received electromagnetic wave signal into a low-frequency baseband signal. The demodulator then transmits the demodulated low-frequency baseband signal to the baseband processor for processing. After processing by the baseband processor, the low-frequency baseband signal is transmitted to the application processor. The application processor outputs sound signals through an audio device (not limited to speaker 270A, receiver 270B, etc.) or displays images or videos through the display screen 294. In some embodiments, the modem processor may be a separate device. In other embodiments, the modem processor may be independent of the processor 210 and may be housed in the same device as the mobile communication module 240 or other functional modules.
[0081] The wireless communication module 250 can provide solutions for wireless communication applications on the electronic device 200, including wireless local area networks (WLAN) (such as wireless fidelity (Wi-Fi) networks), Bluetooth (BT), global navigation satellite system (GNSS), frequency modulation (FM), near field communication (NFC), and infrared (IR) technologies. The wireless communication module 250 can be one or more devices integrating at least one communication processing module. The wireless communication module 250 receives electromagnetic waves via antenna 2, performs frequency modulation and filtering of the electromagnetic wave signals, and sends the processed signal to processor 210. The wireless communication module 250 can also receive signals to be transmitted from processor 210, perform frequency modulation and amplification, and convert them into electromagnetic waves for radiation via antenna 2.
[0082] Electronic device 200 implements display functions through a GPU, a display screen 294, and an application processor. The GPU is a microprocessor for image processing, connected to the display screen 294 and the application processor. The GPU is used to perform mathematical and geometric calculations and for graphics rendering. Processor 210 may include one or more GPUs, which execute program instructions to generate or modify display information.
[0083] Display screen 294 is used to display images, videos, etc. Display screen 294 includes a display panel. The display panel may be a liquid crystal display (LCD), an organic light-emitting diode (OLED), an active-matrix organic light-emitting diode (AMOLED), a flexible light-emitting diode (FLED), a miniature LED, a microLED, a quantum dot light-emitting diode (QLED), etc. In some embodiments, electronic device 200 may include one or N displays 294, where N is a positive integer greater than 1.
[0084] Electronic device 200 can perform shooting functions through ISP, camera 293, video codec, GPU, display screen 294 and application processor.
[0085] The external storage interface 220 can be used to connect an external memory card, such as a Micro SD card, to expand the storage capacity of the electronic device 200.
[0086] Internal memory 221 can be used to store computer executable program code, which includes instructions. Processor 210 executes various functional applications and data processing of electronic device 200 by running the instructions stored in internal memory 221. Internal memory 221 may include a program storage area and a data storage area. The program storage area may store the operating system, at least one application program required for a function (such as sound playback, image playback, etc.), etc. The data storage area may store data created during the use of electronic device 200 (such as audio data, phonebook, etc.). Furthermore, internal memory 221 may include high-speed random access memory and may also include non-volatile memory, such as at least one disk storage device, flash memory device, universal flash storage (UFS), etc.
[0087] Electronic device 200 can implement audio functions such as music playback and recording through audio module 270, speaker 270A, receiver 270B, microphone 270C, headphone jack 270D, and application processor.
[0088] Button 290 includes a power button, volume buttons, etc. Button 290 can be a mechanical button or a touch button.
[0089] Indicator 292 can be an indicator light, which can be used to indicate charging status, power changes, messages, missed calls, notifications, etc.
[0090] The sensor module 280 may include pressure sensors, gyroscope sensors, barometric pressure sensors, magnetic sensors, accelerometers, distance sensors, proximity sensors, fingerprint sensors, temperature sensors, touch sensors, ambient light sensors, bone conduction sensors, etc.
[0091] This application provides a texture enhancement method. An electronic device receives a video stream. For example... Figure 3AAs shown, for each first image in the video stream, the electronic device performs scene detection on the first image, such as comparing the scene of the first image with the scene of the previous frame to determine whether the scene of the first image has changed. If the scene of the first image has changed, it indicates that the texture information of the first image differs significantly from that of the previous frame, and the texture adjustment parameters corresponding to the first image need to be determined. The electronic device can then input the first image into the trained first MobileNetV3 model to determine the texture adjustment parameters, thus achieving accurate determination of the texture adjustment parameters. If the scene of the first image has not changed, it indicates that the texture information of the first image differs only slightly from that of the previous frame. Therefore, the texture adjustment parameters corresponding to the previous frame can be directly used as the texture adjustment parameters for the first image, achieving rapid determination of the texture adjustment parameters. Afterwards, the electronic device can use the texture adjustment parameters to adjust the first image, enhancing its texture clarity and thus improving its quality, which in turn improves the video quality, achieving real-time texture enhancement of the first image.
[0092] The first MobileNetV3 model trained above was obtained using sample images. This trained first MobileNetV3 model can predict the texture adjustment parameters corresponding to the image. The process of obtaining the first trained MobileNetV3 model using sample images will be described in detail below.
[0093] In this embodiment, the first device can acquire a first sample image set, which may include multiple original first sample images. The first device can perform texture degradation processing on the original first sample images to reduce their texture sharpness, obtaining a degraded sample image corresponding to the original first sample image. Then, the first device can use the original first sample image and its corresponding degraded sample image to train a MobileNetV3 model. The trained MobileNetV3 model can predict texture adjustment parameters corresponding to images in a video, thereby using these parameters to improve the texture sharpness and enhance details, thus improving the image's display effect. Specifically, as shown... Figure 3B As shown, the process of training the MobileNetV3 model using the first sample image set can include S301-S305.
[0094] S301. The electronic device acquires a first sample image set. The first sample image set includes multiple original first sample images.
[0095] Among them, the original image of the first sample is of high quality, such as a high-resolution image (or a standard image).
[0096] In some embodiments, if the format of the first sample original image is not RGB, the electronic device may first convert the format of the first sample original image to RGB format for training using the first sample original image in RGB format.
[0097] In some embodiments, the aforementioned first sample image set can be of various types. These types may include collaborative, movie, game, office (e.g., PPT), meeting, short video, and recording types. The methods for creating different types of image sets also differ. For example, a movie-type image set may include multiple (e.g., 100) domestic and international 1080P / 2K high-definition movies, forming a large dataset evenly distributed across four aspects: brightness, texture complexity, motion complexity, and color complexity. In other words, it includes movie frames with different display effects (e.g., texture display effects), ensuring the richness of the movie-type image set and improving the model's generalization ability.
[0098] For example, the image set of the above-mentioned office PPT type can include: PPT content in the images including common content such as text, pictures, tables, and statistical charts.
[0099] For example, the image set of the above-mentioned meeting type may include different facial images, such as facial images of different genders, facial images of different ages, facial images of different regions, etc.
[0100] For example, the image set of the aforementioned game type can include game screens from different game applications, that is, images.
[0101] In some embodiments, since model training actually refers to the process of recovering a high-quality image from a low-quality image, after obtaining a high-quality first sample original image, for each first sample original image, the electronic device can degrade the sharpness of the first sample original image to obtain an image with lower sharpness, that is, a sample degraded image corresponding to the first sample original image. This allows the electronic device to train the model using each first sample original image and its corresponding sample degraded image. It should be understood that the content, i.e., the objects, of the first sample original image and its corresponding sample degraded image are the same; however, compared to the first sample original image, the objects in the sample degraded image corresponding to the first sample original image have lower sharpness, such as objects that were originally sharp becoming blurry. The following will describe a possible process for degrading the sharpness of the first sample original image in conjunction with S302-S303.
[0102] S302. For each of the multiple first sample original images, the electronic device downsamples the first sample original image to obtain the downsampled first sample original image.
[0103] In this case, the size of the original image corresponding to the first sample after downsampling is smaller than the size of the original image of the first sample.
[0104] S303. The electronic device upsamples the downsampled original image of the first sample to obtain a degraded image corresponding to the original image of the first sample. The size of the degraded image corresponding to the original image of the first sample is equal to that of the original image of the first sample.
[0105] In this embodiment, for each first sample original image, the electronic device can first downsample the high-quality first sample original image. Then, the electronic device uses an image interpolation algorithm (such as the Bicubic algorithm) to upsample the downsampled first sample original image, obtaining a low-quality image with the same size as the first sample original image but reduced clarity, i.e., a sample degradation image. This sample degradation image simulates the low-resolution image corresponding to the first sample original image. In this way, the first sample original image and its corresponding sample degradation image constitute a pair of image sets with enhanced texture clarity.
[0106] S304. The electronic device calculates the variance of the original image of the first sample to obtain the complexity corresponding to the original image of the first sample.
[0107] For example, the electronic device can first convert the first sample original image into a grayscale image. Then, the electronic device can use the pixel values in the grayscale image to calculate the variance of the first sample original image. This variance can reflect the dispersion of the color distribution of the first sample original image. The larger the variance, the more drastic the change in color brightness, and the higher the complexity of the first sample original image; the smaller the variance, the more gradual the change in color brightness, and the lower the complexity of the first sample original image.
[0108] S305. The electronic device trains the first MobileNetV3 model based on each original first sample image, the corresponding degraded sample image, and the complexity of each original first sample image, to obtain the trained first MobileNetV3 model. The trained first MobileNetV3 model can predict the texture adjustment parameters corresponding to images in videos under different scenarios.
[0109] Among them, the texture adjustment parameter is used to enhance the texture of the image and improve its clarity.
[0110] In this embodiment, the electronic device can use the variance corresponding to the first sample original image as the label of the first sample original image. Then, the electronic device can input the labeled first sample original image and its corresponding sample degraded image into the first MobileNetV3 model, so that the first MobileNetV3 model learns how to recover the first sample original image from the sample degraded image corresponding to the first sample original image, and determines the mapping relationship between texture complexity and texture adjustment parameters (or texture sharpness intensity), such as generating a mapping curve between texture complexity and texture adjustment parameters.
[0111] It should be understood that for scenes with low texture complexity, the required texture sharpness intensity gradually decreases. If an image has low texture complexity and its texture is adjusted using a large texture sharpness intensity, over-sharpening may occur, leading to unnatural and distorted images. Similarly, for scenes with high texture complexity, excessively high texture sharpness intensity is not necessary; the higher the texture complexity, the lower the required texture sharpness intensity. If an image has high texture complexity and its texture is adjusted using a large texture sharpness intensity, over-sharpening may also occur. Therefore, scenes with medium texture complexity should correspond to higher texture sharpness intensity. To achieve a mapping curve that satisfies this rule, the first MobileNetV3 model can use four tangent quadratic curves to fit a curve, resulting in... Figure 4 The diagram shows the mapping curve between texture complexity and texture sharpness intensity. It should be noted that... Figure 4 The texture complexity ranges from 0 to 1, and is mapped to 0 to α via curves. The parameters at the junctions of curve segments can be adjusted. Figure 4 The parameters at the five junctions shown can change the overall shape of the curve. λ determines the maximum texture sharpness intensity, and α determines the texture complexity corresponding to the maximum texture sharpness intensity. X1 and x2 determine the junction positions of the curve segments, used for fine-tuning of the wireless detail. μ determines the overall average texture sharpness intensity of the wireless signal. The first MobileNetV3 model, based on overall adjustments plus fine-tuning of the detail, can ultimately determine a suitable mapping curve, thus ensuring the accuracy of the mapping between texture complexity and texture sharpness intensity.
[0112] In some embodiments, the structure of the first MobileNetV3 model described above may include a first convolutional layer, a batch normalization (BN) layer and an activation function layer, multiple (e.g., 11) coding blocks, an average pooling layer and a fully connected layer.
[0113] Optionally, the last fully connected layer of the first MobileNetV3 model described above may include three fully connected layers: fully connected layer a, fully connected layer b, and fully connected layer c. Fully connected layer a has 1024 neurons in the relevant feature map at its input and outputs 512 neurons; it has a Dropout rate of 0.5 and uses the ReLU activation function. Fully connected layer b has 512 neurons in the relevant feature map at its input and outputs 256 neurons; it has a Dropout rate of 0.2 and uses the ReLU activation function. Fully connected layer c has 256 neurons in the relevant feature map at its input and outputs 1 neuron; it has a Dropout rate of 0.1 and uses the ReLU activation function.
[0114] During training, the electronic device freezes all layers of the first MobileNetV3 model except for the parameters of the last fully connected layer, training only that fully connected layer. The electronic device can first downsample the input image (such as the original image of the first sample and its corresponding degraded image) to 256x256, set the learning rate to 0.0001, and the number of iterations (epochs) to 100, using the Adam optimizer for training. It should be understood that the electronic device can also train directly using the original image without downsampling the input image, although the training time will be longer.
[0115] In some embodiments, the first MobileNetV3 model described above is merely an example, and other lightweight models can also be used, which are not limited in this application. Examples include MobileNetV1 and MobileNetV2.
[0116] The training process of the first MobileNetV3 model has been described above. After obtaining the trained first MobileNetV3 model, electronic devices using this model can predict texture adjustment parameters for images in the video stream. These parameters are then used to adjust the images, enabling real-time adjustment of the video images, enhancing details, improving clarity, and ultimately improving video quality and user visual experience. The following section will combine... Figure 5 This section details the process of enhancing video image clarity using the first trained MobileNetV3 model. For example, ... Figure 5 As shown, the process is as follows:
[0117] S401, Electronic device acquires video stream.
[0118] The aforementioned video stream includes at least one image. The electronic device acquires the video stream in real time. This video stream can be generated by the electronic device itself, such as when the electronic device is in recording mode and generates a corresponding video stream. Alternatively, the video stream can be sent by other devices, such as when the electronic device is in a collaborative state and displays the content displayed by other devices, and the images in the video stream include the display content of those other devices.
[0119] S402. For each first image in the video stream, the electronic device downsamples the first image to obtain the downsampled first image.
[0120] In some embodiments, for each image (or first image) obtained by the electronic device, the first image can be downsampled, such as downsampling the first image to 256*256. Of course, the first image can also be downsampled to other sizes, such as 512*512. However, compared to other sizes, 256*256 can ensure the accuracy of scene detection while requiring less computing power from the electronic device.
[0121] S403. The electronic device uses the previous image after downsampling of the first image to perform scene detection on the first image, and obtains the scene detection result corresponding to the first image. The scene detection result indicates whether the scene of the first image has changed.
[0122] In this embodiment, the electronic device compares the downsampled first image with the previous downsampled image (or the previous frame image) to determine whether the first image and the previous frame image belong to the same scene. If the scene detection result indicates that the scene of the first image has changed, it means that the texture information difference between the first image and the previous frame image is large, and the texture adjustment parameters corresponding to the previous frame image cannot be used to adjust the current image, that is, to adjust the texture of the first image. Therefore, the electronic device needs to determine new texture adjustment parameters corresponding to the first image, and can execute S404.
[0123] If the scene detection result indicates that the scene of the first image has not changed, it means that the difference in texture information between the first image and the previous image (or the previous frame image) is small. The texture adjustment parameters corresponding to the previous frame image can be used to enhance the current image, that is, the texture information of the first image. The electronic device can execute S405.
[0124] In some embodiments, the electronic device can use the difference between the pixel value of the target channel of the first image and the pixel value of the target channel of the previous frame of the first image to determine whether the first image and the previous frame belong to the same scene, so as to obtain the scene detection result corresponding to the first image. Possible implementation methods for determining the scene detection result will be described below.
[0125] In one implementation, because the HSV format corresponds to different dimensions, such as hue, saturation, and brightness, it allows for checking whether the scene in the first image has changed across different dimensions, ensuring more accurate scene detection results. The following will use the aforementioned target channels, including the H, S, and V channels, as an example to detail how to use HSV to obtain the scene detection results corresponding to the first image.
[0126] First, for the hue (H) dimension, the electronic device can calculate the difference between the pixel value of the H channel of the downsampled first image and the pixel value of the H channel of the downsampled previous frame image, obtaining the H channel difference value corresponding to the first image. This H channel difference value represents the degree of hue difference between the first image and the previous frame image. The larger the H channel difference value, the greater the degree of hue difference, the less content the first image and the previous frame image share, that is, the less identical texture information they contain, and the lower the probability that the first image and the previous frame image belong to the same scene. For example, for each pixel point (or first pixel point) in the downsampled first image, the electronic device can calculate the difference between the pixel value of the H channel of the first pixel point and the pixel value of the H channel of the corresponding pixel point (or second pixel point) in the previous frame image, obtaining the H channel difference value corresponding to the first pixel point. The position of the second pixel point at the corresponding position in the previous frame image is the same as the position of the first pixel point in the first image. Then, the electronic device can calculate the difference of the H channel corresponding to all first pixels in the first image, thus obtaining the difference of the H channel corresponding to the first image. Specifically, the electronic device can use the following formula 1 to calculate the difference of the H channel corresponding to the first image.
[0127]
[0128] Among them, score h This represents the difference in the H channels of the first image, num_piexls represents the total number of pixels in the first image, and current_H represents the difference in the H channels. i Last_H represents the i-th pixel in the first image. i This represents the i-th pixel in the previous frame of the first image.
[0129] For the saturation (S) dimension, the electronic device can calculate the difference between the pixel value of the S channel of the downsampled first image and the pixel value of the S channel of the downsampled previous frame image, obtaining the S channel difference value corresponding to the first image. This S channel difference value represents the degree of saturation difference between the first image and the previous frame image. The larger the S channel difference value, the greater the degree of saturation difference, the less content the first image and the previous frame image share, that is, the less identical texture information they contain, and the lower the probability that the first image and the previous frame image belong to the same scene. For example, for each first pixel in the downsampled first image, the electronic device can calculate the difference between the pixel value of the S channel of that first pixel and the pixel value of the H channel of the corresponding second pixel in the previous frame image, obtaining the S channel difference value corresponding to that first pixel. Then, the electronic device can calculate the S channel difference values corresponding to all first pixels in the first image, obtaining the S channel difference value corresponding to the first image. Specifically, the electronic device can use the following formula 2 to calculate the S channel difference value corresponding to the first image.
[0130]
[0131] Among them, score s The current_S value represents the difference in the S channels of the first image. i Last_S represents the pixel value of the S channel of the i-th pixel in the first image. i This represents the pixel value of the S channel of the i-th pixel in the previous frame of the first image.
[0132] For the luminance (V) dimension, the electronic device can calculate the difference between the pixel value of the V channel of the downsampled first image and the pixel value of the V channel of the downsampled previous frame image, obtaining the V channel difference value corresponding to the first image. This V channel difference value represents the degree of luminance difference between the first image and the previous frame image. The larger the V channel difference value, the greater the degree of luminance difference, the less content the first image and the previous frame image share, that is, the less identical texture information they contain, and the lower the probability that the first image and the previous frame image belong to the same scene. For example, for each first pixel in the downsampled first image, the electronic device can calculate the difference between the pixel value of the V channel of that first pixel and the pixel value of the V channel of the corresponding second pixel in the previous frame image, obtaining the V channel difference value corresponding to that first pixel. Then, the electronic device can calculate the V channel difference value corresponding to all first pixels in the first image, obtaining the V channel difference value corresponding to the first image. Specifically, the electronic device can use the following formula 3 to calculate the V channel difference value corresponding to the first image.
[0133]
[0134] Among them, score v The current_V represents the difference in the V channel of the first image, and the last_V represents the pixel value of the V channel of the i-th pixel in the first image. i This represents the pixel value of the V channel of the i-th pixel in the previous frame of the first image.
[0135] It should be noted that the process of determining the difference values of the H channel, S channel, and V channel corresponding to the first image described above is only an example. Electronic devices may determine the difference values using several pixels in the first image instead of individual pixels. This application does not limit the method of determining the difference values of the H channel, S channel, and V channel corresponding to the first image.
[0136] Then, the electronic device can obtain the scene detection result corresponding to the first image based on the differences in the H channel, S channel, and V channel of the first image. For example, the electronic device can calculate the average of the differences in the H channel, S channel, and V channel of the first image to obtain the scene difference corresponding to the first image. Specifically, the electronic device can use the following formula 4 to calculate the scene difference corresponding to the first image.
[0137]
[0138] Here, score represents the scene difference corresponding to the first image.
[0139] When the average value, i.e., the scene difference corresponding to the first image, is greater than the threshold of 1, it indicates a significant difference in the displayed content between the first image and the previous frame. This means there is a large difference in texture information between the first image and the previous frame, and the scene difference between them differs considerably. Therefore, the first image and the previous frame do not belong to the same scene, and the electronic device can determine that the scene detection result corresponding to the first image indicates a scene change. When the scene difference is less than or equal to the threshold of 1 (e.g., 27), it indicates a small difference in the displayed content between the first image and the previous frame. This means there is a small difference in texture information between the first image and the previous frame, and the scene difference between them is relatively small. Therefore, the first image and the previous frame belong to the same scene, and the electronic device can determine that the scene detection result corresponding to the first image indicates no scene change. Of course, using the average of the differences in the H channel, S channel, and V channel is just one example; other methods can also be used to determine the scene detection result corresponding to the first image. For example, if the difference in the H channel of the first image is greater than threshold 2, the difference in the S channel is greater than threshold 3, and the difference in the V channel is greater than threshold 4, the electronic device determines that the scene detection result for the first image indicates a change in the scene. If the difference in the H channel of the first image is less than or equal to threshold 2, the difference in the S channel is less than or equal to threshold 3, and the difference in the V channel is less than or equal to threshold 4, the electronic device determines that the scene detection result for the first image indicates no change in the scene.
[0140] In the embodiments of this application, such as Figure 6 As shown, the electronic device downsamples (resizes) the first image in the video stream to improve image processing speed, obtaining a downsampled first image. Then, the electronic device converts the downsampled first image to HSV format, obtaining an HSV format first image. Next, the electronic device calculates the difference between the pixel value of the target channel of each first pixel in the HSV format first image and the pixel value of the target channel of the corresponding second pixel in the previous frame of the HSV format image, obtaining the pixel difference of the target channel for each first pixel. Then, the electronic device calculates the average of the pixel differences of the target channels for each first pixel, obtaining the pixel difference of the target channels for the first image. Finally, the electronic device calculates the average of the pixel differences of the target channels for the first image, obtaining the scene detection result for the first image. This scene detection result indicates whether the scene has changed, achieving accurate determination of the scene detection result. Furthermore, by downsampling the first image, the number of pixels that need to be calculated can be reduced, thereby reducing the amount of data processing and improving the efficiency of scene detection result determination.
[0141] The process of determining the previous frame image in HSV format is similar to the process of determining the first image in HSV format described above.
[0142] In some embodiments, the image format in the video stream is generally YUV format. The electronic device can first convert the YUV format of the downsampled first image to RGB format, and then convert the first image in RGB format to the first image in HSV format.
[0143] It should be understood that the above example, using the target channels including H, S, and V channels, illustrates how to obtain the scene detection result corresponding to the first image using HSV. Of course, the target channels can include one or two of the H, S, and V channels. The electronic device can calculate the differences between the target channels corresponding to the first image to determine the detection result for the first image. For example, if the target channels include H and S channels, the electronic device calculates the difference between the H channel and the S channel corresponding to the first image. Then, the electronic device can calculate the average of the differences between the H channel and the S channel corresponding to the first image. If this average is greater than a threshold, the electronic device can determine that the scene detection result for the first image indicates a change in the scene.
[0144] S404. The electronic device inputs the downsampled first image into the trained first MobileNetV3 model to obtain the texture adjustment parameters corresponding to the first image.
[0145] In this embodiment, since the first image and its previous frame do not belong to the same scene, different scenes correspond to different texture adjustment parameters. Therefore, the electronic device can use the trained first MobileNetV3 model to process the first image to obtain texture adjustment parameters that match the scene of the first image, thus achieving accurate prediction of the texture adjustment parameters. Then, the electronic device can use the texture adjustment parameters that match the scene of the first image to adjust the texture information of the first image, thereby achieving sharpening processing of the first image, enhancing the texture clarity of the first image, and avoiding situations where the texture information enhancement is too large or too small. For example, if the first image is distorted due to over-sharpening or the texture clarity of the first image is not effectively improved due to insufficient sharpening, the clarity of the first image is guaranteed.
[0146] For example, the trained first MobileNetV3 model can calculate the texture complexity corresponding to the downsampled first image. Then, the trained first MobileNetV3 model can use the mapping curve between texture complexity and texture sharpness intensity to determine the texture sharpness intensity of the downsampled first image, that is, the texture sharpness intensity corresponding to the first image, and output the texture sharpness intensity corresponding to the first image (i.e., the texture adjustment parameter corresponding to the first image). Here, the texture sharpness intensity corresponding to the first image represents the texture adjustment parameter that matches the texture information of the first image. Adjusting the first image using this texture sharpness intensity will not result in over-adjustment or under-adjustment.
[0147] In some embodiments, since the model typically processes RGB format images, the electronic device can convert the downsampled first image to RGB format. Then, the electronic device can input the downsampled RGB format first image into the trained first MobileNetV3 model.
[0148] In some embodiments, the above-described method of using the trained first MobileNetV3 model to determine the texture adjustment parameters corresponding to the first image is only an example. This application does not limit the way of predicting the texture adjustment parameters corresponding to the image. For example, an electronic device can use the pixel values (such as YUV pixel values) of the pixels in the first image to adjust the preset texture adjustment parameters to obtain texture adjustment parameters that match the first image.
[0149] It should be noted that the downsampling process of the first image in S402 above is optional. The electronic device may also choose not to downsample the first image, but instead directly use the first image for scene detection to determine whether the scene of the first image has changed. If it has changed, the electronic device can input the first image into the trained first MobileNetV3 model to determine the texture adjustment parameters corresponding to the first image. However, this method involves a large amount of data processing and has lower image processing efficiency.
[0150] It should be understood that the first MobileNetV3 model trained above is only one example of an image processing model. Image processing models can also be other lightweight models trained in other ways.
[0151] S405. The electronic device uses the texture adjustment parameters corresponding to the previous image of the first image as the texture adjustment parameters corresponding to the first image.
[0152] In this embodiment of the application, since the first image and the previous frame of the first image belong to the same scene, it indicates that the texture information difference between the first image and the previous frame is small, that is, the difference in texture clarity is small. Therefore, the electronic device can use the texture adjustment parameters corresponding to the previous frame as the texture adjustment parameters corresponding to the first image.
[0153] In some embodiments, the electronic device utilizes texture adjustment parameters, combined with a sharpening algorithm, to enhance the texture of a first image. This involves adjusting the high-frequency information of the first image to enhance its details and improve its clarity. The high-frequency information forms the edges and details of the image; it refers to information in high-frequency regions, which are areas in the image where the grayscale values of pixels change rapidly. These high-frequency regions include information in the image that needs to clearly display details. The following section, in conjunction with steps S406-S407, will describe a possible implementation method for improving the clarity of a first image using texture adjustment parameters.
[0154] S406. The electronic device determines the high-frequency information of the first image.
[0155] S407. For each pixel in the high-frequency information of the first image, the electronic device multiplies the pixel value of the Y channel of the pixel by the texture adjustment parameter to obtain the target first image.
[0156] In some embodiments, the process of determining the high-frequency region of the first image described above may include: such as Figure 7 As shown, firstly, the electronic device performs mean filtering on the Y-channel pixel values of the pixels in the first image to obtain a blurred first image (or a first filtered image). Then, for each pixel in the first image, the Y-channel pixel value is subtracted from the corresponding pixel value in the blurred first image to obtain the high-frequency information (or high-frequency edge information) of the first image. This blurs the non-high-frequency information, such as low-frequency information, i.e., parts with less texture information. The blurred first image then contains the high-frequency information, i.e., the parts where texture needs to be enhanced. For example, the Y-channel pixel value of the first pixel in the first image is subtracted from the corresponding pixel value in the blurred first image to obtain the pixel difference value corresponding to the first pixel in the first image. This process is repeated to obtain the pixel differences corresponding to each pixel in the first image, and these pixel differences constitute the high-frequency information of the first image.
[0157] In some embodiments, a possible implementation of S407 described above may include: the electronic device calculating the product of the high-frequency information of the first image and the texture adjustment parameter λ to obtain the texture adjustment value of the Y channel of each pixel in the first image. For example, the electronic device can multiply the pixel value of the Y channel of each pixel in the high-frequency information of the first image by the texture adjustment parameter, that is, multiply the pixel difference corresponding to each pixel in the first image by the texture adjustment parameter to obtain the texture adjustment value of the Y channel corresponding to each pixel in the first image.
[0158] Next, for each pixel in the first image, the electronic device adds the pixel value of the Y channel (i.e., the original Y channel pixel value) of that pixel to the texture adjustment value at the corresponding position to obtain the texture-enhanced Y channel pixel value of that pixel. For example, the pixel value of the Y channel of the first pixel in the first image is added to the texture adjustment value at the corresponding position, i.e., the texture adjustment value of the Y channel of the first pixel, to obtain the texture-enhanced Y channel pixel value of the first pixel in the first image.
[0159] Then, the electronic device combines the pixel values of the Y channel and the UV channel of the first image after texture enhancement to obtain the texture-enhanced first image, which is the target first image, and thus obtains the texture-enhanced video stream.
[0160] In some embodiments, after obtaining the high-frequency information of the first image, a threshold can be used to determine whether to retain a portion of the high-frequency information of the first image to suppress noise generated by texture enhancement. For example, for the pixel value of the Y channel of each pixel in the high-frequency information of the first image, the electronic device can determine whether the pixel value of the Y channel of that pixel is greater than or equal to a threshold of 5. If it is greater than or equal to the threshold of 5, the electronic device can retain that pixel, meaning the high-frequency information of the first image still includes the pixel value of that pixel. If it is less than the threshold of 5, the electronic device can discard that pixel, meaning the high-frequency information of the first image does not include the pixel value of that pixel.
[0161] In this embodiment, to reduce computational power and improve the speed of image texture sharpness adjustment, the electronic device can perform real-time texture sharpness enhancement processing on the first image in the video stream using the YUV control. Since Y corresponds to brightness and UV corresponds to chromaticity, UV has a relatively small impact on texture sharpness. Therefore, the electronic device can enhance only the pixel values of the Y channel of the first image, ensuring the texture enhancement effect of the first image and reducing data processing volume. Furthermore, the electronic device only performs texture enhancement on the high-frequency information of the first image, while not enhancing non-high-frequency information, i.e., areas with less texture information, to avoid unnecessary texture adjustments. For example, the first image includes a puppy image portion and a background image portion. The background image portion refers to the blue sky, which contains less texture information, so it can be blurred. The puppy image portion includes the puppy's fur, which is high-frequency information; therefore, it is necessary to enhance the texture sharpness of the puppy's fur to make the details of the puppy image portion clearer. Simply put, determining the high-frequency information of the first image involves identifying lines like the puppy's fur, increasing the brightness of these lines, and then combining them with the pixel values of the UV channels of the original first image to obtain a first image with higher clarity.
[0162] In some embodiments, the aforementioned high-frequency information may include noise. For example, if the video stream is a video stream generated from a recorded video, the image in the captured video stream may include white dots, which are noise. Therefore, the electronic device can perform mean filtering on the high-frequency information of the first image again to obtain the filtered high-frequency information of the first image. Then, for the pixel value of the Y channel of each pixel in the high-frequency information of the filtered first image, the electronic device can determine whether the pixel value of the Y channel is less than a threshold of 6. If it is less than the threshold of 6, it indicates that the pixel is likely noise, and the electronic device can discard the pixel, that is, discard the pixel value of the Y channel. If it is greater than or equal to the threshold of 6, the pixel can be retained. Then, for each pixel in the high-frequency information of the filtered first image, the electronic device multiplies the pixel value of the Y channel by a texture adjustment parameter to obtain the target first image.
[0163] In some embodiments, the mean filtering method used in the above-described blurring process is merely an example. Electronic devices can also employ other methods to blur the first image, such as Gaussian blur (or Gaussian filtering), Laplacian filtering, median filtering, etc. The differences in high-frequency information of the first image determined by different blurring methods are relatively small. However, compared to other blurring methods, mean filtering has higher performance and determines the high-frequency information of the first image faster, thereby effectively improving the texture enhancement rate of the first image and ensuring the real-time texture enhancement of the video stream.
[0164] In this embodiment, after obtaining the high-frequency information of the first image, the electronic device adjusts the high-frequency information of the first image using texture adjustment parameters that match the scene of the first image, thereby achieving adaptive texture adjustment and avoiding excessive or insufficient texture enhancement of the first image. Furthermore, to reduce the impact of noise, after obtaining the high-frequency information of the first image, a mean filter is performed again, discarding some pixels in the high-frequency information of the first image based on a threshold. These two noise reduction processes effectively suppress noise generated by texture enhancement.
[0165] In some embodiments, the trained first MobileNetV3 model can predict not only the texture adjustment parameters corresponding to the first image, but also the color adjustment parameters corresponding to the first image. For example, when the first image and the previous frame do not belong to the same scene, the electronic device inputs the downsampled first image into the trained first MobileNetV3 model, which can obtain not only the texture adjustment parameters corresponding to the first image, but also the color adjustment parameters corresponding to the first image. The color adjustment parameters include brightness adjustment parameters, contrast adjustment parameters, and saturation adjustment parameters. Alternatively, when the first image and the previous frame belong to the same scene, the electronic device can directly use the color adjustment parameters corresponding to the previous frame as the color adjustment parameters corresponding to the first image.
[0166] In this embodiment, since the first image and its previous frame do not belong to the same scene, different scenes correspond to different color adjustment parameters and texture adjustment parameters. Therefore, the electronic device can use the trained first MobileNetV3 model to process the downsampled first image to obtain color adjustment parameters and texture adjustment parameters that match the scene of the first image, thus achieving accurate prediction of the color adjustment parameters and texture adjustment parameters. Then, the electronic device can use the color adjustment parameters that match the scene of the first image to adjust the color of the first image, avoiding excessive or insufficient color enhancement, such as excessive or insufficient brightness, contrast, or saturation enhancement, ensuring the color display effect of the first image. Furthermore, the texture adjustment parameters of the first image can be used to adjust the sharpness of the first image.
[0167] It is understandable that the aforementioned first sample image set may also include images with high color quality and images with low color quality. The first device uses the high-quality and low-quality images to train the first MobileNetV3 model, enabling the trained first MobileNetV3 model to predict color adjustment parameters.
[0168] The detailed process of adjusting the sharpness of the first image has been described above. The process of adjusting the color of the first image will be described in detail below.
[0169] In some embodiments, the electronic device uses brightness adjustment parameters to increase the brightness of the first image. Furthermore, the electronic device uses contrast adjustment parameters to increase the contrast of the first image. The electronic device also uses saturation adjustment parameters to increase the contrast of the first image. Since brightness is related to the pixel values of the Y channel of an image, the brightness of the first image can be increased by adjusting the pixel values of the Y channel of pixels in the first image using brightness adjustment parameters.
[0170] The following section will introduce a possible method to improve the brightness of the first image by using brightness adjustment parameters.
[0171] For example, for each pixel in the first image, the electronic device uses the brightness adjustment parameter as the index of the Y channel pixel of that pixel to obtain the pixel value of the Y channel after brightness enhancement.
[0172] In this embodiment of the application, for each pixel in the first image, the electronic device may use Y... L =Y0 L The pixel value of the Y channel of the pixel after brightness enhancement is calculated to achieve the brightness enhancement of the first image.
[0173] Among them, Y L This represents the pixel value of the Y channel after brightness enhancement, where L represents the brightness adjustment parameter, and Y0 represents the original pixel value of the Y channel of the pixel, which is the pixel value of the Y channel of the pixel in the first image.
[0174] In some embodiments, since contrast refers to the degree of difference between brightness values in an image, reflecting the difference in brightness between different areas of the image, it is evident that the contrast of an image is also related to the pixel values of the Y channel of the image. Therefore, the electronic device can also improve the contrast of the first image by adjusting the pixel values of the Y channel of pixels in the first image using contrast adjustment parameters. For example, for each pixel in the first image, the electronic device can obtain the adjusted value of the Y channel of that pixel based on the pixel value of the brightness-enhanced Y channel and the average value of the brightness-enhanced Y channel of the corresponding pixel in the first image, combined with the contrast adjustment parameters. Here, the average value of the brightness-enhanced Y channel of the corresponding pixel in the first image represents the average value of the brightness-enhanced Y channel of all pixels in the first image. For example, the electronic device can calculate the difference (or second difference) between the pixel value of the brightness-enhanced Y channel of that pixel and the average value of the brightness-enhanced Y channel of the corresponding pixel in the first image. Then, the electronic device can determine the adjusted value of the Y channel of that pixel based on the second difference and the contrast adjustment parameters.
[0175] Then, the electronic device obtains the pixel value of the Y channel of the pixel after contrast enhancement based on the average value of the adjustment value of the Y channel of the pixel and the pixel value of the Y channel of the first image after brightness enhancement. By reducing the brightness value of the darker pixels in the first image and increasing the brightness value of the brighter pixels in the first image, the contrast of the first image is improved, thereby enhancing the contrast of the first image.
[0176] The following section will introduce a possible method to improve the brightness of the first image by using brightness adjustment parameters.
[0177] For each pixel in the first image, the electronic device calculates a second difference between the pixel value of the Y channel after brightness enhancement and the average value of the corresponding pixel value of the Y channel after brightness enhancement in the first image.
[0178] The electronic device calculates the product of the second difference and the contrast adjustment parameter to obtain the adjustment value of the Y channel of that pixel.
[0179] The electronic device calculates the sum of the adjustment value of the Y channel of the pixel and the average value of the pixel value of the Y channel after brightness enhancement corresponding to the first image, to obtain the pixel value of the Y channel after contrast enhancement of the pixel.
[0180] In this embodiment of the application, for each pixel in the first image, the electronic device may use Y... C =(Y L-mean0)*C+mean0 calculates the pixel value of the Y channel of the pixel after contrast enhancement, so as to enhance the contrast of the first image.
[0181] Among them, Y C This represents the pixel value of the Y channel after contrast enhancement, C represents the contrast adjustment parameter, and mean0 represents the average value of the pixel value of the Y channel of the first image after contrast enhancement, thereby enhancing the contrast of the first image.
[0182] Understandably, although both contrast and brightness enhancement affect the pixel values of the Y channel, for the same pixel, the Y channel pixel value after contrast enhancement may differ from that after brightness enhancement—it might be larger or smaller. In other words, the Y channel pixel value after contrast enhancement might be smaller than the original Y channel pixel value, meaning the pixel brightness would be darker. However, experimental testing confirms that the overall brightness of the first image is still enhanced. Furthermore, the Y channel pixel value after contrast enhancement might be larger than that after brightness enhancement, meaning the pixel brightness would be brighter. However, experimental testing confirms that the overall brightness of the first image will not be excessively bright.
[0183] In some embodiments, since saturation represents the purity and vividness of a color, and in the YUV color space, the UV channels represent the chromaticity information of an image, i.e., the colors in the image, meaning that saturation is related to the pixel values of the U and V channels of the image, the electronic device can improve the saturation of the first image by adjusting the pixel values of the U and V channels of the pixels in the first image using saturation adjustment parameters. The following will further describe a possible implementation method for improving the saturation of the first image using saturation adjustment parameters.
[0184] The electronic device calculates the product between the pixel value of the U channel of the pixel and the saturation adjustment parameter to obtain the pixel value of the U channel of the pixel after saturation enhancement.
[0185] For each pixel in the first image, the electronic device can employ U S =U0*S, calculates the pixel value of the U channel of this pixel after saturation enhancement. Where U... S This represents the pixel value of the U channel of this pixel after saturation enhancement, where S represents the saturation adjustment parameter. U0 represents the original pixel value of the U channel of this pixel, which is the pixel value of the U channel of this pixel in the first image.
[0186] The electronic device calculates the product between the pixel value of the V channel of the pixel and the saturation adjustment parameter to obtain the pixel value of the V channel of the pixel after saturation enhancement.
[0187] In this embodiment of the application, for each pixel in the first image, the electronic device may employ V S =V0*S, calculates the pixel value of the V channel of this pixel after saturation enhancement. Where V... S V0 represents the pixel value of the V channel after saturation enhancement. V0 represents the original V channel pixel value of the pixel, which is the V channel pixel value of the pixel in the first image.
[0188] The above description is merely an example of how to enhance the saturation of a first image. Electronic devices can also use other methods to enhance the saturation of a first image. For instance, for each pixel in the first image, the electronic device can use the saturation as the exponent of the pixel's U channel to obtain the pixel value of the U channel after saturation enhancement. Furthermore, it can use the saturation as the exponent of the pixel's V channel to obtain the pixel value of the pixel after saturation enhancement, thus achieving saturation enhancement of the first image. However, the saturation enhancement effect described above is better, resulting in better color display of the first image.
[0189] In this embodiment, the electronic device enhances the brightness of the first image using brightness adjustment parameters, enhances the saturation of the first image using saturation adjustment parameters, and enhances the contrast of the first image using contrast adjustment parameters, resulting in a color-enhanced first image. The electronic device can then display this color-enhanced first image. Compared to the original first image, the color-enhanced first image has a brighter overall color, greater contrast between light and dark areas, and more vivid colors, resulting in better color display and improved image color quality. This, in turn, improves video quality and ensures a better user visual experience.
[0190] In this embodiment, when the electronic device performs color enhancement on a video stream, it first performs scene detection on the first image in the video stream to determine whether the current frame image (i.e., the first image) and the previous frame image belong to the same scene. If they belong to the same scene, the electronic device can use the color adjustment parameters corresponding to the previous frame image as the color adjustment parameters corresponding to the first image, without needing to re-predict the color adjustment parameters, thus achieving rapid determination of the color adjustment parameters and improving the processing efficiency of the video image. If they do not belong to the same scene, the electronic device can re-predict the color adjustment parameters corresponding to the first image to obtain color adjustment parameters that match the color scene of the first image. The electronic device can then use the color adjustment parameters that match the color scene of the first image to adjust the color of the first image, achieving adaptive color adjustment. Furthermore, since there is an error in the prediction of color adjustment parameters, that is, for the same scene, the predicted color adjustment parameters may have errors. Therefore, when the current frame and the previous frame belong to the same scene, directly using the color adjustment parameters corresponding to the previous frame to adjust the current frame can avoid the color display effect of two adjacent frames belonging to the same scene being different due to errors, thereby avoiding the occurrence of discontinuity between video frames (such as frame skipping, flickering, etc.).
[0191] In some embodiments, although brightness and contrast are based on the processing of pixel values in the Y channel of the first image, and texture adjustment parameters are also based on the processing of pixel values in the Y channel of the first image, texture adjustment parameters mainly adjust the pixel values in the Y channel of the first image in high-frequency regions, while brightness and contrast adjustment parameters adjust the pixel values in the Y channel of all pixels in the first image. Therefore, after adjusting the pixel values in the Y channel of pixels based on color adjustment parameters and contrast adjustment parameters, the color quality (such as brightness and contrast) and sharpness of the first image as a whole can still be improved.
[0192] For example, when it is necessary to process the pixel values of the channels of pixels in the first image using color adjustment parameters and texture adjustment parameters, for each pixel in the image, the color adjustment parameters can be used simultaneously to determine the color-enhanced pixel value of the corresponding channel of the pixel, and the texture adjustment parameters can be used to determine the sharpness-enhanced pixel value of the Y channel of the pixel. If the channel corresponding to the pixel also includes a Y channel, the target Y channel pixel value of the pixel can be determined based on the color-enhanced Y channel pixel value and the sharpness-enhanced Y channel pixel value of the pixel. For example, the average value between the two can be calculated as the target Y channel pixel value, or the color-enhanced Y channel pixel value or the sharpness-enhanced Y channel pixel value of the pixel can be used as the target Y channel pixel value.
[0193] For example, after adjusting the pixel values of the Y channel of each pixel in the first image based on brightness adjustment parameters, the pixel values of the Y channel of each pixel after brightness enhancement are obtained. The pixel values of the Y channel of pixels in the high-frequency part of the first image are adjusted using texture adjustment parameters, resulting in pixel values of the Y channel of pixels in the high-frequency part after sharpness enhancement. For the same pixel, the average value between the pixel value of the Y channel after brightness enhancement and the pixel value of the Y channel after sharpness enhancement can be calculated to obtain the target Y channel pixel value for that pixel. Although the target Y channel pixel value of the pixel may be smaller than the pixel value of the Y channel after brightness enhancement, texture adjustment targets the high-frequency pixels of the first image, so overall, the brightness of the first image is still enhanced. Furthermore, although the target Y channel pixel value of the pixel may be smaller than the pixel value of the Y channel after sharpness enhancement, the sharpness of the first image can still be improved. In summary, although adjusting the image using both texture and color adjustment parameters may change the pixel value of the Y channel of the high-frequency pixels, thus altering the brightness of those pixels, it will not affect the overall brightness display effect of the first image. Overall, the brightness and sharpness of the first image are still enhanced.
[0194] It should be noted that the above example uses color adjustment parameters including saturation, contrast, and brightness to illustrate the process of enhancing the saturation, contrast, and brightness of a first image to obtain a color-enhanced first image. Of course, color adjustment parameters can also include only one of these three parameters, or even just two. In other words, color adjustment parameters can include at least one of these three parameters. Accordingly, the electronic device can use these color adjustment parameters to enhance the colors corresponding to the first image, obtaining a color-enhanced first image. For example, if the color adjustment parameters include brightness, the electronic device can use the brightness adjustment parameter to enhance the brightness of the first image, obtaining a color-enhanced first image. As another example, if the color adjustment parameters include both brightness and saturation, the electronic device can use the brightness adjustment parameter to enhance the brightness of the first image, obtaining a color-enhanced first image, and then use the saturation adjustment parameter to enhance the saturation of the first image.
[0195] S408, The electronic device displays the first image of the target.
[0196] In this embodiment of the application, after the electronic device obtains the texture-enhanced first image, that is, the target first image, it can display the first image to realize real-time processing and display of the video stream, ensure the clarity of the displayed video image, improve the image quality, thereby improve the video quality and ensure the user's visual experience.
[0197] In addition, after color enhancement of the first image, the color display effect of the target first image was also improved.
[0198] In some embodiments, where the first image in the aforementioned video stream includes a face region, such as a selfie image, the electronic device can determine the high-frequency information of the face image in the first image. Then, for each pixel in the high-frequency information of the face image in the first image, the electronic device multiplies the pixel value of the Y channel of that pixel by a texture adjustment parameter to obtain the target first image. This achieves texture enhancement of the region of interest, ensuring the texture enhancement effect of the image and reducing the amount of data processing, thereby reducing power consumption. Specifically, as shown... Figure 8 As shown, the electronic device can use a segmentation network to segment the face region and background in the first image to obtain the face image and background image of the first image, that is, to obtain the pixel value of the Y channel of the pixel in the face image and the pixel value of the Y channel of the pixel in the background image.
[0199] After obtaining the face image, the electronic device can first use mean filtering to blur the face image, resulting in a blurred face image. Then, the electronic device can determine the high-frequency information of the blurred face image, obtaining its high-frequency information. Next, the pixel value of the Y channel of each pixel in the high-frequency information of the face image is multiplied by a texture adjustment parameter to obtain the texture adjustment value of the Y channel corresponding to each pixel in the high-frequency information, which is the texture adjustment value of the Y channel corresponding to each pixel in the face image. The specific process of determining the high-frequency information of the blurred face image can be referred to the process of determining the high-frequency information of the first image described above, and will not be repeated here.
[0200] Subsequently, for each pixel in the face image of the first image, the electronic device can add the pixel value of the Y channel of that pixel (i.e. the original pixel value of the Y channel) to the texture adjustment value at the corresponding position to obtain the pixel value of the Y channel after texture enhancement, thereby realizing texture enhancement of the high-frequency information of the face part of the first image.
[0201] Subsequently, the electronic device can synthesize the texture-enhanced Y-channel pixel values of the face image in the first image with the UV-channel pixel values of the face image in the first image to obtain a texture-enhanced face image. Furthermore, the electronic device can combine this texture-enhanced face image with the background image to obtain the target first image, achieving texture enhancement of the region of interest in the first image. This enhances facial details while avoiding over-sharpening the background, thus preventing the introduction of excessive background noise. However, this increases data processing volume and power consumption.
[0202] Optionally, the structure of the above segmentation network model is as follows: Figure 9 As shown, the segmentation network model outputs a binary image (or binary mask image) corresponding to the first image. In this binary image, the pixel values of the face portion can be 1, and the pixel values of the background portion are 0, thus achieving face segmentation of the first image and determining the face image of the first image. It should be understood that the above... Figure 9 The segmentation network model shown is just an example. The structure of this segmentation network model can be other structures, as long as it can segment the face part of the image.
[0203] In some embodiments, the aforementioned electronic device can directly utilize a pure deep learning texture enhancement algorithm, that is, a pure deep learning model, to enhance the texture of the first image to obtain the target first image. For example... Figure 10 As shown, the electronic device can input the pixel values of the Y channel of the first image into a pure deep learning model. This model performs relevant processing on the Y channel pixel values (downsampling, double convolution, and upsampling) to obtain the texture-enhanced Y channel pixel values of the first image. Then, the electronic device can synthesize the texture-enhanced Y channel pixel values and the texture-enhanced UV channel pixel values of the first image to obtain the target first image.
[0204] Optionally, image synthesis can also be performed using a pure deep learning model. Accordingly, the electronic device can input the first image into the pure deep learning model, and the pure deep learning model can directly output the target first image to achieve texture enhancement of the first image.
[0205] In some embodiments, the electronic device can not only perform texture enhancement on images in a video in real time, such as performing texture enhancement on the first image after obtaining the first image in the video stream, as described above, the electronic device can also perform texture enhancement on images in a video file that is already stored locally, as long as it can obtain the pixel values (such as the pixel values of the YUV channels) of the pixels in the images in the video.
[0206] It should be noted that for the first image received, that is, the first frame of the first image, the electronic device can directly input it into the trained first MobileNetV3 model to obtain the texture adjustment parameters corresponding to the first frame of the first image, thereby realizing scene detection of the first frame of the first image.
[0207] In some embodiments, the first device and the electronic device described above may be the same device or different devices, and this application does not limit them. In some embodiments, this application provides a computer storage medium including computer instructions, which, when executed on an electronic device, cause the electronic device to perform the method described above.
[0208] In some embodiments, this application provides a computer program product that, when run on an electronic device, causes the electronic device to perform the method described above.
[0209] In the above embodiments, implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented in software, it can be implemented, in whole or in part, as a computer program product. The computer program product includes one or more computer instructions. When these computer instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated.
[0210] It should be understood that the term "embodiment" used throughout the specification means that a specific feature, structure, or characteristic related to an embodiment is included in at least one embodiment of this application. Therefore, various embodiments throughout the specification do not necessarily refer to the same embodiment. Furthermore, these specific features, structures, or characteristics can be combined in any suitable manner in one or more embodiments. It should be understood that in the various embodiments of this application, the sequence numbers of the above processes do not imply a sequential order of execution; the execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.
[0211] It should also be understood that in this application, “when…”, “if” and “if” all refer to the UE or base station taking corresponding actions under certain objective circumstances, and are not time-limited, nor do they require the UE or base station to perform a judgment action, nor do they imply any other limitations.
[0212] Those skilled in the art will understand that the various numerical designations, such as "first" and "second," used in this application are merely for descriptive convenience and are not intended to limit the scope of the embodiments of this application, nor to indicate a sequential order. Elements represented by singular numbers in this application are intended to represent "one or more," not "one and only one," unless otherwise specified. In this application, unless otherwise specified, "at least one" is intended to represent "one or more," and "more" is intended to represent "two or more."
[0213] The term "and / or" in this document merely describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, and B existing alone. A can be singular or plural, and B can be singular or plural. The terms "at least one of..." or "at least one of..." in this document indicate all or any combination of the listed items. For example, "at least one of A, B, and C" can represent: A existing alone, B existing alone, C existing alone, A and B existing simultaneously, B and C existing simultaneously, and A, B, and C existing simultaneously. A can be singular or plural, B can be singular or plural, and C can be singular or plural.
[0214] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application. Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0215] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.
[0216] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0217] In addition, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.
[0218] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0219] The same or similar parts between the various embodiments in this application can be referred to mutually. In the various embodiments of this application, and in the various implementation methods / methods / implementations within each embodiment, unless otherwise specified or logically conflicting, the terminology and / or descriptions between different embodiments and between the various implementation methods / methods / implementations within each embodiment are consistent and can be mutually referenced. The technical features in different embodiments and the various implementation methods / methods / implementations within each embodiment can be combined according to their inherent logical relationships to form new embodiments, implementation methods, methods, or implementation approaches. The above-described embodiments of this application do not constitute a limitation on the scope of protection of this application.
[0220] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of protection of the claims. In conclusion, the above description is merely a preferred embodiment of the technical solution of this application and is not intended to limit the scope of protection of this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of protection of this application.
Claims
1. A texture enhancement method, characterized in that, include: Acquire a video stream; wherein the video stream includes at least one first frame; For each frame of the at least one first image, a scene detection result corresponding to the first image is obtained based on the difference between the pixel value of the first target channel of the first image and the pixel value of the first target channel of the previous frame of the first image, the difference between the pixel value of the second target channel of the first image and the pixel value of the second target channel of the previous frame of the first image, and the difference between the pixel value of the third target channel of the first image and the pixel value of the third target channel of the previous frame of the first image; wherein, the scene detection result corresponding to the first image indicates whether the first image and the previous frame image belong to the same scene; the first target channel is the hue H channel, the second target channel is the saturation S channel, and the third target channel is the luminance V channel; the first target channel, the second target channel, and the third target channel are included in the target channel; Wherein, if the average of the differences corresponding to the first target channel, the second target channel, and the third target channel is less than the first threshold, the scene detection result indicates that the first image and the previous frame image belong to the same scene; If the average of the differences corresponding to the first target channel, the second target channel, and the third target channel is greater than the first threshold, the scene detection result indicates that the first image and the previous frame image do not belong to the same scene. If the scene detection result indicates that the first image and the previous frame image do not belong to the same scene, the texture adjustment parameters corresponding to the first image are determined based on the first image; the texture adjustment parameters corresponding to the first image are texture adjustment parameters that match the scene of the first image. The texture adjustment parameters are obtained based on a trained image training model. During the training of the image training model, a first sample image set is used, which includes multiple original first sample images. The first sample image set includes various types of image sets. The image training model is obtained based on at least one original image in the first sample image set, the corresponding degraded sample image, and the corresponding complexity. The image training model uses four tangent quadratic curves to fit a curve, obtaining a mapping curve between texture complexity and texture sharpness intensity. The corresponding degraded image and its corresponding complexity are obtained as follows: For each of the multiple original first sample images, the original first sample image is downsampled to obtain a downsampled original first sample image; the downsampled original first sample image is upsampled to obtain a degraded image corresponding to the original first sample image; wherein, the size of the degraded image corresponding to the original first sample image is equal to that of the original first sample image; the variance of the original first sample image is calculated to obtain the complexity corresponding to the original first sample image. The method further includes: The first image is subjected to mean filtering to obtain the first filtered image; For each pixel in the first image, the pixel value of the Y channel of that pixel is subtracted from the pixel value of the Y channel of the corresponding pixel in the first filtered image to obtain the high-frequency information of the first image. For each pixel in the high-frequency information of the first image, the pixel value of the Y channel of the pixel is multiplied by the texture adjustment parameter to obtain the target first image; Display the first image of the target.
2. The method according to claim 1, characterized in that, The method further includes: For each pixel in the first image, the difference between the pixel value of the target channel of the pixel and the pixel value of the target channel of the corresponding pixel in the previous frame image is calculated to obtain the difference of the target channel of the pixel in the first image. Calculate the average of the differences between the target channels corresponding to each pixel in the first image to obtain the difference between the pixel value of the target channel in the first image and the pixel value of the target channel in the previous frame image.
3. The method according to claim 2, characterized in that, The number of target channels is one or more; The step of calculating the average of the differences between the target channels corresponding to each pixel in the first image to obtain the difference between the pixel values of the target channels in the first image and the pixel values of the target channels in the previous frame image includes: For each target channel, the average of the differences between the target channels corresponding to each pixel is calculated to obtain the difference between the target channels corresponding to the first image; Calculate the average of the differences between each target channel corresponding to the first image to obtain the difference between the pixel value of the target channel of the first image and the pixel value of the target channel of the previous frame image.
4. The method according to any one of claims 1 to 3, characterized in that, The scene detection result corresponding to the first image is obtained by taking the difference between the pixel value of the target channel of the first image and the pixel value of the target channel of the previous frame of the first image, including: If the difference between the pixel value of the target channel of the first image and the pixel value of the target channel of the previous frame image is less than or equal to a first threshold, the scene detection result corresponding to the first image is determined to indicate that the first image and the previous frame image belong to the same scene. If the difference between the pixel value of the target channel of the first image and the pixel value of the target channel of the previous frame image is greater than the first threshold, the scene detection result corresponding to the first image indicates that the first image and the previous frame image do not belong to the same scene.
5. The method according to any one of claims 1 to 4, characterized in that, The step of adjusting the clarity of high-frequency information in the first image based on the texture adjustment parameters corresponding to the first image includes: Based on the texture adjustment parameters, the pixel values of each pixel in the high-frequency information of the first image are adjusted respectively to obtain the texture adjustment value of each pixel in the high-frequency information; For each pixel in the first image, the texture-enhanced pixel value is obtained based on the pixel value of the pixel and the texture adjustment value corresponding to the pixel at the corresponding position in the high-frequency information.
6. The method according to claim 5, characterized in that, The texture adjustment value of the pixel includes the texture adjustment value of the pixel's brightness Y channel, and the pixel value after texture enhancement includes the pixel value of the Y channel after texture enhancement. The step of adjusting the pixel values of each pixel in the high-frequency information of the first image based on the texture adjustment parameters to obtain the texture adjustment value of each pixel in the high-frequency information includes: The texture adjustment value of the Y channel of each pixel in the high-frequency information is obtained by multiplying the pixel value of the Y channel of each pixel in the high-frequency information with the texture adjustment parameter. For each pixel in the first image, the process of obtaining the texture-enhanced pixel value of the pixel based on the pixel value of the pixel and the texture adjustment value corresponding to the pixel at the corresponding position in the high-frequency information includes: For each pixel in the first image, the pixel value of the Y channel of the pixel and the texture adjustment value of the Y channel of the corresponding pixel in the high-frequency information are calculated to obtain the texture-enhanced Y channel pixel value of the pixel.
7. The method according to any one of claims 1 to 6, characterized in that, The step of determining the texture adjustment parameters corresponding to the first image based on the first image includes: The first image is used as the input parameter of the image processing model, and the image processing model is run to obtain the texture adjustment parameters corresponding to the first image; the image processing model is a lightweight model.
8. The method according to any one of claims 1 to 7, characterized in that, The scene detection result corresponding to the first image is obtained by taking the difference between the pixel value of the target channel of the first image and the pixel value of the target channel of the previous frame of the first image, including: The first image is downsampled to obtain the downsampled first image; The scene detection result corresponding to the first image is obtained based on the difference between the pixel value of the target channel of the downsampled first image and the pixel value of the target channel of the previous frame image after downsampling the first image.
9. An electronic device, characterized in that, The electronic device includes a display screen, a memory, and one or more processors; the display screen, the memory, and the processors are coupled; the display screen is used to display an image generated by the processor, the memory is used to store computer program code, the computer program code including computer instructions; when the processor executes the computer instructions, the electronic device performs the method as described in any one of claims 1 to 8.
10. A computer storage medium, characterized in that, Includes computer instructions that, when executed on an electronic device, cause the electronic device to perform the method as described in any one of claims 1 to 8.
Citation Information
Patent Citations
Scene switching detection method and device, electronic equipment and storage medium
CN110675371A
Image color adjustment method and device based on scene, and computer equipment
CN111062860A
Image quality adjusting method and device
CN114363693A
Photographing method and related device
CN116095513A