Image processing method and device, storage medium and program product

By recognizing and sending facial images to the server in real time during live streaming for unified beautification, the problem of inconsistent beautification effects across different live streaming applications is solved, achieving stability of beautification effects and efficient data transmission.

CN121509689APending Publication Date: 2026-02-10GUANGZHOU KUGOU COMP TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511579497.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-31
Publication Date
2026-02-10

AI Technical Summary

Technical Problem

Across different live streaming applications, it is difficult for streamers to achieve consistent beauty filter effects, resulting in a lack of stability in the beauty filter results.

Method used

By performing image recognition on image frames in real time during live streaming, facial area images are obtained and sent to the server for unified portrait beautification processing. The server executes portrait beautification technology to ensure consistent beautification effects across different applications.

Benefits of technology

It improves the stability and consistency of beautification effects, reduces the bandwidth required for data transmission, and improves data transmission efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121509689A_ABST
    Figure CN121509689A_ABST
Patent Text Reader

Abstract

The invention discloses an image processing method and device, a storage medium and a program product, and relates to the technical field of computers. The method comprises the following steps: acquiring a live video stream of a first anchor account, wherein the live video stream of the first anchor account comprises a first image frame; image recognition is carried out on the first image frame, a first local image corresponding to the first image frame is obtained, and the first local image comprises a face area in the first image frame; sending the first local image to a server; the processed first local image sent by the server is received, and the processed first local image is obtained after the first local image is processed through the portrait beautifying technology; and obtaining the processed first image frame according to the first image frame and the processed first local image. According to the method and the device, the beautifying processing part of the image frame is executed by the server, so that an anchor can achieve the same portrait beautifying effect in different live broadcast application programs, and the stability and the consistency of the beautifying effect are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and in particular to an image processing method, device, storage medium, and program product. Background Technology

[0002] When live streaming on a live streaming platform, hosts usually use beauty filters to enhance their faces in order to achieve a better live streaming effect.

[0003] In related technologies, the host starts a live broadcast on a live broadcast application and adjusts the beauty parameters suitable for the host based on the beauty function provided in the live broadcast application. The host's face is beautified by the beauty parameters adjusted by the host, so that the final beauty effect is suitable for the host.

[0004] However, if a streamer uses different live streaming applications, it is difficult to adjust the same beauty filter effect across different applications, resulting in a lack of consistency in the beauty filter effect. Summary of the Invention

[0005] This application provides an image processing method, apparatus, storage medium, and program product. The technical solutions provided by this application are as follows: According to one aspect of the embodiments of this application, an image processing method is provided, the method comprising: Obtain the live video stream of the first streamer account, wherein the live video stream of the first streamer account includes the first image frame; Image recognition is performed on the first image frame to obtain a first local image corresponding to the first image frame, and the first local image includes the face region in the first image frame; Send the first partial image to the server; The server sends a processed first partial image, wherein the processed first partial image is obtained by processing the first partial image using portrait beautification technology; Based on the first image frame and the processed first local image, the processed first image frame is obtained.

[0006] According to one aspect of the embodiments of this application, an image processing apparatus is provided, the apparatus comprising: The video stream acquisition module is used to acquire the live video stream of the first broadcaster account, wherein the live video stream of the first broadcaster account includes the first image frame; An image recognition module is used to perform image recognition on the first image frame to obtain a first local image corresponding to the first image frame, wherein the first local image includes the face region in the first image frame; An image sending module is used to send the first partial image to the server; An image receiving module is used to receive a processed first partial image sent by the server, wherein the processed first partial image is obtained by processing the first partial image using portrait beautification technology; The image processing module is used to obtain the processed first image frame based on the first image frame and the processed first local image.

[0007] According to one aspect of the embodiments of this application, an image processing method is provided, the method comprising: The system receives a first partial video stream sent by a first client. The first partial video stream includes a first partial image. The first partial image is obtained by the first client through image recognition of a first image frame in the live video stream of the first broadcaster account. The first partial image includes the face region in the first image frame. The first partial image is processed using portrait beautification technology to obtain the processed first partial image; The processed first partial image is sent to the first client.

[0008] According to one aspect of the embodiments of this application, an image processing apparatus is provided, the apparatus comprising: The video stream receiving module is used to receive a first partial video stream sent by a first client. The first partial video stream includes a first partial image. The first partial image is obtained by the first client performing image recognition on a first image frame in the live video stream of the first broadcaster account. The first partial image includes the face region in the first image frame. The image enhancement module is used to process the first partial image using portrait enhancement technology to obtain the processed first partial image; An image sending module is used to send the processed first partial image to the first client.

[0009] According to one aspect of the embodiments of this application, a computer device is provided, the computer device including a processor and a memory, the memory storing a computer program, the computer program being loaded and executed by the processor to implement the above-described image processing method.

[0010] According to one aspect of the embodiments of this application, a computer-readable storage medium is provided, wherein a computer program is stored in the computer-readable storage medium, the computer program being loaded and executed by a processor to implement the above-described image processing method.

[0011] According to one aspect of the embodiments of this application, a computer program product is provided, the computer program product including a computer program, the computer program being loaded and executed by a processor to implement the above-described image processing method.

[0012] The technical solution provided in this application can bring the following beneficial effects: By performing real-time image recognition on the first image frame during the live stream of the first broadcaster's account, obtaining the first partial image corresponding to the first image frame, and sending the first partial image to the server for beautification processing, the lower resolution of the first partial image reduces the bandwidth required for data transmission, thus improving data transmission efficiency. Furthermore, by offloading the portrait beautification processing of the first image frame to the server, the server applies a unified portrait beautification technology to live stream images from different live streaming applications. This ensures that the broadcaster achieves the same portrait beautification effect when broadcasting across different applications. Compared to related technologies that use client-side adjustments of beautification parameters, the technical solution provided in this application improves the stability and consistency of the beautification effect while maintaining its effectiveness. Attached Figure Description

[0013] Figure 1 This is a schematic diagram of a computer system provided in one embodiment of this application; Figure 2 This is a flowchart of an image processing method for a single-person live streaming scenario provided in one embodiment of this application; Figure 3 This is a schematic diagram of the image processing flow in a single-person live streaming scenario provided in one embodiment of this application; Figure 4 This is a flowchart of an image processing method for a multi-person live streaming scenario provided in one embodiment of this application; Figure 5 This is a schematic diagram of the image processing flow in a two-person live streaming scenario provided in one embodiment of this application; Figure 6 This is a schematic diagram of the image stitching process in a two-person live streaming scenario provided in one embodiment of this application; Figure 7 This is a flowchart of an image processing method for a single-person live streaming scenario provided in another embodiment of this application; Figure 8 This is a flowchart of an image processing method for a multi-person live streaming scenario provided in another embodiment of this application; Figure 9 This is a flowchart of an image processing method for a single-person live streaming scenario provided in another embodiment of this application; Figure 10This is a flowchart of an image processing method for a multi-person live streaming scenario provided in another embodiment of this application; Figure 11 This is a block diagram of an image processing apparatus provided in one embodiment of this application; Figure 12 This is a block diagram of an image processing apparatus provided in another embodiment of this application; Figure 13 This is a structural block diagram of a computer device provided in one embodiment of this application. Detailed Implementation

[0014] To make the objectives, technical solutions, and advantages of this application clearer, the embodiments of this application will be described in further detail below with reference to the accompanying drawings.

[0015] Please refer to Figure 1 This illustration shows a schematic diagram of a computer system provided in one embodiment of this application. This computer system can be implemented as a live interactive system. The computer system may include: a broadcaster terminal device 10, a viewer terminal device 20, and a server 30.

[0016] The broadcaster terminal device 10 refers to the terminal device used by the broadcaster, and the viewer terminal device 20 refers to the terminal device used by the viewer. The broadcaster refers to the user participating in the live broadcast, and the viewer refers to the user watching the live broadcast in which the broadcaster participates. The broadcaster terminal device 10 and the viewer terminal device 20 (collectively referred to as "terminal devices") can be electronic devices such as mobile phones, tablets, laptops, desktop computers, game consoles, e-book readers, multimedia playback devices, wearable devices, and in-vehicle terminals.

[0017] The broadcast terminal device 10 includes a first terminal device and a second terminal device. The first terminal device is used to log in with the first broadcaster's account, and the first broadcaster's account is a user account held by the first broadcaster. The second terminal device is used to log in with the second broadcaster's account, and the second broadcaster's account is a user account held by the second broadcaster. A camera is installed on the broadcast terminal device 10 to capture the live broadcast footage.

[0018] Both the broadcaster terminal device 10 and the viewer terminal device 20 run the target application, which is used to implement the functions involved in this application, such as starting a live broadcast, watching a live broadcast, and participating in live broadcast interactions. This application does not limit the type of the target application; for example, the target application can be a live broadcast application, a short video application, a game application, etc. Optionally, the target application can be an application that requires downloading and installation, or it can be an application that can be used instantly; this application does not limit the type of application.

[0019] The number of broadcaster terminal devices 10 can be one or more. Optionally, each broadcaster can start a solo live stream on a broadcaster terminal device 10, in which case the broadcaster's live stream screen will be displayed in the live stream interface of the target application. For example, if the first broadcaster account can start a solo live stream on the first terminal device, then the first broadcaster's live stream screen will be displayed in the live stream interface of the target application. Optionally, each broadcaster can also start a joint live stream with other broadcasters on their respective broadcaster terminal devices 10, in which case the broadcaster's live stream screens of multiple broadcasters will be displayed simultaneously in the live stream interface of the target application. For example, if the first broadcaster account and the second broadcaster account can start a joint live stream on their respective terminal devices, then the live stream screen of the first broadcaster and the second broadcaster will be displayed in a combined live stream in the live stream interface of the target application. The number of viewer terminal devices 20 can be one or more, and viewers can select to watch the live stream of broadcasters they are interested in within the target application.

[0020] Server 30 provides background services to the clients of the target application installed and running on the broadcast terminal device 10 and the viewer terminal device 20. For example, server 30 can be the background server of the aforementioned target application. Server 30 can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms, but is not limited to these. Optionally, server 30 can simultaneously provide background services to the target application in multiple broadcast terminal devices 10 and multiple viewer terminal devices 20. Broadcast terminal devices 10 and viewer terminal devices 20 can communicate with server 30 through the network respectively.

[0021] In this embodiment, the first client of the target application logged in by the first broadcaster account obtains the live video stream of the first broadcaster account, which includes a first image frame. The first client performs image recognition on the first image frame to obtain a first partial image corresponding to the first image frame, the first partial image including the facial region in the first image frame. Then, the first client sends the first partial image to the server, which processes the first partial image using portrait beautification technology to obtain a processed first partial image. The first client receives the processed first partial image sent by the server. Finally, the first client obtains a processed first image frame based on the first image frame and the processed first partial image, and displays the processed first image frame on the first client.

[0022] Please refer to Figure 2This document illustrates a flowchart of an image processing method for a single-person live streaming scenario according to an embodiment of this application. The execution entity for each step of this method can be a first terminal device, specifically referring to the streamer terminal device corresponding to the aforementioned first streamer account. The method may include at least one of the following steps 210-250: Step 210: Obtain the live video stream of the first streamer account, which includes the first image frame.

[0023] The first streamer account is the user account held by the first streamer. A streamer refers to a user participating in the live stream interaction. The first streamer conducts live streaming on the first client of the target application through their first streamer account. The first streamer turns on their camera during the live stream, and the live video stream of the first streamer account refers to the continuous image frames captured by the camera during the live stream. The live video stream of the first streamer account includes the first image frame, which is a real-time image frame captured by the camera. The first client acquires the first image frame in real time, processes the acquired first image frame, and then displays the processed first image frame.

[0024] Specifically, when the first anchor chooses to enter the camera's shooting range, the first main image frame is an image frame captured focusing on the anchor's image; when the first anchor chooses not to enter the camera's shooting range, the first image frame is an image frame captured focusing on any scene the camera is facing. The image processing method involved in this application pertains to the first image frame containing the anchor's image.

[0025] Step 220: Perform image recognition on the first image frame to obtain a first local image corresponding to the first image frame. The first local image includes the face region in the first image frame.

[0026] The first local image is an image composed of local regions in the first image frame, including the face region in the first image frame.

[0027] In some embodiments, face recognition is performed on the first image frame to obtain a face image corresponding to the first image frame, and the face image corresponding to the first image frame is determined as a first local image.

[0028] Face recognition is used to identify the face region in the first image frame to obtain the face image corresponding to the first image frame. The face image includes the region corresponding to the face of the first anchor in the first image frame. Then the first local image includes the complete face region in the first image frame.

[0029] In some embodiments, facial region recognition is performed on the first image frame to obtain a facial region image corresponding to the first image frame, and the facial region image corresponding to the first image frame is determined as a first local image, wherein the facial region image includes key facial regions in the first image frame.

[0030] The face region recognition is used to identify key facial regions in the first image frame to obtain the face region image corresponding to the first image frame. The face region image includes the region corresponding to the key facial position of the first anchor in the first image frame. Then the first local image includes part of the face region in the first image frame.

[0031] For example, key facial regions include, but are not limited to, any key regions on the face of the first anchor, such as the eye region, mouth region, nose region, eyebrow region, and cheek region. Correspondingly, the facial region image corresponding to the first image frame refers to the eye image, mouth image, nose image, eyebrow image, cheek image, etc., corresponding to the first image frame.

[0032] In some embodiments, head recognition is performed on the first image frame to obtain the head image corresponding to the first image frame, and the head image corresponding to the first image frame is determined as the first local image.

[0033] Head recognition is used to identify the head region in the first image frame to obtain the head image corresponding to the first image frame. The head image includes the region corresponding to the head of the first anchor in the first image frame. The first partial image includes the complete face region in the first image frame and other head regions in the first image frame besides the face region, such as the hair region and neck region. For example, the first partial image may include the complete face region and hair region in the first image frame, or it may include the complete face region, hair region and neck region in the first image frame.

[0034] By applying different image recognition methods to the first image frame, a first local image containing different regions of the first image frame can be obtained, enabling image processing of different regions of the first image frame. For example, the face in the first image frame can be processed in a targeted manner, or the eyes in the first image frame can be processed in a targeted manner, etc., so as to improve the performance of image processing and enhance image processing capabilities.

[0035] In some embodiments, image recognition is performed on a first image frame to obtain a first local image and a first other image corresponding to the first image frame. The first other image includes the region in the first image frame other than the first local image. The first image frame includes a first local image and a first other image.

[0036] Optionally, face recognition is performed on the first image frame to obtain a face image and a first other image corresponding to the first image frame. The face image corresponding to the first image frame is determined as a first local image, and the first other image includes other regions in the first image frame other than the face region.

[0037] Optionally, facial region recognition is performed on the first image frame to obtain a facial region image corresponding to the first image frame. The facial region image corresponding to the first image frame is determined as a first local image. The facial region image includes the key facial region in the first image frame, and the first other image includes other regions in the first image frame other than the key facial region.

[0038] Optionally, head recognition is performed on the first image frame to obtain the head image corresponding to the first image frame, and the head image corresponding to the first image frame is determined as the first local image. The first other images include other regions in the first image frame other than the head region.

[0039] Step 230: Send the first partial image to the server.

[0040] In some embodiments, the first partial image and its image information are encoded to obtain encoded data of the first partial image, wherein the image information of the first partial image includes timestamp information of the first partial image; wherein the timestamp information of the first partial image is used to indicate the timestamp of the first image frame in the live video stream of the first broadcaster account; and the encoded data of the first partial image is sent to the server.

[0041] The timestamp information of the first partial image can also be used to indicate the timestamp of the first partial image in the live video stream of the first broadcaster's account.

[0042] In some embodiments, the image information of the first local image further includes the position information of the first local image, which is used to indicate the position of the first local image in the first image frame.

[0043] In some embodiments, the image information of the first partial image further includes facial recognition point information of the first partial image, which indicates the position of at least one facial key point in the first partial image. For example, facial key points may include eyes, eyebrows, nose, mouth, etc., and the facial recognition point information of the first partial image includes the position of facial key points such as eyes, eyebrows, nose, and mouth in the first partial image. The server can perform portrait beautification processing on at least one facial key point in the first partial image based on the facial recognition point information of the first partial image.

[0044] It is important to note that zero-latency coding is used here to encode the first partial image and its image information, resulting in the encoded data of the first partial image. Furthermore, a low-latency transmission protocol is employed to send the encoded data of the first partial image to the server, thereby reducing both encoding and transmission latency.

[0045] By encoding the first partial image and its image information, the first partial image carries its timestamp information. This allows the server to process each first partial image in sequence based on its timestamp information, avoiding the situation where prioritizing the processing of first partial images with later timestamps would cause disorder in the order of the subsequently processed first partial images, thus affecting the live broadcast.

[0046] Step 240: Receive the processed first partial image sent by the server, wherein the processed first partial image is obtained by processing the first partial image using portrait beautification technology.

[0047] The aforementioned portrait enhancement technology is used to enhance the facial area indicated by the first partial image to optimize the portrait image of the first anchor. The portrait enhancement technology includes at least one of the following: portrait whitening technology, portrait skin smoothing technology, portrait reshaping technology, portrait makeup technology, etc. Among them, portrait whitening technology is used to whiten and brighten the skin area in the image. Portrait skin smoothing technology is used to uniformly process the skin area in the image, removing blemishes and adjusting skin tone, optimizing visual effects while ensuring defect repair. Portrait reshaping technology is used to adjust facial features in the image, such as adjusting the shape and size of the eyes, mouth, eyebrows, face shape, etc. Portrait makeup technology is used to apply makeup to the facial area in the image, adding makeup effects to the original face, such as adding eyeshadow, lipstick, nose contouring, blush, etc.

[0048] In general, portrait beautification technology can also be called beauty enhancement technology.

[0049] In some embodiments, a first client sends encoded data of a first partial image to a server. After receiving the encoded data stream of the first partial image from the first client, the server decodes the encoded data of the first partial image to obtain a decoded first partial image and its image information. The decoded first partial image is then processed using portrait beautification technology to obtain a processed first partial image. The server then encodes the processed first partial image and its image information to obtain encoded data of the processed first partial image. The image information of the processed first partial image includes timestamp information, and may also include at least one of location information and facial recognition point information. The timestamp information indicates the timestamp of the processed first partial image in the live video stream of the first broadcaster's account. The location information indicates the position of the processed first partial image in the first image frame. The facial recognition point information indicates the position of at least one facial key point in the processed first partial image within the processed first partial image.

[0050] The server sends the encoded data of the processed first partial image to the first client. After receiving the encoded data of the processed first partial image sent by the server, the first client decodes the encoded data of the processed first partial image to obtain the processed first partial image and the image information of the processed first partial image.

[0051] It is important to note that zero-latency coding is used here to encode the processed first partial image and its image information, resulting in the encoded data of the processed first partial image. Furthermore, a low-latency transmission protocol is employed to send the encoded data of the processed first partial image to the first client, thereby reducing both encoding and transmission latency.

[0052] For example, if the above image processing steps are not performed using zero-latency coding and low-latency transmission protocols, it will take 200ms. If the above image processing steps are performed using zero-latency coding and low-latency transmission protocols, it will take 70ms for the first client to send the encoded data of the first partial image to the server, 70ms for the server to send the encoded data of the processed first partial image to the first client, 30ms for the encoding and decoding process in the first client and the server, and 20ms for the image processing process in the server using portrait beautification technology. The total time is less than the original 200ms, so the encoding latency and transmission latency in the entire image processing process can be reduced.

[0053] Step 250: Obtain the processed first image frame based on the first image frame and the processed first local image.

[0054] In some embodiments, if image recognition is performed on the first image frame to obtain only the first partial image corresponding to the first image frame, then the image information carried by the processed first partial image includes the timestamp information and the position information of the processed first partial image. Step 250 includes sub-step 251.

[0055] Sub-step 251: Based on the timestamp information of the processed first partial image, obtain the first image frame from the live video stream of the first anchor account; based on the position information of the processed first partial image, the processed first partial image, and the first image frame, obtain the processed first image frame.

[0056] Based on the timestamp information of the processed first partial image, a first image frame with the same timestamp information is obtained from the live video stream of the first broadcaster account, thus obtaining a first image frame that matches the processed first partial image. Based on the position information of the processed first partial image, the position of the processed first partial image in the first image frame is determined. Then, based on the position of the processed first partial image in the first image frame, the image region at the corresponding position in the first image frame is replaced by the processed first partial image, thus obtaining the processed first image frame.

[0057] By first obtaining the first image frame with the same timestamp information from the live video stream of the first broadcaster's account based on the timestamp information, image processing errors can be avoided, ensuring matching processing between image frames. Furthermore, by replacing the image region at the corresponding position in the first image frame, the aesthetics of the image after replacement can be guaranteed, ensuring the effectiveness of image processing.

[0058] In some embodiments, if image recognition is performed on the first image frame to obtain a first partial image and a first other image corresponding to the first image frame, then the image information carried by the processed first partial image includes the timestamp information of the processed first partial image. Step 250 includes sub-step 252.

[0059] Sub-step 252: Based on the timestamp information of the processed first partial image, obtain the first image frame from the live video stream of the first anchor account; stitch the first other image and the processed first partial image together to obtain the processed first image frame.

[0060] Based on the timestamp information of the processed first partial image, a first other image with the same timestamp information is obtained from the live video stream of the first broadcaster's account, i.e., a first other image that matches the processed first partial image. The first other image and the processed first partial image are then stitched together to obtain the processed first image frame.

[0061] By acquiring the first other image during image recognition, the first other image and the processed first local image can be directly stitched together, simplifying the image processing steps and avoiding the need to acquire additional positional information of the first local image, which would increase the amount of data required for image processing and affect the efficiency of data encoding and data transmission, thereby improving the efficiency of image processing.

[0062] In some embodiments, to make the stitching of the first other image and the processed first partial image more accurate, the position information of the first partial image can be added during image encoding. In this case, the image information carried by the processed first partial image includes the timestamp information and the position information of the processed first partial image. Therefore, based on the timestamp information of the processed first partial image, a first image frame is obtained from the live video stream of the first broadcaster account; based on the position information of the processed first partial image, the first other image and the processed first partial image are stitched together to obtain the processed first image frame.

[0063] In some embodiments, after displaying the processed first image frame, the first client, in response to an operation on the live stream, displays a reprocessed first image frame. If the operation on the live stream is performed by the user to add filters or effects to the live stream, then the reprocessed first image frame is an image frame with an overlaid filter or effect layer.

[0064] The technical solution provided in this application involves real-time image recognition of the first image frame during the live stream of the first broadcaster's account to obtain a first partial image corresponding to the first image frame. This first partial image is then sent to the server for beautification processing. Since the first partial image has a lower resolution, the bandwidth required for data transmission is reduced, thus improving data transmission efficiency. Furthermore, by delegating the portrait beautification processing of the first image frame to the server, the server applies a unified portrait beautification technology to live stream images from different live streaming applications. This ensures that the broadcaster achieves the same portrait beautification effect when broadcasting in different applications. Compared to related technologies that use client-side adjustments of beautification parameters, the technical solution provided in this application improves the stability and consistency of the beautification effect while maintaining its effectiveness.

[0065] Figure 3 This diagram illustrates the image processing flow in a single-person live streaming scenario. After acquiring a first image frame in real-time, the first client performs image recognition on the first image frame to obtain a first partial image corresponding to the first image frame, and sends this first partial image to the server. The server uses portrait beautification technology to beautify the first partial image, obtaining a processed first partial image, and sends the processed first partial image to the first client. The first client receives the processed first partial image sent by the server, and based on the first image frame and the processed first partial image, obtains a processed first image frame. The user can customize the processed first image frame, such as adding filters or effects, and then the first client streams and displays the further processed first image frame.

[0066] Please refer to Figure 4 This document illustrates a flowchart of an image processing method for a multi-user live streaming scenario according to an embodiment of this application. The executing entity for each step of this method can be a first terminal device. The method may include at least one of the following steps 410-460: Step 410: Obtain the live video stream of the first broadcaster account and the live video stream corresponding to at least one second broadcaster account. The live video stream of the first broadcaster account includes a first image frame, and the live video stream of the second broadcaster account includes a second image frame.

[0067] The second streamer account is a user account held by a second streamer. This second streamer account can be any streamer account other than the first streamer account. The first streamer and at least one second streamer conduct a joint live stream on their respective clients. The second streamer turns on their camera during the live stream. The live video stream of the second streamer account refers to the continuous image frames captured by the camera during the live stream. The live video stream of the second streamer account includes second image frames, which are real-time image frames captured by the camera. After the second client acquires the second image frames in real time, it sends them to the first client for joint display. Simultaneously, the first client sends the first image frames to at least one second client for joint display. The image recognition step for the second image frames is performed by the second client.

[0068] In some embodiments, step 410 includes sub-step 411.

[0069] Sub-step 411: Obtain the live video stream of the first broadcaster account and other image video streams corresponding to at least one second broadcaster account. The other image video streams of the second broadcaster account include the second other image corresponding to the second image frame. The second other image is the region in the second image frame other than the second local image.

[0070] After acquiring the second image frame, the second client performs image recognition on the second image frame to obtain a second partial image and a second other image corresponding to the second image frame. The second other image is then sent to the first client for joint display. Correspondingly, the first client will send a first other image to at least one second client for joint display.

[0071] By acquiring a second, additional image, the amount of data sent from the second client to the first client can be reduced, thereby saving bandwidth required for data transmission and improving data transmission efficiency. Furthermore, acquiring the second, additional image facilitates subsequent direct stitching of the second, additional image and the processed second local image, improving image processing efficiency.

[0072] Step 420: Perform image recognition on the first image frame to obtain a first local image corresponding to the first image frame. The first local image includes the face region in the first image frame.

[0073] Step 430: Send the first partial image to the server.

[0074] The specific execution process of steps 420 and 430 can be referred to the above embodiments, and will not be repeated here.

[0075] Step 440: Receive the processed partial video stream sent by the server. The processed partial video stream includes a processed first partial image and at least one processed second partial image. The processed first partial image is obtained by processing the first partial image using portrait beautification technology, and the processed second partial image is obtained by processing the second partial image corresponding to the second image frame using portrait beautification technology.

[0076] In some embodiments, a first client sends encoded data of a first partial image to a server, and at least one second client sends encoded data of a second partial image to the server. After receiving the encoded data stream of the first partial image from the first client, the server decodes the encoded data of the first partial image to obtain a decoded first partial image and image information of the decoded first partial image. The server also receives encoded data of the second partial images sent by at least one second client, and decodes the encoded data of each of the at least one second partial image to obtain at least one decoded second partial image and image information of at least one decoded second partial image. Based on timestamp information, portrait beautification technology is used to process the decoded first partial image and at least one decoded second partial image with the same timestamp information in parallel to obtain a processed first partial image and at least one processed second partial image. If the image information includes facial recognition point information, the server uses portrait beautification technology to process the facial key points in the decoded first partial image based on the facial recognition point information of the decoded first partial image to obtain a processed first partial image. Based on the facial recognition point information of at least one decoded second local image, facial key points in at least one decoded second local image are processed using portrait beautification technology to obtain at least one processed second local image.

[0077] In some embodiments, the server encodes the processed first partial image, at least one processed second partial image, and processed image information to obtain processed image encoded data. The server sends the processed image encoded data to a first client. After receiving the processed image encoded data sent by the server, the first client decodes the processed image encoded data to obtain the processed first partial image, at least one processed second partial image, and processed image information.

[0078] The processed local video stream includes processed image encoded data, which includes a processed first local image, at least one processed second local image, and processed image information.

[0079] Optionally, the processed image information may include only the timestamp information of the processed first local image and at least one processed second local image, or it may additionally include the position information of the processed first local image and the position information corresponding to at least one processed second local image, as well as at least one of the face recognition point information of the processed first local image and the face recognition point information corresponding to at least one processed second local image.

[0080] In some embodiments, the server stitches together a processed first partial image and at least one processed second partial image to obtain a stitched partial image. Then, the server encodes the stitched partial image and its image information to obtain encoded data of the stitched partial image. The server sends the encoded data of the stitched partial image to a first client. After receiving the encoded data, the first client decodes it to obtain the stitched partial image and its image information.

[0081] The processed local video stream includes encoded data of stitched local images. The encoded data of stitched local images includes the stitched local images and image information of the stitched local images. The stitched local images include a processed first local image and at least one processed second local image.

[0082] Optionally, the image information for stitching the partial images may include only the timestamp information of the stitched partial images, or it may include both the timestamp information and the position information of the stitched partial images. Specifically, the timestamp information of the stitched partial images includes the timestamp information of the processed first partial image and at least one processed second partial image, and the position information of the stitched partial images includes the position information of the processed first partial image and the position information corresponding to each of the at least one processed second partial image.

[0083] Figure 5 The diagram illustrates the image processing flow in a two-person live streaming scenario. After acquiring a first image frame in real-time, the first client performs image recognition on the first image frame to obtain a first partial image corresponding to the first image frame, and sends the first partial image to the server. After acquiring a second image frame in real-time, the second client performs image recognition on the second image frame to obtain a second partial image corresponding to the second image frame, and sends the second partial image to the server. Based on the timestamp information of the first and second partial images, the server uses portrait beautification technology to perform parallel beautification processing on the first and second partial images with the same timestamp information, obtaining processed first and second partial images. The processed first and second partial images are then stitched together to obtain a stitched partial image, which is sent to both the first and second clients.

[0084] By stitching the processed first local image and at least one processed second local image into a single image, data transmission requires only the transmission resources needed for one image. Compared to transmitting multiple images, this reduces the bandwidth required for data transmission, saves data transmission costs, and improves data transmission efficiency.

[0085] In some embodiments, if the processed local video stream includes stitched local images, step 440 is followed by step 470.

[0086] Step 470: The stitched local image is split to obtain a processed first local image and at least one processed second local image.

[0087] After obtaining the stitched partial image, the first client splits the stitched partial image and extracts the processed first partial image and at least one processed second partial image from the stitched partial image.

[0088] By splitting the stitched local image as described above, subsequent image processing can be performed based on the obtained processed first local image and at least one processed second local image, providing the conditions for subsequent image processing steps.

[0089] Step 450: Obtain the processed first image frame based on the first image frame and the processed first local image, and obtain the processed second image frame based on the second image frame and the processed second local image.

[0090] In some embodiments, if image recognition is performed on the first image frame and only the first partial image corresponding to the first image frame is obtained, then the processed first partial image carries the timestamp information and the position information of the processed first partial image. Similarly, if the second image frame is obtained in step 410, then the processed second partial image carries the timestamp information and the position information of the processed second partial image. Step 450 includes sub-step 451.

[0091] Sub-step 451: Based on the timestamp information of the processed first partial image, obtain a first image frame from the live video stream of the first broadcaster account; and based on the timestamp information of the processed second partial image, obtain a second image frame from the live video stream of the second broadcaster account. Based on the position information of the processed first partial image, replace the corresponding image region in the first image frame with the processed first partial image to obtain the processed first image frame; and based on the position information of the processed second partial image, replace the corresponding image region in the second image frame with the processed second partial image to obtain the processed second image frame.

[0092] In some embodiments, if image recognition is performed on the first image frame to obtain a first partial image and a first other image corresponding to the first image frame, the processed first partial image carries timestamp information of the processed first partial image; and if the second image frame is obtained in step 410, the processed second partial image carries timestamp information of the processed second partial image and position information of the processed second partial image. Step 450 includes sub-step 452.

[0093] Sub-step 452: Based on the timestamp information of the processed first partial image, obtain the first image frame from the live video stream of the first broadcaster account; and based on the timestamp information of the processed second partial image, obtain the second image frame from the live video stream of the second broadcaster account. The first other image and the processed first partial image are stitched together to obtain the processed first image frame. Based on the position information of the processed second partial image, the processed second partial image replaces the corresponding image region in the second image frame to obtain the processed second image frame.

[0094] In some embodiments, if image recognition is performed on the first image frame and only the first partial image corresponding to the first image frame is obtained, then the processed first partial image carries the timestamp information and the position information of the processed first partial image. Furthermore, if the image obtained in step 410 is a second other image, then the processed second partial image carries the timestamp information of the processed second partial image. Step 450 includes sub-step 453.

[0095] Sub-step 453 involves obtaining a first image frame from the live video stream of the first broadcaster account based on the timestamp information of the processed first partial image, and obtaining a second image frame from the live video stream of the second broadcaster account based on the timestamp information of the processed second partial image. Based on the position information of the processed first partial image, the image region at the corresponding position in the first image frame is replaced with the processed first partial image to obtain the processed first image frame. Finally, the other second images and the processed second partial image are stitched together to obtain the processed second image frame.

[0096] In some embodiments, if image recognition is performed on the first image frame to obtain a first local image and a first other image corresponding to the first image frame, the processed first local image carries the timestamp information of the processed first local image; and if the second other image is obtained in step 410, the processed second local image carries the timestamp information of the processed second local image. Step 450 includes sub-step 454.

[0097] Sub-step 454: Based on the timestamp information of the processed first partial image, obtain the first image frame from the live video stream of the first broadcaster account; and based on the timestamp information of the processed second partial image, obtain the second image frame from the live video stream of the second broadcaster account. The first other image and the processed first partial image are stitched together to obtain the processed first image frame, and the second other image and the processed second partial image are stitched together to obtain the processed second image frame.

[0098] By acquiring the first and second other images as described above, the processed first and second image frames can be directly stitched together, simplifying the image processing steps, avoiding the waste of image stitching time due to re-positioning based on location information, and improving the efficiency of image processing.

[0099] Step 460: Based on the processed first image frame and at least one processed second image frame, a multi-host image frame is obtained.

[0100] The processed first image frame and at least one processed second image frame are stitched together to obtain a multi-host image frame.

[0101] In some embodiments, after displaying multi-anchor image frames, the first client, in response to an operation on the live stream, displays processed multi-anchor image frames. If the operation on the live stream is performed by the user to add filters or effects to the live stream, then the processed multi-anchor image frames are image frames with overlaid filter or effect layers.

[0102] Figure 6 This diagram illustrates the image stitching process in a two-person live streaming scenario. The server sends a stitched partial image to the client. After receiving the stitched partial image from the server, either the first or second client splits it into a processed first partial image and a processed second partial image. Then, it stitches the first remaining image with the processed first partial image to obtain a processed first image frame, and it stitches the second remaining image with the processed second partial image to obtain a processed second image frame. Finally, the stitched first and second partial images are displayed on the client.

[0103] By receiving the live video stream from the second broadcaster's account sent by the second client on the first client, and receiving the processed partial video stream sent by the server, the server applies a unified portrait beautification technology to the live portrait images from different live streaming applications. This allows the portrait images in the multiple broadcaster image frames on the first client to achieve the same portrait beautification effect, improving the uniformity and coordination of the live stream images when multiple people are live streaming together, thereby improving the live streaming effect.

[0104] Please refer to Figure 7 This document illustrates a flowchart of an image processing method for a single-person live streaming scenario according to another embodiment of this application. The execution entity for each step of this method can be a server. The method may include at least one of the following steps 710-730: Step 710: Receive a first partial video stream sent by the first client. The first partial video stream includes a first partial image. The first partial image is obtained by the first client through image recognition of the first image frame in the live video stream of the first broadcaster account. The first partial image includes the face region in the first image frame.

[0105] In some embodiments, the first local video stream includes encoded data of a first local image, which is obtained by encoding the first local image and its image information. The image information of the first local image includes timestamp information of the first local image. The timestamp information of the first local image is used to indicate the timestamp of the first image frame in the live video stream of the first broadcaster's account.

[0106] In some embodiments, the image information of the first local image further includes the position information of the first local image, which is used to indicate the position of the first local image in the first image frame.

[0107] In some embodiments, the image information of the first partial image further includes facial recognition point information of the first partial image, which is used to indicate the position of at least one facial key point in the first partial image.

[0108] Step 720: The first partial image is processed using portrait beautification technology to obtain the processed first partial image.

[0109] In some embodiments, the encoded data of the first local image is decoded to obtain the decoded first local image; the decoded first local image is then processed using portrait beautification technology to obtain the processed first local image.

[0110] If the image information of the first partial image does not include facial recognition point information, then facial key point recognition needs to be performed on the decoded first partial image to obtain the facial recognition point information. Then, portrait enhancement technology is used to process at least one facial key point in the decoded first partial image to obtain the processed first partial image. If the image information of the first partial image includes facial recognition point information, then portrait enhancement technology can be directly used to process at least one facial key point in the decoded first partial image based on the facial recognition point information to obtain the processed first partial image. By having the first client handle facial key point recognition of the first partial image, secondary recognition by the server can be avoided, thereby improving the image processing efficiency on the server side.

[0111] Step 730: Send the processed first partial image to the first client.

[0112] In some embodiments, the processed first partial image and the image information of the processed first partial image are encoded to obtain encoded data of the processed first partial image; the encoded data of the processed first partial image is sent to the first client.

[0113] The image processing steps on the first client side can refer to the above embodiment, and will not be repeated here.

[0114] By using a unified portrait enhancement technology to process live stream images from different live streaming applications through the server, the same portrait enhancement effect can be achieved when the streamer broadcasts in different live streaming applications, thus improving the stability and consistency of the enhancement effect.

[0115] Please refer to Figure 8 This document illustrates a flowchart of an image processing method for a multi-user live streaming scenario according to another embodiment of this application. The execution entity for each step of this method can be a server. The method may include at least one of the following steps 810-840: Step 810: Receive a first partial video stream sent by a first client and a second partial video stream sent by at least one second client.

[0116] The first partial video stream includes a first partial image, which is obtained by the first client performing image recognition on a first image frame in the live video stream of the first broadcaster account. The first partial image includes the facial region in the first image frame. The second partial video stream includes a second partial image, which is obtained by the second client performing image recognition on a second image frame in the live video stream of the second broadcaster account. The second partial image includes the facial region in the second image frame.

[0117] In some embodiments, the first local video stream includes encoded data of a first local image, which is obtained by encoding the first local image and its image information. The image information of the first local image includes timestamp information of the first local image. The timestamp information of the first local image is used to indicate the timestamp of the first image frame in the live video stream of the first broadcaster's account.

[0118] In some embodiments, the image information of the first local image further includes the position information of the first local image, which is used to indicate the position of the first local image in the first image frame.

[0119] In some embodiments, the image information of the first partial image further includes facial recognition point information of the first partial image, which is used to indicate the position of at least one facial key point in the first partial image.

[0120] In some embodiments, the second local video stream includes encoded data of a second local image, which is obtained by encoding the image information of the second local image and the second local image. The image information of the second local image includes timestamp information of the second local image. The timestamp information of the second local image is used to indicate the timestamp of the second image frame in the live video stream of the second broadcaster's account.

[0121] In some embodiments, the image information of the second local image further includes the position information of the second local image, which is used to indicate the position of the second local image in the second image frame.

[0122] In some embodiments, the image information of the second partial image further includes facial recognition point information of the second partial image, which is used to indicate the location of at least one facial key point in the second partial image.

[0123] Step 820: The first partial image is processed using portrait beautification technology to obtain the processed first partial image, and at least one second partial image is processed using portrait beautification technology to obtain at least one processed second partial image.

[0124] In some embodiments, the encoded data of a first partial image is decoded to obtain a decoded first partial image; portrait enhancement techniques are applied to the decoded first partial image to obtain a processed first partial image. The encoded data of at least one second partial image is decoded to obtain at least one decoded first partial image; portrait enhancement techniques are applied to the at least one decoded second partial image to obtain at least one processed second partial image.

[0125] If the image information of the first partial image does not include facial recognition point information, then facial key point recognition needs to be performed on the decoded first partial image to obtain the facial recognition point information. Then, portrait beautification techniques are used to process at least one facial key point in the decoded first partial image to obtain the processed first partial image. If the image information of the first partial image includes facial recognition point information, then portrait beautification techniques can be directly used to process at least one facial key point in the decoded first partial image based on the facial recognition point information to obtain the processed first partial image.

[0126] If the image information of the second partial image does not include facial recognition point information, then facial key point recognition needs to be performed on the decoded second partial image to obtain the facial recognition point information. Then, portrait enhancement techniques are used to process at least one facial key point in the decoded second partial image to obtain the processed second partial image. If the image information of the second partial image includes facial recognition point information, then portrait enhancement techniques can be directly used to process at least one facial key point in the decoded second partial image based on the facial recognition point information to obtain the processed second partial image.

[0127] Step 830: The processed first local image and at least one processed second local image are stitched together to obtain a stitched local image.

[0128] In some embodiments, the processed first partial image carries image information of the processed first partial image, and the processed second partial image carries image information of the processed second partial image. The image information of the processed first partial image includes timestamp information of the processed first partial image, and the image information of the processed second partial image includes timestamp information of the processed second partial image. Then, processed first partial images with the same timestamp information and at least one processed second partial image are stitched together to obtain a stitched partial image.

[0129] Step 840: Send the stitched partial image to the first client.

[0130] In some embodiments, the image information of the stitched partial image, the processed first partial image, and the processed second partial image are encoded to obtain the encoded data of the stitched partial image; the encoded data of the stitched partial image is then sent to the first client.

[0131] The image processing steps on both the first and second client sides can refer to the above embodiments, and will not be repeated here.

[0132] By having the server apply a uniform portrait enhancement technology to the live stream images from different live streaming applications, the portrait images in multiple anchor image frames on the client can achieve the same portrait enhancement effect, improving the uniformity and coordination of the live stream images when multiple people are live streaming together, thereby improving the live streaming effect.

[0133] Please refer to Figure 9 This document illustrates a flowchart of an image processing method for a single-person live streaming scenario according to another embodiment of this application. The execution entity for each step of this method can be a viewer terminal device. The method may include at least one of the following steps 910-930: Step 910: Receive the live video stream of the first broadcaster account sent by the first client. The live video stream of the first broadcaster account includes the first image frame.

[0134] In some embodiments, other image and video streams of the first broadcaster account sent by the first client are received. The other image and video streams of the first broadcaster account include a first other image corresponding to the first image frame. The first other image is a region in the first image frame other than the first partial image. The first partial image includes the face region in the first image frame.

[0135] Step 920: Receive the processed first partial video stream sent by the server. The processed first partial video stream includes a processed first partial image, which is obtained by processing the first partial image corresponding to the first image frame using portrait beautification technology.

[0136] Step 930: Obtain the processed first image frame based on the first image frame and the processed first local image.

[0137] In some embodiments, a first other image and a processed first partial image are stitched together to obtain a processed first image frame.

[0138] Please refer to Figure 10This document illustrates a flowchart of an image processing method for a multi-person live streaming scenario according to another embodiment of this application. The execution entity for each step of this method can be a viewer terminal device. The method may include at least one of the following steps 1010-1050: Step 1010: Receive the live video stream of the first broadcaster account sent by the first client, and receive the live video stream of the second broadcaster account sent by at least one second client respectively. The live video stream of the first broadcaster account includes a first image frame, and the live video stream of the second broadcaster account includes a second image frame.

[0139] In some embodiments, other image and video streams of the first broadcaster account sent by the first client are received. The other image and video streams of the first broadcaster account include a first other image corresponding to the first image frame. The first other image is a region in the first image frame other than the first partial image. The first partial image includes the face region in the first image frame.

[0140] In some embodiments, other image video streams of a second broadcaster account sent by at least one second client are received. The other image video streams of the second broadcaster account include second other images corresponding to the second image frame. The second other images are regions in the second image frame other than the second local image. The second local image includes the face region in the second image frame.

[0141] Step 1020: Receive the processed partial video stream sent by the server. The processed partial video stream includes stitched partial images, which include a processed first partial image and at least one processed second partial image.

[0142] The first local image after processing is obtained by processing the first local image using portrait beautification technology, and the second local image after processing is obtained by processing the second local image corresponding to the second image frame using portrait beautification technology.

[0143] Step 1030: The stitched local image is split to obtain a processed first local image and at least one processed second local image.

[0144] Step 1040: Obtain the processed first image frame based on the first image frame and the processed first local image, and obtain the processed second image frame based on the second image frame and the processed second local image.

[0145] In some embodiments, a first other image and a processed first partial image are stitched together to obtain a processed first image frame, and a second other image and a processed second partial image are stitched together to obtain a processed second image frame.

[0146] Step 1050: Based on the processed first image frame and at least one processed second image frame, obtain the multi-host image frame.

[0147] The following are embodiments of the apparatus described in this application, which can be used to execute the embodiments of the method described in this application. For details not disclosed in the apparatus embodiments of this application, please refer to the embodiments of the method described in this application.

[0148] Please refer to Figure 11 This diagram illustrates a block diagram of an image processing apparatus according to an embodiment of this application. The apparatus has the function of implementing the above-described image processing method; this function can be implemented in hardware or by hardware executing corresponding software. The apparatus can be the first terminal device described above, or it can be disposed within the first terminal device. Figure 11 As shown, the device 1100 may include: a video stream acquisition module 1110, an image recognition module 1120, an image sending module 1130, an image receiving module 1140, and an image processing module 1150.

[0149] The video stream acquisition module 1110 is used to acquire the live video stream of the first broadcaster account, wherein the live video stream of the first broadcaster account includes the first image frame.

[0150] The image recognition module 1120 is used to perform image recognition on the first image frame to obtain a first local image corresponding to the first image frame, wherein the first local image includes the face region in the first image frame.

[0151] The image sending module 1130 is used to send the first partial image to the server.

[0152] The image receiving module 1140 is used to receive the processed first partial image sent by the server, wherein the processed first partial image is obtained by processing the first partial image using portrait beautification technology.

[0153] The image processing module 1150 is used to obtain the processed first image frame based on the first image frame and the processed first local image.

[0154] In some embodiments, the first image frame includes the first partial image and a first other image, wherein the first other image includes a region in the first image frame other than the first partial image; the image processing module 1150 is configured to: The first other image and the processed first partial image are stitched together to obtain the processed first image frame.

[0155] In some embodiments, the image recognition module 1120 is configured to: Perform face recognition on the first image frame to obtain the face image corresponding to the first image frame, and determine the face image corresponding to the first image frame as the first local image. or, Facial region recognition is performed on the first image frame to obtain the facial region image corresponding to the first image frame. The facial region image corresponding to the first image frame is determined as the first local image. The facial region image includes the key facial regions in the first image frame. or, Head recognition is performed on the first image frame to obtain the head image corresponding to the first image frame, and the head image corresponding to the first image frame is determined as the first local image.

[0156] In some embodiments, the image sending module 1130 is configured to: The first partial image and its image information are encoded to obtain the encoded data of the first partial image. The image information of the first partial image includes the timestamp information of the first partial image. The timestamp information of the first partial image is used to indicate the timestamp of the first image frame in the live video stream of the first broadcaster account. The encoded data of the first partial image is sent to the server.

[0157] In some embodiments, the processed first partial image carries image information of the processed first partial image, the image information of the processed first partial image including timestamp information of the processed first partial image; the image processing module 1150 is configured to: Based on the timestamp information of the processed first local image, the first image frame is obtained from the live video stream of the first broadcaster account; The first image frame is obtained based on the image information of the processed first local image, the processed first local image, and the first image frame.

[0158] In some embodiments, the video stream acquisition module 1110 is configured to: Obtain the live video stream corresponding to at least one second broadcaster account, wherein the live video stream of the second broadcaster account includes a second image frame; The image receiving module 1140 is used for: The server receives a processed partial video stream, which includes a processed first partial image and at least one processed second partial image. The processed second partial image is obtained by processing the second partial image corresponding to the second image frame using portrait beautification technology. The image processing module 1150 is used for: Based on the second image frame and the processed second partial image, the processed second image frame is obtained; Based on the processed first image frame and at least one processed second image frame, a multi-host image frame is obtained.

[0159] In some embodiments, the video stream acquisition module 1110 is configured to: Obtain other image and video streams corresponding to the at least one second broadcaster account, wherein the other image and video streams of the second broadcaster account include second other images corresponding to the second image frame, and the second other images include the region in the second image frame other than the second local image; The image processing module 1150 is used for: The second other image and the processed second partial image are stitched together to obtain the processed second image frame.

[0160] In some embodiments, the processed local video stream includes stitched local images, which include the processed first local image and the at least one processed second local image.

[0161] In some embodiments, the image processing module 1150 is configured to: The stitched partial image is split to obtain the processed first partial image and the at least one processed second partial image.

[0162] Please refer to Figure 12 This diagram illustrates a block diagram of an image processing apparatus according to another embodiment of this application. The apparatus has the function of implementing the above-described image processing method; this function can be implemented in hardware or by hardware executing corresponding software. The apparatus can be the server described above, or it can be located within a server. Figure 12 As shown, the device 1200 may include: a video stream receiving module 1210, an image enhancement module 1220, and an image sending module 1230.

[0163] The video stream receiving module 1210 is used to receive a first partial video stream sent by a first client. The first partial video stream includes a first partial image. The first partial image is obtained by the first client through image recognition of the first image frame in the live video stream of the first broadcaster account. The first partial image includes the face region in the first image frame.

[0164] The image enhancement module 1220 is used to process the first partial image using portrait enhancement technology to obtain the processed first partial image.

[0165] The image sending module 1230 is used to send the processed first partial image to the first client.

[0166] In some embodiments, the first local video stream includes encoded data of the first local image, which is obtained by encoding the first local image and its image information. The image information of the first local image includes timestamp information of the first local image. The timestamp information of the first local image is used to indicate the timestamp of the first image frame in the live video stream of the first broadcaster's account. The image enhancement module 1220 is used to: The encoded data of the first local image is decoded to obtain the decoded first local image; The decoded first partial image is processed using portrait beautification technology to obtain the processed first partial image.

[0167] In some embodiments, the image sending module 1230 is configured to: The processed first local image and its image information are encoded to obtain the encoded data of the processed first local image; The encoded data of the processed first local image is sent to the first client.

[0168] In some embodiments, the video stream receiving module 1210 is configured to: The system receives a second partial video stream sent by at least one second client. The second partial video stream includes a second partial image. The second partial image is obtained by the second client through image recognition of a second image frame in the live video stream of the second broadcaster account. The second partial image includes a face region in the second image frame. The image enhancement module 1220 is used for: At least one second local image is processed using portrait beautification technology to obtain at least one processed second local image; The processed first local image and the at least one processed second local image are stitched together to obtain a stitched local image; The image sending module 1230 is used for: The stitched partial image is sent to the first client.

[0169] In some embodiments, the processed first partial image carries image information of the processed first partial image, and the processed second partial image carries image information of the processed second partial image. The image information of the processed first partial image includes timestamp information of the processed first partial image, and the image information of the processed second partial image includes timestamp information of the processed second partial image. The image enhancement module 1220 is used for: The first processed local image and the at least one processed second local image with the same timestamp information are stitched together to obtain the stitched local image.

[0170] It should be noted that the apparatus provided in the above embodiments is only illustrated by the division of the above functional modules when implementing its functions. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the content structure of the device can be divided into different functional modules to complete all or part of the functions described above. In addition, the apparatus and method embodiments provided in the above embodiments belong to the same concept, and the specific implementation process can be found in the method embodiments, which will not be repeated here.

[0171] Please refer to Figure 13 This diagram illustrates a structural block diagram of a computer device 1300 provided in one embodiment of this application. The computer device 1300 can be any electronic device capable of data calculation, processing, and storage. The computer device 1300 can be used to implement the image processing method provided in the above embodiments.

[0172] Typically, computer device 1300 includes a processor 1301 and a memory 1302.

[0173] Processor 1301 may include one or more processing cores, such as a quad-core processor, an octa-core processor, etc. Processor 1301 may be implemented using at least one hardware form selected from DSP (Digital Signal Processing), FPGA (Field Programmable Gate Array), and PLA (Programmable Logic Array). Processor 1301 may also include a main processor and a coprocessor. The main processor, also known as a CPU (Central Processing Unit), is used to process data in the wake-up state; the coprocessor is a low-power processor used to process data in the standby state. In some embodiments, processor 1301 may integrate a GPU (Graphics Processing Unit), which is responsible for rendering and drawing the content to be displayed on the screen. In some embodiments, processor 1301 may also include an AI (Artificial Intelligence) processor, which is used to handle computational operations related to machine learning.

[0174] The memory 1302 may include one or more computer-readable storage media, which may be non-transitory. The memory 1302 may also include high-speed random access memory and non-volatile memory, such as one or more disk storage devices or flash memory devices. In some embodiments, the non-transitory computer-readable storage media in the memory 1302 are used to store a computer program configured to be executed by one or more processors to implement the image processing method described above.

[0175] Those skilled in the art will understand that Figure 13 The structure shown does not constitute a limitation on the computer device 1300, and may include more or fewer components than shown, or combine certain components, or use different component arrangements.

[0176] In an illustrative embodiment, a computer-readable storage medium is also provided, wherein a computer program is stored in the storage medium, and the computer program implements the above-described image processing method when executed by a processor of a computer device. Optionally, the above-described computer-readable storage medium may be ROM (Read-Only Memory), RAM (Random Access Memory), CD-ROM (Compact Disc Read-Only Memory), magnetic tape, floppy disk, and optical data storage device, etc.

[0177] In an exemplary embodiment, a computer program product is also provided, comprising a computer program stored in a computer-readable storage medium. A processor of a computer device reads the computer program from the computer-readable storage medium and executes the computer program, causing the computer device to perform the image processing method described above.

[0178] It should be understood that "multiple" as used herein refers to two or more. "And / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, or B alone. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. Furthermore, the step numbers described herein are merely illustrative of one possible execution order. In some other embodiments, the steps may not be executed in numerical order, such as two steps with different numbers being executed simultaneously, or two steps with different numbers being executed in the reverse order of the illustration. This application does not limit this.

[0179] The above description is merely an exemplary embodiment of this application and is not intended to limit this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.

Claims

1. An image processing method, characterized in that, The method includes: Obtain the live video stream of the first streamer account, wherein the live video stream of the first streamer account includes the first image frame; Image recognition is performed on the first image frame to obtain a first local image corresponding to the first image frame, and the first local image includes the face region in the first image frame; Send the first partial image to the server; The server sends a processed first partial image, wherein the processed first partial image is obtained by processing the first partial image using portrait beautification technology; Based on the first image frame and the processed first local image, the processed first image frame is obtained.

2. The method according to claim 1, characterized in that, The first image frame includes the first partial image and the first other image, wherein the first other image includes the region in the first image frame other than the first partial image; The step of obtaining the processed first image frame based on the first image frame and the processed first partial image includes: The first other image and the processed first partial image are stitched together to obtain the processed first image frame.

3. The method according to claim 1 or 2, characterized in that, The step of performing image recognition on the first image frame to obtain the first local image corresponding to the first image frame includes: Perform face recognition on the first image frame to obtain the face image corresponding to the first image frame, and determine the face image corresponding to the first image frame as the first local image. or, Facial region recognition is performed on the first image frame to obtain the facial region image corresponding to the first image frame. The facial region image corresponding to the first image frame is determined as the first local image. The facial region image includes the key facial regions in the first image frame. or, Head recognition is performed on the first image frame to obtain the head image corresponding to the first image frame, and the head image corresponding to the first image frame is determined as the first local image.

4. The method according to any one of claims 1 to 3, characterized in that, Sending the first partial image to the server includes: The first partial image and its image information are encoded to obtain the encoded data of the first partial image. The image information of the first partial image includes the timestamp information of the first partial image. The timestamp information of the first partial image is used to indicate the timestamp of the first image frame in the live video stream of the first broadcaster account. The encoded data of the first partial image is sent to the server.

5. The method according to claim 4, characterized in that, The processed first partial image carries image information of the processed first partial image, and the image information of the processed first partial image includes the timestamp information of the processed first partial image; The step of obtaining the processed first image frame based on the first image frame and the processed first partial image includes: Based on the timestamp information of the processed first local image, the first image frame is obtained from the live video stream of the first broadcaster account; The first image frame is obtained based on the image information of the processed first local image, the processed first local image, and the first image frame.

6. The method according to any one of claims 1 to 5, characterized in that, The method further includes: Obtain the live video stream corresponding to at least one second broadcaster account, wherein the live video stream of the second broadcaster account includes a second image frame; Receiving the processed first partial image sent by the server includes: The server receives a processed partial video stream, which includes a processed first partial image and at least one processed second partial image. The processed second partial image is obtained by processing the second partial image corresponding to the second image frame using portrait beautification technology. The method further includes: Based on the second image frame and the processed second partial image, the processed second image frame is obtained; Based on the processed first image frame and at least one processed second image frame, a multi-host image frame is obtained.

7. The method according to claim 6, characterized in that, The step of obtaining the live video stream corresponding to at least one second streamer account includes: Obtain other image and video streams corresponding to the at least one second broadcaster account, wherein the other image and video streams of the second broadcaster account include second other images corresponding to the second image frame, and the second other images include the region in the second image frame other than the second local image; The step of obtaining the processed second image frame based on the second image frame and the processed second partial image includes: The second other image and the processed second partial image are stitched together to obtain the processed second image frame.

8. The method according to claim 6 or 7, characterized in that, The processed local video stream includes stitched local images, which include the processed first local image and the at least one processed second local image.

9. The method according to claim 8, characterized in that, After receiving the processed partial video stream sent by the server, the method further includes: The stitched partial image is split to obtain the processed first partial image and the at least one processed second partial image.

10. An image processing method, characterized in that, The method includes: The system receives a first partial video stream sent by a first client. The first partial video stream includes a first partial image. The first partial image is obtained by the first client through image recognition of a first image frame in the live video stream of the first broadcaster account. The first partial image includes the face region in the first image frame. The first partial image is processed using portrait beautification technology to obtain the processed first partial image; The processed first partial image is sent to the first client.

11. The method according to claim 10, characterized in that, The first local video stream includes encoded data of the first local image. The encoded data of the first local image is obtained by encoding the first local image and the image information of the first local image. The image information of the first local image includes the timestamp information of the first local image. The timestamp information of the first local image is used to indicate the timestamp of the first image frame in the live video stream of the first broadcaster account. The process of processing the first partial image using portrait beautification technology to obtain the processed first partial image includes: The encoded data of the first local image is decoded to obtain the decoded first local image; The decoded first partial image is processed using portrait beautification technology to obtain the processed first partial image.

12. The method according to claim 11, characterized in that, Sending the processed first partial image to the first client includes: The processed first local image and its image information are encoded to obtain the encoded data of the processed first local image; The encoded data of the processed first local image is sent to the first client.

13. The method according to any one of claims 10 to 12, characterized in that, The method further includes: The system receives a second partial video stream sent by at least one second client. The second partial video stream includes a second partial image. The second partial image is obtained by the second client through image recognition of a second image frame in the live video stream of the second broadcaster account. The second partial image includes a face region in the second image frame. At least one second local image is processed using portrait beautification technology to obtain at least one processed second local image; The processed first local image and the at least one processed second local image are stitched together to obtain a stitched local image; Sending the processed first partial image to the first client includes: The stitched partial image is sent to the first client.

14. The method according to claim 13, characterized in that, The processed first partial image carries the image information of the processed first partial image, and the processed second partial image carries the image information of the processed second partial image. The image information of the processed first partial image includes the timestamp information of the processed first partial image, and the image information of the processed second partial image includes the timestamp information of the processed second partial image. The step of stitching together the processed first local image and the at least one processed second local image to obtain a stitched local image includes: The first processed local image and the at least one processed second local image with the same timestamp information are stitched together to obtain the stitched local image.

15. A computer device, characterized in that, The computer device includes a processor and a memory, the memory storing a computer program, the computer program being loaded and executed by the processor to implement the image processing method as described in any one of claims 1 to 9, or to implement the image processing method as described in any one of claims 10 to 14.

16. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, which is loaded and executed by a processor to implement the image processing method as described in any one of claims 1 to 9, or to implement the image processing method as described in any one of claims 10 to 14.

17. A computer program product, characterized in that, The computer program product includes a computer program that is loaded and executed by a processor to implement the image processing method as described in any one of claims 1 to 9, or to implement the image processing method as described in any one of claims 10 to 14.