Live-streaming picture adjustment method and related device

Through object detection and image processing technology, the target objects in the live broadcast screen are adjusted and fused, and the problem of preset objects cannot be highlighted in the prior art is solved, and the effect of attracting audience attention without changing the picture size is achieved.

WO2025092188A9PCT designated stage expired Publication Date: 2025-07-03TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/115090
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-10-31
Filing Date
2024-08-28
Publication Date
2025-07-03

AI Technical Summary

Technical Problem

The existing live broadcast technology cannot effectively highlight the preset objects in the live broadcast screen, resulting in audience users being unable to pay attention to it, affecting the live broadcast effect.

Method used

By detecting, segmenting and adjusting the target objects in the live broadcast screen, and re-fusion them into the live broadcast screen, the preset adjustment strategy is used to highlight the preset objects, including technical means such as size and position adjustment, image fusion, etc.

Benefits of technology

Without changing the size of the live broadcast screen, the preset objects in the live broadcast screen can be highlighted to attract the attention of viewers and improve the live broadcast effect.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024115090_03072025_PF_FP_ABST
    Figure CN2024115090_03072025_PF_FP_ABST
Patent Text Reader

Abstract

Disclosed in the present application are a live-streaming picture adjustment method and a related device. The method comprises: performing object detection on picture content of a first live-streaming picture, so as to obtain candidate detection boxes respectively corresponding to a plurality of preset objects within the first live-streaming picture; for a target detection box among the plurality of candidate detection boxes, separating out a target object in the target detection box from the first live-streaming picture, so as to obtain a second live-streaming picture and a first object image of the target object; according to a preset adjustment strategy used for highlighting one or more of the plurality of preset objects, adjusting the first object image to obtain a second object image; and fusing the second object image and the second live-streaming picture to obtain a third live-streaming picture. By means of detection, separation and adjustment of target objects within live-streaming pictures and re-fusion of same into the live-streaming pictures, the method achieves the display effect of highlighting preset objects within the live-streaming pictures without changing the size of the live-streaming pictures, so as to attract users' attention to the preset objects within the live-streaming pictures, thus improving the live-streaming effect.
Need to check novelty before this filing date? Find Prior Art

Description

A method and related device for adjusting live broadcast images

[0001] This application claims priority to the Chinese patent application filed with the China Patent Office on October 31, 2023, with application number 202311438267.4 and application name “A method for adjusting live broadcast images and related devices”, the entire contents of which are incorporated by reference into this application. Technical Field

[0002] The present application relates to the field of computer technology, and in particular to a technology for adjusting live broadcast images. Background Art

[0003] With the rapid development of live streaming technology, it has become ubiquitous in our daily lives and work. Live streaming users can use live streaming software on their terminals to capture live footage and send it to a server, which then sends it to viewer terminals for viewing.

[0004] In related technologies, methods for adjusting live broadcast images include extracting the host's outline image from the live broadcast image and beautifying the host's outline image to make the host look better; or replacing the live broadcast background other than the host in the live broadcast image to make the live broadcast background more interesting.

[0005] However, the above methods only beautify the main objects such as the anchor in the live broadcast screen, or replace the background and other areas in the live broadcast screen, but cannot highlight the preset objects such as the anchor or objects in the live broadcast screen, resulting in failure to attract audience users to focus on the preset objects in the live broadcast screen, thereby affecting the live broadcast effect.

[0006] Summary of the Invention

[0007] In order to solve the above technical problems, the present application provides an adjustment method and related devices based on the live broadcast screen, which can achieve the display effect of highlighting the preset objects in the live broadcast screen by detecting, segmenting, adjusting the target objects in the live broadcast screen and reintegrating them into the live broadcast screen without changing the size of the live broadcast screen, so as to attract audience users to focus on the preset objects in the live broadcast screen, thereby improving the live broadcast effect.

[0008] The embodiments of this application disclose the following technical solutions:

[0009] In one aspect, an embodiment of the present application provides an adjustment method based on a live broadcast screen, the method being executed by a computer device, the method comprising:

[0010] Performing object detection on the screen content of the first live screen to obtain candidate detection frames corresponding to a plurality of preset objects in the first live screen;

[0011] segmenting a target object in the target detection frame from the first live broadcast screen based on the target detection frames in the plurality of candidate detection frames, and obtaining a second live broadcast screen and a first object image of the target object, where the second live broadcast screen refers to the live broadcast screen excluding the first object image of the target object in the first live broadcast screen;

[0012] Adjusting the first object image according to a preset adjustment strategy to obtain a second object image; the preset adjustment strategy is used to highlight one or more of the plurality of preset objects;

[0013] The second object image and the second live broadcast picture are fused to obtain a third live broadcast picture.

[0014] On the other hand, an embodiment of the present application provides an adjustment device based on a live picture, the device being deployed on a computer device, the device comprising: a detection unit, a segmentation unit, an adjustment unit, and a fusion unit;

[0015] The detection unit is configured to perform object detection on the content of the first live broadcast screen to obtain candidate detection frames corresponding to a plurality of preset objects in the first live broadcast screen;

[0016] The segmentation unit is configured to segment a target object in the target detection frame from the first live broadcast screen based on the target detection frames in the plurality of candidate detection frames, and obtain a second live broadcast screen and a first object image of the target object, where the second live broadcast screen refers to the live broadcast screen excluding the first object image of the target object in the first live broadcast screen;

[0017] The adjustment unit is configured to adjust the first object image according to a preset adjustment strategy to obtain a second object image; the preset adjustment strategy is configured to highlight one or more of the plurality of preset objects;

[0018] The fusion unit is configured to perform image fusion on the second object image and the second live broadcast picture to obtain a third live broadcast picture.

[0019] In another aspect, an embodiment of the present application provides a computer device, comprising a processor and a memory.

[0020] The memory is used to store a computer program and transmit the computer program to the processor;

[0021] The processor is configured to execute the method described in any one of the preceding aspects according to instructions in the computer program.

[0022] On the other hand, an embodiment of the present application provides a computer-readable storage medium, which is used to store a computer program. When the computer program is run on a computer device, the computer device executes the method described in any one of the above aspects.

[0023] On the other hand, an embodiment of the present application provides a computer program product, including a computer program, which, when executed on a computer device, enables the computer device to execute the method described in any one of the aforementioned aspects.

[0024] As can be seen from the above technical solution, first, based on the screen content of the first live broadcast screen, object detection is performed to obtain multiple candidate detection frames corresponding to multiple preset objects in the first live broadcast screen; based on the target detection frames in the multiple candidate detection frames, the target object in the target detection frame is segmented from the first live broadcast screen to obtain a second live broadcast screen and a first object image of the target object; this step can accurately segment the first live broadcast screen into the first object image of the target object to be adjusted and the second live broadcast screen to be fused through methods such as object detection and image segmentation. Then, according to a preset adjustment strategy for highlighting one or more of the multiple preset objects, the first object image is adjusted to obtain a second object image; the second object image and the second live broadcast screen are fused to obtain a third live broadcast screen; this step changes the first object image of the target object into a second object image through methods such as image adjustment and image fusion based on the preset adjustment strategy, and fuses it into the second live broadcast screen to obtain a third live broadcast screen, so that the third live broadcast screen can highlight one or more of the multiple preset objects. Based on this, this method can achieve the display effect of highlighting the preset objects in the live broadcast screen by detecting, segmenting, adjusting the target objects in the live broadcast screen and reintegrating them into the live broadcast screen without changing the size of the live broadcast screen, so as to attract the audience to focus on the preset objects in the live broadcast screen, thereby improving the live broadcast effect. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technical members in this field, other drawings can be obtained based on these drawings without paying any creative labor.

[0026] FIG1 is a schematic diagram of a system architecture of a method for adjusting a live broadcast screen according to an embodiment of the present application;

[0027] FIG2 is a flow chart of an adjustment method based on a live broadcast screen provided in an embodiment of the present application;

[0028] FIG3 is a schematic diagram of multiple candidate detection frames corresponding to multiple preset objects in a first live broadcast screen provided by an embodiment of the present application;

[0029] FIG4 is a schematic diagram of a target detection frame among multiple candidate detection frames in a first live broadcast picture provided by an embodiment of the present application;

[0030] FIG5 is a schematic diagram of a second live broadcast screen and a first object image of a target object provided by an embodiment of the present application;

[0031] FIG6 is a schematic diagram of adjusting a first object image of a target object to a second object image according to an embodiment of the present application;

[0032] FIG7 is a schematic diagram of a third live broadcast screen provided in an embodiment of the present application;

[0033] FIG8 is a diagram showing the specific steps of a method for adjusting a live broadcast screen according to an embodiment of the present application;

[0034] FIG9 is a schematic diagram of a method of moving a first object image of a target object in a first live broadcast image to a second object image provided by an embodiment of the present application;

[0035] FIG10 is a schematic diagram of a second live broadcast screen including an area to be filled provided by an embodiment of the present application;

[0036] FIG11 is a flow chart of determining a detection frequency according to an embodiment of the present application;

[0037] FIG12 is a structural diagram of an adjustment device based on a live screen according to an embodiment of the present application;

[0038] FIG13 is a structural diagram of a terminal provided in an embodiment of the present application;

[0039] FIG14 is a structural diagram of a server provided in an embodiment of the present application. DETAILED DESCRIPTION

[0040] The embodiments of the present application are described below with reference to the accompanying drawings.

[0041] At present, the method of adjusting the live broadcast screen can be to extract the outline image of the anchor from the live broadcast screen and beautify the outline image of the anchor to obtain a better-looking anchor; or, to replace the live broadcast background except the outline image of the anchor in the live broadcast screen to obtain a more interesting live broadcast background.

[0042] However, after research, it was found that the above methods only beautify the main objects such as the anchor in the live broadcast screen, or replace the background and other areas in the live broadcast screen, but are unable to highlight the preset objects such as the anchor or objects in the live broadcast screen, resulting in the inability to attract audience users to focus on the preset objects in the live broadcast screen, thereby affecting the live broadcast effect.

[0043] The embodiment of the present application provides an adjustment method based on a live broadcast screen, which can achieve a display effect of highlighting a preset object in the live broadcast screen by detecting, segmenting, adjusting the target object in the live broadcast screen and reintegrating it into the live broadcast screen without changing the size of the live broadcast screen, so as to attract audience users to focus on the preset object in the live broadcast screen, thereby improving the live broadcast effect.

[0044] Next, the system architecture of the adjustment method based on the live screen will be introduced. Referring to Figure 1, Figure 1 is a schematic diagram of the system architecture of a method for adjusting the live screen provided in an embodiment of the present application, wherein the system architecture includes a terminal 100, which is used to execute the adjustment method based on the live screen.

[0045] The terminal 100 performs object detection on the screen content of the first live screen, and obtains candidate detection frames corresponding to a plurality of preset objects in the first live screen.

[0046] As an example, the first live screen is live screen 1, and the multiple preset objects in live screen 1 include M objects, where M is a positive integer. The terminal 100 can perform object detection on the screen content of live screen 1 to obtain M candidate detection frames corresponding to the M objects in live screen 1.

[0047] The terminal 100 segments the target object in the target detection frame from the first live broadcast picture according to the target detection frames in the multiple candidate detection frames, and obtains the second live broadcast picture and the first object image of the target object.

[0048] As an example, based on the above example, the target detection frame among the M candidate detection frames is the i-th detection frame, i∈M, and the terminal 100 segments the target object in the i-th detection frame from the live broadcast screen 1 as the i-th object for the i-th detection frame, and obtains the second live broadcast screen as the live broadcast screen 2 and the first object image of the i-th object as the object image 1; wherein, the live broadcast screen 2 refers to the live broadcast screen 1 from which the object image 1 is segmented, and the live broadcast screen 2 includes M-1 preset objects among the M objects except the i-th object.

[0049] The terminal 100 adjusts the first object image according to a preset adjustment strategy to obtain a second object image; the preset adjustment strategy is used to highlight one or more of a plurality of preset objects.

[0050] As an example, the preset adjustment strategy includes one or more of a size adjustment strategy and a position adjustment strategy. Based on the above example, the terminal 100 adjusts the object image 1 according to one or more of the size adjustment strategy and the position adjustment strategy to obtain the second object image as object image 2.

[0051] The terminal 100 performs image fusion on the second object image and the second live picture to obtain a third live picture.

[0052] As an example, based on the above example, the object image 2 and the live broadcast screen 2 are fused to obtain a third live broadcast screen 3, which includes M-1 preset objects and the i-th object shown in the object image 2 (adjusted object image 1), a total of M objects.

[0053] That is, through methods such as object detection and image segmentation, the first live broadcast screen can be accurately segmented into a first object image of the target object to be adjusted and a second live broadcast screen to be fused; through methods such as image adjustment and image fusion based on a preset adjustment strategy, the first object image of the target object is changed into a second object image, and then fused into the second live broadcast screen to obtain a third live broadcast screen, so that the third live broadcast screen can highlight one or more of the multiple preset objects. Based on this, the method can achieve the display effect of highlighting the preset objects in the live broadcast screen by detecting, segmenting, adjusting the target objects in the live broadcast screen and re-integrating them into the live broadcast screen without changing the size of the live broadcast screen, so as to attract viewers to focus on the preset objects in the live broadcast screen, thereby improving the live broadcast effect.

[0054] It should be noted that in the embodiments of the present application, the computer device may be a server or a terminal, and the method provided in the embodiments of the present application may be executed by the terminal or the server alone, or by the terminal and the server in combination. The embodiment corresponding to FIG1 is mainly described by taking the terminal executing the method provided in the embodiments of the present application as an example.

[0055] Furthermore, when the method provided in the embodiment of the present application is executed solely by the server, its execution method is similar to the embodiment corresponding to FIG1 , primarily with the terminal being replaced by the server. Furthermore, when the method provided in the embodiment of the present application is executed by a terminal and a server in cooperation, steps that need to be displayed on the front-end interface can be executed by the terminal, while steps that require background computing and do not need to be displayed on the front-end interface can be executed by the server.

[0056] The terminal may be, but is not limited to, a smartphone, tablet computer, laptop computer, desktop computer, intelligent voice interaction device, vehicle-mounted terminal, or aircraft. The server may be, but is not limited to, an independent physical server, a server cluster or distributed system consisting of multiple physical servers, or a cloud server providing cloud computing services. The terminal and server may be connected directly or indirectly via wired or wireless communication, which is not limited in this application. For example, the terminal and server may be connected via a network, which may be a wired or wireless network.

[0057] The method provided in the embodiment of the present application involves artificial intelligence (AI) technology, and automatically implements an adjustment method based on a live broadcast screen based on AI technology.

[0058] In addition, the embodiments of the present application can be applied to various scenarios, including but not limited to cloud technology, artificial intelligence, smart transportation, audio and video, assisted driving, etc.

[0059] Next, the method for adjusting the live screen provided by the embodiment of the present application will be described in detail with reference to the accompanying drawings, using the method provided by the terminal executing the embodiment of the present application as an example. Referring to FIG. 2 , FIG. 2 is a flow chart of a method for adjusting the live screen provided by the embodiment of the present application, the method comprising:

[0060] S201: Performing object detection on the picture content of a first live broadcast picture to obtain candidate detection frames corresponding to a plurality of preset objects in the first live broadcast picture.

[0061] In related technologies, it is possible to extract the host's outline image from the live broadcast screen and beautify the host's outline image to make the host look better; or to replace the live broadcast background other than the host's outline image to create a more interesting live broadcast background. However, after research, it was found that the above methods only beautify the main subject such as the host in the live broadcast screen, or replace the background and other areas in the live broadcast screen, but fail to highlight the preset subject such as the host or objects in the live broadcast screen. As a result, it is impossible to attract the audience to focus on the preset subject in the live broadcast screen, thus affecting the live broadcast effect.

[0062] Therefore, in an embodiment of the present application, in order to highlight a preset object in a live broadcast image, it is first necessary to detect multiple preset objects in the live broadcast image so that the preset objects can be subsequently highlighted in the live broadcast image. Based on this, a first live broadcast image is obtained, and object detection is performed on the image content of the first live broadcast image to obtain multiple candidate detection frames corresponding to the multiple preset objects in the first live broadcast image. The preset objects can be objects captured or displayed during the live broadcast, such as the host or various objects.

[0063] In practical applications, S201 may be performing object detection on the screen content of the first live screen by using an object detection algorithm to obtain multiple candidate detection frames corresponding to multiple preset objects in the first live screen. Wherein, the object detection algorithm may be a target detection algorithm, such as a Fast Region-based Convolutional Neural Network (Fast R-CNN) algorithm. The specific implementation process of S201 is as follows: performing feature extraction on the screen content of the first live screen by using the Fast R-CNN algorithm to obtain image features of the first live screen; performing object-based region extraction on the screen content of the first live screen according to the image features of the first live screen to obtain region images corresponding to multiple preset objects in the first live screen; performing object detection on the multiple region images in the first live screen to obtain predicted detection frames corresponding to the multiple preset objects in the first live screen; performing regression processing on the multiple predicted detection frames in the first live screen to obtain candidate detection frames corresponding to the multiple preset objects in the first live screen.

[0064] S201 processes the first live screen through object detection, and can accurately detect the candidate detection frame of each preset object in the first live screen, thereby providing effective and accurate detection data for subsequently accurately segmenting the first live screen into the first object image of the target object to be adjusted and the second live screen to be fused.

[0065] As an example of S201, refer to Figure 3, which is a schematic diagram of multiple candidate detection frames corresponding to multiple preset objects in a first live broadcast screen provided by an embodiment of the present application, wherein Figure 3 (a) indicates that the first live broadcast screen is live broadcast screen 1, and the multiple preset objects in live broadcast screen 1 include 4 objects, namely, anchor, object 1, object 2 and object 3; Figure 3 (b) indicates that the terminal can perform object detection on the screen content of live broadcast screen 1 to obtain the candidate detection frame corresponding to the anchor in live broadcast screen 1 as the anchor detection frame, the candidate detection frame corresponding to object 1 as the object 1 detection frame, the candidate detection frame corresponding to object 2 as the object 2 detection frame and the candidate detection frame corresponding to object 3 as the object 3 detection frame, namely, the 4 dotted boxes in Figure 3 (b).

[0066] S202: Segment a target object in the target detection frame from the first live broadcast image according to the target detection frames in the plurality of candidate detection frames, and obtain a second live broadcast image and a first object image of the target object.

[0067] In an embodiment of the present application, in order to highlight preset objects such as anchors or objects in the live broadcast screen, after detecting multiple preset objects such as anchors and objects in the live broadcast screen, it is also necessary to select a target object from the multiple preset objects and segment the target object from the live broadcast screen, so that the preset object can be highlighted in the live broadcast screen by adjusting the target object later.

[0068] Based on this, after executing S201 to detect multiple candidate detection frames corresponding to multiple preset objects in the first live broadcast image, it is necessary to select a target detection frame from the multiple candidate detection frames, segment the target object in the target detection frame from the first live broadcast image, and obtain a second live broadcast image and a first object image of the target object. The second live broadcast image refers to the first live broadcast image from which the first object image of the target object is segmented, that is, the second live broadcast image refers to the other live broadcast images in the first live broadcast image except the first object image of the target object; the second live broadcast image then includes the other objects in the multiple preset objects except the target object.

[0069] In practical applications, S202 may be in response to a selection operation for a target detection frame from a plurality of candidate detection frames, determining a target detection frame from the plurality of candidate detection frames, or in response to detecting that the live speech corresponding to the first live screen includes a target object, determining a target detection frame corresponding to the target object from the plurality of candidate detection frames; segmenting the target object in the target detection frame from the first live screen using an image segmentation algorithm based on the target detection frame, thereby obtaining a second live screen and a first object image of the target object. The image segmentation algorithm may be an instance segmentation algorithm, such as a Mask Region-based Convolutional Neural Network (Mask R-CNN) algorithm. The specific implementation process of S202 is as follows: adding a segmentation branch to the Fast R-CNN algorithm for detecting the contour of the object, performing object-based contour detection based on the image features of the target detection frame and the target detection frame in the first live screen using the Mask R-CNN algorithm, thereby obtaining a target contour image of the target object; performing image segmentation on the first live screen based on the target contour image, thereby obtaining a target object image of the target object; and performing image separation on the target object and the background region in the target object image, thereby obtaining a first object image of the target object.

[0070] S202 processes the first live screen by image segmentation, and can accurately segment the first live screen into a first object image of the target object to be adjusted and a second live screen to be fused, thereby providing targeted and accurate image data for highlighting the preset object in the second live screen by subsequently adjusting the target object.

[0071] As an example of S202, based on the above-mentioned S201 example, refer to Figure 4, which is a schematic diagram of a target detection frame among multiple candidate detection frames in a first live broadcast screen provided by an embodiment of the present application; based on the above-mentioned Figure 3, the terminal responds to the selection operation of the object 2 detection frame among the host detection frame, object 1 detection frame, object 2 detection frame and object 3 detection frame in the live broadcast screen 1 shown in Figure 4 (a), and determines that the target detection frame is the object 2 detection frame from the host detection frame, object 1 detection frame, object 2 detection frame and object 3 detection frame, that is, the bold frame in Figure 4 (b).

[0072] Refer to Figure 5, which is a schematic diagram of a second live broadcast screen and a first object image of a target object provided in an embodiment of the present application; based on the above-mentioned Figure 4, the terminal segments object 2 in the object 2 detection frame from the live broadcast screen 1, and obtains the second live broadcast screen represented by (a) in Figure 5 as live broadcast screen 2 and the first object image of the target object represented by (b) in Figure 5 as object 2 image 1; wherein, live broadcast screen 2 refers to the live broadcast screen 1 from which object 2 image 1 is segmented, that is, live broadcast screen 2 refers to other live broadcast screens in the live broadcast screen 1 except object 2 image 1, and live broadcast screen 2 includes the anchor, object 1 and object 3.

[0073] S203: Adjust the first object image according to a preset adjustment strategy to obtain a second object image; the preset adjustment strategy is used to highlight one or more of a plurality of preset objects.

[0074] In an embodiment of the present application, in order to highlight preset objects such as an anchor or an object in a live broadcast screen, after segmenting a target object from multiple preset objects from the live broadcast screen, it is also necessary to adjust the target object based on the principle of highlighting one or more of the multiple preset objects so that the preset object can be highlighted in the live broadcast screen later.

[0075] Based on this, after executing S202 to split the first live broadcast screen into the second live broadcast screen and the first object image of the target object, it is necessary to adjust the first object image according to a preset adjustment strategy for highlighting one or more of the multiple preset objects, thereby obtaining the adjusted first object image as the second object image. The preset adjustment strategy for highlighting one or more of the multiple preset objects may be for highlighting the target object among the multiple preset objects, or may be for highlighting other preset objects among the multiple preset objects except the target object.

[0076] S203 processes the first object image of the target object through an image adjustment method based on a preset adjustment strategy, so that the first object image of the target object is accurately changed into a second object image. The second object image is integrated into the second live screen to highlight one or more of the multiple preset objects, providing accurate and effective image data for the subsequent highlighting of the preset objects in the second live screen.

[0077] As an example of S203, based on the above S203 example, refer to Figure 6, which is a schematic diagram provided by an embodiment of the present application of adjusting the first object image of the target object in the first live broadcast screen to the second object image. Based on the above Figure 5, according to the preset adjustment strategy for highlighting one or more of a plurality of preset objects, the object 2 image 1 represented by (a) in Figure 6 is adjusted to obtain the second object image represented by (b) in Figure 6 as object 2 image 2.

[0078] S204: Perform image fusion on the second object image and the second live broadcast picture to obtain a third live broadcast picture.

[0079] In the embodiment of the present application, in order to achieve the display effect of highlighting a preset object in the live broadcast image and thereby attracting the audience's attention to the preset object in the live broadcast image, after adjusting the target object, it is necessary to reintegrate the adjusted target object into the live broadcast image to achieve the effect of highlighting the preset object, such as the host or object, in the live broadcast image. Based on this, after executing S203 and adjusting the first object image of the target object to the second object image according to the preset adjustment strategy for highlighting one or more of the multiple preset objects, it is necessary to fuse the second object image with the second live broadcast image to obtain a third live broadcast image.

[0080] The S204 processes the second live broadcast screen and the second object image by image fusion, and fuses the second object image obtained by changing the first object image of the target object into the second live broadcast screen to obtain a third live broadcast screen, so that the third live broadcast screen can highlight one or more of the multiple preset objects, achieving a display effect of highlighting the preset objects in the third live broadcast screen, so as to attract audience users to focus on the preset objects in the third live broadcast screen, thereby improving the live broadcast effect.

[0081] As an example of S204, based on the above S203, refer to Figure 7, which is a schematic diagram of a third live broadcast screen provided in an embodiment of the present application. Based on the above Figures 5 and 6, the object 2 image 2 represented by (b) in Figure 6 and the live broadcast screen 2 represented by (a) in Figure 5 are integrated to obtain the third live broadcast screen represented by Figure 7, which is live broadcast screen 3. The live broadcast screen 3 includes the host, object 1 and object 3, as well as the object 2 shown in the object 2 image 2 (adjusted object 2 image 1).

[0082] In summary, referring to Figure 8, Figure 8 is a specific step diagram of an adjustment method based on a live screen provided by an embodiment of the present application, and the specific steps include: the first step, object detection, that is, performing object detection on the screen content of the first live screen through an object detection algorithm to obtain multiple candidate detection frames corresponding to multiple preset objects in the first live screen; the second step, object selection, that is, in response to the selection operation of the target detection frame in the multiple candidate detection frames, determining the target detection frame from the multiple candidate detection frames; the third step, image segmentation, that is, using the image segmentation algorithm to segment the target object in the target detection frame from the first live screen according to the target detection frame, to obtain a second live screen and a first object image of the target object; the fourth step, image adjustment, that is, adjusting the first object image according to a preset adjustment strategy for highlighting one or more of the multiple preset objects to obtain a second object image; the fifth step, image fusion, that is, performing image fusion on the second object image and the second live screen to obtain a third live screen.

[0083] As can be seen from the above technical solution, first, based on the screen content of the first live broadcast screen, object detection is performed to obtain multiple candidate detection frames corresponding to multiple preset objects in the first live broadcast screen; based on the target detection frames in the multiple candidate detection frames, the target object in the target detection frame is segmented from the first live broadcast screen to obtain a second live broadcast screen and a first object image of the target object; this step can accurately segment the first live broadcast screen into the first object image of the target object to be adjusted and the second live broadcast screen to be fused through methods such as object detection and image segmentation. Then, according to a preset adjustment strategy for highlighting one or more of the multiple preset objects, the first object image is adjusted to obtain a second object image; the second object image and the second live broadcast screen are fused to obtain a third live broadcast screen; this step changes the first object image of the target object into a second object image through methods such as image adjustment and image fusion based on the preset adjustment strategy, and fuses it into the second live broadcast screen to obtain a third live broadcast screen, so that the third live broadcast screen can highlight one or more of the multiple preset objects. Based on this, this method can achieve the display effect of highlighting the preset objects in the live broadcast screen by detecting, segmenting, adjusting the target objects in the live broadcast screen and reintegrating them into the live broadcast screen without changing the size of the live broadcast screen, so as to attract the audience to focus on the preset objects in the live broadcast screen, thereby improving the live broadcast effect.

[0084] In an embodiment of the present application, based on the principle of highlighting one or more of the plurality of preset objects, when executing S203, adjusting the target object so that the preset object is subsequently highlighted in the live broadcast image may involve adjusting the size of the target object, adjusting the position of the target object, or adjusting both the size and position of the target object so that the preset object is subsequently highlighted in the live broadcast image; the preset adjustment strategy may be a size adjustment strategy, a position adjustment strategy, or both a size adjustment strategy and a position adjustment strategy, etc. Therefore, the present application provides a possible implementation method, wherein the preset adjustment strategy includes one or more of a size adjustment strategy and a position adjustment strategy.

[0085] In the embodiment of the present application, when the above S203 is specifically implemented, when the preset adjustment strategy is a size adjustment strategy, the size adjustment strategy can be a size upsampling strategy, a size downsampling strategy, a size enlargement model, or a size reduction model; different size adjustment strategies correspond to different implementation methods of S203, as shown below:

[0086] One implementation method is: the resizing strategy is a size upsampling strategy. For the first object image of the target object, the size upsampling strategy can generate new pixel values ​​based on adjacent pixel values ​​in the first object image, and the new pixel values ​​will be inserted into the gaps between the existing pixel values ​​of the first object image, increasing the number of pixel values ​​of the first object image to enlarge the size of the first object image and improve the resolution of the first object image, thereby obtaining the enlarged first object image as the second object image. Therefore, the present application provides a possible implementation method. When the preset adjustment strategy is the resizing strategy, the resizing strategy is a size upsampling strategy. S203 is specifically S2031 (not shown in the figure): the first object image is enlarged according to the size upsampling strategy to obtain the second object image. The size of the second object image is larger than the size of the first object image, and the resolution of the second object image is larger than the resolution of the first object image.

[0087] Among them, the size upsampling strategy can be a bilinear interpolation algorithm or a bicubic interpolation algorithm. The bilinear interpolation algorithm refers to calculating a new pixel value through 4 (2×2) adjacent pixel values ​​in the image, and the bicubic interpolation algorithm refers to calculating a new pixel value through 16 (4×4) adjacent pixel values ​​in the image.

[0088] S2031 processes the first object image of the target object through a size upsampling strategy, so that the first object image of the target object is accurately and controllably enlarged to a second object image. The second object image is integrated into the second live broadcast screen to highlight the target object among multiple preset objects, providing accurate and effective image data for subsequently highlighting the target object among multiple preset objects in the second live broadcast screen.

[0089] As an example of S2031, based on the above-mentioned S202 example, the object 2 image 1 represented by (a) in Figure 6 is enlarged through a bilinear interpolation algorithm or a bicubic interpolation algorithm to obtain a second object image as object 2 image 3. The size of object 2 image 3 is larger than the size of object 2 image 1, and the resolution of object 2 image 3 is greater than the resolution of object 2 image 1.

[0090] Another implementation method is: the resizing strategy is a downsampling strategy. For the first object image of the target object, the downsampling strategy can generate new pixel values ​​based on the adjacent pixel values ​​in the first object image, and use the new pixel values ​​to replace the adjacent pixel values, thereby reducing the number of pixel values ​​in the first object image, thereby reducing the size of the first object image and lowering the resolution of the first object image, and obtaining the reduced first object image as the second object image. Therefore, the present application provides a possible implementation method. When the preset adjustment strategy is the resizing strategy, the resizing strategy is a downsampling strategy. S203 is specifically S2032 (not shown in the figure): the first object image is reduced in size according to the downsampling strategy to obtain the second object image. The size of the second object image is smaller than the size of the first object image, and the resolution of the second object image is smaller than the resolution of the first object image.

[0091] The downsampling strategy may also be a bilinear interpolation algorithm or a bicubic interpolation algorithm.

[0092] S2032 processes the first object image of the target object through a size downsampling strategy, so that the first object image of the target object is accurately and controllably reduced to a second object image. The second object image is integrated into the second live broadcast screen to highlight other objects except the target object among multiple preset objects, providing accurate and effective image data for subsequently highlighting other objects except the target object among multiple preset objects in the second live broadcast screen.

[0093] As an example of S2032, based on the above S203 example, based on the above S203 example, through the bilinear interpolation algorithm or the bicubic interpolation algorithm, the object 2 image 1 represented by (a) in Figure 6 is reduced to obtain the second object image represented by (b) in Figure 6 as object 2 image 2, the size of which is smaller than the size of object 2 image 1, and the resolution of which is smaller than the resolution of object 2 image 1.

[0094] Another implementation method is: the resizing strategy is a resizing model, and the training process of the resizing model is to pre-acquire a large number of first sample images and second sample images, the resolution of the first sample image is smaller than the resolution of the corresponding second sample image, that is, the size of the first sample image is smaller than the size of the corresponding second sample image; based on the large number of first sample images and second sample images, a first preset model is trained to obtain a resizing model; the resizing model learns the mapping relationship between the first sample image and the second sample image, that is, the mapping relationship between the low-resolution image and the high-resolution image.

[0095] Based on this, for the first object image of the target object, the first object image is enlarged using a size magnification model, that is, the size of the first object image is enlarged according to the mapping relationship between the first sample image and the second sample image, and the enlarged first object image is obtained as the second object image. Therefore, the present application provides a possible implementation method. When the preset adjustment strategy is a size adjustment strategy, the size adjustment strategy is a size magnification model. S203 is specifically S2033 (not shown in the figure): the first object image is enlarged according to the size magnification model to obtain the second object image; the size magnification model is used to enlarge the size of the first object image according to the mapping relationship between the first sample image and the second sample image, and the resolution of the first sample image is smaller than the resolution of the second sample image. The size of the second object image is larger than the size of the first object image.

[0096] The size enlargement model may be a super-resolution convolutional neural network (SRCNN) model, which is used to convert an input low-resolution image into a high-resolution image.

[0097] S2033 processes the first object image of the target object through the size magnification model, so that the first object image of the target object is accurately and intelligently magnified into a second object image. The second object image is integrated into the second live broadcast screen to highlight the target object among multiple preset objects, providing accurate and effective image data for subsequently highlighting the target object among multiple preset objects in the second live broadcast screen.

[0098] As an example of S2033, based on the above-mentioned S202 example, the object 2 image 1 represented by (a) in Figure 6 is enlarged by SRCNN to obtain a second object image as object 2 image 4. The size of object 2 image 4 is larger than the size of object 2 image 1, and the resolution of object 2 image 4 is greater than the resolution of object 2 image 1.

[0099] Another implementation method involves: the resizing strategy is a size reduction model. The training process for this size reduction model involves pre-acquiring a large number of third and fourth sample images, where the resolution of the third sample images is greater than the resolution of the corresponding fourth sample images, i.e., the size of the third sample images is greater than the size of the corresponding fourth sample images. Based on the large number of third and fourth sample images, a second preset model is trained to obtain a size reduction model. The size reduction model learns the mapping relationship between the third and fourth sample images, i.e., the mapping relationship between high-resolution images and low-resolution images. Based on this, for a first object image of the target object, the first object image is reduced in size using the size reduction model. i.e., the size of the first object image is reduced according to the mapping relationship between the third and fourth sample images, and the reduced first object image is obtained as the second object image. Therefore, the present application provides a possible implementation method. When the preset adjustment strategy is a size adjustment strategy, the size adjustment strategy is a size reduction model. S203 is specifically S2034 (not shown in the figure): the first object image is reduced according to the size reduction model to obtain a second object image; the size reduction model is used to reduce the size of the first object image according to the mapping relationship between the third sample image and the fourth sample image, and the resolution of the third sample image is greater than the resolution of the fourth sample image. The size of the second object image is smaller than the size of the first object image.

[0100] The size reduction model may be an Enhanced Super-Resolution Generative Adversarial Networks (ESRGAN) model, which is used to convert an input high-resolution image into a low-resolution image.

[0101] S2034 processes the first object image of the target object through a size reduction model, so that the first object image of the target object is accurately and intelligently reduced to a second object image. The second object image is integrated into the second live broadcast screen to highlight other objects except the target object among multiple preset objects, and provide accurate and effective image data for subsequently highlighting other objects except the target object among multiple preset objects in the second live broadcast screen.

[0102] As an example of S2034, based on the above S203 example, the second object image represented by (b) in FIG6 is obtained by reducing the object 2 image 1 represented by (a) in FIG6 through ESRGAN, which is the object 2 image 2. The size of the object 2 image 2 is smaller than the size of the object 2 image 1, and the resolution of the object 2 image 2 is smaller than the resolution of the object 2 image 1.

[0103] In the embodiment of the present application, when the preset adjustment strategy is a position adjustment strategy, in the specific implementation of S203 above, with respect to a first object image of the target object, the first object image is moved according to the position adjustment strategy to obtain a second object image at a position different from that of the first object image. Therefore, the present application provides a possible implementation method, where, when the preset adjustment strategy is a position adjustment strategy, S203 is specifically S2035 (not shown in the figure): the first object image is moved according to the position adjustment strategy to obtain a second object image; the position of the second object image is different from that of the first object image.

[0104] S2035 processes the first object image of the target object through a position adjustment strategy, so that the first object image of the target object is accurately moved to the second object image. The second object image is integrated into the second live screen to highlight the target object among multiple preset objects or other objects among multiple preset objects except the target object, providing accurate and effective image data for the subsequent highlighting of the preset object in the second live screen.

[0105] Position movement may refer to movement of the coordinates of the central pixel point. As an example of S2035, based on the above-mentioned S201 example, the terminal responds to the selection operation of the object 1 detection frame among the host detection frame, object 1 detection frame, object 2 detection frame, and object 3 detection frame in the live screen 1 shown in FIG4 (a), determines that the target detection frame is the object 1 detection frame from the host detection frame, object 1 detection frame, object 2 detection frame, and object 3 detection frame, segments object 1 in the object 1 detection frame from the live screen 1, and obtains object 1 image 1 in which the first object image of the target object is object 1; see FIG9, FIG9 is a schematic diagram provided by an embodiment of the present application of moving the first object image of the target object in the first live screen to a second object image, and moving the object 1 image 1 shown in FIG9 (a) through the position adjustment strategy to obtain the second object image as object 1 image 2, and the position of the object 1 image 2, that is, the center pixel coordinates (xt, yt), is different from the position of the object 1 image 1, that is, the center pixel coordinates (xs, ys).

[0106] Based on the above description, S202-S203 can be: according to the target object in the live voice corresponding to the first live screen, determine the target detection frame corresponding to the target object from multiple candidate detection frames; segment the target object in the target detection frame from the first live screen according to the target detection frame through the image segmentation algorithm, and obtain the second live screen and the first object image of the target object; enlarge the first object image according to the size upsampling strategy to obtain the second object image; or, enlarge the first object image according to the size enlargement model to obtain the second object image; or, move the first object image according to the position adjustment strategy to obtain the second object image.

[0107] This method can automatically detect whether the live voice corresponding to the first live screen includes a target object among multiple preset objects. If so, the first live screen can be processed by image segmentation to accurately segment the first live screen into a first object image of the target object to be adjusted and a second live screen to be fused; the first object image of the target object is processed by a size upsampling strategy, so that the first object image of the target object is accurately and controllably enlarged to the second object image; or, the first object image of the target object is processed by a size magnification model, so that the first object image of the target object is accurately and intelligently enlarged to the second object image, or, the first object image of the target object is processed by a position adjustment strategy, so that the first object image of the target object is accurately moved to the second object image; the second object image is fused into the second live screen to highlight the target object among the multiple preset objects, so that the target object among the multiple preset objects can be highlighted in the second live screen subsequently.

[0108] Specifically, this method automatically detects the target object within the pre-set objects described by the corresponding live audio in the live video feed, automatically enlarging or adjusting the target object's position within the live video feed, thereby highlighting the target object within the live video feed. In a product live broadcast scenario, this target object could be the product being introduced in the live audio feed corresponding to the live video feed.

[0109] As an example of S202-S203, based on the above Figure 3, in response to detecting that the live voice corresponding to the live broadcast screen 1 includes the target object being object 2, the terminal determines that the target detection frame is the object 2 detection frame from the anchor detection frame, the object 1 detection frame, the object 2 detection frame and the object 3 detection frame, that is, the bold frame in (b) in Figure 4; based on the above Figure 4, the terminal segments object 2 in the object 2 detection frame from the live broadcast screen 1, and obtains the second live broadcast screen represented by (a) in Figure 5 as live broadcast screen 2 and the first object image of the target object represented by (b) in Figure 5 as object 2 image 1. Based on Figure 5 above, the bilinear interpolation algorithm or the bicubic interpolation algorithm is used to enlarge the object 2 image 1 represented by Figure 6 (a) to obtain the second object image as object 2 image 3; or, the object 2 image 1 represented by Figure 6 (a) is enlarged by SRCNN to obtain the second object image as object 2 image 4; or, the object 2 image 1 represented by Figure 6 (a) is moved by the position adjustment strategy to obtain the second object image as object 2 image 5, which can be located at the center of the live broadcast screen 2. In the product live broadcast scenario, object 2 is the target product being introduced in the live broadcast voice corresponding to the live broadcast screen 1.

[0110] In the embodiment of the present application, when S203 is specifically implemented above, if the preset adjustment strategy includes a size adjustment strategy and a position adjustment strategy, for a first object image of the target object, the size adjustment strategy is used to scale the first object image to obtain an intermediate object image having a different size from the first object image; and the position adjustment strategy is used to move the intermediate object image to obtain a second object image having a different position from the intermediate object image. Therefore, the present application provides a possible implementation method, where the preset adjustment strategy is a position adjustment strategy, S203 includes the following S2036-S2037 (not shown in the figure):

[0111] S2036: Scale the first object image according to the size adjustment strategy to obtain an intermediate object image; the size of the intermediate object image is different from the size of the first object image.

[0112] S2037: Move the intermediate object image according to the position adjustment strategy to obtain a second object image; the position of the second object image is different from that of the intermediate object image.

[0113] The examples of S2036-S2037 refer to the above examples and are not repeated here.

[0114] In an embodiment of the present application, when the above-mentioned S204 is specifically implemented, since the second live broadcast screen refers to the first live broadcast screen from which the first object image of the target object is segmented, that is, the second live broadcast screen refers to other live broadcast screens except the first object image of the target object in the first live broadcast screen, and the second object image is the adjusted first object image; therefore, the first object image is adjusted to the second object image, for example, the first object image is reduced or shrunk to the second object image, or the first object image is moved to the second object image, and an area to be filled that does not match the surrounding pixels appears in the second live broadcast screen, and the area to be filled also needs to be filled with content that matches the surrounding pixels, so that the second object image and the filled second live broadcast screen can be subsequently fused to obtain the third live broadcast screen.

[0115] Based on this, when determining that the second live screen includes a to-be-filled area through the second object image and the second live screen, first, it is necessary to determine the surrounding similar area of ​​the to-be-filled area in the second live screen as a reference filling area; then, fill the to-be-filled area in the second live screen according to the regional content of the reference filling area, and obtain the filled second live screen as the fourth live screen; finally, fuse the second object image and the fourth live screen to obtain the third live screen. Therefore, this application provides a possible implementation method, S204 includes the following S2041-S2043 (not shown in the figure):

[0116] S2041: If it is determined according to the second object image and the second live broadcast image that the second live broadcast image includes an area to be filled, a similar area around the area to be filled in the second live broadcast image is determined as a reference filling area.

[0117] S2042: Filling the area to be filled in the second live broadcast picture with content according to the area content of the reference filling area to obtain a fourth live broadcast picture.

[0118] In actual applications, filling the area to be filled in the second live screen according to the area content of the reference filling area, and obtaining the filled second live screen as the fourth live screen means: filling the texture, color, shape and other features of the area content of the reference filling area into the area to be filled in the second live screen, and obtaining the filled second live screen as the fourth live screen.

[0119] S2043: Perform image fusion on the second object image and the fourth live broadcast picture to obtain a third live broadcast picture.

[0120] When the second live broadcast screen includes an area to be filled, S2041-S2043 fills the area to be filled with the area content of similar areas around the area to be filled to obtain a fourth live broadcast screen without the area to be filled, thereby fusing the second object image to obtain the third live broadcast screen, which can avoid the appearance of the area to be filled that does not match the surrounding pixels in the third live broadcast screen, thereby improving the display effect of the third live broadcast screen.

[0121] As an example of S2041-S2043, see FIG10, which is a schematic diagram of a second live screen including an area to be filled provided by an embodiment of the present application. Based on the above adjustment of the image 1 of the object 2 represented by (a) in FIG6 to obtain the image 2 of the object 2 represented by (b) in FIG6, the second live screen includes an area to be filled, that is, the area filled with the oblique line pattern in FIG10. The similar area around the area to be filled in the live screen 2 is determined as the reference filling area; the area to be filled in the live screen 2 is filled according to the area content of the reference filling area, and the filled live screen 2 is obtained as the fourth live screen, that is, live screen 4; the object image 2 and the live screen 4 are fused to obtain the live screen 3.

[0122] In addition, in an embodiment of the present application, based on the above-mentioned S2041-S2042, in the fourth live picture obtained from the second live picture, there is an unnatural transition problem between the filled area to be filled and the surrounding pixels; in order to solve this problem, it is also necessary to smooth the filled area to be filled in the fourth live picture based on the second live picture, and obtain the smoothed fourth live picture as the fifth live picture; correspondingly, when S2043 is specifically implemented, it is necessary to fuse the second object image and the fifth live picture to obtain the third live picture. Therefore, the present application provides a possible implementation method, S2043 is specifically S1 and S2044 (not shown in the figure): S1, smooth the filled area to be filled in the fourth live picture based on the second live picture to obtain the fifth live picture; S2044, fuse the second object image and the fifth live picture to obtain the third live picture.

[0123] Here, S1 may be smoothing the filled area to be filled in the fourth live picture according to the second live picture using the Boisson fusion algorithm to obtain the fifth live picture. The Boisson fusion algorithm is used to smooth the source image in the fused image according to the gradient field of the source image and the gradient field of the target image when fusing the source image into the target image.

[0124] S1 and S2044 obtain a fifth live picture by smoothing the filled area to be filled in the fourth live picture, so that the transition between the filled area to be filled in the fifth live picture and the surrounding pixels is more natural, thereby fusing the second object image to obtain a third live picture, thereby avoiding unnatural pixel transition in the third live picture, and further improving the display effect of the third live picture.

[0125] As an example of S1 and S2044, based on the live broadcast picture 4 obtained in the above S2041-S2043 examples, the filled area to be filled in the live broadcast picture 4 is smoothed based on the live broadcast picture 2 to obtain the smoothed live broadcast picture 4 as the fifth live broadcast picture, that is, the live broadcast picture 5; correspondingly, the object image 2 and the live broadcast picture 5 are fused to obtain the live broadcast picture 3.

[0126] Furthermore, in the embodiment of the present application, based on the above-described S204, in the third live picture obtained by fusing the second object image with the second live picture, there is an unnatural transition between the second object image and the second live picture. To address this issue, it is necessary to smooth the second object image in the third live picture based on the second live picture in the third live picture, obtaining the smoothed third live picture as the sixth live picture. Therefore, the present application provides a possible implementation method, which further includes S2 (not shown in the figure): smoothing the second object image in the third live picture based on the second live picture in the third live picture to obtain the sixth live picture.

[0127] The S2 smoothes the second object image in the third live screen to make the transition between the second object image in the sixth live screen and the second live screen more natural, and replaces the third live screen with the sixth live screen. On the basis of achieving the display effect of highlighting the preset object in the live screen to attract the audience to focus on the preset object in the live screen, the display effect of the sixth live screen is improved to further enhance the live broadcast effect.

[0128] As an example of S2, based on the above-mentioned S204 example, the image 2 of the object 2 in the live screen 3 is smoothed based on the live screen 2 in the live screen 3, and the smoothed live screen 3 is obtained as the sixth live screen, that is, the live screen 6.

[0129] The above S2 can be implemented in at least the following two ways:

[0130] One implementation method is: using the Boisson fusion algorithm to smooth the second object image in the third live screen according to the second live screen in the third live screen, to obtain a sixth live screen; because the Boisson fusion algorithm is used to smooth the source image in the fused image according to the gradient field of the source image and the gradient field of the target image when fusing the source image to the target image; therefore, taking the second object image as the source image and the second live screen as the target image, first, it is necessary to obtain the first gradient field of the second object image in the third live screen and the second gradient field of the second live screen in the third live screen; then, based on the first gradient field and the second gradient field, the second object image in the third live screen is smoothed to obtain the smoothed third live screen as the sixth live screen. That is, the present application provides a possible implementation method, S2 includes the following S21-S22 (not shown in the figure):

[0131] S21: Acquire a first gradient field of a second object image in a third live picture, and a second gradient field of a second live picture in the third live picture.

[0132] S22: performing a smoothing process on the second object image in the third live picture according to the first gradient field and the second gradient field to obtain a sixth live picture.

[0133] As an example of S21-S22, based on the above S2 example, the first gradient field of the image 2 of object 2 in the live screen 3 is obtained as gradient field 1, and the second gradient field of the live screen 2 in the live screen 3 is obtained as gradient field 2. Based on the gradient field 1 and the gradient field 2, the image 2 of object 2 in the live screen 3 is smoothed, and the smoothed live screen 3 is obtained as the sixth live screen, that is, live screen 6.

[0134] Another implementation method is to obtain a large number of source images and target images in advance, and the source image is to be fused into the target image; based on the large number of source images and target images, train a third preset model to obtain an image smoothing model; the image smoothing model learns the mapping relationship between the source image and the target image. Based on this, the second object image is used as the source image and the second live screen is used as the target image. The second object image in the third live screen is smoothed according to the mapping relationship between the second object image and the second live screen by the image smoothing model, and the smoothed third live screen is obtained as the sixth live screen. Therefore, the present application provides a possible implementation method, S2 is specifically S23 (not shown in the figure): according to the mapping relationship between the second object image and the second live screen, the second object image in the third live screen is smoothed by the image smoothing model to obtain the sixth live screen.

[0135] As an example of S23, based on the above-mentioned S2 example, the image 2 of object 2 in the live screen 3 is smoothed according to the mapping relationship between the image 2 of object 2 in the live screen 3 and the live screen 2 in the live screen 3 through the image smoothing model, and the smoothed live screen 3 is obtained as the sixth live screen, that is, live screen 6.

[0136] In an embodiment of the present application, if the above-mentioned adjustment method based on the live broadcast screen is executed in real time, it is necessary to execute a large number of calculation steps and consume a large amount of computing resources. To save computing resources, a certain detection frequency can be set. According to the set detection frequency, object detection is performed on the screen content of the first live broadcast screen to obtain candidate detection frames corresponding to multiple preset objects in the first live broadcast screen. Therefore, the present application provides a possible implementation method, S201 is specifically S2011 (not shown in the figure): object detection is performed on the screen content of the first live broadcast screen according to the detection frequency to obtain candidate detection frames corresponding to multiple preset objects in the first live broadcast screen.

[0137] S2011 processes the first live broadcast picture through object detection according to the detection frequency, and can intermittently and accurately detect the candidate detection frame of each preset object in the first live broadcast picture. On the basis of saving computing resources, it provides effective and accurate detection data for subsequently accurately segmenting the first live broadcast picture into the first object image of the target object to be adjusted and the second live broadcast picture to be fused.

[0138] As an example of S2011, the detection frequency is tar_freq. Based on the above S201 example, according to tar_freq, the object detection is performed for the picture content of the live screen 1 to obtain the anchor detection frame, object 1 detection frame, object 2 detection frame and object 3 detection frame in the live screen 1.

[0139] Among them, the detection frequency can be dynamically updated based on whether the preset object in the first live screen of the two frames before and after moves. In actual applications, any preset object among the multiple preset objects is taken as the first object, and i is a positive integer; when the first object in the first live screen of the i+1th frame moves relative to the first object in the first live screen of the i-th frame, it is necessary to perform extremely frequent object detection to obtain multiple candidate detection frames corresponding to the multiple preset objects in the first live screen, and the maximum frequency is used as the detection frequency; when the multiple preset objects in the first live screen of the i+1th frame do not move relative to the multiple preset objects in the first live screen of the i-th frame, the detection frequency can be reduced, and there is no need to perform extremely frequent object detection to obtain multiple candidate detection frames corresponding to the multiple preset objects in the first live screen, and the difference frequency between the detection frequency and the preset frequency is used as the detection frequency, and when the detection frequency is the minimum frequency, there is no need to further reduce the detection frequency. Therefore, the present application provides a possible implementation method, and the step of obtaining the detection frequency includes the following S3 or S4:

[0140] S3: If the first object in the first live broadcast picture of the (i+1)th frame moves relative to the first object in the first live broadcast picture of the (i)th frame, update the detection frequency according to the maximum frequency, where i is a positive integer.

[0141] S4: If the multiple preset objects in the first live broadcast picture of the (i+1)th frame do not move relative to the multiple preset objects in the first live broadcast picture of the i-th frame, update the detection frequency according to the difference between the detection frequency and the preset frequency; the updated detection frequency is greater than or equal to the minimum frequency.

[0142] S3 and S4 dynamically update the detection frequency based on whether the preset object in the two preceding and subsequent frames of the first live screen moves, and on the basis of saving computing resources, when the preset object in the two preceding and subsequent frames of the first live screen moves, the object detection is performed extremely frequently at the maximum detection frequency to obtain multiple candidate detection frames corresponding to the multiple preset objects in the first live screen; when the preset object in the two preceding and subsequent frames of the first live screen does not move, the detection frequency of the multiple candidate detection frames corresponding to the multiple preset objects in the first live screen is reduced.

[0143] As an example of S3 and S4, based on the above-mentioned S201 example, refer to Figure 11. Figure 11 is a flowchart for determining the detection frequency provided in an embodiment of the present application, and the process refers to: obtaining the detection frequency; judging whether the first object in the first live broadcast picture of the i+1th frame moves relative to the first object in the first live broadcast picture of the i-th frame, and if so, updating the detection frequency according to the maximum frequency; if not, that is, the multiple preset objects in the first live broadcast picture of the i+1th frame do not move relative to the multiple preset objects in the first live broadcast picture of the i-th frame, updating the detection frequency according to the difference frequency between the detection frequency and the preset frequency; the updated detection frequency is greater than or equal to the minimum frequency.

[0144] Among them, the implementation method of determining that the first object in the first live broadcast picture of the i+1th frame moves relative to the first object in the first live broadcast picture of the i-th frame refers to: determining that the positions of a part of the pixel points of the first object in the first live broadcast picture of the i+1th frame have changed relative to the multiple pixel points of the first object in the first live broadcast picture of the i-th frame. Based on this, first, obtain a plurality of coordinate differences between the coordinates of the multiple pixel points of the first object in the first live broadcast picture of the i+1th frame and the coordinates of the multiple pixel points of the first object in the first live broadcast picture of the i-th frame; then, determine whether a preset number of coordinate differences in the multiple coordinate differences are greater than the preset difference. If so, it can be determined that the first object in the first live broadcast picture of the i+1th frame moves relative to the first object in the first live broadcast picture of the i-th frame; if not, it means that the first object in the first live broadcast picture of the i+1th frame does not move relative to the first object in the first live broadcast picture of the i-th frame; wherein the preset number is configured according to actual needs. Therefore, the present application provides a possible implementation method, and the steps of determining that the first object in the first live broadcast picture of the i+1th frame moves relative to the first object in the first live broadcast picture of the i-th frame include the following S5-S6:

[0145] S5: Obtain multiple coordinate differences between the coordinates of multiple pixel points of the first object in the (i+1)th frame of the first live picture and the coordinates of multiple pixel points of the first object in the (i)th frame of the first live picture.

[0146] S6: If a preset number of coordinate differences among the plurality of coordinate differences are greater than a preset difference, it is determined that the first object in the (i+1)th frame of the first live picture moves relative to the first object in the (i)th frame of the first live picture.

[0147] In S5-S6, the positions of a portion of pixel points of the first object in the first live broadcast picture of the i+1th frame relative to the corresponding portion of pixel points of the first object in the first live broadcast picture of the i-th frame change, and it is determined that the first object in the first live broadcast picture of the i+1th frame moves relative to the first object in the first live broadcast picture of the i-th frame. This can accurately judge whether the preset objects in the two previous and next frames of the first live broadcast picture move, and provide an accurate update basis for the subsequent dynamic update detection frequency.

[0148] As an example of S5-S6, a preset difference value is Diff_Threshold. N coordinate differences are obtained between the N pixel coordinates of the first object in the first live broadcast frame (i+1) and the N pixel coordinates of the first object in the first live broadcast frame (i). The coordinate difference refers to the spatial distance between the pixel coordinates of the first object in the first live broadcast frame (i+1) and the pixel coordinates of the first object in the first live broadcast frame (i). A determination is made as to whether N / 16 of the N coordinate differences are greater than Diff_Threshold, where N is a positive integer and a multiple of 16. If so, it is determined that the first object in the first live broadcast frame (i+1) has moved relative to the first object in the first live broadcast frame (i). Diff_Threshold may be 1 / 100 of the width of the first live broadcast frame.

[0149] In addition, in an embodiment of the present application, when executing S203, based on the principle of highlighting one or more of a plurality of preset objects, the target object is adjusted so that the preset object can be highlighted in the live broadcast screen later. Alternatively, the target object can be replaced to highlight the preset object in the live broadcast screen. In this case, the preset adjustment strategy can also be a replacement adjustment strategy. For the first object image of the target object, the first object image is replaced with a preset replacement image different from the first object image through the replacement adjustment strategy, and the preset replacement image is used as the second object image. Therefore, the present application provides a possible implementation method, in which the preset adjustment strategy also includes a replacement adjustment strategy. Specifically, S203 is S2038 (not shown in the figure): the first object image is replaced according to the preset replacement image through the replacement adjustment strategy to obtain the second object image.

[0150] S2038 processes the first object image of the target object through a replacement adjustment strategy, so that the first object image of the target object is replaced with a preset replacement image as the second object image. The second object image is integrated into the second live screen to highlight other objects except the target object among multiple preset objects, and provide accurate and effective image data for subsequently highlighting other objects except the target object among multiple preset objects in the second live screen.

[0151] As an example of S2038, based on the above S203 example, the object 2 image 1 shown in FIG6 (a) is replaced with a preset replacement image through a replacement adjustment strategy, and the second object image obtained is the preset replacement image.

[0152] In addition, in an embodiment of the present application, after executing S202 to divide the first live broadcast screen into the second live broadcast screen and the first object image of the target object, the first object image can also be deleted, and the second live broadcast screen can be filled with content to obtain a seventh live broadcast screen; so that other objects except the target object among multiple preset objects can be highlighted in the seventh live broadcast screen, achieving a display effect of highlighting the preset objects in the live broadcast screen, so as to attract audience users to focus on the preset objects in the live broadcast screen, thereby improving the live broadcast effect.

[0153] It should be noted that, based on the implementation methods provided in the above aspects, this application can also be further combined to provide more implementation methods.

[0154] Based on the adjustment method based on the live screen provided in the embodiment corresponding to FIG2 , the embodiment of the present application further provides an adjustment device based on the live screen. Referring to FIG12 , FIG12 is a structural diagram of an adjustment device based on the live screen provided in the embodiment of the present application. The adjustment device based on the live screen 1200 includes: a detection unit 1201, a segmentation unit 1202, an adjustment unit 1203, and a fusion unit 1204.

[0155] The detection unit 1201 is configured to perform object detection on the content of the first live broadcast image to obtain candidate detection frames corresponding to a plurality of preset objects in the first live broadcast image.

[0156] a segmentation unit 1202 configured to segment a target object in the target detection frame from the first live broadcast image based on the target detection frames in the plurality of candidate detection frames, and obtain a second live broadcast image and a first object image of the target object, where the second live broadcast image refers to the live broadcast image excluding the first object image of the target object in the first live broadcast image;

[0157] An adjusting unit 1203 is configured to adjust the first object image according to a preset adjustment strategy to obtain a second object image; the preset adjustment strategy is configured to highlight one or more of a plurality of preset objects;

[0158] The fusion unit 1204 is configured to perform image fusion on the second object image and the second live picture to obtain a third live picture.

[0159] In a possible implementation, the preset adjustment strategy includes one or more of a size adjustment strategy and a position adjustment strategy.

[0160] In a possible implementation, when the preset adjustment strategy is a size adjustment strategy, the size adjustment strategy is a size upsampling strategy, and the adjustment unit 1203 is specifically configured to:

[0161] Enlarging the first object image according to a size upsampling strategy to obtain a second object image; or,

[0162] The size adjustment strategy is a size downsampling strategy, and the adjustment unit 1203 is specifically used to:

[0163] The first object image is downsized according to a downsampling strategy to obtain a second object image.

[0164] In a possible implementation, when the preset adjustment strategy is a size adjustment strategy, the size adjustment strategy is a size enlargement model, and the adjustment unit 1203 is specifically configured to:

[0165] The first object image is enlarged according to the size enlargement model to obtain a second object image; the size enlargement model is used to enlarge the size of the first object image according to the mapping relationship between the first sample image and the second sample image, and the resolution of the first sample image is smaller than the resolution of the second sample image; or,

[0166] The size adjustment strategy is a size reduction model, and the adjustment unit 1203 is specifically used to:

[0167] The first object image is reduced in size according to the size reduction model to obtain a second object image; the size reduction model is used to reduce the size of the first object image according to a mapping relationship between the third sample image and the fourth sample image, and the resolution of the third sample image is greater than the resolution of the fourth sample image.

[0168] In a possible implementation, when the preset adjustment strategy is a position adjustment strategy, the adjustment unit 1203 is specifically configured to:

[0169] The first object image is moved according to the position adjustment strategy to obtain a second object image; the position of the second object image is different from that of the first object image.

[0170] In a possible implementation, the fusion unit 1204 is specifically configured to:

[0171] If it is determined according to the second object image and the second live broadcast image that the second live broadcast image includes the area to be filled, a similar area around the area to be filled in the second live broadcast image is determined as a reference filling area;

[0172] Filling the to-be-filled area in the second live picture with content according to the area content of the reference filling area to obtain a fourth live picture;

[0173] The second object image and the fourth live picture are fused to obtain a third live picture.

[0174] In a possible implementation, the apparatus further includes: a smoothing unit;

[0175] a smoothing unit, configured to smooth the filled area to be filled in the fourth live picture according to the second live picture, to obtain a fifth live picture;

[0176] The fusion unit 1204 is specifically configured to:

[0177] The second object image and the fifth live picture are fused to obtain a third live picture.

[0178] In a possible implementation, the smoothing unit is further configured to:

[0179] According to the second live picture in the third live picture, a smoothing process is performed on the second object image in the third live picture to obtain a sixth live picture.

[0180] In a possible implementation, the smoothing unit is specifically configured to:

[0181] Acquire a first gradient field of the second object image in the third live picture, and a second gradient field of the second live picture in the third live picture;

[0182] The second object image in the third live picture is smoothed according to the first gradient field and the second gradient field to obtain a sixth live picture.

[0183] In a possible implementation, the smoothing unit is specifically configured to:

[0184] According to the mapping relationship between the second object image and the second live picture, the second object image in the third live picture is smoothed by an image smoothing model to obtain a sixth live picture.

[0185] In a possible implementation, the detection unit 1201 is specifically configured to:

[0186] Object detection is performed on the picture content of the first live broadcast picture according to the detection frequency to obtain candidate detection frames corresponding to multiple preset objects in the first live broadcast picture.

[0187] In a possible implementation, the apparatus further includes: an updating unit;

[0188] Update unit, specifically used for:

[0189] If the first object in the first live broadcast picture of the (i+1)th frame moves relative to the first object in the first live broadcast picture of the (i)th frame, the detection frequency is updated according to the maximum frequency, the first object is any preset object among multiple preset objects, and i is a positive integer;

[0190] If the multiple preset objects in the first live broadcast picture of the i+1th frame do not move relative to the multiple preset objects in the first live broadcast picture of the i-th frame, the detection frequency is updated according to the difference between the detection frequency and the preset frequency; the updated detection frequency is greater than or equal to the minimum frequency.

[0191] In a possible implementation, the apparatus further includes: a determining unit;

[0192] Identify units, specifically for:

[0193] Obtaining a plurality of coordinate differences between the coordinates of a plurality of pixel points of the first object in the first live picture of the (i+1)th frame and the coordinates of a plurality of pixel points of the first object in the first live picture of the (i)th frame;

[0194] If a preset number of coordinate differences among the plurality of coordinate differences is greater than a preset difference, it is determined that the first object in the (i+1)th frame of the first live picture moves relative to the first object in the (i)th frame of the first live picture.

[0195] In a possible implementation, the preset adjustment strategy further includes a replacement adjustment strategy; the adjustment unit 1203 is specifically configured to:

[0196] The first object image is replaced according to a preset replacement image using a replacement adjustment strategy to obtain a second object image.

[0197] As can be seen from the above technical solution, based on the content of the first live broadcast screen, object detection is performed to obtain multiple candidate detection frames corresponding to multiple preset objects in the first live broadcast screen. Based on the target detection frames in the multiple candidate detection frames, the target object in the target detection frames is segmented from the first live broadcast screen to obtain a second live broadcast screen and a first object image of the target object. Through the detection unit and the segmentation unit, the first live broadcast screen can be accurately segmented into the first object image of the target object to be adjusted and the second live broadcast screen to be fused. According to a preset adjustment strategy for highlighting one or more of the multiple preset objects, the first object image is adjusted to obtain a second object image. The second object image and the second live broadcast screen are fused to obtain a third live broadcast screen. Through the adjustment unit and the fusion unit based on the preset adjustment strategy, the first object image of the target object is changed to a second object image and fused into the second live broadcast screen to obtain a third live broadcast screen, so that the third live broadcast screen can highlight one or more of the multiple preset objects. Based on this, the method can achieve the display effect of highlighting the preset objects in the live broadcast screen by detecting, segmenting, adjusting, and re-integrating the target objects into the live broadcast screen without changing the size of the live broadcast screen, thereby attracting viewers to focus on the preset objects in the live broadcast screen and improving the live broadcast effect.

[0198] The embodiment of the present application also provides a computer device, which can be a terminal. See Figure 13, which is a structural diagram of a terminal provided by the embodiment of the present application. Taking the terminal as a smartphone as an example, the smartphone includes components such as a radio frequency (RF) circuit 1310, a memory 1320, an input unit 1330, a display unit 1340, a sensor 1350, an audio circuit 1360, a wireless fidelity (WiFi) module 1370, a processor 1380, and a power supply 13120. The input unit 1330 may include a touch panel 1331 and other input devices 1332, the display unit 1340 may include a display panel 1341, and the audio circuit 1360 may include a speaker 1361 and a microphone 1362. Those skilled in the art will understand that the smartphone structure shown in Figure 13 does not constitute a limitation on smartphones, and may include more or fewer components than shown, or combine certain components, or arrange components differently.

[0199] The memory 1320 can be used to store software programs and modules. The processor 1380 executes the various functional applications and data processing of the smartphone by running the software programs and modules stored in the memory 1320. The memory 1320 may mainly include a program storage area and a data storage area. The program storage area may store an operating system, at least one application required for a function (such as a sound playback function, an image playback function, etc.), etc.; the data storage area may store data created based on the use of the smartphone (such as audio data, a phone book, etc.). In addition, the memory 1320 may include a high-speed random access memory and may also include a non-volatile memory, such as at least one disk storage device, a flash memory device, or other volatile solid-state storage device.

[0200] Processor 1380 is the control center of the smartphone, connecting all components of the smartphone using various interfaces and circuits. It executes software programs and / or modules stored in memory 1320 and accesses data stored in memory 1320 to perform various smartphone functions and process data. Optionally, processor 1380 may include one or more processing units. Preferably, processor 1380 integrates an application processor and a modem processor. The application processor primarily handles the operating system, user interface, and application programs, while the modem processor primarily handles wireless communications. It is understood that the modem processor may not be integrated into processor 1380.

[0201] In this embodiment, the processor 1380 in the smartphone may execute the methods provided in various optional implementations of the above embodiments.

[0202] The computer device provided in the embodiment of the present application can also be a server. See Figure 14, which is a structural diagram of a server provided in the embodiment of the present application. The server 1400 may have relatively large differences due to different configurations or performances, and may include one or more processors, such as a central processing unit (CPU) 1422, and a memory 1432, one or more storage media 1430 (such as one or more massive storage devices) storing application programs 1442 or data 1444. Among them, the memory 1432 and the storage medium 1430 can be temporary storage or permanent storage. The program stored in the storage medium 1430 may include one or more modules (not shown in the figure), and each module may include a series of instruction operations on the server. Furthermore, the central processing unit 1422 can be configured to communicate with the storage medium 1430 to execute a series of instruction operations in the storage medium 1430 on the server 1400.

[0203] The server 1400 may also include one or more power supplies 1426, one or more wired or wireless network interfaces 1450, one or more input and output interfaces 1458, and / or one or more operating systems 1441, such as Windows Server 2003. TM , Mac OS X TM , Unix TM ,Linux TM , FreeBSD TM etc.

[0204] In this embodiment, the central processing unit 1422 in the server 1400 can execute the methods provided in various optional implementations of the above embodiments.

[0205] According to one aspect of the present application, a computer-readable storage medium is provided, which is used to store a computer program. When the computer program is run on a computer device, the computer device executes the methods provided in various optional implementations of the above embodiments.

[0206] According to one aspect of the present application, a computer program product is provided, comprising a computer program stored in a computer-readable storage medium. A processor of a computer device reads the computer program from the computer-readable storage medium and executes the computer program, causing the computer device to perform the methods provided in various optional implementations of the above-described embodiments.

[0207] The descriptions of the processes or structures corresponding to the above figures have different emphases. For parts that are not described in detail in a certain process or structure, please refer to the relevant descriptions of other processes or structures.

[0208] The terms "first", "second" etc. in the specification of the present application and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequential order. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein, for example, can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, for example, processes, methods, systems, products or equipment comprising a series of steps or units are not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or that are inherent to these processes, methods, products or equipment.

[0209] In the several embodiments provided in this application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of the units is merely a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interfaces, devices or units, which can be electrical, mechanical or other forms.

[0210] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.

[0211] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.

[0212] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: various media that can store computer programs, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a RAM, a magnetic disk or an optical disk.

[0213] As described above, the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the above embodiments, ordinary technical members in this field should understand that they can still modify the technical solutions recorded in the above embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present application.

Claims

1. A method for adjusting a live video screen, which is executed by a computer device, and the method includes: Performing object detection on the screen content of a first live video screen to obtain candidate detection frames respectively corresponding to multiple preset objects in the first live video screen; According to a target detection frame among the multiple candidate detection frames, segmenting the target object in the target detection frame from the first live video screen to obtain a second live video screen and a first object image of the target object, where the second live video screen refers to the live video screen in the first live video screen except for the first object image of the target object; Performing adjustment processing on the first object image according to a preset adjustment strategy to obtain a second object image; the preset adjustment strategy is used to highlight one or more of the multiple preset objects; Performing image fusion on the second object image and the second live video screen to obtain a third live video screen.

2. The method according to claim 1, wherein the preset adjustment strategy includes one or more of a size adjustment strategy and a position adjustment strategy.

3. The method according to claim 2, when the preset adjustment strategy is the size adjustment strategy, the size adjustment strategy is a size upsampling strategy, and the performing adjustment processing on the first object image according to the preset adjustment strategy to obtain a second object image is specifically: Performing an enlarging process on the first object image according to the size upsampling strategy to obtain the second object image; or, The size adjustment strategy is a size downsampling strategy, and the performing adjustment processing on the first object image according to the preset adjustment strategy to obtain a second object image is specifically: Performing a reducing process on the first object image according to the size downsampling strategy to obtain the second object image.

4. The method according to claim 2, when the preset adjustment strategy is the size adjustment strategy, the size adjustment strategy is a size enlargement model, and the performing adjustment processing on the first object image according to the preset adjustment strategy to obtain a second object image is specifically: Performing an enlarging process on the first object image according to the size enlargement model to obtain the second object image; the size enlargement model is used to enlarge the size of the first object image according to the mapping relationship between a first sample image and a second sample image, and the resolution of the first sample image is less than the resolution of the second sample image; or, The size adjustment strategy is a size reduction model, and the performing adjustment processing on the first object image according to the preset adjustment strategy to obtain a second object image is specifically: Performing a reducing process on the first object image according to the size reduction model to obtain the second object image; the size reduction model is used to reduce the size of the first object image according to the mapping relationship between a third sample image and a fourth sample image, and the resolution of the third sample image is greater than the resolution of the fourth sample image.

5. The method according to claim 2, when the preset adjustment strategy is the position adjustment strategy, the performing adjustment processing on the first object image according to the preset adjustment strategy to obtain a second object image is specifically: Perform position movement on the first object image according to the position adjustment strategy to obtain the second object image; the position of the second object image is different from that of the first object image.

6. The method according to any one of claims 1-5, wherein the image fusion of the second object image and the second live broadcast screen to obtain a third live broadcast screen includes: If it is determined according to the second object image and the second live broadcast screen that the second live broadcast screen includes an area to be filled, determine the surrounding similar area of the area to be filled in the second live broadcast screen as a reference filling area; Perform content filling on the area to be filled in the second live broadcast screen according to the area content of the reference filling area to obtain a fourth live broadcast screen; Perform image fusion on the second object image and the fourth live broadcast screen to obtain the third live broadcast screen.

7. The method according to claim 6, wherein the performing image fusion on the second object image and the fourth live broadcast screen to obtain the third live broadcast screen is specifically: Perform smoothing processing on the filled area to be filled in the fourth live broadcast screen according to the second live broadcast screen to obtain a fifth live broadcast screen; Perform image fusion on the second object image and the fifth live broadcast screen to obtain the third live broadcast screen.

8. The method according to any one of claims 1-7, after obtaining the third live broadcast screen, the method further includes: Perform smoothing processing on the second object image in the third live broadcast screen according to the second live broadcast screen in the third live broadcast screen to obtain a sixth live broadcast screen.

9. The method according to claim 8, wherein the performing smoothing processing on the second object image in the third live broadcast screen according to the second live broadcast screen in the third live broadcast screen to obtain a sixth live broadcast screen includes: Obtain the first gradient field of the second object image in the third live broadcast screen and the second gradient field of the second live broadcast screen in the third live broadcast screen; Perform smoothing processing on the second object image in the third live broadcast screen according to the first gradient field and the second gradient field to obtain the sixth live broadcast screen.

10. The method according to claim 8, wherein the performing smoothing processing on the second object image in the third live broadcast screen according to the second live broadcast screen in the third live broadcast screen to obtain a sixth live broadcast screen is specifically: Perform smoothing processing on the second object image in the third live broadcast screen through an image smoothing model according to the mapping relationship between the second object image and the second live broadcast screen to obtain the sixth live broadcast screen.

11. The method according to any one of claims 1-10, wherein the performing object detection on the content of the first live broadcast screen to obtain candidate detection frames corresponding to multiple preset objects in the first live broadcast screen is specifically: Perform object detection on the content of the first live broadcast screen according to the detection frequency to obtain candidate detection frames corresponding to multiple preset objects in the first live broadcast screen.

12. The method according to claim 11, wherein the step of obtaining the detection frequency comprises: If a first object in the (i + 1)-th frame of the first live video moves relative to the first object in the i-th frame of the first live video, updating the detection frequency according to the maximum frequency, where the first object is any one of the plurality of preset objects, and i is a positive integer; If the plurality of preset objects in the (i + 1)-th frame of the first live video do not move relative to the plurality of preset objects in the i-th frame of the first live video, updating the detection frequency according to the difference frequency between the detection frequency and the preset frequency; the updated detection frequency is greater than or equal to the minimum frequency.

13. The method according to claim 12, wherein the step of determining that a first object in the (i + 1)-th frame of the first live video moves relative to the first object in the i-th frame of the first live video comprises: Obtaining a plurality of coordinate differences between the pixel point coordinates of the first object in the (i + 1)-th frame of the first live video and the pixel point coordinates of the first object in the i-th frame of the first live video; If a preset number of the plurality of coordinate differences are greater than a preset difference, determining that the first object in the (i + 1)-th frame of the first live video moves relative to the first object in the i-th frame of the first live video.

14. The method according to claim 1 or 2, wherein the preset adjustment strategy further comprises a replacement adjustment strategy; the step of adjusting the first object image according to the preset adjustment strategy to obtain a second object image specifically comprises: Performing replacement processing on the first object image according to a preset replacement image through the replacement adjustment strategy to obtain the second object image.

15. An adjustment device based on a live broadcast screen, the device is deployed on a computer device, and the device includes: A detection unit, a segmentation unit, an adjustment unit, and a fusion unit; The detection unit is configured to perform object detection on the content of the first live video to obtain candidate detection frames corresponding to the plurality of preset objects in the first live video; The segmentation unit is configured to segment the target object in the target detection frame from the first live video according to the target detection frame among the plurality of candidate detection frames to obtain a second live video and a first object image of the target object, where the second live video refers to the live video in the first live video other than the first object image of the target object; The adjustment unit is configured to perform adjustment processing on the first object image according to a preset adjustment strategy to obtain a second object image; The preset adjustment strategy is used to highlight one or more of the plurality of preset objects; The fusion unit is configured to perform image fusion on the second object image and the second live video to obtain a third live video.

16. A computer device, the computer device comprising a processor and a memory: The memory is configured to store a computer program and transmit the computer program to the processor; The processor is configured to execute the method according to any one of claims 1-14 according to the instructions in the computer program.

17. A computer-readable storage medium for storing a computer program, which, when run on a computer device, causes the computer device to execute the method according to any one of claims 1-14.

18. A computer program product comprising a computer program, which, when run on a computer device, causes the computer device to execute the method according to any one of claims 1-14.