Live broadcast goods-carrying commodity pushing method based on artificial intelligence

Through an artificial intelligence-based method, the head, clothing and body area images of virtual anchors and combined with voice synthesis model, the trust and generation cost problems of virtual anchors in the sales scene are solved, and efficient and real-time virtual live broadcast push is achieved.

CN120075484AActive Publication Date: 2025-05-30MEIFU NET (GUANGZHOU) TECH CO LTD
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202510185846.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-20
Publication Date
2025-05-30
Estimated Expiration
2045-02-20

AI Technical Summary

Technical Problem

There is a gap in the naturalness of action and language expression ability of existing virtual anchors, which leads to low audience trust when selling clothing and beauty products, and the generation cost of virtual anchors is high.

Method used

Through an artificial intelligence-based method, virtual anchor images are generated by fusion of images in the head area, clothing area and body size area, and live video streams are generated in real time with voice synthesis model.

Benefits of technology

It realizes the display of the upper-body effect of clothing products by virtual anchors, and reduces the demand for graphics computing resources generated by virtual anchors, and improves the real-time and efficiency of live streaming.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120075484A_ABST
    Figure CN120075484A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of live broadcast goods carrying, in particular to a live broadcast goods carrying commodity pushing method based on artificial intelligence, and the method comprises the steps: generating a virtual anchor image of a current frame based on the fusion of a head region image, a clothing region image and a body type region image; splicing the virtual anchor image of the current frame into a live broadcast room background image to obtain a virtual live broadcast video frame; acquiring to-be-pushed clothing data, inputting the to-be-pushed clothing data into the voice synthesis model, and predicting to obtain live voice data; aligning the pushed voice data with the virtual live video frame, and generating a live video stream in real time; and sending the live broadcast video stream to a server for live broadcast stream pushing. According to the method, the virtual anchor image is divided into the body type area, the clothing area and the head area, the clothing area image and the head area image are spliced above the image layer of the body type area image, the complete virtual anchor image is generated, and the virtual anchor image is quickly generated for live broadcast stream pushing.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of live streaming with goods, and particularly to a method for pushing live streaming with goods based on artificial intelligence. Background Art

[0002] With the rapid development of the Internet economy, online live streaming is changing people's lifestyles, and live streaming with goods, as a new e-commerce marketing model, has also started to develop rapidly.

[0003] Compared with traditional e-commerce, the operating cost of live streaming with goods is actually higher. Because live streaming with goods requires hiring a live streaming team, and the live streaming team takes a commission as a reward in the completed transactions. In order to reduce the operating cost of live streaming with goods, many e-commerce platforms have started to use virtual video live streaming technology for live streaming with goods. The virtual anchor generated based on the virtual video live streaming technology can automatically generate corresponding voices and actions according to the input text.

[0004] At present, there is a certain gap between virtual anchors and real anchors in terms of the naturalness of movements and language expression abilities. In complex live streaming with goods scenarios, the audience's trust in virtual anchors is not high, such as when promoting clothing and beauty products. When promoting clothing and beauty products, real anchors will display the wearing effects of the products.

[0005] Existing virtual anchors need to complete 3D modeling of the virtual anchor and the product before the live stream. Among them, modeling and streaming the texture and gloss of clothing require a large amount of graphics computing resources, increasing the generation cost of virtual anchors.

[0006] Therefore, in order to enable virtual anchors to display the wearing effects of clothing products and conduct real-time virtual live streaming, this application provides a method for pushing live streaming with goods based on artificial intelligence. Summary of the Invention

[0007] To overcome the problems existing in the related art, this application provides a method for pushing live streaming with goods based on artificial intelligence, including: S1. Generating a virtual anchor image of the current frame by fusing the head region image, the clothing region image, and the body shape region image; the clothing region image is obtained by image segmentation of a captured real clothing image, the body shape region image is obtained by image segmentation of a captured real human body image, and the head region image is a human image for generating the head of the virtual anchor; S2. Splicing the virtual anchor image of the current frame into the live streaming room background image to obtain a virtual live streaming video frame; S3. Obtaining the clothing data to be pushed, and inputting the clothing data to be pushed into a speech synthesis model to predict and obtain live streaming voice data; S4. Align the pushed voice data and the virtual live video frames to generate a live video stream in real time; S5. Send the live video stream to the server side for live streaming.

[0008] In one implementation, step S1 specifically includes: S101. Collect real person images in real time; the real person in the real person image wears a motion capture suit, and the surface of the motion capture suit is provided with N rectangular grids, and each of the rectangular grids is provided with a grid identification code; N is an integer greater than or equal to 2; S102. Perform contour recognition on the real person image of the current frame, and segment the head region image and the body shape region image; S103. Obtain the clothing region image; S104. Cover the clothing region image above the layers of each of the rectangular grids according to the grid identification code; S105. Obtain the head region image, and cover the head region image above the layer of the body shape region image to obtain the virtual anchor image.

[0009] In one implementation, in step S104, it specifically includes: S1041. Calculate the area of the rectangular grid of the body shape region image; S1042. Determine the rectangular grid with the largest area as the central grid; S1043. Extract the grid identification code of the central grid; S1044. Search for the corresponding human body offset angle according to the grid identification code; S1045. Determine the image regions corresponding to the central grid and the clothing region image according to the human body offset angle; S1046. Cover the clothing region image above the layer of the central grid; S1047. Perform layer covering on the rectangular grids around the central grid until all the rectangular grids are filled.

[0010] In one implementation, in step S102, it specifically includes: Extract and recognize the contour features of the input real person image through a pre-trained YOLOv8 model, and segment the head region image and the body shape region image along the contour of the real person.

[0011] In one implementation, step S103 specifically includes: Take 360-degree photos around the center line of the clothing to be pushed to obtain the clothing area image, which is a panoramic view of the clothing to be pushed. In one embodiment, step S103 specifically includes: Cut the clothing to be pushed open, flatten it and then take a photo to obtain the clothing area image.

[0012] In one embodiment, step S3 specifically includes: The voice synthesis model is built based on a large language model, which generates live streaming text from the input clothing parameter text, and then predicts live voice data through the live streaming text.

[0013] In one embodiment, step S5 specifically includes: Perform frame classification on the commodity push video stream through an encoder; Compress and encapsulate each video frame and send it to the server side.

[0014] The technical solution provided by this application may include the following beneficial effects: In this application, the virtual anchor image is divided into a body area, a clothing area, and a head area. Images of the three areas are obtained respectively, and then the clothing area image and the head area image are spliced above the layer of the body area image to generate a complete virtual anchor image. Finally, the virtual live video frames and the live voice data are spliced into virtual live video frames, and a live video stream is output in real time on the server side. When generating the virtual anchor image in this application, the virtual anchor image is quickly generated through the images of the 3 areas, and the virtual anchor image is quickly generated for live streaming push.

[0015] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit this application. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] By describing the exemplary embodiments of the present application in more detail in conjunction with the accompanying drawings, the above and other objects, features, and advantages of the present application will become more obvious. Among them, in the exemplary embodiments of the present application, the same reference numerals generally represent the same components.

[0017] Figure 1 It is a flowchart of the live streaming commodity push method shown in the embodiments of this application; Figure 2 is Figure 1 a flowchart of step S1 of the live streaming commodity push method shown; Figure 3 is Figure 2 a flowchart of step S104 of the live streaming commodity push method shown. Detailed implementation manners

[0018] The preferred implementation manners of the present application will be described in more detail below with reference to the accompanying drawings. Although the preferred implementation manners of the present application are shown in the drawings, it should be understood that the present application can be implemented in various forms and should not be limited by the implementation manners set forth herein. On the contrary, these implementation manners are provided to make the present application more thorough and complete, and to convey the scope of the present application to those skilled in the art completely.

[0019] The terms used in the present application are only for the purpose of describing specific embodiments and are not intended to limit the present application. The singular forms "a", "the" and "said" used in the present application and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term "and / or" used herein refers to and includes any or all possible combinations of one or more of the associated listed items.

[0020] It should be understood that although the terms "first", "second", "third", etc. may be used in the present application to describe various information, such information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other. For example, without departing from the scope of the present application, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more of such features. In the description of the present application, "a plurality of" means two or more unless otherwise specifically defined. Embodiment

[0021] Existing virtual anchors need to complete the 3D modeling of virtual anchors and goods before the live broadcast. Among them, modeling and streaming the texture and gloss of clothing requires a large amount of graphics computing resources, increasing the generation cost of virtual anchors.

[0022] Therefore, in order to enable the virtual anchor to display the upper body effect of clothing goods and conduct real-time virtual live broadcast, the embodiment of the present application provides a method for pushing live broadcast goods based on artificial intelligence, as Figure 1 shown, including the following steps: S1. Generate a virtual anchor image of the current frame based on the fusion of a head region image, a clothing region image, and a body shape region image; the clothing region image is obtained by image segmentation of a captured real clothing image, the body shape region image is obtained by image segmentation of a captured real person image, and the head region image is a person image for generating the head of the virtual anchor; It can be understood that the head region image is used to generate the head of the virtual anchor image, the clothing region image is used to generate the clothing of the virtual anchor image, and the body shape region image is used to generate the torso of the virtual anchor image.

[0023] Since the action postures of existing digital humans are relatively rigid, and predicting the posture changes of digital humans through a neural network model requires a large amount of computing power, which cannot meet the real-time requirements of live streaming.

[0024] Therefore, in the embodiments of the present application, by collecting images of real human bodies as the bottom layer, and filling the head region image and the clothing region image above the bottom layer, a virtual anchor image with coordinated limbs can be quickly generated.

[0025] Specifically, the head region image is pre-stored in the system, and the head region image can be a photo of a real person's head or a photo of a digital human's head. When generating a virtual anchor, the head photo in the system is selected as the head region image.

[0026] Specifically, the head region image, the clothing region image, and the body shape region image are spliced through a pre-trained image synthesis model to obtain a virtual anchor image.

[0027] S2. Splice the virtual anchor image of the current frame into the live broadcast background image to obtain a virtual live video frame; S3. Obtain the clothing data to be pushed, and input the clothing data to be pushed into the speech synthesis model to predict and obtain the live broadcast voice data; In the embodiments of the present application, the clothing data to be pushed is the parameters of the clothing. The speech synthesis model is built based on a large language model, generates live broadcast text for e-commerce from the input clothing parameter text, and then predicts the live broadcast voice data through the live broadcast text for e-commerce.

[0028] S4. Align the pushed voice data and the virtual live video frame to generate a live video stream in real time; Specifically, the virtual anchor in the virtual live video frame is driven by the pushed voice data, and the virtual anchor can adjust the head expression according to different voice features. It can be understood that driving a digital human to generate a video stream through a voice signal or a text signal is achieved through known technologies.

[0029] S5. Send the live video stream to the server side for live streaming.

[0030] Specifically, in step S5, the commodity push video stream is frame-classified through an encoder; each video frame is compressed and encapsulated and sent to the server side.

[0031] In the embodiments of the present application, the virtual anchor image is divided into a body area, a clothing area, and a head area. Images of the three areas are obtained respectively, and then the clothing area image and the head area image are spliced above the layer of the body area image to generate a complete virtual anchor image. Finally, the virtual live video frame and the live voice data are spliced into a virtual live video frame, and the live video stream is output in real time on the server side.

[0032] In the embodiments of the present application, when generating the virtual anchor image, the virtual anchor image is quickly generated through the images of the 3 areas, and the quickly generated virtual anchor image is used for live streaming. Embodiment

[0033] The clothing demonstration effect of the existing virtual anchors is not good, and the real wearing effect of the clothing cannot be reflected. The existing virtual try-on technology is based on 3D models, and the workload of modeling each piece of clothing is large.

[0034] Based on the live streaming goods promotion method in Embodiment 1, in order to achieve the clothing demonstration effect of the virtual anchor through a planar image.

[0035] The embodiments of the present application also provide a live streaming goods promotion method based on artificial intelligence, as Figure 1 shown, including the following steps: S1. Generate the virtual anchor image of the current frame based on the fusion of the head area image, the clothing area image, and the body area image; S2. Splice the virtual anchor image of the current frame into the live broadcast background image to obtain a virtual live video frame; S3. Obtain the clothing data to be promoted, input the clothing data to be promoted into the voice synthesis model, and predict to obtain the live voice data; S4. Align the promoted voice data and the virtual live video frame to generate a live video stream in real time; S5. Send the live video stream to the server side for live streaming.

[0036] Further, in step S1, as Figure 2 shown, specifically includes: S101. Real-time collect real person images; the real person in the real person images wears a motion capture clothing, and the surface of the motion capture clothing is provided with N rectangular grids, and each of the rectangular grids is provided with a grid identification code; N is an integer greater than or equal to 2; S102. Perform contour recognition on the real person image of the current frame, and segment to obtain the head area image and the body area image; S103. Obtain the clothing area image; S104. Cover the clothing area image above the layers of each of the rectangular grids according to the grid identification code.

[0037] S105. Obtain the head area image, cover the head area image above the layer of the body shape area image to obtain the virtual anchor image.

[0038] Specifically, in step S101, it is necessary to photograph a real person wearing a motion capture clothing.

[0039] Since N rectangular grids are printed on the surface of the motion capture clothing, each rectangular grid contains a separate grid identification code, and the grid identification codes in each column correspond to the person offset angle in the data table. Therefore, when the real person offsets relative to the camera, by identifying the grid identification code on the body shape area image, the person offset angle of the real person can be determined.

[0040] In step S102, when generating the virtual anchor image, layer filling is performed on the head area image and the body shape area image respectively.

[0041] Furthermore, in step S102, the contour features of the input real person image are extracted and recognized through a pre-trained YOLOv8 model, and the head area image and the body shape area image are segmented along the contour of the real person.

[0042] It can be understood that the YOLOv8 model in the embodiment of the present application is pre-trained based on an existing data set and can predict the human body contour in the image.

[0043] In one implementation manner of step S103, the clothing area image is a panoramic view of the clothing to be pushed. Specifically, a 360-degree photograph is taken around the center line of the clothing to be pushed to obtain the clothing area image.

[0044] In another implementation manner of step S103, the clothing to be pushed is cut open, flattened and then photographed to obtain the clothing area image.

[0045] Specifically, in step S104, as Figure 3 shown, it includes the following steps: S1041. Calculate the area of the rectangular grid of the body shape area image; S1042. Determine the rectangular grid with the largest area as the central grid; S1043. Extract the grid identification code of the central grid; S1044. Look up the corresponding person offset angle according to the grid identification code; S1045. Determine the image regions corresponding to the central grid and the clothing region image according to the person offset angle; S1046. Overlay the clothing region image above the layer of the central grid; S1047. Perform layer overlay along the central grid to the surrounding rectangular grids until all the rectangular grids are filled.

[0046] Further, in step S1047, it further includes: Calculate the affine transformation matrix between the Nth rectangular grid and the initial rectangular grid according to the feature point matching algorithm; Project the clothing region image on the Nth rectangular grid through the affine transformation matrix.

[0047] Through two-dimensional image processing, the virtual anchor generated in the embodiments of the present application can display clothing at different angles. Since 3D modeling is not required, the graphic calculation amount can be reduced, thereby realizing real-time live streaming.

[0048] The solutions of the present application have been described in detail above with reference to the accompanying drawings. In the above embodiments, the descriptions of the various embodiments have their own emphases. For the parts not described in detail in a certain embodiment, reference may be made to the relevant descriptions of other embodiments. Those skilled in the art should also know that the actions and modules involved in the specification are not necessarily essential to the present application.

[0049] In addition, it can be understood that the steps in the method of the embodiments of the present application can be adjusted, combined, and deleted according to actual needs, and the modules in the device of the embodiments of the present application can be combined, divided, and deleted according to actual needs.

[0050] In addition, the method according to the present application can also be implemented as a computer program or a computer program product, which includes computer program code instructions for executing some or all of the above steps of the method of the present application.

[0051] Alternatively, the present application can also be implemented as a non-transitory machine-readable storage medium (or computer-readable storage medium, or machine-readable storage medium), on which executable code (or computer program, or computer instruction code) is stored. When the executable code (or computer program, or computer instruction code) is executed by a processor of an electronic device (or an electronic device, a server, etc.), the processor is caused to execute some or all of the steps of the above method according to the present application.

[0052] Those skilled in the art will also understand that the various exemplary logical blocks, modules, circuits, and algorithm steps described in connection with the applications herein can be implemented as electronic hardware, computer software, or a combination of both.

[0053] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems and methods according to various embodiments of the present application. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a portion of code that contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions noted in the blocks may occur in a different order than that noted in the accompanying drawings. For example, two consecutive blocks may actually be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented by a dedicated hardware-based system that performs the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.

[0054] The embodiments of the present application have been described above. The above description is exemplary, not exhaustive, and is not limited to the disclosed embodiments. Many modifications and variations are obvious to those of ordinary skill in the art in the technical field without departing from the scope and spirit of the described embodiments. The choice of terms used herein is intended to best explain the principles of the embodiments, the practical application, or the improvement of technologies in the market, or to enable other ordinary skill in the art in the technical field to understand the embodiments disclosed herein.

Claims

1. A method for pushing goods through live streaming based on artificial intelligence, characterized in that: include: S1, generating a virtual anchor image of the current frame based on the fusion of the head area image, the clothing area image and the body area image; The clothing area image is obtained by performing image segmentation on a real clothing image taken, the body shape area image is obtained by performing image segmentation on a real human body image taken, and the head area image is a character image for generating a virtual anchor's head; S2, splicing the virtual anchor image of the current frame into the background image of the live broadcast room to obtain a virtual live broadcast video frame; S3, obtaining clothing data to be pushed, inputting the clothing data to be pushed into a speech synthesis model, and predicting live speech data; S4, aligning the pushed voice data with the virtual live video frame to generate a live video stream in real time; S5. Send the live video stream to the server for live streaming.

2. According to the method for pushing goods through live streaming based on artificial intelligence in claim 1, it is characterized in that: Step S1 specifically includes: S101, real-time acquisition of a real person image; the real person in the real person image wears a motion capture suit, the surface of the motion capture suit is provided with N rectangular grids, each of the rectangular grids is provided with a grid identification code; N is an integer greater than or equal to 2; S102, performing contour recognition on the real person image of the current frame, and segmenting to obtain the head area image and the body area image; S103, acquiring the clothing area image; S104, covering the clothing area image on the layer of each rectangular grid according to the grid identification code; S105: Acquire the head region image, and overlay the head region image on the layer of the body region image to obtain the virtual anchor image.

3. According to the method for pushing goods through live streaming based on artificial intelligence in claim 2, it is characterized in that: In step S104, it specifically includes: S1041, calculating the rectangular grid area of ​​the body shape region image; S1042, determining the rectangular grid with the largest area as the central grid; S1043, extracting the grid identification code of the central grid; S1044, searching for a corresponding character offset angle according to the grid identification code; S1045, determining the image area corresponding to the central grid and the clothing area image according to the character offset angle; S1046, overlaying the clothing area image on the layer of the central grid; S1047, performing layer covering along the central grid to the surrounding rectangular grids until all the rectangular grids are filled.

4. According to the method for pushing goods through live streaming based on artificial intelligence in claim 2, it is characterized in that: In step S102, it specifically includes: The input real person image is subjected to contour feature extraction and recognition through a pre-trained YOLOv8 model, and the head area image and the body area image are segmented along the contour of the real person.

5. According to the method of pushing goods through live streaming based on artificial intelligence in claim 2, it is characterized in that: Step S103 specifically includes: A 360-degree shooting is performed around the center line of the garment to be pushed to obtain the garment area image, which is a panoramic view of the garment to be pushed.

6. A method for pushing goods through live streaming based on artificial intelligence according to claim 2, characterized in that: Step S103 specifically includes: The garments to be pushed are cut open, flattened and photographed to obtain the garment area image.

7. A method for pushing goods through live streaming based on artificial intelligence according to claim 1, characterized in that: Step S3 specifically includes: The speech synthesis model is built based on a large language model, which generates live streaming text based on the input clothing parameter text, and then predicts the live streaming voice data based on the live streaming text.

8. The method for pushing goods through live streaming based on artificial intelligence according to claim 1, characterized in that: Step S5 specifically includes: Performing frame classification on the product push video stream by using an encoder; Each video frame is compressed and packaged and sent to the server.

Citation Information

Patent Citations

  • Body scanning and movement capturing method based on clothes feature points

    CN104766345A

  • Virtual clothing try-on method and device, terminal equipment and storage medium

    CN111508079A

  • Virtual try-on method and device based on artificial intelligence, server and storage medium

    CN111784845A

  • Augmented reality fitting method

    CN112348647A

  • Virtual anchor processing method and device, computing equipment and storage medium

    CN117315102A