Method and system for displaying products in live broadcast mode
By combining large language model and zero-sample object detection technology, the control robotic arm and rotary device conduct product display in three-dimensional space, solving the problem of single display effect in the existing technology, and realizing multi-angle and multi-dimensional automated display.
Patent Information
- Application Number
- CN202510867600.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-26
- Publication Date
- 2025-08-15
AI Technical Summary
The existing live product display solutions have problems such as a single perspective, subjective dependence, or single plane rotation, and it is impossible to achieve multi-dimensional and multi-directional three-dimensional display.
The target requirements are obtained through the barrage or voice acquisition module on the server side, combined with the large language model and the zero-sample object detection model, the control robotic arm and automatic rotation device are displayed in three-dimensional space to realize automatic positioning of any angle and area.
Multi-angle display in three-dimensional space is realized, and the audience or anchor can dominate the display area. The system has high scalability and positioning accuracy, and is suitable for various product categories.
Smart Images

Figure CN120499409A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of live broadcast display technology, and in particular to a method and system for live broadcast display of products. Background Art
[0002] With the rapid development of internet live streaming platforms and the continuous expansion of their user base, the marketing model of selling products through live streaming is becoming increasingly popular. In this marketing model, the sales host needs to use cameras to display the products for sale to the audience watching the live broadcast to improve sales conversion rate. In the existing technology, live streaming product display solutions mainly include:
[0003] 1. Place the product in front of the camera for fixed display;
[0004] 2. The sales host holds the product and adjusts the display angle relative to the camera.
[0005] 3. Place the product on a rotating display device and display it using a camera.
[0006] However, all of the above display solutions have obvious disadvantages:
[0007] Regarding the above solution 1, the static display solution can only present a single fixed perspective of the product;
[0008] Regarding the above solution 2, although the handheld display solution can achieve multi-angle display, it is completely dependent on the subjective operation of the sales anchor. The display effect lacks objectivity and the audience cannot control the angle of the displayed product.
[0009] Regarding the above-mentioned solution 3, although the rotating display solution can be operated automatically, its mechanical structure only supports rotational movement in a single plane, and cannot achieve multi-dimensional and multi-directional three-dimensional display, which seriously restricts the product display effect. Summary of the Invention
[0010] The present invention aims to provide a method for live display of products, comprising the following steps:
[0011] S1. Obtain target requirements through the bullet comment acquisition module or voice acquisition module on the server side;
[0012] S2. Based on the target requirements, obtain preset prompt words from the Prompt module, and combine the target requirements and the preset prompt words into prompt words to be retrieved; call the large language model located on the remote server, and input the prompt words to be retrieved into the large language model to obtain a list of target objects to be detected;
[0013] S3. Using a controller module on the server side, control a robotic arm and an automatic rotating device on the product shooting side, causing the robotic arm to enter a preset level view posture and the automatic rotating device to start rotating; the robotic arm holds a camera device, and the automatic rotating device has a product placed on it;
[0014] S4. Use a camera to photograph the rotating product and send the photographed image to a zero-shot object detection model for detection to determine whether there is an image that matches the first object in the target object list and the target object is facing the camera. If so, jump to step S6; otherwise, go to step S5.
[0015] S5. Determine whether the automatic rotating device has completed one circle. If so, change the preset posture of the robotic arm to a downward-looking posture and proceed to step S6). Otherwise, return to step S4.
[0016] S6. Control the automatic rotating device to stop rotating, upload the current frame captured by the camera device and the prompt word to the multimodal large language model, and determine whether the position of the product in the current frame is correct based on the prompt word;
[0017] If correct, delete the object from the target object list and proceed to step S7);
[0018] Otherwise, return a message indicating that the current target positioning has failed, and proceed to step S7);
[0019] S7. Determine whether the target object still exists in the target object list. If so, return to step S3); otherwise, return to step S1.
[0020] Furthermore, in step S1), when the target requirement is obtained through the bullet screen acquisition module, step S2) can be directly entered.
[0021] When the target requirement is obtained through the voice acquisition module, the host's voice input needs to be converted into text through the voice-to-text module, and then step S2 is entered).
[0022] Furthermore, in step S3), the server communicates remotely with the product shooting end via a Bluetooth protocol module.
[0023] Furthermore, in step S4), the method for determining an image using a zero-sample object detection model comprises the following steps:
[0024] S4.1. Using a zero-shot object detection model, continuously search for the target object in the image stream and obtain the object bounding box. When the target object is detected in the bounding box, proceed to step S4.2.
[0025] S4.2. Calculate the confidence of each bounding box and select the bounding box with the highest confidence as the current bounding box area A i, then record the area of the bounding box of the previous frame of the current bounding box as A i-1 , calculate the absolute value of the difference in area of the target bounding box at adjacent moments ΔA=|A i -A i-1 |;
[0026] Determine whether ΔA<τ1 holds. If so, proceed to step S4.3). Otherwise, set i=1 and return to step S4.1). The initial value of i is 1.
[0027] S4.3. Determine whether i = n1. If so, set i = 1 and proceed to S4.4) to start tracking. Otherwise, set i = i + 1 and return to step S4.1) to continue searching.
[0028] S4.4. Reduce the speed of the automatic rotation device;
[0029] S4.5. Calculate the absolute value of the difference in the area of the target bounding box at adjacent moments and determine whether ΔA > τ3. If so, set i = 1 and return to S4.1. Otherwise, proceed to step S4.6.
[0030] S4.6. Determine whether ΔA < τ2. If so, proceed to step S4.7. Otherwise, set i = 1 and return to step S4.5.
[0031] S4.7. Determine whether i=n2. If so, set i=1 and proceed to S4.8) to start the confirmation state. Otherwise, set i=i+1 and return to step S4.5).
[0032] S4.8. Calculate the absolute value of the difference in the area of the target bounding box at adjacent moments and determine whether ΔA < τ4 holds. If so, proceed to S4.9. Otherwise, set i = 1 and return to step S4.5.
[0033] S4.9. Determine whether i=n3. If so, the target position is preliminarily located and the automatic rotation device is controlled to stop rotating. Otherwise, set i=i+1 and return to step S4.8).
[0034] Another object of the present invention is to provide a system for implementing a live broadcast product display method, including a product shooting end and a service end.
[0035] The product shooting end includes a main control, a mechanical arm, a camera device and an automatic rotation device.
[0036] The main control is a circuit device based on an embedded chip that can burn programs.
[0037] The mechanical arm is used to clamp the camera device and control the shooting angle of the camera device.
[0038] The camera device is a device with camera function and network streaming function.
[0039] The automatic rotating device includes a motor and a rotating platform. The motor is located below the rotating platform and is used to drive the rotating platform to rotate.
[0040] The server includes a user interface module, a function module, a device management module, a controller module and a Bluetooth protocol module.
[0041] The user interface module is used to display system status data in real time.
[0042] The functional modules include a bullet screen acquisition module, a voice acquisition module, a voice-to-text module, a zero-sample target detection module and a large model calling module.
[0043] The barrage acquisition module is used to obtain barrage data from the live broadcast platform server in real time.
[0044] The voice acquisition module is used to obtain the host's voice input.
[0045] The speech-to-text module is used to convert the acquired speech into text.
[0046] The zero-shot target detection module is used to call the zero-shot target detection model to perform zero-shot target detection on the target object.
[0047] The large model calling module is used to remotely call a large language model or a multimodal large language model.
[0048] The device management module is used to manage the status of the underlying devices of the system and provide system status data to the user interface module.
[0049] The controller module is used to remotely control the product shooting end.
[0050] The Bluetooth protocol module is used for remote communication with the product shooting end.
[0051] Furthermore, the server is a device that can run Windows, Linux, or MacOS operating systems.
[0052] Furthermore, the system status data includes the connection status, operation status, operation angle of the robot arm, the connection status of the camera device, and the connection status, operation status, operation speed, and operation direction of the automatic rotation device.
[0053] Furthermore, the robotic arm and the automatic rotation device communicate using a CAN bus.
[0054] Furthermore, the live broadcast platform is a live broadcast platform that can provide user nicknames, user barrages, and gift information interfaces.
[0055] The technical effects of the present invention are undoubted, and the beneficial effects of the present invention are as follows:
[0056] 1. The present invention can display any area of the product in three dimensions, namely, horizontal axis, vertical axis and vertical axis, in three-dimensional space.
[0057] 2. This invention combines the latest speech-to-text technology, zero-sample target detection technology and large language models to locate the detection target in the middle area of the camera image. The system can identify and locate targets of any category to adapt to a variety of different products.
[0058] 3. Compared with the above-mentioned existing solution 2 in which only the host can control the angle of the displayed product, the present invention allows the audience to control the display area of the displayed product through barrage or the host to control the display area of the displayed product through voice.
[0059] 4. Compared to the existing solution 3, which only has a single display function, this invention can flexibly expand its functions and has greater scalability. For example: 1) Because the robotic arm and the turntable communicate using the CAN bus, multiple turntables can be added, eliminating the need for frequent manual replacement of products on the turntable. BRIEF DESCRIPTION OF THE DRAWINGS
[0060] Figure 1 This is a schematic diagram of the system structure according to Example 10 of the present invention;
[0061] Figure 2 This is a structural diagram of a product shooting terminal according to embodiment 10 of the present invention;
[0062] Figure 3 This is a schematic diagram of the server program structure according to Example 10 of the present invention;
[0063] Figure 4 A flowchart of detecting and locating a target by a server according to embodiment 10 of the present invention;
[0064] Figure 5 This is a flow chart of the server confirming a preliminary positioning target according to embodiment 10 of the present invention. DETAILED DESCRIPTION
[0065] The present invention will be further described below with reference to the following examples, but it should not be understood that the scope of the present invention is limited to the following examples. Without departing from the above technical ideas of the present invention, various substitutions and modifications can be made according to common technical knowledge and customary means in the art, and all should be included in the scope of protection of the present invention.
[0066] Example 1:
[0067] A method for live broadcasting of a product, comprising the following steps:
[0068] S1. Obtain target requirements through the bullet screen acquisition module 312 or the voice acquisition module 313 on the server side;
[0069] S2. Based on the target requirements, obtain preset prompt words from the Prompt module, and combine the target requirements and the preset prompt words into prompt words to be retrieved; call the large language model located on the remote server, and input the prompt words to be retrieved into the large language model to obtain a list of target objects to be detected;
[0070] S3. Using the controller module 331 on the server side, the robotic arm 202 and the automatic rotating device 204 on the product shooting side are controlled, so that the robotic arm 202 enters a preset level view posture and the automatic rotating device 204 starts to rotate; the robotic arm 202 holds the camera 203, and the automatic rotating device 204 has the product placed on it;
[0071] S4. Use the camera 203 to shoot the rotating product, and send the shot image to the zero-sample target detection model for detection to determine whether there is an image that matches the first object in the target object list, and the target object is facing the camera (203). If so, jump to step S6), otherwise, go to step S5);
[0072] S5. Determine whether the automatic rotating device 204 has completed one rotation. If so, change the preset posture of the robotic arm 202 to a top-down posture and proceed to step S6). Otherwise, return to step S4.
[0073] S6. Control the automatic rotation device 204 to stop rotating, upload the current frame captured by the camera device 203 and the prompt word to the multimodal large language model, and determine whether the position of the product in the current frame is correct based on the prompt word;
[0074] If correct, delete the object from the target object list and proceed to step S7);
[0075] Otherwise, return a message indicating that the current target positioning has failed, and proceed to step S7);
[0076] S7. Determine whether the target object still exists in the target object list. If so, return to step S3); otherwise, return to step S1.
[0077] Example 2:
[0078] The main structure of this embodiment is the same as that of embodiment 1. Furthermore, in step S1), when the target requirement is obtained by the bullet screen obtaining module 312, step S2) can be directly entered.
[0079] When the target requirement is obtained through the voice acquisition module 313, the host's voice input needs to be converted into text through the voice-to-text module 314, and then the process goes to step S2).
[0080] Example 3:
[0081] The main structure of this embodiment is the same as any one of Embodiments 1 to 2. Furthermore, in step S3), the server communicates remotely with the product shooting end via the Bluetooth protocol module 341.
[0082] Example 4:
[0083] The main structure of this embodiment is the same as any one of Embodiments 1 to 3. Furthermore, in step S4), the method for determining an image using a zero-sample object detection model includes the following steps:
[0084] S4.1. Using a zero-shot object detection model, continuously search for the target object in the image stream and obtain the object bounding box. When the target object is detected in the bounding box, proceed to step S4.2.
[0085] S4.2. Calculate the confidence of each bounding box and select the bounding box with the highest confidence as the current bounding box area A i , then record the area of the bounding box of the previous frame of the current bounding box as A i-1 , calculate the absolute value of the difference in area of the target bounding box at adjacent moments ΔA=|A i -A i-1 |;
[0086] Determine whether ΔA<τ1 holds. If so, proceed to step S4.3). Otherwise, set i=1 and return to step S4.1). The initial value of i is 1.
[0087] S4.3. Determine whether i = n1. If so, set i = 1 and proceed to S4.4) to start tracking. Otherwise, set i = i + 1 and return to step S4.1) to continue searching.
[0088] S4.4. Reduce the rotation speed of the automatic rotation device 204;
[0089] S4.5. Calculate the absolute value of the difference in the area of the target bounding box at adjacent moments and determine whether ΔA > τ3. If so, set i = 1 and return to S4.1. Otherwise, proceed to step S4.6.
[0090] S4.6. Determine whether ΔA < τ2. If so, proceed to step S4.7. Otherwise, set i = 1 and return to step S4.5.
[0091] S4.7. Determine whether i=n2. If so, set i=1 and proceed to S4.8) to start the confirmation state. Otherwise, set i=i+1 and return to step S4.5).
[0092] S4.8. Calculate the absolute value of the difference in the area of the target bounding box at adjacent moments and determine whether ΔA < τ4 holds. If so, proceed to S4.9. Otherwise, set i = 1 and return to step S4.5.
[0093] S4.9. Determine whether i=n3. If so, the target position is preliminarily located and the automatic rotation device 204 is controlled to stop rotating. Otherwise, set i=i+1 and return to step S4.8).
[0094] Example 5:
[0095] The main structure of this embodiment is the same as any one of embodiments 1 to 4. Furthermore, the system for applying the method includes a product shooting end and a service end. The product shooting end includes a main control 201, a mechanical arm 202, a camera device 203 and an automatic rotation device 204.
[0096] The main control 201 is a circuit device based on an embedded chip that can burn programs.
[0097] The robotic arm 202 is used to clamp the camera device 203 and control the shooting angle of the camera device 203 .
[0098] The camera device 203 is a device with camera functions and network streaming functions.
[0099] The automatic rotating device 204 includes a motor and a rotating platform. The motor is located below the rotating platform and is used to drive the rotating platform to rotate.
[0100] The server includes a user interface module 301 , a function module 311 , a device management module 321 , a controller module 331 and a Bluetooth protocol module 341 .
[0101] The user interface module 301 is used to display system status data in real time.
[0102] The functional module 311 includes a bullet comment acquisition module 312 , a voice acquisition module 313 , a voice-to-text module 314 , a prompt module, a zero-sample target detection module 315 and a large model calling module 316 .
[0103] The barrage acquisition module 312 is used to obtain barrage data from the live broadcast platform server in real time.
[0104] The voice acquisition module 313 is used to acquire the host's voice input.
[0105] The speech-to-text module 314 is used to convert the acquired speech into text.
[0106] The zero-shot target detection module 315 is used to call a zero-shot target detection model to perform zero-shot target detection on the target object.
[0107] The large model calling module 316 is used to remotely call a large language model or a multimodal large language model.
[0108] The device management module 321 is used to manage the status of the underlying devices of the system and provide system status data to the user interface module.
[0109] The controller module 331 is used to remotely control the product shooting end.
[0110] The Bluetooth protocol module 341 is used for remote communication with the product shooting end.
[0111] Example 6:
[0112] The main structure of this embodiment is the same as any one of Embodiments 1 to 5. Furthermore, the server is a device that can run Windows, Linux, or MacOS operating systems.
[0113] Example 7:
[0114] The main structure of this embodiment is the same as any one of Embodiments 5 to 6. Furthermore, the system status data includes the connection status, operation status, operation angle of the robotic arm 202, the connection status of the camera device 203, and the connection status, operation status, operation speed, and operation direction of the automatic rotation device 204.
[0115] Example 8:
[0116] The main structure of this embodiment is the same as any one of Embodiments 1 to 7. Furthermore, the robotic arm 202 and the automatic rotation device 204 communicate using a CAN bus.
[0117] Example 9:
[0118] The main structure of this embodiment is the same as any one of Embodiments 5 to 8. Furthermore, the live broadcast platform is a live broadcast platform that can provide user nicknames, user barrages, and gift information interfaces.
[0119] Example 10:
[0120] The main structure of this embodiment is the same as any one of Embodiments 1 to 9. Furthermore, the present invention provides a method and device for live broadcasting of product display, which expands upon the existing live broadcasting product display solution and adds a robotic arm and a service end to the automatic rotating device.
[0121] By adding a robotic arm, the horizontal axis of the product shooting angle is controlled by an automatic rotation device, and the vertical and vertical axes of the product shooting angle are controlled by the robotic arm, so that any area of the product can be displayed in three-dimensional space.
[0122] Adding a server can communicate with the product shooting terminal wirelessly, and the server controls the automatic operation of the product shooting terminal and controls any shooting area of the displayed product in the three-dimensional space.
[0123] Furthermore, the present invention integrates a speech-to-text model, a zero-sample target detection model, and a large language model remote API on the server side, which can perform zero-sample target detection on the target product and move the detection target to the front of the camera.
[0124] Ultimately, users can display products through barrage or anchors can control the system through voice.
[0125] See also Figure 1 , is a schematic diagram of the system structure of the present invention:
[0126] The system includes: a live broadcast platform, a service end, and a product shooting end.
[0127] The live broadcast platform is any live broadcast platform that can provide user nicknames, user barrages, and gift information interfaces.
[0128] The server is a device that can run Windows, Linux, or MacOS operating systems.
[0129] See also Figure 2 , is a structural diagram of the shooting end of the product of the present invention:
[0130] The product shooting end includes: a main control 201, a mechanical arm 202, a camera device 203, and an automatic rotation device 204.
[0131] The main control 201 is any circuit device based on an embedded chip that can burn a program.
[0132] The robot arm 202 is any general-purpose robot arm device.
[0133] A camera device is a device with camera and network streaming functions.
[0134] The automatic rotating device 204 is composed of a motor and a rotating platform. The motor rotates under the rotating platform to drive the rotating platform to rotate.
[0135] See also Figure 3 , is a schematic diagram of the server program structure of the present invention:
[0136] The server program includes:
[0137] User interface module 301, for displaying system status data in real time;
[0138] Functional module 311 includes:
[0139] The bullet screen acquisition module 312 (connected to the bullet screen API of the live broadcast platform) is used to obtain bullet screen data from the live broadcast platform server in real time;
[0140] The voice acquisition module 313 (i.e., the system microphone) is used to obtain the host's voice input;
[0141] Speech-to-text module 314 (using the SenseVoice model), used to convert speech input into text;
[0142] A zero-shot object detection module 315 (using the OWL v2 model) is used to perform zero-shot object detection on the target object;
[0143] Large model calling module 316 (connecting to the large model platform API) is used to remotely call the large language model;
[0144] The device management module 321 is used to manage the status of the underlying devices of the system and provide system status data to the user interface module;
[0145] Controller module 331, used to remotely control the product shooting end;
[0146] The Bluetooth protocol module 341 is used for remote communication with the product shooting end.
[0147] See also Figure 4 , is a flow chart of the present invention's server detecting and locating a target, comprising the following steps:
[0148] Step 401: The system obtains input.
[0149] Step 402 , determining whether the system input is voice or text, if it is text, proceed to step 404 , if not text, proceed to step 403 .
[0150] Step 403 , convert the input voice into text and proceed to step 404 .
[0151] Step 404: Fill the text into the prompt word.
[0152] Step 405: Send the prompt word to the large language model and wait for a response.
[0153] Step 406: After receiving the large language model, wait for a response and output a list of target objects to be detected.
[0154] Step 407: The robotic arm enters a preset detection posture and the rotating device starts rotating.
[0155] In step 408 , the zero-shot object detection model starts detecting the first object in the target object list.
[0156] Step 409 , based on the target preliminary positioning algorithm of 501 - 514 , determine whether the object is preliminarily positioned, if yes, proceed to step 410 , if not, proceed to step 413 .
[0157] Step 410: Send the current camera frame and the prompt word to the multimodal large language model and wait for a response.
[0158] Step 411: Receive the output of the multimodal large language model and determine whether the position is correct. If the position is correct, proceed to step 412; if the position is incorrect, proceed to step 408.
[0159] Step 412, determine whether the target object list is empty. If it is empty, the process ends; if it is not empty, proceed to step 408.
[0160] Step 413 , determining whether the turntable has completed one rotation, if so, proceed to step 414 , otherwise proceed to step 408 .
[0161] Step 414: Set the robotic arm to a preset downward-looking posture.
[0162] Step 415: Send the current camera frame and the prompt word to the multimodal large language model and wait for a response.
[0163] Step 416: Receive the output of the multimodal large language model and determine whether the position is correct. If the position is correct, proceed to step 418; if the position is incorrect, proceed to step 417.
[0164] Step 417 , returns a message indicating that the current target positioning has failed, and proceeds to step 418 .
[0165] Step 418, determine whether the target object list is empty. If it is empty, the process ends; if it is not empty, proceed to step 408.
[0166] See also Figure 5 , is a flow chart of the present invention for confirming a preliminary positioning target by the server, including the following steps:
[0167] In step 501 , the system is in a search state and continuously detects a target object until the target object is detected and enters step 502 .
[0168] Step 502: Select the bounding box with the highest confidence.
[0169] Step 503: Calculate the absolute value of the difference between the bounding box area of this time and the previous time ΔA=|A t -A t-1 |.
[0170] Step 504 : If the absolute value of the area difference ΔA is less than τ1 , proceed to step 502 ; otherwise, return to step 501 .
[0171] Step 505: If the condition of step 504 is met for n1 consecutive times, proceed to step 506; otherwise, return to step 501.
[0172] Step 506: Confirm it as a suspected target and enter the target tracking state.
[0173] Step 507: Calculate the absolute value of the difference between the bounding box area of this time and the last time ΔA=|A t -A t-1 |.
[0174] Step 508 : If the absolute value of the area difference ΔA>τ3, proceed to step 509 ; otherwise, return to step 501 .
[0175] Step 509 : If the area difference ΔA is less than τ 2 , then go to step 510 ; otherwise, go back to step 506 .
[0176] Step 510: If the condition of step 509 is met for n2 consecutive times, proceed to step 511; otherwise, return to step 506.
[0177] Step 511: Confirm the target as initially located and enter the confirmation state. (Make sure the target is in the center of the camera image)
[0178] Step 512: Calculate the absolute value of the difference between the bounding box area of this time and the last time ΔA=|A t -A t-1 |.
[0179] Step 513: If the absolute value of the area difference ΔA < τ4, proceed to step 514; otherwise, return to step 511.
[0180] Example 11:
[0181] The main structure of this embodiment is the same as any one of Embodiments 1 to 10. Furthermore, this system adopts a layered design concept and is divided into four functional levels from top to bottom to form a complete intelligent interactive display control system.
[0182] The top layer is the cloud-based intelligent service layer, which is responsible for providing powerful natural language understanding and processing capabilities, exchanging data with lower-level systems through the network, and providing advanced intelligent support for the entire system.
[0183] The second layer is the upper computer application layer, serving as the system's core processing unit and integrating multiple functional modules. The master control unit serves as the central coordinator, connecting and scheduling the work of each module. The Prompt module generates optimized prompts and interacts with cloud services via the LLM API to obtain semantic understanding results. The Danmu acquisition module is responsible for receiving Danmu messages, while the Speech-to-Text module is responsible for speech recognition and converting user speech into text. The zero-shot object detection model is responsible for detecting and localizing target objects. These modules work together to implement the entire process from user input to intelligent processing.
[0184] The third layer is the embedded control layer, responsible for hardware control and communication. The main controller maintains a Bluetooth connection with the host computer to receive control commands. It also communicates with the camera via the UART protocol, forwarding commands from the host computer in real time. The system uses the CAN bus to control the stepper motor driver, ensuring precise mechanical motion control.
[0185] The lowest layer is the physical execution layer, encompassing both the mechanical and physical components of the system. The robotic arm adjusts the vertical axis's angle, while the stepper motor precisely rotates under the control of the main controller to adjust the horizontal axis's angle. This layer ultimately translates the software system's intelligent decisions into actual actions in the physical world.
[0186] The four layers are tightly connected via standardized interfaces. Data flows from top to bottom, forming instructions, and then feeds back execution results and device status from bottom to top, forming a closed-loop feedback control system. This layered architecture design makes the system modular, highly scalable, and easy to maintain. Each layer can be independently upgraded and optimized while maintaining the stability and consistency of the overall system functionality.
[0187] Example 12:
[0188] The main structure of this embodiment is the same as any one of Embodiments 1 to 11. Furthermore, the operation process of this system is mainly divided into two stages: input processing and target positioning, forming a complete interactive control closed loop.
[0189] After the system boots up, it first enters the input acquisition phase. Here, the system receives user input and processes it differently depending on the input type: if text input (such as a bullet comment) is detected, it proceeds directly to subsequent processing; if voice input is detected, the SenseVoice module is called to perform speech recognition and convert it into text. Regardless of the method used, all input is ultimately converted to a standardized text format.
[0190] After receiving the text, the system enters the prompt processing phase, which structures and optimizes the raw input, extracts key information, and constructs standardized prompt words. These optimized prompt words are then fed into the Large Language Model (LLM) for deep semantic understanding and intent analysis. The LLM outputs processed detection target descriptions, providing more accurate search targets for the zero-shot object detection model.
[0191] At the same time, the system control unit enters a preset posture and prepares to start searching for the target. The OWL-ViT v2 zero-shot object detection model receives the optimized description from the LLM and begins searching for matching target objects in the current field of view.
[0192] If the initial positioning is unsuccessful, the system will continue to search using OWL-ViT v2. Once the target location is preliminarily determined, the system will capture the current frame of the camera and further call the multimodal large language model to confirm the location through customized prompt words. This step provides an additional verification mechanism, greatly improving the accuracy of the system.
[0193] If the VLM confirms the location is correct, the system completes the target location task and ends the current process. If the location confirmation fails, it returns to OWL-ViT v2 to continue searching for a more suitable target. This dual verification mechanism ensures extremely high positioning accuracy and effectively avoids misidentification issues.
[0194] The entire operational process forms a complete intelligent closed loop, from user input to understanding and processing to visual search. Each module works closely together to form an efficient and accurate interactive object positioning system. The system's dual-branch design ensures flexible input processing and accurate object positioning, fully demonstrating the system's intelligent nature.
[0195] Example 13:
[0196] The main structure of this embodiment is the same as any one of Embodiments 1 to 12. Furthermore, this system adopts a hierarchical module design to realize a complete functional link from user interaction to device control.
[0197] 1.1 System Architecture Overview
[0198] The host computer software adopts a five-layer architecture: user interface layer, functional module layer, device management layer, controller layer, and communication protocol layer. This layered design clarifies the responsibilities of each system component and standardizes interfaces, significantly improving software maintainability and scalability.
[0199] The system data flow follows the processing chain of "user input → semantic understanding → target detection → device control → state feedback". Each layer interacts through standardized interfaces to form a closed-loop information processing process.
[0200] 1.2 User Interface and Functional Modules
[0201] The user interface layer includes a UI updater and three main controllers (overview page, settings page, and mode controller), providing users with an intuitive operation experience and system status feedback.
[0202] The functional module layer, serving as the core of the system's intelligent processing, includes a bullet comment acquisition module, a speech acquisition module, a speech-to-text module, a zero-shot object detection module, and a large model call module. Together, these modules support the entire processing process, from natural language input to visual object localization.
[0203] 1.3 Equipment Management and Control Implementation
[0204] The device management layer coordinates data exchange between upper and lower layers, maintains device status information, and implements a unified device access interface. The controller layer connects software logic and hardware operations, including camera controllers, OBS controllers, joint controllers, and Bluetooth controllers, ensuring the accuracy and reliability of device operations.
[0205] Example 14:
[0206] The main structure of this embodiment is the same as any of Examples 1-13. Furthermore, the core target localization algorithm of this system utilizes a state machine-based adaptive detection framework, combined with a zero-shot target detection model and a multimodal verification mechanism, to achieve high-precision and robust target localization. The overall algorithm architecture includes four main steps: preliminary detection, state tracking, stability analysis, and multimodal confirmation, forming a complete target detection and verification closed loop.
[0207] During the initial detection phase, the system uses OWL-ViT v2 as a zero-shot object detection model, processing the real-time image stream captured by the camera and performing zero-shot object detection based on natural language descriptions. Unlike traditional methods, this model does not require prior training on specific categories and can directly understand text descriptions and locate corresponding objects, significantly improving the system's flexibility and applicability.
[0208] 1.1 Three-stage state machine design
[0209] The core of the target location algorithm is a three-stage state machine design, including search, tracking and confirmation states:
[0210] Searching state (SEARCHING): The system actively controls the rotating stage to perform a circumferential scan to find potential matching targets. State transition condition: When ΔA<τ1 and it is satisfied for n1 consecutive times, it switches to the tracking state, where ΔA=|A i -A i-1 | represents the difference in the target bounding box area at adjacent moments, τ1 = 400, n1 = 3.
[0211] Tracking: The system slows down and observes the target more closely. State transition conditions: When ΔA < τ2 and this condition is met for n2 consecutive times, the system transitions to the Confirmation state. When ΔA > τ3, the system reverts to the Search state, where τ2 = 400, n2 = 3, and τ3 = 1000.
[0212] CONFIRMING: The system performs final position confirmation. State transition condition: When ΔA < τ4 and is met for n3 consecutive times, target confirmation is triggered, where τ4 = 300 and n3 = 3.
[0213] This progressive state transition design significantly improves the stability and accuracy of the detection process, effectively avoiding positioning errors caused by noise or short-term false detections.
[0214] To ensure that the positioning accuracy meets the requirements, the system implements an innovative two-layer verification and confirmation mechanism:
[0215] Internal verification of the state machine: Primary verification is achieved through area stability analysis, which can be expressed as:
[0216]
[0217] Multimodal LLM secondary confirmation: Use a multimodal large language model for advanced semantic verification, the expression is V2 = M(I t ,D), where I t is the current image frame, D is the target text description, M is the multimodal model function, and the output value is a binary result, that is, confirmation or error. Only when V1 = 1 and V2 = confirmation, the system finally confirms the target location.
[0218] Through this carefully designed two-layer verification mechanism, the system achieves extremely high target positioning accuracy in complex environments, meeting the high-precision requirements of product display.
Claims
1. A method for live broadcasting of products, characterized in that: The following steps are involved: S1. Obtaining target requirements through the bullet screen acquisition module (312) or the voice acquisition module (313) of the server; S2. Based on the target requirement, obtain the preset prompt word from the Prompt module, and combine the target requirement and the preset prompt word into the prompt word to be retrieved; Call the large language model located on the remote server and input the prompt words to be searched into the large language model to obtain a list of target objects to be detected; S3, using the controller module (331) on the service end to control the mechanical arm (202) and the automatic rotation device (204) on the product shooting end, so that the mechanical arm (202) enters a preset horizontal posture and the automatic rotation device (204) starts to rotate; the mechanical arm (202) clamps the camera device (203), and the product is placed on the automatic rotation device (204); S4, using the camera device (203) to shoot the rotating product, and sending the shot image to the zero-sample target detection model for detection, to determine whether there is an image that matches the first object in the target object list, and the target object is facing the camera device (203), if so, jump to step S6), otherwise, go to step S5); S5, determining whether the automatic rotating device (204) has completed one rotation; if so, changing the preset posture of the robotic arm (202) to a downward-looking posture and proceeding to step S6); otherwise, returning to step S4); S6, controlling the automatic rotating device (204) to stop rotating, uploading the current frame and the prompt word captured by the camera device (203) to the multimodal large language model, and judging whether the position of the product in the current frame is correct based on the prompt word; If correct, delete the object from the target object list and proceed to step S7); Otherwise, return a message indicating that the current target positioning has failed, and proceed to step S7); S7. Determine whether the target object still exists in the target object list. If so, return to step S3); otherwise, return to step S1.
2. A method for live broadcasting of products according to claim 1, characterized in that: In step S1), when the target requirement is obtained through the bullet screen acquisition module (312), step S2) can be directly entered; When the target requirement is obtained through the voice acquisition module (313), the host's voice input is converted into text through the voice-to-text module (314), and then the process goes to step S2).
3. The method for live broadcasting and displaying products according to claim 1, characterized in that: In step S3), the server communicates remotely with the product shooting end via the Bluetooth protocol module (341).
4. The method for live broadcasting and displaying products according to claim 1, characterized in that: In step S4), the method for determining an image using a zero-shot object detection model comprises the following steps: S4.
1. Using a zero-shot object detection model, continuously search for the target object in the image stream and obtain the object bounding box. When the target object is detected in the bounding box, proceed to step S4.
2. S4.
2. Calculate the confidence of each bounding box and select the bounding box with the highest confidence as the current bounding box area A i , then record the area of the bounding box of the previous frame of the current bounding box as A i-1 , calculate the absolute value of the difference in area of the target bounding box at adjacent moments ΔA=|A i -A i-1 |; Determine whether ΔA<τ1 holds. If so, proceed to step S4.3). Otherwise, set i=1 and return to step S4.1). The initial value of i is 1. S4.
3. Determine whether i = n1. If so, set i = 1 and proceed to S4.4) to start tracking. Otherwise, set i = i + 1 and return to step S4.1) to continue searching. S4.
4. Reduce the rotation speed of the automatic rotation device (204); S4.
5. Calculate the absolute value of the difference in the area of the target bounding box at adjacent moments and determine whether ΔA > τ3. If so, set i = 1 and return to S4.
1. Otherwise, proceed to step S4.
6. S4.
6. Determine whether ΔA < τ2. If so, proceed to step S4.
7. Otherwise, set i = 1 and return to step S4.
5. S4.
7. Determine whether i=n2. If so, set i=1 and proceed to S4.8) to start the confirmation state. Otherwise, set i=i+1 and return to step S4.5). S4.
8. Calculate the absolute value of the difference in the area of the target bounding box at adjacent moments and determine whether ΔA < τ4 holds. If so, proceed to S4.
9. Otherwise, set i = 1 and return to step S4.
5. S4.9, determine whether i=n3. If so, the target position is initially located and the automatic rotation device (204) is controlled to stop rotating. Otherwise, set i=i+1 and return to step S4.8).
5. A system according to any one of claims 1 to 4, wherein: Including product shooting end and service end; The product shooting end includes a main control (201), a mechanical arm (202), a camera device (203) and an automatic rotation device (204); The main control (201) is a circuit device based on an embedded chip capable of burning a program; The mechanical arm (202) is used to clamp the camera device (203) and control the shooting angle of the camera device (203); The camera device (203) is a device with camera function and network streaming function; The automatic rotating device (204) comprises a motor and a rotating platform, wherein the motor is located below the rotating platform and is used to drive the rotating platform to rotate; The server includes a user interface module (301), a function module (311), a device management module (321), a controller module (331) and a Bluetooth protocol module (341); The user interface module (301) is used to display system status data in real time; The functional module (311) includes a bullet screen acquisition module (312), a voice acquisition module (313), a voice-to-text module (314), a zero-sample target detection module (315), and the large model calling module (316); The bullet screen acquisition module (312) is used to obtain bullet screen data from the live broadcast platform server in real time; The voice acquisition module (313) is used to acquire the host's voice input; The speech-to-text module (314) is used to convert the acquired speech into text; The zero-sample target detection module (315) is used to call the zero-sample target detection model to perform zero-sample target detection on the target object; The large model calling module (316) is used to remotely call a large language model or a multimodal large language model; The management module (321) is used to manage the status of the system's underlying equipment and provide system status data to the user interface module; The controller module (331) is used to remotely control the product shooting end; The Bluetooth protocol module (341) is used for remote communication with the product shooting terminal.
6. The system for applying the method of live broadcasting to display products according to claim 5, characterized in that: The server is a device that can run Windows, Linux, or MacOS operating systems.
7. The method for live broadcasting and displaying products according to claim 5, characterized in that: The system status data includes the connection status, operation status, and operation angle of the robot arm (202), the connection status of the camera device (203), and the connection status, operation status, operation speed, and operation direction of the automatic rotation device (204).
8. The method for live broadcasting and displaying products according to claim 5, characterized in that: The mechanical arm (202) and the automatic rotation device (204) communicate with each other using a CAN bus.
9. The method for live broadcasting and displaying products according to claim 5, characterized in that: The live broadcast platform is a live broadcast platform that can provide user nicknames, user barrages, and gift information interfaces.