Image tracking-based virtual anchor multi-camera view automatic switching method and system
By constructing a live streaming process state machine and multimodal feature abstraction, combined with dynamic threshold adjustment, the system accurately distinguishes between valid command actions and unconscious habitual actions, solving the problem of unstable camera switching in the automatic switching of multi-camera perspectives for virtual anchors, and improving the professionalism and user experience of live streaming.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SHANLIN ARITHMETIC (SHANGHAI) TECHNOLOGY CO LTD
- Filing Date
- 2026-01-12
- Publication Date
- 2026-05-29
AI Technical Summary
Existing technologies cannot effectively distinguish between unconscious habitual actions and effective instructions with clear directorial intent in the automatic switching of multi-camera perspectives for virtual anchors, resulting in frequent and chaotic camera switching, reducing the continuity of live broadcasts and user experience.
By constructing a live streaming process state machine, combining a posture estimation model, a temporal action classification model, and a multi-level keyword trigger library, the system identifies the driver's actions and voice commands, uses an intent arbitrator to predict the comprehensive intent confidence, and dynamically adjusts the threshold to filter out invalid actions, thereby achieving precise camera switching.
It effectively reduces the rate of accidental camera triggers, improves the smoothness of camera transitions and the consistency of intent, and ensures the professionalism and stability of live streaming.
Smart Images

Figure CN122120397A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of live streaming technology, specifically to a method and system for automatic switching of multi-camera perspectives for virtual anchors based on image tracking. Background Technology
[0002] With the deep integration of virtual reality and live streaming technologies, live streaming e-commerce featuring virtual digital humans (hereinafter referred to as "virtual anchors") has become an emerging model in the e-commerce field. In this model, a real person typically drives the virtual anchor's image to showcase and explain products in real time using cameras, sensors, and other devices. To enhance the expressiveness and professionalism of the live stream and simulate the multi-camera effect of traditional television broadcasts, multiple virtual cameras with different angles and perspectives are set up in the virtual scene, and automatic switching between shots is achieved.
[0003] Currently, common virtual anchor multi-camera perspective automatic switching solutions mainly rely on real-time tracking and motion recognition of the driver's image. The typical implementation method is as follows: real-time capture of the driver's skeletal key points or body posture through human pose estimation algorithms; pre-setting a series of mapping relationships between actions and virtual camera positions (for example, when the "spread arms" action is detected, switch to a panoramic shot; when the "pointing" action is detected, switch to a close-up shot); after the preset action is detected to be triggered, the corresponding shot switching command is executed.
[0004] However, in the impromptu, highly interactive live-streaming e-commerce process, the aforementioned existing technology solutions based on single image tracking often result in numerous unconscious, habitual movements (such as adjusting clothing, scratching the head, or briefly resting the chin on the hand). The mechanical mapping logic of existing technologies, lacking the ability to understand the intent behind these movements, cannot effectively distinguish between these unconscious, ineffective actions and the clear, intended commands from the director. This misjudgment of the movement's intent leads to frequent and chaotic camera cuts, disrupting the continuity of the live stream and degrading the viewing experience.
[0005] Therefore, there is an urgent need for a method for automatic switching of multiple camera angles for virtual anchors that can accurately understand the live streamer's intentions and effectively filter out invalid interference, in order to solve the aforementioned prominent problems faced by existing technologies and improve the overall quality and user experience of virtual anchor live streams. Summary of the Invention
[0006] The purpose of this invention is to provide a method, system, electronic device, readable storage medium, and computer program product for automatic switching of multiple camera perspectives for virtual anchors based on image tracking, so as to solve the problems mentioned in the background art.
[0007] This invention provides a method for automatic switching of multi-camera perspectives for virtual anchors based on image tracking, comprising the following steps: Step S10: Initialize and maintain a live streaming process state machine, obtain its output current live streaming state, including at least the product explanation state and the interactive Q&A state; at the same time, receive the video stream and audio stream from the driver; Step S20: Temporal human key point data is obtained from the video stream through the pose estimation model and input into the temporal action classification model to obtain a primary action semantic vector representing the physical action category and confidence level. Step S30: Identify the speech-text stream from the audio stream, and match the speech-text stream using a multi-level keyword triggering library to obtain an instruction strength vector; Step S40: Input the primary action semantic vector, the instruction intensity vector and the current live broadcast state together into the intent arbitrator to predict the comprehensive intent confidence. Step S50: When the comprehensive intent confidence exceeds the dynamic threshold, the current action is determined to be a valid camera switching instruction, and the corresponding virtual camera position is switched according to the primary action semantic vector; wherein, the dynamic threshold is adaptively adjusted according to the historical switching frequency.
[0008] This invention also provides an image tracking-based multi-camera perspective automatic switching system for virtual anchors, the system comprising: The acquisition unit is used to: initialize and maintain a live streaming process state machine, obtain its output current live streaming state, including at least the product explanation state and the interactive Q&A state; and simultaneously receive the video stream and audio stream from the driver. The action semantic analysis unit is used to: obtain temporal human key point data from the video stream through the pose estimation model, and input it into the temporal action classification model to obtain a primary action semantic vector representing the physical action category and confidence level; The instruction strength analysis unit is used to: identify the speech-text stream from the audio stream, and match the speech-text stream through a multi-level keyword triggering library to obtain an instruction strength vector; The intent analysis unit is used to: input the primary action semantic vector, the instruction intensity vector and the current live broadcast state into the intent arbitrator, and predict the comprehensive intent confidence. The perspective switching unit is used to: determine that the current action is a valid camera switching instruction when the comprehensive intent confidence exceeds the dynamic threshold, and trigger the corresponding virtual camera position switching according to the primary action semantic vector; wherein, the dynamic threshold is adaptively adjusted according to the historical switching frequency.
[0009] The present invention also provides an electronic device including a processor and a memory, the memory storing a program or instructions executable on the processor, the program or instructions, when executed by the processor, implementing the steps of the method as described in any of the preceding claims.
[0010] The present invention also provides a readable storage medium on which a program or instructions are stored, which, when executed by a processor, implement the steps of the method as described in any of the preceding claims.
[0011] The present invention also provides a computer program product stored in a storage medium, the program product being executed by at least one processor to implement the steps of the method as described in any of the preceding claims.
[0012] This invention achieves scene awareness by constructing a live streaming process state machine. Combined with multimodal feature abstraction and a dynamic weight arbitration mechanism, it accurately distinguishes between valid command actions and unconscious habitual actions based on intelligent intent understanding, thereby effectively reducing the rate of accidental camera triggers. Simultaneously, this invention also automatically raises the trigger threshold during periods of concentrated action by dynamically adjusting the threshold based on historical switching frequencies, effectively suppressing switching oscillations and enabling virtual anchor camera switching to achieve a smoothness and intent consistency similar to that of a professional director. Attached Figure Description
[0013] Figure 1 This is a flowchart illustrating an image tracking-based method for automatic switching of multiple camera perspectives for virtual anchors, as disclosed in an embodiment of the present invention. Figure 2 This is a functional schematic diagram of the method scheme disclosed in the embodiments of the present invention; Figure 3 This is a schematic diagram of the structure of a virtual anchor multi-camera perspective automatic switching system based on image tracking, as disclosed in an embodiment of the present invention. Detailed Implementation
[0014] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0015] Please see Figure 1 , Figure 2 This invention provides a method 100 for automatic switching of multi-camera perspectives for virtual anchors based on image tracking, comprising the following steps: Step S10: Initialize and maintain a live streaming process state machine, obtain its output current live streaming state, including at least the product explanation state and the interactive Q&A state; at the same time, receive the video stream and audio stream from the driver; In this step, the live streaming process state machine is used to continuously track and determine the different stages of the live streaming process, including at least the product explanation state and the interactive Q&A state. In the product explanation state, the core task of the live stream is to showcase product details and explain product features; while in the interactive Q&A state, the focus shifts to responding to audience questions and creating an interactive atmosphere. The state machine transitions and maintains between these states according to preset rules (such as specific keyword triggers, timers, or manual switching signals). It can be understood that the live streaming process state machine is constructed and driven by a rule-based finite state machine and a lightweight, highly efficient event detection algorithm.
[0016] Simultaneously, video streams from the camera and audio streams from the microphone are received in real time, and these two signals constitute the multimodal data source for subsequent analysis.
[0017] Step S20: Temporal human key point data is obtained from the video stream through the pose estimation model and input into the temporal action classification model to obtain a primary action semantic vector representing the physical action category and confidence level. In this step, based on the video stream, the temporal key point coordinate data of various parts of the driver's body are extracted in real time using a pose estimation model (such as MediaPipe Pose) to form a digital description of the human posture.
[0018] The aforementioned temporal keypoint data is input into a temporal action classification model, such as a lightweight network based on Temporal Convolutional Network or LSTM. This temporal action classification model can analyze pose changes within a short temporal window, identify predefined physical action categories such as "pointing," "opening arms," and "picking up an item," and output a primary action semantic vector. It is understood that this primary action semantic vector includes not only the identified action type encoding but also the confidence level of the temporal action classification model in the identification result.
[0019] Step S30: Identify the speech-text stream from the audio stream, and match the speech-text stream using a multi-level keyword triggering library to obtain an instruction strength vector; In this step, the received audio stream undergoes real-time speech recognition (e.g., using a speech-to-text engine like Whisper) to convert it into a speech-text stream. Then, this text stream is matched in real-time using a predefined multi-level keyword trigger library. This multi-level keyword trigger library contains at least a set of strong guiding instructions (such as "Look here," "Pay attention to this detail") and a set of weakly related instructions (such as "This," "Over there"). The matching result is quantified into an instruction strength vector, which characterizes the strength of the director's intention and the specific instruction type contained in the current speech.
[0020] Step S40: Input the primary action semantic vector, the instruction intensity vector and the current live broadcast state together into the intent arbitrator to predict the comprehensive intent confidence. In this step, the intent arbitrator, as an intelligent decision-making module, receives the primary action semantic vector from step S20, the instruction strength vector from step S30, and the current live broadcast status from step S10.
[0021] The intent arbitrator has a pre-defined dynamic weight allocation strategy that dynamically adjusts the decision weights of different input vectors based on the current live stream status. For example, in a product demonstration mode, "pointing" actions related to product display are given higher weight; while in an interactive Q&A mode, the weight of voice commands is significantly increased. By comprehensively analyzing the weighted and fused information, the intent arbitrator calculates a comprehensive intent confidence score, which quantifies the overall probability that the driver's action at the current moment constitutes a valid camera switching command.
[0022] Step S50: When the comprehensive intent confidence exceeds the dynamic threshold, the current action is determined to be a valid camera switching instruction, and the corresponding virtual camera position is switched according to the primary action semantic vector; wherein, the dynamic threshold is adaptively adjusted according to the historical switching frequency.
[0023] In this step, the comprehensive intent confidence obtained above is compared with a dynamic threshold, which can be adaptively adjusted according to the historical switching frequency, thereby increasing the trigger threshold during periods of high activity, preventing excessive switching, and ensuring the stability of the system output.
[0024] When the overall intent confidence exceeds the adaptively adjusted dynamic threshold, the current action of the driver is ultimately determined as a valid shot switching instruction. This rigorous conditional judgment mechanism ensures that only actions with clear directorial intent and high confidence supported by multimodal information can trigger a switch, filtering out unconscious, habitual actions. Upon successful determination, the preset virtual camera position (such as close-up shots, panoramic shots, etc.) mapped to the action type represented by the primary action semantic vector is triggered to complete the switch, thereby achieving an intelligent, smooth, and professional directing effect.
[0025] This invention achieves scene awareness by constructing a live streaming process state machine. Combined with multimodal feature abstraction and a dynamic weight arbitration mechanism, it accurately distinguishes between valid command actions and unconscious habitual actions based on intelligent intent understanding, thereby effectively reducing the rate of accidental camera triggers. Simultaneously, this invention also automatically raises the trigger threshold during periods of concentrated action by dynamically adjusting the threshold based on historical switching frequencies, effectively suppressing switching oscillations and enabling virtual anchor camera switching to achieve a smoothness and intent consistency similar to that of a professional director.
[0026] As an example, temporal human keypoint data is obtained from the video stream using a pose estimation model and input into a temporal action classification model to obtain a primary action semantic vector representing the physical action category and confidence level, including: Step S201: Extract the two-dimensional or three-dimensional coordinate sequence of multiple key points, including the driver's head, hands, elbows, and shoulders, from the video stream in real time using a pose estimation model to form temporal human key point data. In this step, to extract effective motion information from the video stream, this invention employs a pose estimation model (such as the mature MediaPipe Pose model) to process the video stream frame by frame. This pose estimation model can accurately detect and output the coordinate data of multiple key points of the driver's body in two-dimensional image space or three-dimensional physical space. These key points include at least the joints of the head, hands, elbows, and shoulders. The coordinate data is then arranged chronologically to form a temporal sequence of human key point data. It should be noted that three-dimensional coordinate data is preferentially used to improve the accuracy of motion recognition.
[0027] Step S202: Input the temporal human key point data into a temporal action classification model based on a temporal convolutional network. The temporal action classification model analyzes the motion trajectory of key points within a preset short temporal window, identifies multiple physical action categories, and outputs a primary action semantic vector containing action type encoding and corresponding confidence for each identified action.
[0028] In this step, the aforementioned temporal human keypoint data is input into a temporal action classification model for action recognition. This temporal action classification model is preferably built upon a Temporal Convolutional Network (TCN). Through its unique causal convolution and dilated convolution structures, the TCN can efficiently process temporal data and capture long-term dependencies over time. Compared to models such as Recurrent Neural Networks (RNNs), it maintains high accuracy while possessing stronger parallel computing capabilities and more stable gradients, making it more suitable for the real-time live streaming scenarios involved in this invention.
[0029] The temporal action classification model uses a pre-defined short temporal window (e.g., corresponding to the most recent 30 frames of data in one second) as the analysis unit, comprehensively analyzing the motion trajectory, relative position changes, and velocity characteristics of all keypoints within this window. Through its internally trained network parameters, the model can identify various predefined physical action categories, such as "pointing" (whose trajectory is characterized by the continuous movement of hand keypoints in a certain direction), "opening arms" (whose trajectory is characterized by the simultaneous outward extension of both arm keypoints), "picking up an item," "waving," and "nodding," etc. For each identified action, a structured data set, namely a primary action semantic vector, is output.
[0030] The primary action semantic vector includes at least: the encoded action type, represented by a numeric ID or a one-hot vector, indicating the specific action; and the confidence level ([0,1]) corresponding to the action, reflecting the degree of understanding of the temporal action classification model regarding the recognition result.
[0031] As an example, a multi-level keyword triggering library includes at least a strong guidance instruction set and a weak association instruction set; wherein, the strong guidance instruction set includes clearly directional broadcast instruction keywords, and the weak association instruction set includes context-dependent indicative keywords; The step of matching the speech-text stream using a multi-level keyword trigger library to obtain the instruction strength vector includes: Step S301: Synchronously match the real-time recognized speech-text stream with the strong guidance instruction set and the weak association instruction set; In this step, the speech-text stream obtained in real time through the speech recognition engine is continuously analyzed. To accurately extract the director's intent from the speech information, this invention establishes a multi-level keyword trigger library as a matching benchmark. This trigger library contains at least two levels: a strong guidance instruction set and a weak association instruction set.
[0032] A multi-threaded parallel processing approach is adopted to synchronously match the real-time acquired speech-text stream with keywords in two instruction sets. In specific implementation, the matching process uses a high-efficiency string matching algorithm based on Trie trees to ensure that the text stream scanning and keyword recognition are completed within milliseconds, meeting the real-time requirements of live streaming scenarios.
[0033] In step S302, if a keyword in the strong guidance instruction set is matched, an instruction strength vector with a high weight is generated; if only a keyword in the weak association instruction set is matched, an instruction strength vector with a low weight is generated; wherein, the instruction strength vector is used to characterize the explicitness and intensity of the voice instruction.
[0034] In this step, instruction strength vectors with corresponding weights are generated based on the different matching results. When keywords in the strong guiding instruction set are matched in the speech text stream, such as phrases with clear directing intent like "Look here" or "Pay attention to this detail," a high-weight instruction strength vector is generated. This high-weight vector reflects that the current speech instruction has a strong directing intent, and the judgment can be made without relying on other modal information.
[0035] When only keywords from the weakly related instruction set are matched in the speech-text stream, such as words like "this" or "over there" whose intent needs to be understood in context, a low-weight instruction strength vector is generated. This low-weight vector indicates that the director's intent for the current speech instruction is weak and requires comprehensive judgment in conjunction with visual action information.
[0036] Understandably, the instruction strength vector is represented in the form of a numerical feature vector, which includes information such as the type of matched keyword, its frequency of occurrence, and the calculated weight.
[0037] This implementation method, based on a multi-level keyword triggering and intent intensity quantification mechanism, achieves fine differentiation of voice directing instructions from presence to strength, effectively identifies and strengthens the decision weight of strong guiding instructions, while weakening the interference of vague expressions, thereby effectively suppressing false triggers caused by habitual spoken language and significantly improving the parsing accuracy of voice instructions.
[0038] As an example, the primary action semantic vector, the instruction intensity vector, and the current live streaming state are input together into the intent arbitrator to predict the comprehensive intent confidence, including: Step S401: Real-time acquisition of live stream bullet comments and extraction of semantic information containing requests for camera switching, generating a bullet comment interference intensity vector; In this step, during live streaming, bullet comments often contain numerous requests for immediate camera switching (such as "close-up" or "see product"), but these requests often conflict with the presenter's presentation rhythm and professional judgment. To address this, this invention quantifies the intensity of bullet comment interference and assigns negative weights, effectively identifying and suppressing misleading signals caused by audience group behavior, thus preventing camera switching from being excessively influenced by unprofessional opinions. This mechanism preserves the presenter's control over the live streaming process while enabling intelligent directorial decision-making through negative feedback adjustment, thereby maintaining the professionalism and stability of decision-making in complex interactive environments.
[0039] First, the degree of interference from the bullet comment environment on camera switching decisions is quantitatively assessed. Through a multi-level processing flow, semantic information containing camera switching requests is identified and extracted from the original bullet comment stream. The specific processing flow is roughly as follows: A dedicated dictionary containing industry terminology is established, and a three-level semantic filtering mechanism is used to identify explicit switching commands, negative feedback requests, and product display needs. The intent intensity of the selected bullet comments is quantified, and dynamic weight adjustments are made based on factors such as user value, time density, and state consistency, ultimately generating a standardized bullet comment interference intensity vector. This vector numerically represents the real-time interference intensity of the current bullet comment environment.
[0040] Step S402: Input the primary action semantic vector, the instruction strength vector, the barrage interference strength vector, and the current live broadcast status together into the intent arbitrator based on the gated recurrent unit network; In this step, a deep neural network based on gated recurrent units (GRUs) is constructed as an intent arbitrator, capable of processing temporal data. The input vectors mentioned above, after feature alignment and dimensionality unification, are jointly input into the GRU network. Specifically, the primary action semantic vector provides information on the driver's body movements, the instruction intensity vector provides the clarity of the voice instruction, the barrage interference intensity vector provides environmental interference signals, and the current live stream state provides the context for decision-making. The GRU network uses its gating mechanism to initially fuse this multimodal information.
[0041] In step S403, the intent arbitrator dynamically adjusts the weight allocation of each input vector according to the current live broadcast state through its internal state-aware gating mechanism, wherein the barrage interference intensity vector is given a negative weight coefficient; and, based on the temporal memory capability of the gated recurrent unit network, it performs temporal integration on the weighted multimodal vector sequence, filters out instantaneous interference signals, and outputs the comprehensive intent confidence.
[0042] In this step, the state-aware gating mechanism dynamically adjusts the importance weights of each input vector based on the current live stream state. The degree of dependence on each input vector varies depending on the live stream state: in the product explanation state, more attention is paid to the action semantic vector, while in the interactive Q&A state, more emphasis is placed on the instruction intensity vector. It should be noted that the barrage interference intensity vector is given a negative weight coefficient to suppress false triggers.
[0043] The temporal memory capability of the GRU network enables it to analyze multimodal information sequences in consecutive time steps. Through update and reset gate mechanisms, it selectively retains important historical information and filters out transient interference signals. This temporal integration ensures the stability of decision-making and avoids misjudgments caused by fluctuations in single-frame signals. Finally, the GRU outputs a comprehensive intent confidence score between 0 and 1, which comprehensively reflects the overall probability that the driver's actions constitute a valid shot-switching instruction in the current multimodal environment.
[0044] This invention introduces a barrage interference intensity vector and assigns it a negative weight, enabling the system to resist audience interference and effectively suppress false triggers caused by non-professional demands from the audience. Combined with the temporal modeling capabilities of the GRU network, it achieves dynamic weight allocation and temporal integration of multimodal information, ensuring accurate identification of the anchor's true intentions in complex live streaming environments and reducing the false trigger rate of camera switching due to external interference.
[0045] As an example, the live stream's bullet comments are captured in real time, and semantic information containing requests for camera switching is extracted to generate a bullet comment interference intensity vector, including: Step S4011: Collect and preprocess bullet screen data in real time, and identify bullet screens containing explicit switching instructions, negative feedback and product display requests through multi-level semantic filtering; In this step, during the preprocessing stage, a dictionary specifically for the live streaming industry is constructed, containing industry terms such as "close-up shot" and "wide-angle shot." The first level of semantic filtering is based on an instruction pattern library, which contains the following three types of instructions: Explicit switching commands: such as direct commands like "cut shot", "rotate screen", "change camera position"; Shot size control instructions: such as professional terms like "close-up," "wide shot," "close-up," and "long shot"; Directional instructions: such as "look to the left", "take a picture of the right", "show me the side view", etc.
[0046] The second level of sentiment analysis uses an attention-based classification model to identify visual discomfort expressions such as "too blurry" and "eye-straining," as well as content quality complaints such as "boring" and "uninteresting." The third level of product association analysis tracks the live stream's product list in real time, identifying specific product display needs such as "try on the lipstick" and "take a picture of the back of your phone."
[0047] Step S4012: Quantify the intent intensity of the filtered bullet comments, and assign basic intensity score, emotion adjustment score and relevance score respectively. Dynamically adjust the scores based on user value, time density and state consistency. In this step, the base intensity score is divided according to the command type. For example, a direct switching command scores 0.9 points, a shot control command scores 0.7 points, and a directional indication scores 0.5 points. The emotion adjustment score is based on the intensity grading of emotion words. For example, words like "strong dissatisfaction" are weighted 1.5 times, and words like "general advice" are weighted 1.2 times. The relevance score is determined by calculating the semantic similarity between the bullet screen content and the product features. A perfect match scores 1.0 points, and a partial match scores 0.6 points.
[0048] Based on this, the following adjustment mechanism will be further implemented: User value weighting: User influence is assessed based on historical user interaction data (including fan level, consumption history, viewing time, etc.), and users are divided into different weight levels of 1.0-3.0 times. The bullet comments of high-value users have a stronger influence. Temporal density detection: Statistically count the frequency of similar semantic bullet comments within a unit of time, and exponentially decay repeated bullet comments that appear densely in a short period of time to avoid weight distortion caused by group spamming; State consistency check: The request in the bullet comments is compared with the current live broadcast state (including camera framing, shooting angle, displayed products, etc.) in real time. When the request in the bullet comments conflicts with the current state (such as the bullet comments requesting a close-up but the current state is already a close-up), the weight of that bullet comment is reduced.
[0049] Step S4013: Extract the weighted barrage intensity sequence using a sliding time window, calculate the peak intensity, average intensity, and intensity change gradient of the weighted barrage intensity sequence, and generate the barrage interference intensity vector by linear weighted combination.
[0050] In this step, time-series analysis techniques are used to synthesize the final interference intensity vector. Specifically, a sliding time window of, for example, 30 seconds is maintained, and the weighted barrage intensity sequence within the window is updated and analyzed every second.
[0051] Calculate the following three key indicators: 1) Peak intensity: The 90th percentile value algorithm is used to take the intensity value of the top 10% of the bullet screen intensity in the window, which effectively avoids the influence of extreme outliers; 2) Average intensity: A time decay weighted algorithm is used to establish a time decay function, in which the weight of the barrage in the most recent 5 seconds is set to 1.0, the weight of the barrage in 6-10 seconds is 0.8, the weight of the barrage in 11-15 seconds is 0.6, and so on, to ensure that the recent barrage has a higher weight. 3) Intensity change gradient: The intensity sequence within the window is subjected to linear regression analysis using the least squares method to calculate its slope value. A positive value indicates that the degree of interference is intensifying, while a negative value indicates that the degree of interference is easing.
[0052] Finally, a feature fusion formula, such as interference intensity = 0.5 × peak intensity + 0.3 × average intensity + 0.2 × intensity change gradient, is used to quantify the real-time interference of the bullet screen environment on camera switching decisions, providing a precise suppression signal for intent arbitration. Integrating the interference intensity from multiple sliding time windows constitutes the bullet screen interference intensity vector.
[0053] Please see Figure 3 This invention also provides an image tracking-based multi-camera perspective automatic switching system for virtual anchors, the system comprising: The acquisition unit 2011 is used to: initialize and maintain a live streaming process state machine, obtain its output current live streaming state, including at least the product explanation state and the interactive Q&A state; and simultaneously receive the video stream and audio stream from the driver. The action semantic analysis unit 2012 is used to: obtain temporal human key point data from the video stream through the pose estimation model, and input it into the temporal action classification model to obtain a primary action semantic vector representing the physical action category and confidence level; The instruction strength analysis unit 2013 is used to: identify the speech-text stream from the audio stream, and match the speech-text stream through a multi-level keyword trigger library to obtain an instruction strength vector; The intent analysis unit 2014 is used to: input the primary action semantic vector, the instruction intensity vector and the current live broadcast state into the intent arbitrator, and predict the comprehensive intent confidence. The perspective switching unit 2015 is used to: determine that the current action is a valid camera switching instruction when the comprehensive intent confidence exceeds the dynamic threshold, and trigger the corresponding virtual camera position switching according to the primary action semantic vector; wherein, the dynamic threshold is adaptively adjusted according to the historical switching frequency.
[0054] As an example, the action semantic analysis unit 2012 is used for: The pose estimation model extracts the two-dimensional or three-dimensional coordinate sequence of multiple key points, including the driver's head, hands, elbows, and shoulders, from the video stream in real time to form temporal human key point data. The temporal human key point data is input into a temporal action classification model based on a temporal convolutional network. The temporal action classification model analyzes the motion trajectory of key points within a preset short temporal window, identifies multiple physical action categories, and outputs a primary action semantic vector containing action type encoding and corresponding confidence for each identified action.
[0055] As an example, a multi-level keyword triggering library includes at least a strong guidance instruction set and a weak association instruction set; wherein, the strong guidance instruction set includes clearly directional broadcast instruction keywords, and the weak association instruction set includes context-dependent indicative keywords; The instruction strength analysis unit 2013 is used for: The real-time recognized speech-text stream is synchronously matched with the strong guidance instruction set and the weak association instruction set; If a keyword in the strongly associated instruction set is matched, an instruction strength vector with a high weight is generated; if only a keyword in the weakly associated instruction set is matched, an instruction strength vector with a low weight is generated; wherein, the instruction strength vector is used to characterize the explicitness and intensity of the voice instruction.
[0056] As an example, the intent analysis unit 2014 is used for: Real-time acquisition of live stream bullet comments and extraction of semantic information containing camera switching requests to generate bullet comment interference intensity vector; The primary action semantic vector, the instruction strength vector, the barrage interference strength vector, and the current live broadcast status are input together into the intent arbitrator based on a gated recurrent unit network. The intent arbitrator dynamically adjusts the weight allocation of each input vector according to the current live broadcast state through its internal state-aware gating mechanism, wherein the barrage interference intensity vector is given a negative weight coefficient; and, based on the temporal memory capability of the gated recurrent unit network, it performs temporal integration on the weighted multimodal vector sequence, filters out instantaneous interference signals, and outputs the comprehensive intent confidence.
[0057] As an example, the intent analysis unit 2014 is also used for: Real-time collection and preprocessing of bullet screen data; identification of bullet screens containing explicit switching instructions, negative feedback and product display requests through multi-level semantic filtering. The filtered bullet comments are quantified in terms of intent intensity, and assigned a base intensity score, an emotion adjustment score, and a relevance score respectively. The scores are then dynamically weighted and adjusted based on user value, time density, and state consistency. A weighted barrage intensity sequence is extracted using a sliding time window. The peak intensity, average intensity, and intensity change gradient of the weighted barrage intensity sequence are calculated, and the barrage interference intensity vector is generated by linear weighted combination.
[0058] The present invention also provides an electronic device including a processor and a memory, the memory storing a program or instructions executable on the processor, the program or instructions being executed by the processor to implement the steps of the method as described in any of the foregoing embodiments.
[0059] The present invention also provides a readable storage medium on which a program or instructions are stored, which, when executed by a processor, implement the steps of the method as described in any of the foregoing embodiments.
[0060] The present invention also provides a computer program product stored in a storage medium, the program product being executed by at least one processor to implement the steps of the method as described in any of the foregoing embodiments.
[0061] The above description is only a part of the embodiments of this application and does not limit the patent scope of this application. All equivalent structural transformations made under the technical concept of this application and using the contents of the specification and drawings of this application, or direct / indirect applications in other related technical fields, are included in the patent protection scope of this application.
Claims
1. A method for automatic switching of multi-camera perspectives for virtual anchors based on image tracking, characterized in that, The methods and steps include the following: Step S10: Initialize and maintain a live streaming process state machine, obtain its output current live streaming state, including at least the product explanation state and the interactive Q&A state; at the same time, receive the video stream and audio stream from the driver; Step S20: Temporal human key point data is obtained from the video stream through the pose estimation model and input into the temporal action classification model to obtain a primary action semantic vector representing the physical action category and confidence level. Step S30: Identify the speech-text stream from the audio stream, and match the speech-text stream using a multi-level keyword triggering library to obtain an instruction strength vector; Step S40: Input the primary action semantic vector, the instruction intensity vector and the current live broadcast state together into the intent arbitrator to predict the comprehensive intent confidence. Step S50: When the comprehensive intent confidence exceeds the dynamic threshold, the current action is determined to be a valid camera switching instruction, and the corresponding virtual camera position is switched according to the primary action semantic vector; wherein, the dynamic threshold is adaptively adjusted according to the historical switching frequency.
2. The method for automatic switching of multi-camera perspectives for virtual anchors based on image tracking according to claim 1, characterized in that: Temporal human keypoint data is obtained from the video stream using a pose estimation model and input into a temporal action classification model to obtain a primary action semantic vector representing the physical action category and confidence level, including: Step S201: Extract the two-dimensional or three-dimensional coordinate sequence of multiple key points, including the driver's head, hands, elbows, and shoulders, from the video stream in real time using a pose estimation model to form temporal human key point data. Step S202: Input the temporal human key point data into a temporal action classification model based on a temporal convolutional network. The temporal action classification model analyzes the motion trajectory of key points within a preset short temporal window, identifies multiple physical action categories, and outputs a primary action semantic vector containing action type encoding and corresponding confidence for each identified action.
3. The method for automatic switching of multi-camera perspectives for virtual anchors based on image tracking according to claim 2, characterized in that: The multi-level keyword triggering library includes at least a strong guidance instruction set and a weak association instruction set; wherein, the strong guidance instruction set includes clearly directional broadcast instruction keywords, and the weak association instruction set includes context-dependent indicative keywords; The step of matching the speech-text stream using a multi-level keyword trigger library to obtain the instruction strength vector includes: Step S301: Synchronously match the real-time recognized speech-text stream with the strong guidance instruction set and the weak association instruction set; In step S302, if a keyword in the strong guidance instruction set is matched, an instruction strength vector with a high weight is generated; if only a keyword in the weak association instruction set is matched, an instruction strength vector with a low weight is generated; wherein, the instruction strength vector is used to characterize the explicitness and intensity of the voice instruction.
4. The method for automatic switching of multi-camera perspectives for virtual anchors based on image tracking according to claim 1, characterized in that: The primary action semantic vector, the instruction intensity vector, and the current live streaming state are input together into the intent arbitrator to predict the comprehensive intent confidence, including: Step S401: Real-time acquisition of live stream bullet comments and extraction of semantic information containing requests for camera switching, generating a bullet comment interference intensity vector; Step S402: Input the primary action semantic vector, the instruction strength vector, the barrage interference strength vector, and the current live broadcast status together into the intent arbitrator based on the gated recurrent unit network; In step S403, the intent arbitrator dynamically adjusts the weight allocation of each input vector according to the current live broadcast state through its internal state-aware gating mechanism, wherein the barrage interference intensity vector is given a negative weight coefficient; and, based on the temporal memory capability of the gated recurrent unit network, it performs temporal integration on the weighted multimodal vector sequence, filters out instantaneous interference signals, and outputs the comprehensive intent confidence.
5. The method for automatic switching of multi-camera perspectives for virtual anchors based on image tracking according to claim 4, characterized in that: Real-time acquisition of live stream comments and extraction of semantic information containing camera switching requests are used to generate a comment interference intensity vector, including: Step S4011: Collect and preprocess bullet screen data in real time, and identify bullet screens containing explicit switching instructions, negative feedback and product display requests through multi-level semantic filtering; Step S4012: Quantify the intent intensity of the filtered bullet comments, and assign basic intensity score, emotion adjustment score and relevance score respectively. Dynamically adjust the scores based on user value, time density and state consistency. Step S4013: Extract the weighted barrage intensity sequence using a sliding time window, calculate the peak intensity, average intensity, and intensity change gradient of the weighted barrage intensity sequence, and generate the barrage interference intensity vector by linear weighted combination.
6. A virtual anchor multi-camera perspective automatic switching system based on image tracking, characterized in that: The system includes: The acquisition unit is used to: initialize and maintain a live streaming process state machine, obtain its output current live streaming state, including at least the product explanation state and the interactive Q&A state; and simultaneously receive the video stream and audio stream from the driver. The action semantic analysis unit is used to: obtain temporal human key point data from the video stream through the pose estimation model, and input it into the temporal action classification model to obtain a primary action semantic vector representing the physical action category and confidence level; The instruction strength analysis unit is used to: identify the speech-text stream from the audio stream, and match the speech-text stream through a multi-level keyword triggering library to obtain an instruction strength vector; The intent analysis unit is used to: input the primary action semantic vector, the instruction intensity vector and the current live broadcast state into the intent arbitrator, and predict the comprehensive intent confidence. The perspective switching unit is used to: determine that the current action is a valid camera switching instruction when the comprehensive intent confidence exceeds the dynamic threshold, and trigger the corresponding virtual camera position switching according to the primary action semantic vector; wherein, the dynamic threshold is adaptively adjusted according to the historical switching frequency.
7. The image tracking-based virtual anchor multi-camera perspective automatic switching system according to claim 6, characterized in that: The action semantic analysis unit is used for: The pose estimation model extracts the two-dimensional or three-dimensional coordinate sequence of multiple key points, including the driver's head, hands, elbows, and shoulders, from the video stream in real time to form temporal human key point data. The temporal human key point data is input into a temporal action classification model based on a temporal convolutional network. The temporal action classification model analyzes the motion trajectory of key points within a preset short temporal window, identifies multiple physical action categories, and outputs a primary action semantic vector containing action type encoding and corresponding confidence for each identified action.
8. An electronic device, characterized in that, It includes a processor and a memory, the memory storing a program or instructions that can run on the processor, the program or instructions being executed by the processor to implement the steps of the method as described in any one of claims 1 to 5.
9. A readable storage medium, characterized in that, The readable storage medium stores a program or instructions that, when executed by a processor, implement the steps of the method as described in any one of claims 1 to 5.
10. A computer program product, characterized in that, The program product is stored in a storage medium and is executed by at least one processor to implement the steps of the method as claimed in any one of claims 1 to 5.