Optimization system and method for dynamic tracking of a camera based on topology

By combining motion capture, dynamic rate calculation, and camera switching modules through a topology-based camera dynamic tracking optimization system, the problem of existing camera tracking systems being unable to adapt to changes in user behavior is solved, achieving efficient and accurate camera sensitivity matching and resource optimization.

CN120529185BActive Publication Date: 2025-11-07广东公信智能会议股份有限公司
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510919785.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-04
Publication Date
2025-11-07
Estimated Expiration
2045-07-04

AI Technical Summary

Technical Problem

Existing camera tracking systems cannot intelligently select the appropriate sensitivity camera based on the dynamic changes in user behavior, resulting in poor tracking performance and wasted resources.

Method used

A topology-based camera dynamic tracking optimization system is adopted. The system acquires user action sets through the action acquisition module, quantifies the behavior dynamic rate using the dynamic rate calculation module, constructs the camera module topology, and intelligently switches cameras between different topology layers through the camera switching module to achieve adaptive matching of sensitivity.

Benefits of technology

It enables adaptive selection of appropriate camera sensitivity based on changes in user behavior, improving tracking accuracy and optimizing resource allocation to ensure efficient and stable camera tracking in different dynamic scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120529185B_ABST
    Figure CN120529185B_ABST
Patent Text Reader

Abstract

The application provides a camera dynamic tracking optimization system and method based on a topological structure, and belongs to the field of camera dynamic tracking. The system comprises a motion acquisition module, a dynamic rate calculation module, a topological construction module and a camera switching module. The motion acquisition module detects a sound source user and acquires a sound source user motion set; the dynamic rate calculation module calculates the behavior dynamic rate of the sound source user motion set by using a dynamic rate function; the topological construction module constructs a camera module topological structure comprising a first type and a second type of camera module topological layer, wherein the tracking sensitivity of each camera module in the first type of camera module topological layer is less than that in the second type of camera module topological layer; and the camera switching module determines a switched camera module by optimization in the corresponding camera module topological layer according to the comparison result of the behavior dynamic rate and a preset dynamic rate threshold. The tracking accuracy is improved and resource allocation is optimized by intelligently matching the camera tracking sensitivity according to the user behavior dynamic rate.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of camera dynamic tracking, and in particular to a camera dynamic tracking optimization system and method based on a topological structure. BACKGROUND

[0002] In video conferencing, intelligent monitoring and other applications, camera dynamic tracking is a core technology for realizing user behavior capture and picture optimization. Existing camera tracking usually adopts a fixed tracking strategy, that is, all cameras track the target user with a fixed sensitivity parameter.

[0003] However, in actual applications, user behavior has obvious dynamic variation characteristics. For example, in a meeting scenario, when the user has a heated discussion or moves quickly, the user's head turns and the body swings, and other movements change dramatically. In a static speech or thinking state, the user's movement changes relatively slowly. The camera tracking in the prior art cannot identify and adapt to the difference in dynamic variation of behavior. Specifically, when the user's behavior changes dramatically, tracking with a low-sensitivity camera is prone to tracking lag or defocusing, resulting in poor tracking effect. When the user's behavior changes slowly, tracking with a high-sensitivity camera will produce unnecessary frequent switching and adjustment, causing resource waste and picture jitter. Therefore, the camera tracking system in the prior art cannot intelligently select a camera with appropriate sensitivity according to the degree of dynamic variation of user behavior, resulting in poor tracking effect and resource waste. SUMMARY

[0004] The present application provides a camera dynamic tracking optimization system and method based on a topological structure to solve the technical problem that the camera tracking system in the prior art cannot intelligently select a camera with appropriate sensitivity according to the degree of dynamic variation of user behavior, resulting in poor tracking effect and resource waste.

[0005] The technical solution of the present application to solve the above technical problem is as follows:

[0006] In a first aspect, the present application provides a topological structure-based camera dynamic tracking optimization system, comprising a motion acquisition module, a dynamic rate calculation module, a topological structure construction module, and a camera switching module. The motion acquisition module is configured to detect a sound source user, acquire the motion of the sound source user, and obtain a sound source user motion set, wherein the sound source user motion set comprises a head offset angular velocity, a shoulder displacement amplitude, and a leg swing frequency. The dynamic rate calculation module is configured to calculate a behavior dynamic rate of the sound source user motion set by using a dynamic rate function, wherein the behavior dynamic rate is an index representing the degree of change of user behavior. The topological structure construction module is configured to construct a camera module topological structure, wherein the camera module topological structure comprises a first type of camera module topological layer and a second type of camera module topological layer, and the tracking sensitivity of each camera module in the first type of camera module topological layer is less than the tracking sensitivity of each camera module in the second type of camera module topological layer. The camera switching module is configured to determine a first switching camera module in the second type of camera module topological layer when the behavior dynamic rate is greater than or equal to a preset dynamic rate threshold, or determine a first switching camera module in the first type of camera module topological layer when the behavior dynamic rate is less than the preset dynamic rate threshold.

[0007] In a second aspect, the present application provides a topological structure-based camera dynamic tracking optimization method, comprising: detecting a sound source user, acquiring the motion of the sound source user, and obtaining a sound source user motion set, wherein the sound source user motion set comprises a head offset angular velocity, a shoulder displacement amplitude, and a leg swing frequency; calculating a behavior dynamic rate of the sound source user motion set by using a dynamic rate function, wherein the behavior dynamic rate is an index representing the degree of change of user behavior; constructing a camera module topological structure, wherein the camera module topological structure comprises a first type of camera module topological layer and a second type of camera module topological layer, and the tracking sensitivity of each camera module in the first type of camera module topological layer is less than the tracking sensitivity of each camera module in the second type of camera module topological layer; determining a first switching camera module in the second type of camera module topological layer when the behavior dynamic rate is greater than or equal to a preset dynamic rate threshold, or determining a first switching camera module in the first type of camera module topological layer when the behavior dynamic rate is less than the preset dynamic rate threshold.

[0008] The present application has the following advantages:

[0009] The sound source user is detected by the action acquisition module, the action of the sound source user is acquired, the action set of the sound source user is acquired, the action set of the sound source user includes a head offset angular velocity, a shoulder displacement amplitude and a leg swing frequency, and thus the basic data of the user behavior is obtained; the behavior dynamic rate of the action set of the sound source user is calculated by using the dynamic rate calculation module by using a dynamic rate function, the behavior dynamic rate is an index for representing the change degree of the user behavior, and thus the change degree of the current behavior of the user is quantified; the camera module topology structure is constructed by the topology construction module, the camera module topology structure includes a first type of camera module topology layer and a second type of camera module topology layer, the tracking sensitivity of each camera module in the first type of camera module topology layer is less than the tracking sensitivity of each camera module in the second type of camera module topology layer, and the tracking resources are differentiated by the hierarchical design for different dynamic requirements; when the behavior dynamic rate is greater than or equal to a preset dynamic rate threshold, the first switching camera module is determined by optimization in the second type of camera module topology layer by the camera switching module, or when the behavior dynamic rate is less than the preset dynamic rate threshold, the first switching camera module is determined by optimization in the first type of camera module topology layer, and the adaptive tracking of the camera with appropriate sensitivity is intelligently matched according to the dynamic change degree of the user behavior.

[0010] Through the above technical solution, based on the user behavior action acquisition and dynamic rate quantification, combined with the hierarchical topology structure and the intelligent matching mechanism, the dynamic tracking optimization of the camera with appropriate sensitivity is adaptively selected according to the change degree of the user behavior, and the technical effects of improving the tracking accuracy and optimizing the resource allocation are achieved. BRIEF DESCRIPTION OF DRAWINGS

[0011] Figure 1 The structure diagram of the camera dynamic tracking optimization system based on the topology structure provided by the application is shown.

[0012] Figure 2 The flowchart of the camera dynamic tracking optimization method based on the topology structure provided by the application is shown.

[0013] In the drawings, the components represented by the numbers are as follows:

[0014] The action acquisition module 11, the dynamic rate calculation module 12, the topology construction module 13 and the camera switching module 14. DETAILED DESCRIPTION

[0015] The technical solutions in the embodiments of the application will be clearly and completely described below with reference to the drawings in the embodiments of the application. Obviously, the described embodiments are only part of the embodiments of the application, not all the embodiments. Based on the embodiments in the application, all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the application.

[0016] In the description of this invention, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of the stated features. In the description of this invention, "a plurality of" means two or more, unless otherwise explicitly specified.

[0017] In the description of this invention, the term "for example" is used to mean "used as an example, illustration, or description." Any embodiment described as "for example" in this invention is not necessarily to be construed as being more preferred or advantageous than other embodiments. The following description is provided to enable any person skilled in the art to make and use the invention. Details are set forth in the following description for purposes of explanation. It should be understood that those skilled in the art will recognize that the invention can be made without using these specific details. In other instances, well-known structures and processes will not be described in detail to avoid obscuring the description of the invention with unnecessary detail. Therefore, the invention is not intended to be limited to the embodiments shown, but is consistent with the broadest scope of the principles and features disclosed herein.

[0018] Example 1, as Figure 1 As shown, this embodiment of the invention provides a camera dynamic tracking optimization system based on topology, including a motion acquisition module 11, a dynamic rate calculation module 12, a topology construction module 13, and a camera switching module 14.

[0019] The motion acquisition module 11 is used to detect the sound source user, acquire the motion of the sound source user, and obtain the motion set of the sound source user, which includes the head offset angular velocity, shoulder displacement amplitude and leg swing frequency.

[0020] Specifically, the sound source user refers to the user currently speaking or generating audio signals, who may be present in meetings, speeches, discussions, or other scenarios requiring camera tracking. The motion capture module 11 receives audio signals through multiple high-sensitivity microphones configured and distributed at different locations. It analyzes the time difference, phase difference, and amplitude difference of the audio signals received by each microphone to determine the three-dimensional spatial location of the sound source, obtaining the audio localization result. Then, combined with visual localization, the audio localization result is matched with the personnel position information captured by the camera to determine the current sound source user.

[0021] Subsequently, the action collection module 11 collects the actions of the sound source user through the camera device, thereby obtaining the action set of the sound source user. Specifically, image frames corresponding to the sound source user are continuously acquired at a preset interval time, for example, the interval time is thirty-three point three milliseconds, corresponding to a video frame rate of thirty frames per second. This frequency can effectively capture the subtle changes of human actions. Then, through the difference calculation and feature comparison between adjacent image frames, the action collection of the sound source user is realized, including the head offset angular velocity, the shoulder displacement amplitude, and the leg swing frequency, thereby forming the action set of the sound source user.

[0022] Among them, for the head offset angular velocity, the face key points of the sound source user are detected in the two adjacent image frames respectively, and the head offset angular velocity is calculated through the spatial distribution of the face key points. First, the face region of the sound source user is detected in the previous image frame, the accurate coordinates of the feature points such as the eye corner, the nose tip, and the mouth corner are identified through the face key point detection, and the attitude angle of the head relative to the camera coordinate system is calculated through the spatial distribution of these feature points; then, in the next image frame, the above process is repeated to obtain the attitude angle of the head relative to the camera coordinate system in this image frame; subsequently, the difference of the head attitude angles of the two image frames is compared, and the head offset angular velocity is determined combined with the known interval time.

[0023] For the shoulder displacement amplitude, the shoulder key points of the sound source user are detected in the two adjacent image frames respectively, and the shoulder displacement amplitude is calculated through the shoulder position change. First, the left and right shoulder key points of the sound source user are located through the human skeleton detection in the previous image frame, and the spatial coordinates of the shoulder center point relative to the camera coordinate system are calculated; then, in the next image frame, the above process is repeated to obtain the spatial coordinates of the shoulder center point relative to the camera coordinate system in this image frame; subsequently, the spatial distance between the shoulder center points in the two image frames is calculated, and the shoulder displacement amplitude is determined combined with the known interval time.

[0024] For the leg swing frequency, the lower limb key points of the sound source user are detected in the continuous multiple image frames, and the leg swing frequency is calculated through the leg motion trajectory analysis. First, the joint positions such as the knee and the ankle of the sound source user are continuously located through the lower limb key point detection in the continuous image frame sequence, and the motion trajectories of these key points relative to the camera coordinate system are recorded; then, the time series data of the leg motion is constructed within a preset time window (usually five to ten seconds); subsequently, the time series data is analyzed in the frequency domain, the main periodic characteristics of the leg motion are extracted, and the leg swing frequency is determined.

[0025] Through the above real-time acquisition process, the action acquisition module 11 can continuously output the action feature parameters of the sound source user, forming a sound source user action set containing the head offset angular velocity, shoulder displacement amplitude and leg swing frequency, to provide accurate data basis for subsequent camera dynamic tracking control.

[0026] The dynamic rate calculation module 12 is configured to calculate a behavior dynamic rate of the sound source user action set using a dynamic rate function, the behavior dynamic rate being an index representing the degree of change in user behavior.

[0027] Specifically, the dynamic rate calculation module 12 calculates the behavior dynamic rate representing the degree of change in user behavior based on the action parameters in the sound source user action set through a pre-defined dynamic rate function. The dynamic rate calculation module 12 converts the discrete action parameters (i.e. the head offset angular velocity, shoulder displacement amplitude and leg swing frequency in the sound source user action set) output by the action acquisition module 11 into a unified behavior dynamic rate index, providing a quantitative basis for subsequent camera control decisions.

[0028] The behavior dynamic rate is a comprehensive numerical index for quantitatively describing the behavior activity and change amplitude of the sound source user within a specific time period. The behavior dynamic rate is calculated by weighting the head offset angular velocity, shoulder displacement amplitude and leg swing frequency, and can reflect the behavior change trend of the user from static to dynamic and from stable to active. The higher the behavior dynamic rate, the more dramatic the change in user behavior, requiring more sensitive and rapid tracking response by the camera.

[0029] The dynamic rate calculation module 12 uses a pre-defined dynamic rate function to process the sound source user action set. The dynamic rate function takes the head offset angular velocity, shoulder displacement amplitude and leg swing frequency as input variables and calculates a comprehensive behavior dynamic rate through weighted linear combination. First, the action parameters (head offset angular velocity, shoulder displacement amplitude and leg swing frequency) are standardized to eliminate the dimensional differences between different parameters; then, the action parameters are weighted and summed according to pre-defined weight coefficients to obtain the behavior dynamic rate.

[0030] The calculation process of the dynamic rate function takes into account the contribution of different body part movements to the overall behavior assessment. The head offset angular velocity reflects the head transfer movement of the sound source user, the shoulder displacement amplitude reflects the activity state and posture adjustment of the upper body, and the leg swing frequency indicates the activity degree of the lower limbs. Through reasonable weight configuration, the dynamic rate function can accurately assess the overall behavior change degree of the user, providing a reliable reference for intelligent control of camera tracking.

[0031] Through the dynamic rate calculation module 12, the behavior dynamic rate of the sound source user can be output in real time, providing support for determining appropriate camera tracking and camera switching.

[0032] a topology construction module 13, configured to construct a camera module topology structure, the camera module topology structure comprising a first type of camera module topology layer and a second type of camera module topology layer, a tracking sensitivity of each camera module in the first type of camera module topology layer being less than a tracking sensitivity of each camera module in the second type of camera module topology layer.

[0033] In particular, the topology construction module 13 is responsible for establishing the network topology of all cameras in the camera system, organizing and configuring each camera module in a hierarchical management manner to achieve differentiated tracking control for different behavior dynamic rate scenarios. The topology construction module 13 divides all cameras into two different topology levels according to the performance characteristics and application requirements of the camera module, forming a hierarchical camera module topology structure.

[0034] The camera module topology structure is a camera network architecture formed after all cameras are classified and organized according to functional characteristics and technical parameters. The camera module topology structure adopts a double-layer design, including a first type of camera module topology layer and a second type of camera module topology layer. The camera modules in each topology layer have similar technical characteristics and application positioning, and the different topology layers have different functional divisions.

[0035] The first type of camera module topology layer contains camera modules with relatively low tracking sensitivity, which are mainly used to handle scenes with relatively slow user behavior changes. The camera modules in this layer usually have a wide field of view coverage, relatively slow response speed and low tracking accuracy, and are suitable for monitoring needs when the user is in a relatively static or low dynamic state. The cameras in the first type of camera module topology layer can provide stable image quality and basic tracking functions, meeting the requirements of conventional monitoring scenarios.

[0036] The second type of camera module topology layer contains camera modules with high tracking sensitivity, which are specially used to cope with high dynamic scenes with dramatic changes in user behavior. The camera modules in this layer have faster response speed, higher tracking accuracy and stronger dynamic adaptability, and can quickly capture and track the rapid movement of the user. The cameras in the second type of camera module topology layer are usually equipped with high-speed pan-tilt control systems to ensure good tracking effect in high dynamic environments. Therefore, the tracking sensitivity of each camera module in the first type of camera module topology layer is less than the tracking sensitivity of each camera module in the second type of camera module topology layer.

[0037] Through the construction of the camera module topology structure, intelligent selection can be performed between the first type of camera module topology layer and the second type of camera module topology layer according to the behavior dynamic rate index calculated in real time, the most suitable camera tracking is provided for user behaviors of different dynamic degrees, and thus the tracking precision is improved and resource allocation is optimized.

[0038] The camera switching module 14 is configured to determine the first switching camera module in the second type of camera module topology layer when the behavior dynamic rate is greater than or equal to the preset dynamic rate threshold, or determine the first switching camera module in the first type of camera module topology layer when the behavior dynamic rate is less than the preset dynamic rate threshold.

[0039] Specifically, the camera switching module 14 is responsible for judging the behavior dynamic rate output by the dynamic rate calculation module 12 through a preset dynamic rate threshold, and performing optimization in the corresponding camera module topology layer to determine the camera module most suitable for the current scene for switching, so as to realize intelligent selection and switching of the camera based on the behavior state of the sound source user.

[0040] The camera switching module 14 classifies and judges the behavior state of the sound source user in a threshold comparison manner. The preset dynamic rate threshold is a critical value pre-configured to distinguish the dynamic degree of the user behavior. When the behavior dynamic rate is greater than or equal to the preset dynamic rate threshold, it indicates that the sound source user is in a high dynamic state, the behavior changes dramatically, and a high-sensitivity camera (a camera module in the second type of camera module topology layer) needs to be used for tracking. When the behavior dynamic rate is less than the preset dynamic rate threshold, it indicates that the sound source user is in a low dynamic state, the behavior changes relatively slowly, and a regular-sensitivity camera (a camera module in the first type of camera module topology layer) can be used for monitoring.

[0041] For a high dynamic scene, when the behavior dynamic rate is greater than or equal to the preset dynamic rate threshold, the camera switching module 14 performs optimization operation in the second type of camera module topology layer to select the camera module most suitable for the current sound source user as the first switching camera module. Since the second type of camera module topology layer includes camera modules with high tracking sensitivity, these camera modules have fast response and high-precision tracking capabilities and can effectively cope with the rapid movement and frequent changes of the user. The camera switching module 14 determines the most suitable camera module as the first switching camera module by comprehensively evaluating the current state, coverage range, image quality and load condition of each camera module in the second type of topology layer.

[0042] For low dynamic scenes, when the behavior dynamic rate is less than a preset dynamic rate threshold, the camera switching module 14 performs optimization operation in the first type of camera module topology layer to select the camera module that is optimal for the current sound source user as the first switched camera module. Although the camera devices in the first type of camera module topology layer have relatively low tracking sensitivity, they can provide stable monitoring effect, while having lower power consumption and system overhead. The camera switching module 14 selects the most suitable camera module as the first switched camera module by evaluating the performance parameters and applicability of each camera module in the first type of topology layer, to realize energy-efficient tracking control.

[0043] The first switched camera module is the target camera device determined by the camera switching module 14 through optimization, which will take over the current camera module to track the sound source user. The optimization process considers different application requirements of the first type of camera module topology layer and the camera modules in the first type of camera module topology layer, and selects the most suitable first switched camera module for the current sound source user according to multiple evaluation indexes. The multiple evaluation indexes include, but are not limited to, the distance between the camera module and the sound source user, the angle of view coverage, the current load state, the image definition, and the switching cost, etc., so as to determine the camera module with the highest comprehensive score as the switching target.

[0044] Through the above intelligent judgment and hierarchical optimization mechanism based on the behavior dynamic rate, the camera switching module 14 can realize accurate selection of camera devices, ensure that the camera module most suitable for the current behavior state of the sound source user is always used for tracking, thereby improving the tracking effect and optimizing the resource utilization efficiency.

[0045] Further, the embodiment of the present application also includes a weight configuration module, which is configured to:

[0046] identify the current conference scene, and configure adaptive weights of the dynamic rate function according to the conference scene, including head movement weight, shoulder movement weight, and leg movement weight;

[0047] define the dynamic rate function using the adaptive weights, wherein the dynamic rate function includes a weighted linear combination of the head offset angular velocity, the shoulder displacement amplitude, and the leg swing frequency, and a nonlinear correction term.

[0048] In a preferred embodiment, the camera dynamic tracking optimization system further includes a weight configuration module responsible for adaptively adjusting the dynamic rate function according to different conference scenes to improve the accuracy and applicability of the behavior dynamic rate calculation. Through scene recognition and weight optimization, the weight configuration module realizes personalized configuration for specific conference scenes, ensuring that the dynamic rate function can accurately reflect the important features of user behavior in different conference scenes.

[0049] Firstly, the weight configuration module identifies the current conference scenario by analyzing multi-dimensional features such as environmental information, user number, spatial layout, and activity type, etc. It identifies the specific scenario type, such as presentation scenario, discussion scenario, workshop scenario, etc. The behavior patterns and focus of users in different scenarios are significantly different, and corresponding weight configuration strategies are needed. Then, based on the current conference scenario, the weight configuration module sets adaptive weights for the dynamic rate function according to the preset configuration rules. The adaptive weights include head motion weight, shoulder motion weight, and leg motion weight, which correspond to the importance of head offset angular velocity, shoulder displacement amplitude, and leg swing frequency in dynamic rate calculation, respectively. For example, in the presentation scenario, the head motion (such as eye movement, expression change) of the sound source user has higher importance for tracking effect, so the head motion weight is set to 0.6, while the shoulder motion weight and leg motion weight are set to 0.2 and 0.2, respectively. In the discussion scenario, the shoulder motion (such as hand gesture expression, body leaning forward) of the sound source user can better reflect its participation level and speaking intention, so the shoulder motion weight is increased to 0.5, and the head motion weight and leg motion weight are set to 0.3 and 0.2, respectively. In the workshop scenario, the leg motion (such as standing up and walking, position changing) is more frequent and important, so the leg motion weight is set to 0.5, and the head motion weight and shoulder motion weight are 0.2 and 0.3, respectively.

[0050] After that, the weight configuration module uses the adaptive weights corresponding to the current conference scenario to define the specific form of the dynamic rate function. The dynamic rate function adopts the basic structure of weighted linear combination, which combines the head offset angular velocity, shoulder displacement amplitude, and leg swing frequency according to the corresponding weight coefficients. At the same time, the dynamic rate function also contains a nonlinear correction term to handle the mutual influence and nonlinear relationship between the action parameters, improving the accuracy and robustness of dynamic rate calculation. The nonlinear correction term can better capture complex user behavior patterns by introducing cross terms, exponential terms, or other nonlinear function forms.

[0051] In addition, the weight configuration module also has a dynamic adjustment function based on historical behavior data. The weight configuration module analyzes the historical activity level and behavior characteristics of the sound source user to further optimize the adaptive weights. For example, for users with high activity level, the leg motion weight is multiplied by a gain coefficient of 1.5 times to better capture their frequent movement behavior; for users with low activity level, the head motion weight is multiplied by a gain coefficient of 1.3 times, focusing on their small attention changes.

[0052] Through the above adaptive weight configuration mechanism, the weight configuration module can ensure that the dynamic rate function can provide accurate and reliable behavior evaluation results in different scenarios and user characteristics, providing more accurate data support for the decision-making of the camera switching module.

[0053] Further, the embodiment of the present application further comprises a weight updating module, which is used to obtain real-time effectiveness detection after the sound source user action set is acquired, and specifically comprises:

[0054] determining whether any action point in the head offset angular velocity, shoulder displacement amplitude and leg swing frequency exists continuous frame occlusion, using an attenuation coefficient to perform weight attenuation on the data item existing occlusion, updating the adaptive weight;

[0055] and determining whether any action point in the head offset angular velocity, shoulder displacement amplitude and leg swing frequency exists action mutation, using a gain coefficient to perform weight gain on the data item existing action mutation, updating the adaptive weight.

[0056] In a preferred embodiment, the embodiment of the present application further comprises a weight updating module, which is responsible for real-time effectiveness detection on the sound source user action set obtained in the action acquisition process, and dynamically adjusts the adaptive weight based on the detection result, to ensure the accuracy and reliability of the dynamic rate calculation. The weight updating module monitors the quality condition of the action data in the sound source user action set in succession, identifies and processes data abnormal conditions in time, and maintains the stable operation of the camera tracking.

[0057] Specifically, the weight updating module starts the effectiveness detection process immediately after the action acquisition module 11 acquires the sound source user action set. The detection process is for the three action parameters of the head offset angular velocity, the shoulder displacement amplitude and the leg swing frequency. The effectiveness detection includes two main aspects of occlusion detection and action mutation detection, to ensure that the data input to the dynamic rate calculation module 12 has good reliability.

[0058] In the occlusion detection aspect, the weight updating module determines whether any action point in the head offset angular velocity, the shoulder displacement amplitude and the leg swing frequency exists continuous frame occlusion. Continuous frame occlusion refers to that in a plurality of continuous image frames, a certain key body part is occluded by other objects or persons, resulting in that the corresponding action parameter cannot be accurately detected and calculated. When it is detected that a certain action parameter exists continuous frame occlusion, the reliability of the parameter is significantly reduced, and the parameter should not occupy too high weight in the dynamic rate calculation. At this time, the weight updating module uses a preset attenuation coefficient to perform weight attenuation processing on the data item existing occlusion, reduces the influence degree of the parameter in the dynamic rate function by multiplying the original weight of the action parameter by an attenuation coefficient less than 1, thereby reducing the interference of unreliable data on the final result. For example, when it is detected that the head of the sound source user is occluded by the podium or other persons in continuous 5 image frames, the head action weight is reduced from the original 0.6 to 0.6x0.7=0.42, and the attenuation coefficient is 0.7, thereby reducing the influence of the inaccurate head offset angular velocity on the behavior dynamic rate calculation.

[0059] In the action mutation detection aspect, the weight updating module judges whether any of the head offset angular velocity, shoulder displacement amplitude and leg swing frequency has the action mutation phenomenon. The action mutation refers to that a certain action parameter has an abnormal sharp change in a short time, which is usually manifested as a sharp jump or non-continuous change of the value. Such mutation can reflect important changes of the user behavior, such as sudden turning, quickly standing up or emergency action, and has important indicative significance for camera tracking. When detecting that a certain action parameter has the action mutation, the weight updating module performs weight gain processing on the data item by using a preset gain coefficient, improves the weight of the parameter in the dynamic rate function by multiplying the original weight of the action parameter by a gain coefficient greater than 1, so that the system can respond more sensitively to important behavior changes. For example, when detecting that the leg swing frequency of the sound source user suddenly increases from 0.5 Hz to 2.1 Hz in a short time, it indicates that the user can be ready to stand up or quickly move, and the leg action weight is increased from the original 0.2 to 0.2 x 1.8 = 0.36, and the gain coefficient is 1.8, so that the system can respond more timely to the dynamic changes of the user.

[0060] Subsequently, the weight updating module updates the adaptive weight in real time based on the above detection results. The updating process adopts a dynamic adjustment strategy, and according to the current detected data quality condition and abnormal situation, the head action weight, shoulder action weight and leg action weight are correspondingly corrected. The updated adaptive weight will be transmitted to the dynamic rate calculation module 12 in real time, ensuring that the subsequent behavior dynamic rate calculation can be based on the latest weight configuration, thereby improving the accuracy and timeliness of the calculation results.

[0061] Through the above real-time effectiveness detection and weight updating mechanism, the weight updating module can effectively cope with various abnormal situations that can occur in the action acquisition process, and ensure that good tracking performance and decision accuracy can still be maintained in complex environments.

[0062] Further, the camera module topology structure includes a first type of camera module topology layer and a second type of camera module topology layer, and there is a topology connection relationship between the plurality of camera modules in each topology layer. The plurality of camera modules in each topology layer are controlled in linkage through switching priorities.

[0063] In a preferred embodiment, the camera module topology structure adopts a hierarchical architecture design, and by establishing the topology connection relationship and linkage control mechanism between the camera modules, not only a physical connection framework of the camera equipment is provided, but also a cooperation mechanism at the logical control level is established, realizing unified management and intelligent scheduling of the entire camera, and ensuring that each camera module can respond to the tracking demand in coordination.

[0064] Specifically, the camera module topology structure includes a first type of camera module topology layer and a second type of camera module topology layer. A complete topology connection relationship is established in each type of camera module topology layer. The topology connection relationship refers to the communication link and data exchange channel established between multiple camera modules in the same topology layer. Through these connection relationships, each camera module can share key parameters such as state information, position data, and load conditions in real time. In the first type of camera module topology layer, each low-sensitivity camera module forms a cooperative network through network connection, and can coordinate the coverage range and working state with each other. Similarly, in the second type of camera module topology layer, each high-sensitivity camera module also establishes a corresponding connection relationship, forming a cooperative network of high-performance tracking devices.

[0065] The multiple camera modules in each camera module topology layer realize the linkage control function through a preset switching priority. The switching priority is a priority order determined comprehensively according to the performance parameters, installation position, coverage range, and current state of the camera module, and is used to guide the decision-making process of the camera switching module 14 when selecting a target camera module. In the first type of camera module topology layer, the camera module with higher priority usually has a better coverage position, more stable image quality, or lower current load, and thus is preferentially selected in a low-dynamic scene. In the second type of camera module topology layer, the priority setting mainly considers the response speed, tracking accuracy, and dynamic adaptation ability of the camera module, to ensure that the most suitable high-performance device can be selected in a high-dynamic scene. The linkage control mechanism ensures that multiple camera modules in the same camera module topology layer can work coordinately, avoiding resource conflicts and repeated coverage. When a camera module is selected as the first switching camera module and starts to perform a tracking task, other camera modules in the camera module topology layer will adjust their working states accordingly to prepare for possible subsequent switching. Meanwhile, the camera modules in the adjacent area can also be activated in advance according to the moving track and predicted position of the sound source user, to realize predictive resource scheduling and seamless tracking switching.

[0066] Through the above topology connection relationship and switching priority mechanism, the camera modules in each camera module topology structure can form an efficient cooperative network, realize intelligent linkage control, and provide continuous, stable, and high-quality camera tracking services for the sound source user.

[0067] Further, the camera switching module 14 is further configured to:

[0068] predict a moving area of the sound source user according to the sound source user action set, and obtain a candidate camera module under the moving area in the second type of camera module topology layer;

[0069] The multi-objective scoring function is used to score each of the candidate camera modules, to obtain a scoring index set, and a first switching camera module is determined from the candidate camera modules according to the scoring index set, wherein the multi-objective scoring includes a view angle range, a zoom response speed, an image definition score, and a current load state;

[0070] A switching communication protocol is established between the first switching camera module and the current camera module, and the sound source user is photographed by the first switching camera module according to the switching communication protocol.

[0071] In a preferred embodiment, when the behavior dynamic rate is greater than or equal to a preset dynamic rate threshold, the camera switching module 14 performs an optimization determination process in the second type of camera module topology layer to determine the first switching camera module, so as to accurately select and smoothly switch the optimal camera module, thereby ensuring that the camera tracking in a high dynamic scene can quickly respond to changes in user behavior and provide high-quality tracking effects.

[0072] First, the camera switching module 14 analyzes the future movement trend of the user according to the sound source user action set. The camera switching module 14 uses the head offset angular velocity, shoulder displacement amplitude, and leg swing frequency as three action parameters to calculate the possible movement direction and target area of the sound source user through motion trajectory analysis and behavior pattern recognition. The prediction process takes into account the user's current position, motion speed, acceleration, and historical movement patterns to calculate one or more possible movement areas. Subsequently, the camera switching module 14 filters out camera modules that can cover these movement areas in the second type of camera module topology layer to form candidate camera modules. The candidate camera module refers to a high-sensitivity camera device that has a field of view coverage capability within the predicted movement area range.

[0073] Next, the camera switching module 14 uses a multi-objective scoring function to evaluate the comprehensive performance of each camera module in the candidate camera module. The multi-objective scoring function is an evaluation system that considers multiple performance indicators to quantitatively compare the applicability and degree of excellence of different camera modules. The multi-objective scoring function includes four core scoring dimensions, namely the view angle range, the zoom response speed, the image definition score, and the current load state. The view angle range is used to evaluate the coverage degree and field of view breadth of the camera module for the target area; the zoom response speed reflects the fast response ability of the camera module to adjust the focal length and view angle; the image definition measures the imaging quality and resolution level of the camera module under the current environmental conditions; and the current load state considers the workload and availability of the camera module. By quantitatively scoring the performance of each candidate camera module in these four dimensions, a scoring index set containing the numerical values of each indicator is obtained.

[0074] Subsequently, based on the set of scoring indicators, the camera switching module 14 determines a first switching camera module among the candidate camera modules. This process integrates the values of each candidate camera module in each scoring dimension of the set of scoring indicators into an overall score through weighted summation or other multi-objective decision-making methods, and selects the camera module with the highest overall score as the optimal switching target, as the first switching camera module. The first switching camera module is the most suitable camera equipment for the current high dynamic scene after comprehensive evaluation, with advantages such as fast response, high-quality imaging, and good coverage capability.

[0075] After determining the first switching camera module, the camera switching module 14 establishes a switching communication protocol between the first switching camera module and the currently working camera module. The switching communication protocol is a communication specification and data exchange standard for coordinating the handover between the two camera modules, including state information synchronization, tracking parameter transfer, image data switching, and control right transfer. Through the switching communication protocol, the current camera module transfers key data such as the position information of the sound source user, the tracking state, and the image parameters to the first switching camera module, ensuring the continuity and consistency of the tracking process. Then, the camera switching module 14 performs camera module switching operation according to the switching communication protocol, and transfers the tracking task from the current camera module to the first switching camera module. At the same time, the switching process adopts seamless switching, ensuring that there is no tracking interruption or image loss phenomenon during the switching period, and the first switching camera module immediately takes over the camera tracking task of the sound source user, providing continuous and stable tracking service for high dynamic scenes.

[0076] Further, the camera switching module 14 is also used for:

[0077] After determining the first switching camera module in the second type of camera module topology layer, the adaptive communication protocol of the first switching camera module is obtained, and the adaptive communication protocol includes the IP network protocol and the RS232 / RS485 serial port protocol;

[0078] According to the adaptive communication protocol, control instructions are issued to the first switching camera module, the camera data set of the sound source user is obtained, and the camera data set of the sound source user is stored in the data center storage module.

[0079] In a preferred embodiment, when the camera switching module 14 completes the optimization determination of the first switching camera module in the second type of camera module topology layer, the camera switching module 14 continues to perform the communication protocol configuration and data acquisition and storage functions, ensuring effective connection with the selected first switching camera module and reliable acquisition of data.

[0080] Firstly, the camera switching module 14 acquires the adaptive communication protocol of the first switching camera module. The adaptive communication protocol refers to the communication standard and interface specification required for data communication and control interaction with a specific camera module. Since different models and manufacturers of camera modules may use different communication methods, the corresponding communication protocol needs to be identified and configured to ensure compatibility. The adaptive communication protocol mainly includes two types: IP network protocol and RS232 / RS485 serial port protocol. Among them, the IP network protocol is suitable for modern camera modules based on Ethernet connection, supports high-speed data transmission and complex control instructions, has good scalability and remote access capability. The RS232 / RS485 serial port protocol is suitable for traditional serial communication camera devices, and realizes basic control and data transmission functions through serial port connection, with stable communication and strong anti-interference capability.

[0081] Subsequently, the camera switching module 14 issues corresponding control instructions to the first switching camera module according to the acquired adaptive communication protocol. The control instructions include the start command of the camera module, tracking parameter setting, image acquisition configuration and data transmission requirements, etc. For camera modules using IP network protocol, send standardized control messages through network sockets, including target position coordinates, tracking accuracy requirements, image resolution settings, etc. For camera modules using RS232 / RS485 serial port protocol, send control byte sequences according to the corresponding serial port communication format to realize basic control functions such as pan-tilt rotation, focal length adjustment and shooting start.

[0082] After the control instructions are successfully issued and the first switching camera module responds, the camera switching module 14 begins to acquire the camera data set of the sound source user. The camera data set refers to the complete data set generated by the first switching camera module after tracking and shooting the sound source user, including high-definition image sequence, video stream data, timestamp information, camera parameter record and tracking state data, etc. At the same time, the camera switching module 14 transmits and stores the acquired camera data set of the sound source user to the data center storage module. The data center storage module is a centralized data management unit responsible for unified storage, index management and backup protection of camera data. The storage process uses standardized data format and file naming rules to ensure data retrievability and integrity. At the same time, generate corresponding metadata information for each camera data set, including acquisition time, camera module identifier, sound source user information and data quality evaluation, etc., to facilitate subsequent data query and analysis application.

[0083] Through the above communication protocol configuration, control instruction issuance and data acquisition and storage process, the camera switching module 14 realizes complete working interface with the first switching camera module, ensuring smooth execution of camera tracking tasks and effective management of data.

[0084] Further, the topology construction module 13 is also used for:

[0085] acquiring the inputted conference personnel information;

[0086] grouping the inputted conference personnel information to obtain a plurality of conference groups, and configuring a plurality of camera module topologies for the plurality of conference groups, wherein each camera module topology comprises a first type of camera module topology layer and a second type of camera module topology layer;

[0087] integrating the plurality of camera module topologies to obtain an integrated camera module topology, and synchronously photographing the plurality of conference groups by the integrated camera module topology through a camera synchronization module.

[0088] In a preferred embodiment, the topology construction module 13 also has a multi-group camera management function, which can handle complex multi-person conference scenarios. Through group division and topology integration, it realizes the coordinated management and synchronous camera control of multiple conference groups, thereby expanding the application range of the embodiment and making it suitable for camera tracking needs in complex scenarios such as large conferences, seminars, and multi-group discussions.

[0089] Specifically, first, the topology construction module 13 acquires pre-inputted conference personnel information. Conference personnel information refers to the basic data and related attributes of all personnel participating in the conference, including personnel name, position role, seat location, participation in topics, and behavior characteristics. This conference personnel information is obtained through a conference management system, a personnel registration system, or manual input, providing basic data support for subsequent group division and camera configuration. Conference personnel information can also include historical behavior patterns, speaking frequency, and activity preferences, which help to make more accurate group analysis and camera strategy formulation.

[0090] Subsequently, based on the acquired conference personnel information, the topology construction module 13 performs group division to group all participants into a plurality of conference groups. The group division process considers multiple dimensions, including spatial location distribution of personnel, role function relationship, topic participation, and interaction frequency. For example, personnel can be divided into different discussion groups according to seat areas, or divided into speaker groups, listener groups, and moderator groups according to functional roles. Each conference group represents a relatively independent camera tracking system with specific behavior patterns and tracking needs.

[0091] Then, the topology construction module 13 configures a corresponding camera module topology structure for each determined conference group, obtaining multiple camera module topology structures. Each camera module topology structure adopts a double-layer architecture design, including a first type of camera module topology layer and a second type of camera module topology layer, which are respectively used to process low dynamic and high dynamic behavior scenes of users in the group. The camera module topology structures of different conference groups may differ in the number of camera modules, coverage range, and performance configuration, etc., to adapt to the specific needs of each group. For example, the topology structure of the presenter group may be configured with more high-sensitivity camera modules, while the topology structure of the audience group focuses on large-scale coverage monitoring.

[0092] After completing the topology structure configuration of each conference group, the topology construction module 13 integrates the multiple camera module topology structures to form a unified integrated camera module topology structure, ensuring that the camera tracking of each conference group can work cooperatively and avoiding resource conflicts and overlapping coverage. The integrated camera module topology structure is a composite network architecture containing multiple camera module topology structures, with global resource scheduling and unified management capabilities. The integrated camera module topology structure realizes the synchronous camera function for multiple conference groups, coordinates the synchronous data acquisition and unified storage management of each conference group, and ensures that the camera activities of different conference groups remain consistent in time and space, thereby simultaneously tracking multiple conference groups and realizing parallel multi-target camera tracking, providing comprehensive visual recording and analysis support for complex conference scenarios.

[0093] Through the above multi-group management and integrated topology architecture, the topology construction module 13 expands the application capability of camera tracking, enabling it to handle large-scale, multi-level conference camera tracking requirements and providing support for conference camera tracking.

[0094] Embodiment two, as shown in the figure, the present application embodiment also provides a kind of camera dynamic tracking optimization method based on topology structure, comprising: Figure 2

[0095] detecting sound source users, collecting actions of the sound source users, obtaining a set of sound source user actions, which includes head offset angular velocity, shoulder displacement amplitude and leg swing frequency;

[0096] calculating the behavior dynamic rate of the set of sound source user actions using a dynamic rate function, which is an index representing the degree of change in user behavior;

[0097] constructing a camera module topology structure, which includes a first type of camera module topology layer and a second type of camera module topology layer, and the tracking sensitivity of each camera module in the first type of camera module topology layer is less than that in the second type of camera module topology layer; ​

[0098] determining a first switching camera module in the second type of camera module topology layer when the behavior dynamic rate is greater than or equal to a preset dynamic rate threshold, or determining a first switching camera module in the first type of camera module topology layer when the behavior dynamic rate is less than the preset dynamic rate threshold.

[0099] Further, the embodiments of the present application further include:

[0100] identifying a current conference scene, and configuring adaptive weights of the dynamic rate function according to the conference scene, including head movement weight, shoulder movement weight, and leg movement weight;

[0101] defining the dynamic rate function using the adaptive weights, the dynamic rate function including a weighted linear combination of the head angular velocity offset, the shoulder displacement amplitude, and the leg swing frequency, and a nonlinear correction term.

[0102] Further, after obtaining the sound source user action set, real-time effectiveness detection is performed, specifically including:

[0103] determining whether any action point in the head angular velocity offset, the shoulder displacement amplitude, and the leg swing frequency exists continuous frame occlusion, using a decay coefficient to weight decay the data item with occlusion, and updating the adaptive weights;

[0104] and determining whether any action point in the head angular velocity offset, the shoulder displacement amplitude, and the leg swing frequency exists action mutation, using a gain coefficient to weight gain the data item with action mutation, and updating the adaptive weights.

[0105] Further, the camera module topology structure includes a first type of camera module topology layer and a second type of camera module topology layer, and there is a topology connection relationship between multiple camera modules in each topology layer, and multiple camera modules in each topology layer are linked and controlled through switching priority.

[0106] Further, when the behavior dynamic rate is greater than or equal to a preset dynamic rate threshold, a first switching camera module is determined in the second type of camera module topology layer, including:

[0107] predicting a movement area of the sound source user according to the sound source user action set, and obtaining a candidate camera module under the movement area in the second type of camera module topology layer;

[0108] performing multi-objective scoring on each camera module in the candidate camera module through a multi-objective scoring function to obtain a scoring index set, and determining a first switching camera module in the candidate camera module using the scoring index set, wherein multi-objective scoring includes view angle range, zoom response speed, image clarity score, and current load state.

[0109] establish a switching communication protocol between the first switching camera module and a current camera module, and switch to the first switching camera module to perform camera shooting on the sound source user according to the switching communication protocol.

[0110] Further, after the first switching camera module is determined by optimization in the second type of camera module topology layer, an adaptive communication protocol of the first switching camera module is obtained, and the adaptive communication protocol includes an IP network protocol and an RS232 / RS485 serial port protocol.

[0111] According to the adaptive communication protocol, a control instruction is sent to the first switching camera module, camera data sets of the sound source user are obtained, and the camera data sets of the sound source user are stored in a data center storage module.

[0112] Further, the camera module topology structure includes:

[0113] Meeting participant information inputted is obtained.

[0114] The inputted meeting participant information is divided into groups, a plurality of meeting groups are obtained, and a plurality of camera module topology structures are configured for the plurality of meeting groups, wherein each camera module topology structure includes a first type of camera module topology layer and a second type of camera module topology layer.

[0115] The plurality of camera module topology structures are integrated and connected to obtain an integrated camera module topology structure, and the integrated camera module topology structure performs synchronous camera shooting on the plurality of meeting groups through a camera synchronization module.

[0116] It should be noted that in the above embodiments, the description of each embodiment has its own emphasis, and the parts not described in detail in a certain embodiment can be referred to the related description of other embodiments.

[0117] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can be in the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can be in the form of a computer program product implemented on one or more computer usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer usable program code.

[0118] The computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process such that the instructions which execute on the computer or other programmable apparatus provide steps for implementing the functions specified in the flowchart block or blocks. Figure 1 one or more flowcharts and / or blocks in the flowcharts and / or blocks of the flowcharts. Figure 1 one or more flowcharts and / or blocks in the flowcharts and / or blocks of the flowcharts.

[0119] These computer program instructions can also be stored in a computer- readable memory that can direct a computer or other programmable data processing apparatus to function in a particular manner, such that the instructions stored in the computer-readable memory produce an article of manufacture including instructions which implement the function specified in the flowchart block or blocks. Figure 1 one or more flowcharts and / or blocks in the flowcharts and / or blocks of the flowcharts. Figure 1 one or more flowcharts and / or blocks in the flowcharts and / or blocks of the flowcharts.

[0120] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process such that the instructions which execute on the computer or other programmable apparatus provide steps for implementing the functions specified in the flowchart block or blocks. Figure 1 one or more flowcharts and / or blocks in the flowcharts and / or blocks of the flowcharts. Figure 1 one or more flowcharts and / or blocks in the flowcharts and / or blocks of the flowcharts.

[0121] While the preferred embodiments of the application have been described, additional variations and modifications can be made to the embodiments described and shown, without departing from the spirit and scope of the application.

[0122] It will be apparent to those skilled in the art that various modifications and variations can be made to the present application without departing from the spirit or scope of the application. Thus, it is intended that the present application cover the modifications and variations of this application and its equivalent.

Claims

1. A topology-based camera dynamic tracking optimization system, characterized in that, The system comprises: An action collection module for detecting a sound source user, collecting actions of the sound source user, and obtaining a sound source user action set, wherein the sound source user action set comprises head offset angular velocity, shoulder displacement amplitude, and leg swing frequency; A dynamic rate calculation module for calculating a behavior dynamic rate of the sound source user action set by using a dynamic rate function, wherein the behavior dynamic rate is an index representing the degree of change in user behavior; A topology construction module for constructing a camera module topology structure, wherein the camera module topology structure comprises a first type of camera module topology layer and a second type of camera module topology layer, the tracking sensitivity of each camera module in the first type of camera module topology layer is less than the tracking sensitivity of each camera module in the second type of camera module topology layer, there is a topology connection relationship between the plurality of camera modules in each topology layer, and the plurality of camera modules in each topology layer are controlled in linkage by switching priorities; A camera switching module for determining a first switching camera module in the second type of camera module topology layer when the behavior dynamic rate is greater than or equal to a preset dynamic rate threshold, or determining a first switching camera module in the first type of camera module topology layer when the behavior dynamic rate is less than the preset dynamic rate threshold; When the behavior dynamic rate is greater than or equal to the preset dynamic rate threshold, the method for determining the first switching camera module in the second type of camera module topology layer comprises: predicting a movement area of the sound source user according to the sound source user action set, and obtaining a candidate camera module under the movement area in the second type of camera module topology layer; performing multi-objective scoring on each camera module in the candidate camera module by using a multi-objective scoring function to obtain a scoring index set, and determining the first switching camera module in the candidate camera module according to the scoring index set, wherein the multi-objective scoring comprises view angle range, zoom response speed, image definition score, and current load state; establishing a switching communication protocol between the first switching camera module and a current camera module, and switching to the first switching camera module to perform camera shooting on the sound source user according to the switching communication protocol.

2. The system of claim 1, wherein, The system further comprises a weight configuration module, which is configured to: identify a current conference scene, and configure adaptive weights of the dynamic rate function according to the conference scene, including head action weight, shoulder action weight, and leg action weight; define the dynamic rate function by using the adaptive weights, wherein the dynamic rate function comprises a weighted linear combination of the head offset angular velocity, the shoulder displacement amplitude, and the leg swing frequency, and a nonlinear correction term.

3. The system of claim 2, wherein, The system further comprises a weight updating module, which is configured to perform effectiveness detection in real time after obtaining the sound source user action set, and specifically comprises: determining whether any action point in the head offset angular velocity, the shoulder displacement amplitude, and the leg swing frequency is continuously blocked, performing weight attenuation on the data item with blocking by using an attenuation coefficient, and updating the adaptive weights. And determine whether any of the head offset angular velocity, shoulder displacement amplitude and leg swing frequency motion point exists motion mutation, using gain coefficient to weight gain the data item which exists motion mutation, update the adaptive weight.

4. The system of claim 1, wherein, The camera switching module is further used for: After the first switching camera module is determined by optimization in the second type of camera module topology layer, an adaptive communication protocol of the first switching camera module is obtained, and the adaptive communication protocol includes an IP network protocol and an RS232 / RS485 serial port protocol. According to the adaptive communication protocol, a control instruction is issued to the first switching camera module, the camera data set of the sound source user is obtained, and the camera data set of the sound source user is stored in the data center storage module.

5. The system of claim 1, wherein, The topology construction module is further used for: Obtaining the entered conference personnel information; Group division is performed on the entered conference personnel information, a plurality of conference groups are obtained, and a plurality of camera module topologies are configured for the plurality of conference groups, wherein each camera module topology includes a first type of camera module topology layer and a second type of camera module topology layer. The plurality of camera module topologies are integrated and connected to obtain an integrated camera module topology, and the integrated camera module topology synchronously photographs the plurality of conference groups through a camera synchronization module.

6. A method for dynamic tracking optimization of a camera based on topology, characterized in that, The method comprises: Detecting a sound source user, collecting the motion of the sound source user, and obtaining a sound source user motion set, wherein the sound source user motion set includes a head offset angular velocity, a shoulder displacement amplitude and a leg swing frequency; A behavior dynamic rate of the sound source user motion set is calculated using a dynamic rate function, and the behavior dynamic rate is an index representing the degree of change of user behavior. A camera module topology structure is constructed, the camera module topology structure includes a first type of camera module topology layer and a second type of camera module topology layer, the tracking sensitivity of each camera module in the first type of camera module topology layer is less than the tracking sensitivity of each camera module in the second type of camera module topology layer, there is a topology connection relationship between the plurality of camera modules in each topology layer, and the plurality of camera modules in each topology layer are linked and controlled through switching priority. When the behavior dynamic rate is greater than or equal to a preset dynamic rate threshold, a first switching camera module is determined by optimization in the second type of camera module topology layer, or when the behavior dynamic rate is less than the preset dynamic rate threshold, a first switching camera module is determined by optimization in the first type of camera module topology layer. When the behavior dynamic rate is greater than or equal to a preset dynamic rate threshold, a first switching camera module is determined by optimization in the second type of camera module topology layer, including: According to the sound source user motion set, the moving area of the sound source user is predicted, and a candidate camera module under the moving area is obtained in the second type of camera module topology layer. Each of the candidate camera modules is scored by a multi-objective scoring function to obtain a scoring index set, and a first switching camera module is determined from the candidate camera modules based on the scoring index set, wherein the multi-objective scoring includes a field of view range, a zoom response speed, an image definition score, and a current load state; A switching communication protocol is established between the first switching camera module and the current camera module, and the sound source user is photographed by the first switching camera module based on the switching communication protocol.

Citation Information

Patent Citations

  • Operation device, tracking system, operation method, and program

    CN107211090A

  • Photography automatic identification lens switching system

    CN107483827A