A visual intercom system for personnel command management

By constructing a line-of-sight gravity field model and heterogeneous modal scheduling, and dynamically controlling the data transmission mode, the problems of image blurring and monitoring blind spots in video intercom systems under bandwidth constraints were solved, and a smooth switching between high-definition video streams in key areas and global situational awareness was achieved.

CN121462713BActive Publication Date: 2026-03-27XIAN XUYANG COMM EQUIP CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-01-06
Publication Date
2026-03-27

AI Technical Summary

Technical Problem

In large-scale personnel command and management scenarios, existing video intercom systems cannot effectively guarantee real-time high-definition video transmission in key areas when network bandwidth resources are limited, resulting in blurry or choppy images. At the same time, bandwidth resources in non-focused areas are wasted or connections are interrupted, causing monitoring blind spots and missed alarms.

Method used

The potential energy parameter acquisition module acquires the event entropy values ​​of the commander's focus and the front-end devices, constructs a line-of-sight gravity field model to calculate the potential energy value, adopts a heterogeneous modal scheduling module to dynamically control the data transmission mode, and combines a virtual-real fusion rendering module to realize hierarchical rendering and virtual avatar reconstruction at the command end, ensuring the transmission of high-definition video streams in key areas and semantic features in non-key areas.

Benefits of technology

It enables lossless real-time video streaming in critical areas under bandwidth-constrained conditions, maintains overall situational awareness, avoids monitoring blind spots, ensures zero-latency delivery of tactical information, and provides a smooth visual experience during network fluctuations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121462713B_ABST
    Figure CN121462713B_ABST
Patent Text Reader

Abstract

The application relates to the fields of multimedia communication and intelligent command and dispatch technology, in particular to a visual intercom system for personnel command management, which comprises a potential parameter acquisition module, a line-of-sight potential calculation module, a heterogeneous modal scheduling module, a virtual-real fusion rendering module and a preset 3D virtual avatar model. The potential parameter acquisition module is used for collecting the physical state mutation degree of the environment where the equipment is located. The line-of-sight potential calculation module is used for calculating the real-time potential value of each front-end acquisition equipment. The heterogeneous modal scheduling module is configured to dynamically control the data transmission mode of each front-end acquisition equipment according to the comparison result of the real-time potential value and the preset potential threshold value. The virtual-real fusion rendering module performs hierarchical rendering on the command terminal display interface. The received semantic feature data drives the preset 3D virtual avatar model to perform action reconstruction. The application realizes lossless transmission of real-time video streams and effectively solves the problem of blurred or stalled key tactical pictures under resource constraints.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of multimedia communication and intelligent command and dispatch technology, in particular to a visual intercom system for personnel command and management. BACKGROUND

[0002] In the current large-scale personnel command and management application scenario, the command center needs to face the concurrent access requests of a large number of front-end collection devices, and the actual communication environment is often in a state of limited or fluctuating network bandwidth resources;

[0003] In order to guarantee the real-time return of the picture, the existing visual intercom system generally adopts a rigid transmission mode of average allocation of code rate, that is, the transmission bandwidth is allocated to all front-end nodes without distinction. This processing method ignores the unevenness of information value in the command scene, and when the total network bandwidth is insufficient, the area that the command personnel pay attention to or the area where a sudden high-risk event occurs cannot obtain enough coding resources, resulting in picture blur, lag or even delay; at the same time, for a large number of background areas in a non-attention state, either the valuable bandwidth resources are wasted, or the connection is directly interrupted under the extreme flow limiting strategy, resulting in the complete loss of the situational awareness ability of the edge node, and simply relying on manual polling is easy to produce monitoring blind area, causing key police information to be missed; therefore, how to construct a dynamic resource scheduling mechanism that combines subjective attention graph and objective environmental mutation degree under the condition of scarce bandwidth resources, which can not only ensure the high-fidelity zero-delay delivery of key tactical information, but also maintain the basic situational awareness of the whole domain at a very low bandwidth cost, has become a technical problem to be solved. SUMMARY

[0004] To solve the above technical problems, the present application provides a visual intercom system for personnel command and management, in particular, the technical scheme of the present application comprises:

[0005] A potential parameter acquisition module is configured to acquire the focus point coordinates of the command end input device in real time, and synchronously receive the event entropy values uploaded by each front-end collection device, the event entropy values representing the physical state mutation degree of the environment where the front-end collection device is located;

[0006] A line-of-sight potential calculation module is configured to build a line-of-sight gravity field model, calculate the line-of-sight distance between the focus point coordinates and the mapping positions of each front-end collection device on the display interface, and based on the weighted calculation of the line-of-sight distance and the event entropy value, solve the real-time potential value of each front-end collection device;

[0007] a heterogeneous modal scheduling module, configured to dynamically control data transmission modalities of the front-end collection devices according to a comparison result of the real-time potential energy value and a preset potential energy threshold; if the real-time potential energy value is greater than the preset potential energy threshold, a full-pixel transmission instruction is generated; if the real-time potential energy value is less than or equal to the preset potential energy threshold, a semantic feature transmission instruction is generated;

[0008] a virtual-real fusion rendering module, configured to perform hierarchical rendering on a command terminal display interface in response to data transmitted by the front-end collection devices; for a region executing the full-pixel transmission instruction, a real-time video stream is rendered; for a region executing the semantic feature transmission instruction, a pre-set 3D virtual avatar model is driven to perform action reconstruction based on received semantic feature data, and the reconstructed virtual image and the real-time video stream are superimposed and displayed in the same user interface coordinate system.

[0009] Preferably, the potential energy parameter collection module comprises:

[0010] a focus capture unit, configured to collect a gaze point coordinate of the commander on a display screen as the focus point coordinate at a preset sampling frequency through a connected eye movement tracking sensor or a cursor control device;

[0011] an entropy value quantification unit, configured to receive multi-dimensional sensor data uploaded by the front-end collection devices, wherein the multi-dimensional sensor data comprises audio decibel data, acceleration mutation data and biological sign data, and the multi-dimensional sensor data is weighted and normalized by using a preset weight coefficient to generate the event entropy value.

[0012] Preferably, the line-of-sight potential energy calculation module performs the following processing:

[0013] determining a current display coordinate of each of the front-end collection devices on the command terminal display interface;

[0014] calculating an Euclidean distance between the focus point coordinate and the current display coordinate;

[0015] processing the Euclidean distance based on a function relationship that the value is inversely proportional to the distance to obtain a basic attention degree, and weighting and summing the basic attention degree and the event entropy value based on a preset proportion coefficient to obtain the real-time potential energy value, wherein the smaller the Euclidean distance or the greater the event entropy value, the higher the real-time potential energy value.

[0016] Preferably, the heterogeneous modal scheduling module performs the following control logic when generating the semantic feature transmission instruction:

[0017] sending a video stream blocking signal to the corresponding front-end collection device to control it to stop encoding and transmission of image pixel data;

[0018] Activating the end-side computing unit built-in the front-end acquisition device, performing structured analysis on the real-time collected picture, extracting human body skeleton key point data and geographical position coordinate data of the target object;

[0019] Establishing a low-bandwidth transmission channel, sending the human body skeleton key point data and the geographical position coordinate data as semantic feature data to the command end, while maintaining the transmission of the event entropy value or the multi-dimensional sensor data used to generate the event entropy value.

[0020] Preferably, the heterogeneous modal scheduling module performs the following control logic when generating the full-pixel transmission instruction:

[0021] Sending a high-bandwidth occupation permission to the corresponding front-end acquisition device to control it to enable a high-definition video encoder;

[0022] Establishing a high-bandwidth transmission channel to send real-time video stream data without dimension reduction processing to the command end;

[0023] When the total network bandwidth is limited, preferentially allocating network resources to the front-end acquisition device executing the full-pixel transmission instruction, and dynamically compressing the channel bandwidth executing the semantic feature transmission instruction.

[0024] Preferably, the virtual-real fusion rendering module comprises:

[0025] A model mapping unit configured to pre-store a 3D virtual avatar model bound to each of the front-end acquisition devices;

[0026] A skeleton driving unit configured to analyze the human body skeleton key point data in the received semantic feature data, map the human body skeleton key point data to the skeleton nodes of the 3D virtual avatar model, and real-time solve and drive the 3D virtual avatar model to produce limb movements synchronized with the front-end target object;

[0027] A scene synthesis unit configured to render the driven 3D virtual avatar model into a corresponding background area of the command end display interface according to the geographical position coordinate data in the semantic feature data.

[0028] Preferably, the system further comprises:

[0029] A smooth switching module configured to monitor the state jump of the real-time potential energy value across the preset potential energy threshold;

[0030] If it is detected that the state jumps from low potential energy to high potential energy, the rendering display of the 3D virtual avatar model is maintained before the first frame of real-time video stream is received, and after the video stream buffering is completed, a fade-in switching from virtual image to real video is performed;

[0031] If the state is detected to jump from high potential energy to low potential energy, semantic feature extraction is immediately started, and the rendering picture of the 3D virtual avatar model is spliced at the next frame of the video stream truncation.

[0032] Preferably, the system further comprises:

[0033] A threshold adaptive adjustment module configured to monitor the total available bandwidth load of the current network environment in real time;

[0034] If the total available bandwidth load is lower than the preset congestion warning line, the preset potential energy threshold is increased to reduce the number of front-end acquisition devices entering the full-pixel transmission mode, so that the non-focus area enters the semantic feature transmission mode.

[0035] If the total available bandwidth load returns to the preset normal range, the preset potential energy threshold is restored to the initial default value.

[0036] Compared with the prior art, the present application has the following beneficial effects:

[0037] 1. The present application constructs a potential energy model based on the line-of-sight gravitational field and event entropy value, breaking the rigid mode of traditional average allocation of code rate; by converting bandwidth allocation into energy allocation, the high potential energy area such as the commander's focus point or high-risk event point is given priority to obtain network resources, realizing lossless transmission of real-time video stream, and effectively solving the problem of blurred or stuttering key tactical pictures under resource limitation;

[0038] 2. The present application adopts a heterogeneous modal scheduling strategy, only transmits semantic feature data such as skeleton and position, and drives the virtual avatar model to reconstruct the action by using the computing power of the command end; this way compresses the transmission bandwidth to the semantic level, greatly saves network resources, maintains basic perception of the position and posture of the edge node, and avoids the monitoring blind area caused by direct disconnection;

[0039] 3. The present application introduces a multi-dimensional sensor entropy value quantization mechanism, giving the front-end device the initiative to seek help; even in the case where the commander is not paying attention, the dramatic mutation of the environmental physical state can also force the node potential energy to be increased by the explosive entropy value, automatically triggering the high-bandwidth transmission channel; this design makes up for the limitations of simply relying on subjective line of sight, ensuring that high-risk events can break through the distance limit and be captured in real time;

[0040] 4. The visual intercom system for personnel command management fills the video stream buffer empty window period with local virtual image through virtual-real fusion rendering and smooth switching technology, eliminating the black screen or flicker during modal switching; cooperating with the bandwidth load adaptive adjustment threshold mechanism, the system not only dynamically adjusts the access quantity to prevent congestion shock according to network fluctuations, but also ensures the continuity and smoothness of the visual experience of the commander when frequently switching focus points. BRIEF DESCRIPTION OF DRAWINGS

[0041] The application will be further explained in connection with the accompanying drawings and embodiments:

[0042] Figure 1 is a structural diagram of the system of the application. DETAILED DESCRIPTION

[0043] To make the objectives, technical solutions and advantages of the application clearer, the application will be further described in detail below with specific embodiments.

[0044] Example 1:

[0045] Please refer to Figure 1 A visual intercom system for personnel command management, comprising:

[0046] A potential parameter acquisition module configured to acquire a focus point coordinate of a command terminal input device in real time, and synchronously receive an event entropy value uploaded by each front-end acquisition device, the event entropy value representing a degree of physical state mutation of an environment in which the front-end acquisition device is located;

[0047] A line-of-sight potential calculation module configured to construct a line-of-sight gravitational field model, calculate a line-of-sight distance between the focus point coordinate and a mapping position of each front-end acquisition device on a display interface, and based on a weighted calculation of the line-of-sight distance and the event entropy value, solve a real-time potential value of each front-end acquisition device;

[0048] A heterogeneous modal scheduling module configured to dynamically control a data transmission mode of each front-end acquisition device according to a comparison result of the real-time potential value and a preset potential threshold; if the real-time potential value is greater than the preset potential threshold, a full-pixel transmission instruction is generated; if the real-time potential value is less than or equal to the preset potential threshold, a semantic feature transmission instruction is generated;

[0049] A virtual-real fusion rendering module configured to perform hierarchical rendering on a command terminal display interface in response to data transmitted by each front-end acquisition device; for a region executing the full-pixel transmission instruction, a real-time video stream is rendered; for a region executing the semantic feature transmission instruction, a preset 3D virtual avatar model is driven to perform action reconstruction based on received semantic feature data, and the reconstructed virtual image and the real-time video stream are superimposed and displayed in the same user interface coordinate system.

[0050] This embodiment details a data transmission system introducing a physical field model, aiming to convert a bandwidth allocation problem in a large-scale command scene into an energy allocation problem in a potential field;

[0051] The system establishes a subjective and objective dual perception input through the potential parameter acquisition module, wherein the focus point coordinate is derived from gaze point tracking data of a command personnel on a terminal screen, and its physical meaning is a current tactical attention center of a commander The event entropy value is derived from the normalized calculation of the environmental data by the front-end device, and its physical meaning is a dimensionless value representing the degree of mutation of the physical state of the on-site environment ;

[0052] The line-of-sight potential energy calculation module constructs a non-uniform information gravity field, defines the commander's attention point as a gravity singularity, and defines the high-risk event at the front end as a charge amount to calculate a real-time potential energy value This value quantifies the urgency of the node to obtain bandwidth resources; on this basis, the heterogeneous modal scheduling module executes a binary decision logic:

[0053] In response to , the system determines that the node is a high-attention node and generates a full-pixel transmission instruction to maintain a high-definition video stream;

[0054] In response to , the system determines that the node is a background node and generates a semantic feature transmission instruction to cut off the video stream and only transmit a very low-bandwidth structured data;

[0055] The virtual-real fusion rendering module does not display a black screen, but instead calls a GPU to drive a 3D Avatar based on semantic data to reconstruct actions and achieve seamless visual superposition;

[0056] This embodiment realizes a bandwidth on-demand allocation mechanism based on the commander's perception intention and the degree of environmental mutation by constructing a line-of-sight gravity field model; in the scenario where the command center faces a large number of front-end concurrent access and the network bandwidth resource is limited, this system breaks the rigid mode of the traditional monitoring picture average allocation code rate, ensuring that the high-potential energy area such as the commander's attention point or the sudden event point obtains lossless real-time video stream, while the low-potential energy area maintains basic situation awareness such as position and posture at a very low bandwidth cost;

[0057] This dynamic scheduling strategy of heterogeneous modal effectively solves the deadlock problem between the lack of bandwidth resources and the demand for global situation awareness in a large-scale visual intercom system, ensuring the zero-delay delivery of key tactical information.

[0058] Embodiment 2:

[0059] The potential energy parameter acquisition module includes:

[0060] A focus capture unit configured to acquire, by a connected eye tracking sensor or a cursor control device, a gaze point coordinate of the commander on a display screen as an attention focus coordinate at a preset sampling frequency;

[0061] The entropy value quantization unit is configured to receive multi-dimensional sensor data uploaded by the front-end acquisition device, the multi-dimensional sensor data including audio decibel data, acceleration mutation data and biological sign data, and to perform weighted normalization processing on the multi-dimensional sensor data by using preset weight coefficients to generate event entropy values.

[0062] The embodiment specifically defines the physical implementation of the potential energy parameter acquisition module;

[0063] The focus capture unit uses an infrared line-of-sight tracking sensor or a high-precision cursor control device to capture the gaze point coordinates of the command personnel on the screen at a preset sampling frequency such as 60 Hz The gaze point coordinates of the command personnel on the screen are continuously acquired , and the focus coordinates are output after being processed by a smoothing filter, aiming to eliminate data noise caused by eyeball microtremor;

[0064] Meanwhile, the entropy value quantization unit deployed at the front end receives multi-dimensional sensor data, including audio decibel data collected by an omnidirectional microphone, acceleration mutation data representing a fall or impact collected by an IMU, and biological sign data such as heart rate and blood oxygen collected by a smart bracelet;

[0065] The unit uses preset weight coefficients to perform weighted normalization processing on the above heterogeneous data to generate event entropy values , and the calculation formula is as follows:

[0066]

[0067] wherein, represents the index of the sensor type; is the real-time acquisition value or the deviation value after preprocessing from the first sensor, and the physical meaning is the instantaneous intensity of the current environmental parameter or the degree of deviation from the normal state;

[0068] is the weight coefficient set according to the task scene, and the physical meaning is the importance weight of the sensor data in the current tactical background; for example, in the intense battle scene mode, the weight of the audio decibel data is set to , the weight of the acceleration mutation data is set to , and the weight of the biological sign data is set to , so as to give priority to the personnel life state; the set value of the weight coefficient can be pre-calibrated by historical data statistical analysis or expert scoring method;

[0069] and is derived from the sensor specification book or preset reference, and the physical meaning is the lower limit and upper limit of the value for normalization calculation, such as 0 or the minimum value;

[0070] The total number of sensor types derived from system integration, the physical meaning is the total number of data dimensions participating in the entropy value calculation;

[0071] The calculation process maps the complex environmental physical quantity into a normalized value, and the larger the value represents the more severe the mutation of the environmental state;

[0072] In addition, in order to ensure the physical meaning robustness of the calculation result, the system also performs the following boundary constraint logic: for biological sign data with bidirectional abnormal characteristics, such as excessively high or low heart rate, the absolute deviation of the data from the preset physiological reference value is calculated in advance, and the deviation value is taken as Substitute the above formula, at this time the corresponding Set to 0; At the same time, the normalized result is truncated in the interval of 0 to 1 to prevent calculation overflow caused by sensor value drift; And the preset weight coefficient satisfies the normalization condition To ensure the dimensional consistency of the total entropy value ;

[0073] The embodiment integrates multi-dimensional sensor data and introduces an entropy value quantization mechanism, giving the front-end acquisition device the initiative to seek help; In the case where the commander does not actively pay attention to a certain area, that is, the line-of-sight distance is far away, if an explosion, a person falls or an abnormal vital sign occurs in the area, the rapidly increasing event entropy value can force the overall potential energy value of the node to increase; This design compensates for the blind spot of simply relying on the commander's subjective line of sight for resource scheduling, ensuring that in a complex and variable battlefield or emergency rescue environment, high-risk emergencies can break through the line-of-sight distance limit and automatically obtain a high-priority transmission channel, thereby minimizing the risk of missing critical alerts.

[0074] Embodiment 3:

[0075] The line-of-sight potential energy calculation module performs the following processing:

[0076] Determine the current display coordinates of each front-end acquisition device on the display interface of the command end; Calculate the Euclidean distance between the focus coordinates and the current display coordinates;

[0077] Process the Euclidean distance based on the inverse function relationship between the value and the distance to obtain the basic attention degree, and weight sum the basic attention degree and the event entropy value based on a preset proportion coefficient to obtain the real-time potential energy value, wherein the smaller the Euclidean distance or the larger the event entropy value, the higher the real-time potential energy value.

[0078] This embodiment details the mathematical model construction and calculation logic of the real-time potential energy value;

[0079] The system performs the coordinate mapping step to obtain the current display coordinates of each front-end acquisition device on the display interface of the command end , usually the geometric center of the video window;

[0080] The system calculates the focus point coordinates The Euclidean distance between the current display coordinates of the first front-end device and the focus point coordinates , the formula is ;

[0081] In order to eliminate the difference in the numerical magnitude between the pixel distance and the dimensionless entropy value, the system performs distance normalization processing; define as the diagonal pixel length of the command terminal display interface, the system calculates the normalized distance , the real-time potential value calculation formula is:

[0082]

[0083] Among them, since and The value range is constrained in the interval [0, 1], the focus gain coefficient and the mutation weight coefficient Can take the same order of magnitude of the value, such as , so as to ensure that the line of sight gravity term and the event entropy term have comparable weight contribution in the total potential , avoid the imbalance of the order of magnitude leading to a single factor dominating the scheduling logic of the system;

[0084] Among them, derived from real-time calculation results, the physical meaning is the information acquisition urgency of the first front-end device at time;

[0085] : derived from system preset constant, the physical meaning is the focus gain coefficient, used to adjust the influence strength of the line of sight gravity field;

[0086] : derived from the preset minimum positive number, the physical meaning is the distance smoothing factor, which aims to prevent numerical singularity;

[0087] : derived from system preset constant, the physical meaning is the mutation weight coefficient, used to adjust the contribution rate of event entropy to the total potential;

[0088] In this model, the first term simulates the gravitational field, and the second term simulates the potential energy of the particle itself; it should be noted that considering the significant difference in numerical magnitude between the pixel unit and the dimensionless , the focus gain coefficient is a dimensionless value, since the aforementioned steps have already normalized the pixel distance converted to normalized , here It is desirable to approach 1.0 dimensionless value to ensure that the line of sight gravity term and event potential energy term have comparable weight contributions in the total potential energy , avoiding the imbalance in the order of magnitude leading to a single factor dominant system scheduling logic;

[0089] The potential energy calculation model constructed in this embodiment accurately simulates the nonlinear superposition effect of human cognitive attention mechanism and environmental saliency; through the inverse function relationship, it ensures that the area closer to the commander's line of sight center obtains exponentially growing basic attention, and the weighted summation mechanism ensures that high-entropy events can directly increase the potential energy independent of the line of sight distance; This algorithm design realizes the mathematical unification of subjective attention demand and objective environmental crisis, ensuring that the system resource allocation strategy not only conforms to the commander's operation intuition, but also maintains high sensitivity to sudden crises on the battlefield edge.

[0090] Embodiment 4:

[0091] When the heterogeneous modal scheduling module generates semantic feature transmission instructions, the following control logic is executed:

[0092] Send a video stream blocking signal to the corresponding front-end acquisition device to control it to stop encoding and transmitting image pixel data; activate the built-in end-side computing power unit of the front-end acquisition device to perform structured analysis on the real-time collected pictures, and extract human body skeleton key point data and geographic position coordinate data of the target object;

[0093] Establish a low-bandwidth transmission channel, and send the human body skeleton key point data and the geographic position coordinate data as semantic feature data to the command end, while maintaining the transmission of event entropy values or multi-dimensional sensor data used to generate event entropy values.

[0094] This embodiment describes in detail the data flow logic in the low potential energy mode, which aims to process non-attention areas at the edge of the commander's line of sight or in a calm state; in response to the determination result that the potential energy value is lower than the threshold, the heterogeneous modal scheduling module sends a video stream blocking signal to the front end, instructing the camera to stop feeding raw data to the encoder or stopping H.264 code stream output at the physical layer;

[0095] The system activates the built-in end-side computing power unit of the front-end device, such as NPU, to run lightweight computer vision algorithms to perform structured analysis on real-time pictures;

[0096] Extract human body skeleton key point data such as coordinates of 17 joint nodes, and geographic position coordinate data such as GPS longitude and latitude;

[0097] The system establishes a low-bandwidth transmission channel, such as using the MQTT protocol, to package and send the above semantic feature data; in this process, the system deliberately maintains the event entropy value or related multi-dimensional sensor data to ensure continuous monitoring of environmental mutations;

[0098] The semantic feature transmission strategy adopted in this embodiment compresses the megabit-level bandwidth Mbps required by traditional video monitoring to a semantic-level bandwidth Kbps that is thousands of times smaller; in an extreme command environment where network bandwidth is extremely limited or congested, this scheme enables the system to maintain dozens of times the number of concurrent connections of front-end nodes compared to traditional schemes, and by preserving key semantic information such as personnel location and posture, it avoids the risk of complete disconnection in non-attention areas, achieving continuous tracking of the global situation under ultra-narrowband conditions.

[0099] Embodiment 5:

[0100] When generating full-pixel transmission instructions, the heterogeneous modal scheduling module executes the following control logic: sends a high-bandwidth occupation permit to the corresponding front-end acquisition device to control it to enable a high-definition video encoder;

[0101] A high-bandwidth transmission channel is established to send real-time video stream data that has not been dimensionally reduced to the command end; when the total network bandwidth is limited, network resources are preferentially allocated to front-end acquisition devices executing full-pixel transmission instructions, and the channel bandwidth for executing semantic feature transmission instructions is dynamically compressed.

[0102] This embodiment describes in detail the resource tilting logic in the high-potential mode, which is applied to the commander's gaze focus or high-risk areas;

[0103] In response to the determination that the system sends a high-bandwidth occupation permit to the front end to control it to enable a high-definition video encoder such as 4K / 60fps;

[0104] A high-bandwidth transmission channel, such as an RTSP stream, is established to transmit real-time video streams that have not been dimensionally reduced; in particular, in a competitive scenario where the total network bandwidth is limited, the heterogeneous modal scheduling module implements a QoS strategy at the network layer to preferentially allocate OFDM subcarrier or time slot resources to devices executing full-pixel transmission instructions; at the same time, the system dynamically compresses the channel bandwidth for executing semantic feature transmission instructions, such as reducing the upload frequency of skeletal point data from 30Hz to 10Hz, sacrificing the refresh rate of edge nodes to ensure the transmission quality of core nodes;

[0105] The embodiment establishes a dynamic QoS strategy of sacrificing the edge and protecting the core; by actively reducing the data throughput of the non-concerned area when the network bandwidth is limited, the system releases valuable channel resources, ensures that the focus picture of the commander's current concern or the high-risk picture of the sudden event always maintains the movie-level clarity and smoothness; this strategy ensures that the commander can still obtain high-fidelity visual information of the key area under poor communication conditions, so as to make accurate tactical judgments and avoid the key picture from being stuck or blurred due to network congestion.

[0106] Embodiment 6:

[0107] The virtual-real fusion rendering module comprises:

[0108] The model mapping unit is configured to pre-store 3D virtual avatar models bound with each front-end acquisition device;

[0109] The skeleton driving unit is configured to analyze the human body skeleton key point data in the received semantic feature data, map the human body skeleton key point data to the skeleton nodes of the 3D virtual avatar model, and real-time solve and drive the 3D virtual avatar model to generate limb movements synchronized with the front-end target object;

[0110] The scene synthesis unit is configured to render the driven 3D virtual avatar model into the corresponding background area of the command terminal display interface according to the geographic position coordinate data in the semantic feature data.

[0111] The embodiment details the rendering mechanism for realizing semantic flow visualization;

[0112] The model mapping unit pre-loads the 3D virtual avatar model bound with the front-end personnel ID in the local video memory of the command terminal; when receiving the semantic data, the skeleton driving unit as an interface layer analyzes the human body skeleton key point data and maps these two-dimensional or three-dimensional coordinates to the skeleton nodes of the virtual avatar model;

[0113] Real-time solving by using inverse kinematics IK algorithm drives the virtual avatar to generate limb movements completely synchronized with the front-end real target;

[0114] The scene synthesis unit renders the dynamic virtual avatar to the corresponding background area of the command terminal display interface, such as a specific coordinate point of the virtual sand table or GIS map, according to the geographic position coordinate data, to complete the reconstruction from data to image;

[0115] Further, in order to solve the spatial misplacement problem caused by inconsistent transmission delay of heterogeneous data sources, i.e. virtual-real asynchronization, the system is built with a time axis alignment unit; the semantic feature data and real-time video stream data uploaded by each front-end acquisition device carry unified GPS time stamps or NTP network time stamps; the scene synthesis unit is provided with an anti-jitter buffer window, and the received semantic feature data is compensated for delay based on the time stamp, so that the rendering timing is strictly consistent with the playing timing of the background real-time video stream, ensuring that the 3D virtual avatar and the real target in the video picture achieve millisecond-level time alignment in the motion track;

[0116] The embodiment makes use of the powerful local graphic rendering capability of the command terminal to compensate for the bandwidth short board of the communication link, and realizes the visual enhancement effect of transmission data and presentation image; for the commander, although the non-attention area presents a virtual avatar, the motion posture and geographical position are all driven by real-time data, which ensures the authenticity and timeliness of the tactical information; this virtual-real integrated presentation method can still intuitively display the tactical actions of personnel such as squatting, raising hands and running under ultra-low bandwidth, and maintains the intuitive perception ability of the commander to the overall situation.

[0117] Embodiment 7:

[0118] The system further comprises:

[0119] The smooth switching module is configured to monitor the state jump of the real-time potential value across the preset potential threshold;

[0120] If it is detected that the state jumps from low potential to high potential, the rendering display of the 3D virtual avatar model is maintained before the first frame of real-time video stream is received, and after the video stream buffering is completed, the fade-in switching from virtual image to real video is performed;

[0121] If it is detected that the state jumps from high potential to low potential, the system preferentially sends a semantic feature extraction start instruction to the front-end acquisition device, and maintains the transmission of full-pixel video stream until the first valid semantic feature data packet is received by the command terminal; after confirming the connection of the semantic data stream, the first frame of semantic feature data is taken as the initial rendering state of the 3D virtual avatar model, and a video stream blocking signal is sent synchronously, and the smooth transition from real-time video to virtual image is performed, so as to avoid picture jump or model posture zero error caused by time sequence difference during data source switching.

[0122] The embodiment constructs a smooth switching module to solve the visual mutation problem during switching of heterogeneous modalities;

[0123] The system monitors the potential value in real time across the threshold the state jumps from low potential energy to high potential energy, i.e., the commander's line of sight moves in or an emergency occurs, considering the physical time delay existing in establishing a video connection, i.e., an I-frame buffering period, the system forces to keep rendering the 3D virtual avatar model before receiving the first frame of complete video stream, filling the waiting empty window period;

[0124] After the video stream is decoded and ready, the fade-in switching is performed; otherwise, in response to the state jumping from high potential energy to low potential energy, i.e., the line of sight moves away, the system starts semantic feature extraction and truncates the video stream immediately in order to immediately release the bandwidth; in the next frame of the video truncation, the rendering engine uses the last frame of the skeleton pose as the initial state to seamlessly connect the rendering of the 3D virtual avatar model, ensuring picture continuity;

[0125] The smooth switching mechanism provided by the embodiment eliminates the black screen flicker or buffer rotation phenomenon commonly seen in traditional streaming media systems when the code stream is switched; by using the locally rendered virtual image to fill the time delay gap caused by network transmission, the system provides the commander with a smooth and continuous visual experience as if from a global perspective; this design ensures that the user's cognitive flow is not interrupted in the commander's operation of frequently switching attention points, and realizes seamless fusion of heterogeneous data sources in the visual presentation layer.

[0126] Embodiment 8:

[0127] The system further comprises:

[0128] The threshold adaptive adjustment module is configured to monitor the total available bandwidth load of the current network environment in real time; if the total available bandwidth load is lower than a preset congestion warning line, the preset potential energy threshold is increased to reduce the number of front-end acquisition devices entering the full-pixel transmission mode, so that the non-attention area enters the semantic feature transmission mode; if the total available bandwidth load recovers to a preset normal range, the preset potential energy threshold is restored to the initial default value;

[0129] The embodiment introduces a threshold adaptive adjustment module with hysteresis characteristics, which gives the system dynamic adaptation capability to network environment fluctuations and eliminates control oscillation; the module clearly quantifies the total available bandwidth load, defines as the remaining bandwidth proportion of the current network, and the calculation formula is:

[0130]

[0131] wherein, is the link physical total bandwidth, is the real-time throughput;

[0132] The system presets two asymmetric thresholds: a congestion warning line of 20% and a recovery safety line of 40% , and satisfies to construct a hysteresis interval;

[0133] The specific control logic is executed as follows:

[0134] Congestion response: when monitoring , the system calculates the congestion severity factor , and dynamically adjusts the potential energy threshold according to the factor, and the initial default threshold is , then the threshold is calculated as follows:

[0135]

[0136] wherein is a preset aggressive coefficient such as 2.0; the formula shows that the more severe the congestion, i.e. , the greater the threshold adjustment, thereby forcing more edge attention, i.e. potential energy value between the original threshold and the new threshold, the front-end device switches to the semantic feature transmission mode, achieving millisecond-level bandwidth release;

[0137] State retention: when , the system keeps the current unchanged to prevent the Ping-Pong effect caused by the immediate fall of the threshold due to temporary release of bandwidth, i.e. the repeated switching between video and semantic modes;

[0138] Restoration and fall: only when , it is determined that the network load has been restored to the normal range, and the system executes the linear decay logic:

[0139]

[0140] wherein is a dynamic recovery step, the value of which is proportional to the surplus degree of the current bandwidth, and the calculation formula is , wherein is a preset recovery rate coefficient; this calculation logic ensures that when the network bandwidth is greatly surplus, the potential energy threshold can quickly fall to restore the quality; when the bandwidth is only hovering on the edge of the safety line, the threshold slowly decreases with a small step, thereby effectively avoiding the risk of inducing network congestion shock caused by the rapid fall of the threshold; derived from the threshold state of the last control cycle from the system cache, which is a historical memory variable for maintaining the continuity of adjustment;

[0141] The embodiment realizes a negative feedback adjustment mechanism based on the principle of Schmidt trigger, significantly enhances the robustness of the system, and through the introduction of the residual bandwidth proportion index and the asymmetric double threshold control, the system can not only automatically cut off the tail to survive according to the congestion degree when the network is extremely congested, accurately calculate the number of edge nodes to be sacrificed, but also effectively avoid mode repeated switching oscillation in the critical state, and ensure the communication stability and smoothness of the operation experience of the command system.

[0142] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present application rather than limit the present application. Although the present application has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present application can be modified or replaced by equivalents without departing from the spirit and scope of the technical solutions of the present application.

Claims

1. A video intercom system for personnel command and management, characterized in that: include: The potential energy parameter acquisition module is configured to acquire the coordinates of the focus of attention of the command terminal input device in real time, and synchronously receive the event entropy value uploaded by each front-end acquisition device. The event entropy value represents the degree of physical state change of the environment in which the front-end acquisition device is located. The line-of-sight potential energy calculation module is configured to construct a line-of-sight gravity field model, calculate the line-of-sight distance between the coordinates of the focus of attention and the mapped positions of each of the front-end acquisition devices on the display interface, and calculate the real-time potential energy value of each of the front-end acquisition devices based on a weighted calculation of the line-of-sight distance and the event entropy value. The heterogeneous modal scheduling module is configured to dynamically control the data transmission mode of each of the front-end acquisition devices based on the comparison result between the real-time potential energy value and the preset potential energy threshold; if the real-time potential energy value is greater than the preset potential energy threshold, a full-pixel transmission instruction is generated; if the real-time potential energy value is less than or equal to the preset potential energy threshold, a semantic feature transmission instruction is generated. The virtual-real fusion rendering module is configured to respond to the data transmitted by each of the front-end acquisition devices and perform hierarchical rendering on the command terminal display interface; for the area where the full-pixel transmission command is executed, the real-time video stream is rendered; for the area where the semantic feature transmission command is executed, the preset 3D virtual avatar model is driven to reconstruct the action based on the received semantic feature data, and the reconstructed virtual image is superimposed and displayed with the real-time video stream in the same user interface coordinate system. The line-of-sight potential energy calculation module performs the following processing: Determine the current display coordinates of each of the aforementioned front-end acquisition devices on the command terminal display interface. ; Calculate the coordinates of the focus of interest Euclidean distance between the current displayed coordinates The formula is ; definition Calculate the normalized distance for the diagonal pixel length of the command terminal display interface. ; Based on the normalized distance and the event entropy value, the real-time potential energy value is calculated. The formula for calculating the real-time potential energy value is as follows: ,in, This is the real-time potential energy value. The event entropy value, This is the attention gain coefficient. For distance smoothing factor, These are the mutation weighting coefficients; The smaller the Euclidean distance or the larger the event entropy value, the higher the real-time potential energy value.

2. The video intercom system for personnel command and management according to claim 1, characterized in that: The potential energy parameter acquisition module includes: The focus capture unit is configured to acquire the coordinates of the commander's gaze point on the display screen at a preset sampling frequency via a connected eye-tracking sensor or cursor control device, and use them as the coordinates of the focus of attention. The entropy quantization unit is configured to receive multi-dimensional sensor data uploaded by the front-end acquisition device. The multi-dimensional sensor data includes audio decibel data, acceleration mutation data, and biological characteristic data. The unit performs weighted normalization processing on the multi-dimensional sensor data using preset weighting coefficients to generate the event entropy value.

3. A video intercom system for personnel command and management according to claim 1, characterized in that: When generating the semantic feature transmission instruction, the heterogeneous modality scheduling module executes the following control logic: Send a video stream blocking signal to the corresponding front-end acquisition device to control it to stop encoding and transmitting image pixel data; The built-in edge computing unit of the front-end acquisition device is activated to perform structured analysis on the real-time acquired images and extract the key point data of the human skeleton and geographical coordinate data of the target object. A low-bandwidth transmission channel is established to send the key point data of the human skeleton and the geographical coordinate data as semantic feature data to the command end, while maintaining the transmission of the event entropy value or the multi-dimensional sensor data used to generate the event entropy value.

4. A video intercom system for personnel command and management according to claim 1, characterized in that: When generating the full-pixel transmission instruction, the heterogeneous modal scheduling module executes the following control logic: Send a high bandwidth usage permission to the corresponding front-end acquisition device to control it to enable the high-definition video encoder; Establish a high-bandwidth transmission channel to send real-time video stream data without dimensionality reduction processing to the command center; When the total network bandwidth is limited, network resources are prioritized for front-end acquisition devices that execute full-pixel transmission commands, and channel bandwidth for executing semantic feature transmission commands is dynamically compressed.

5. A video intercom system for personnel command and management according to claim 1, characterized in that: The virtual-real fusion rendering module includes: The model mapping unit is configured to pre-store 3D virtual avatar models bound to each of the aforementioned front-end acquisition devices; The skeleton driving unit is configured to parse the human skeleton key point data in the received semantic feature data, map the human skeleton key point data to the skeleton nodes of the 3D virtual avatar model, and calculate and drive the 3D virtual avatar model to generate limb movements synchronized with the front-end target object in real time. The scene synthesis unit is configured to render the driven 3D virtual avatar model onto the corresponding background area of ​​the command terminal display interface based on the geographic location coordinate data in the semantic feature data.

6. A video intercom system for personnel command and management according to claim 1, characterized in that: The system also includes: A smooth switching module is configured to monitor state transitions where the real-time potential energy value crosses the preset potential energy threshold. If a change in state from low potential energy to high potential energy is detected, the rendering and display of the 3D virtual avatar model is maintained before the first frame of real-time video stream is received. After the video stream buffering is completed, a fade-in switch from virtual image to real video is performed. If a state transition from high potential energy to low potential energy is detected, semantic feature extraction is immediately initiated, and the rendering screen of the 3D virtual avatar model is connected to the next frame after the video stream is truncated.

7. A video intercom system for personnel command and management according to claim 1, characterized in that: The system also includes: The threshold adaptive adjustment module is configured to monitor the total available bandwidth load of the current network environment in real time. If the total available bandwidth load is lower than the preset congestion warning line, the preset potential energy threshold is increased to reduce the number of front-end acquisition devices entering the full pixel transmission mode, so that non-interested areas enter the semantic feature transmission mode. If the total available bandwidth load recovers to the preset normal range, the preset potential energy threshold will be restored to the initial default value.

Citation Information

Patent Citations

  • Information processing apparatus, information processing method, and information processing system

    CN116137954A

  • Indoor tumble detection method, system and equipment based on computer vision and medium

    CN117115905A