Behavior monitoring system based on video image real-time transmission

By employing real-time semantic analysis and dynamic coding strategies, the problems of resource waste and privacy violations in behavior monitoring systems under normal circumstances have been solved. High-fidelity transmission and reduced false alarm rates have been achieved at critical moments, thereby improving the reliability and privacy protection of the monitoring system.

CN121728243APending Publication Date: 2026-03-24BEIJING CHUANGPU TECH CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-02-25
Publication Date
2026-03-24

AI Technical Summary

Technical Problem

Existing behavior monitoring systems typically lead to wasted network bandwidth and storage resources, privacy violations, and difficulty in ensuring high-fidelity, low-latency transmission of critical images at critical moments. They are also susceptible to false alarms due to changes in lighting and interference from non-target objects.

Method used

The video acquisition unit acquires the raw video stream, the encoding control unit performs real-time semantic analysis, dynamically generates encoding control strategies, and the transmission execution unit adjusts the network service quality to achieve switching between high-fidelity encoding and privacy-preserving low-bandwidth encoding. High-risk behavior footage is prioritized for transmission, and audio-assisted analysis is combined to improve accuracy.

Benefits of technology

While conserving network resources and protecting privacy under normal circumstances, it ensures clear and smooth transmission of images during critical moments, reduces false alarm rates, provides effective alarm and backtracking functions, and improves the reliability and privacy of the monitoring system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121728243A_ABST
    Figure CN121728243A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of intelligent video monitoring, in particular to a behavior monitoring system based on video image real-time transmission. The system comprises a video acquisition unit, a coding control unit and a transmission execution unit. The system analyzes an original video stream in real time to generate a dynamic coding strategy; the core of the method is that semantic perception high-fidelity priority transmission is executed when a high-risk behavior is identified according to a behavior risk assessment result, and a privacy low-bandwidth conventional transmission mode is adaptively switched when a safety behavior is identified; according to the invention, the conversion from continuous high-definition monitoring to on-demand privacy protection is realized, and the supervision effectiveness, bandwidth saving and privacy security are considered.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of intelligent video surveillance and multimedia communication technology, specifically to a behavior monitoring system based on real-time video image transmission. Background Technology

[0002] In the current field of behavioral safety monitoring, monitoring equipment typically continuously collects video streams and transmits them in full over a network, or uses traditional motion detection mechanisms based on pixel changes. To ensure the effectiveness of remote monitoring, existing solutions generally rely on the indiscriminate encoding and transmission of raw images. Although this solution has basic real-time monitoring capabilities, it lacks deep perception and differentiation of scene semantics. In normal safe scenarios, the continuous transmission of high-definition images leads to a serious waste of network bandwidth and storage resources, and also causes a continuous invasion of the privacy of the monitored object. In addition, in network congestion or weak network environments, existing technologies cannot prioritize the transmission quality of high-risk critical images such as falls and foreign object ingestion. Furthermore, traditional detection methods are easily affected by changes in light and shadow and interference from non-target objects, resulting in a large number of false alarms. This leads to the dilemma of missing core information, image lag, or invalid alarms at critical moments for guardians. Therefore, how to dynamically adjust encoding and transmission strategies based on semantic analysis to ensure high-fidelity, low-latency transmission of critical images in critical moments while significantly reducing normal resource consumption and protecting privacy has become an urgent technical problem to be solved. Summary of the Invention

[0003] To address the aforementioned technical problems, this invention provides a behavior monitoring system based on real-time video image transmission. Specifically, the technical solution of this invention includes: Video acquisition unit, encoding control unit, and transmission execution unit: The video acquisition unit is configured to acquire and output raw video stream data of the monitored area. The encoding control unit is configured to perform real-time analysis of the raw video stream data and dynamically generate video encoding control strategies based on the analysis results; the video encoding control strategies include the following two types: In response to the analysis results indicating high-risk behavior, a first control command is generated to trigger a semantically aware high-fidelity encoding and priority transmission process on the raw video stream data; In response to the analysis results indicating security behavior, a second control command is generated to trigger the execution of privacy-preserving low-bandwidth encoding and regular transmission processes on the raw video stream data; The transmission execution unit is configured to receive the video encoded stream processed by the encoding control unit and execute the network service quality policy corresponding to the first control instruction or the second control instruction to complete the video image transmission.

[0004] Preferably, the specific process by which the encoding control unit performs real-time analysis of the raw video stream data is as follows: Extract the sequence of key points of human skeleton in video frames and identify temporal action patterns; The timing action pattern is evaluated according to preset rules. If it matches the preset high-risk behavior pattern library, it is judged as a high-risk behavior. Otherwise, it is considered a safe behavior.

[0005] The preferred high-fidelity encoding and priority transmission process based on semantic awareness is as follows: The human body region is determined based on the sequence of key points of the human skeleton, and this region is marked as the foreground region of interest. Assign a first quantization parameter to the foreground region of interest that is lower than that to the background region, and increase the encoding frame rate; Insert instantaneous decoding refresh frames into the current encoding sequence and add forward error correction redundancy.

[0006] Preferably, the privacy-enhancing low-bandwidth encoding and the conventional transmission process are specifically one of the following: A dynamic skeletal graphic is reconstructed based on the sequence of key points in the human skeleton, and then used to replace the original video footage for encoding and transmission; or, The raw video stream data is globally blurred before being encoded and transmitted. At the same time, the video encoding frame rate is adjusted to a preset low threshold.

[0007] Preferably, the specific process by which the transmission execution unit executes the network service quality policy is as follows: When the video encoded stream corresponds to the first control command, the data packet is marked as the highest quality of service priority and the maximum available bandwidth is requested; When the video encoded stream corresponds to the second control command, the data packet is marked as a best-effort quality of service level, and redundant bandwidth is released.

[0008] Preferably, the encoding control unit is further configured to perform audio-assisted analysis: Synchronously analyze the collected audio data; When a behavior is determined to be safe, but the audio characteristics exceed a preset threshold or match a preset dangerous audio template, the analysis result will be forcibly corrected to a high-risk behavior, and a first control command will be generated.

[0009] Preferably, it also includes a terminal presentation unit, configured as follows: Receive and decode the video encoded stream sent by the transmission execution unit; When the decoded result is a skeleton graphic or a blurred image processed by a privacy-enhanced low-bandwidth encoding process, it is displayed normally. When the decoded video is processed through a high-fidelity encoding process and carries a high-risk behavior indicator, a pop-up alarm will be triggered and high-definition real-time footage and cached clips will be played.

[0010] Compared with the prior art, the present invention has the following beneficial effects: 1. This system performs semantic analysis on the raw video stream through the encoding control unit. When it determines that the behavior is safe, such as sleeping or sitting still, it automatically switches to a privacy-preserving low-bandwidth encoding process. By transmitting only the reconstructed skeleton dynamic graphics or the globally blurred image and reducing the frame rate to a low threshold, the system can significantly reduce the bit rate during non-critical periods, thus saving network bandwidth and storage resources. At the same time, this abstract image display method retains the sense of online confirmation for the guardian while avoiding continuous high-definition spying on the details of life, effectively protecting the privacy of the monitored object. 2. When this system detects high-risk behaviors such as falls or swallowing foreign objects, it triggers a high-fidelity encoding strategy based on semantic awareness. By extracting key points of the human skeleton to determine the region of interest in the foreground, the system can assign more refined quantization parameters to the core action area, while reducing the image quality of the background area. Combined with real-time decoding and refresh frame insertion and forward error correction redundancy technology, it maximizes the core information entropy. With the highest quality of service priority strategy of the transmission execution unit, the system can forcibly seize available bandwidth under adverse conditions such as network congestion or weak signal, ensuring clear, smooth and lag-free images in critical moments, solving the pain point of existing technologies that are prone to information loss at critical moments. 3. Compared to traditional motion detection based on pixel changes, this system utilizes a lightweight edge-side model to extract key points of the human skeleton and identify temporal action patterns, significantly reducing the false alarm rate caused by changes in lighting, pet activity, or movement of non-target objects. Furthermore, the system integrates an audio-assisted analysis mechanism, using timestamp alignment technology to synchronously detect ambient sounds. When visual analysis misjudges a situation as safe but detects dangerous audio features such as screams or impacts, the judgment is forcibly corrected. This multimodal fusion mechanism effectively compensates for the line-of-sight obstruction or blind spots inherent in single-view surveillance, significantly improving the reliability of security monitoring. 4. This system implements hierarchical interaction logic through the terminal presentation unit. Under normal circumstances, the terminal only displays low-power abstract graphics to avoid false alarms causing psychological disturbance to guardians. However, at the moment of danger, the system immediately triggers a strong reminder pop-up and automatically plays high-definition real-time footage containing cached clips from before the event. This mechanism not only ensures the effectiveness of the alarm but also enables guardians to understand the cause and process of the accident through the retrospective function, such as the specific posture of the fall or the swallowed object, so as to take correct and targeted rescue measures quickly. Attached Figure Description

[0011] The present invention will be further explained below with reference to the accompanying drawings and embodiments: Figure 1 This is a structural diagram of the system of the present invention. Detailed Implementation

[0012] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to specific embodiments.

[0013] Example 1: Please see Figure 1 A behavior monitoring system based on real-time video image transmission includes a video acquisition unit, an encoding and control unit, and a transmission execution unit. The video acquisition unit is configured to acquire and output raw video stream data of the monitored area. The encoding control unit is configured to perform real-time analysis of the raw video stream data and dynamically generate video encoding control strategies based on the analysis results; the video encoding control strategies include the following two types: In response to the analysis results indicating high-risk behavior, a first control command is generated to trigger a semantically aware high-fidelity encoding and priority transmission process on the raw video stream data; In response to the analysis results indicating security behavior, a second control command is generated to trigger the execution of privacy-preserving low-bandwidth encoding and regular transmission processes on the raw video stream data; The transmission execution unit is configured to receive the video encoded stream processed by the encoding control unit and execute the network service quality policy corresponding to the first control instruction or the second control instruction to complete the video image transmission.

[0014] This embodiment provides a behavior monitoring system based on real-time video image transmission; the system aims to solve the problems of existing monitoring technologies being unable to guarantee the transmission of critical risk images in weak network environments, as well as the bandwidth waste and privacy violations caused by routine monitoring. The system in this embodiment mainly includes a video acquisition unit, an encoding control unit, and a transmission execution unit; The video acquisition unit is configured to acquire and output raw video stream data of the monitored area. In this embodiment, raw video stream data refers to uncompressed RGB image sequences directly acquired by front-end camera sensors such as CMOS sensors. This unit can integrate an infrared illumination module or a depth sensor to support data acquisition in low-light environments; its purpose is to provide high-fidelity, full-data input for subsequent semantic analysis. The encoding control unit is configured to perform real-time analysis of the raw video stream data and dynamically generate video encoding control strategies based on the analysis results. In this embodiment, the unit acts as the semantic director of the system, and its core logic is to establish a mapping relationship between behavioral risks and video coding parameters; the video coding control strategy refers to a set of instructions that includes encoder configuration parameters such as quantization parameter QP, frame rate FPS, and GOP structure. It should be noted that regardless of the encoding strategy used, the video acquisition unit always maintains the original high frame rate, such as 30fps or 60fps, to ensure the continuity and real-time performance of the data analysis by the encoding control unit. This strategy has the following two modes: First control command: In response to analysis results indicating high-risk behaviors such as falling, climbing, or swallowing foreign objects, it triggers a semantically aware high-fidelity encoding and priority transmission process on the raw video stream data; this process aims to maximize the information entropy density of critical events. The second control command, in response to analysis results indicating safe behaviors such as sleeping, reading, or sitting quietly, is used to trigger a privacy-enhancing low-bandwidth encoding and regular transmission process on the raw video stream data; this process aims to minimize redundant data transmission and protect the privacy of the monitored object. The transmission execution unit is configured to receive the video encoded stream processed by the encoding control unit and execute the network service quality policy corresponding to the first control instruction or the second control instruction to complete the video image transmission. In this embodiment, the unit is located at the network transport layer and is responsible for adjusting the sending priority of data packets according to the semantic instructions of the upper layer; Through the above architecture, the system realizes a paradigm shift from pixel-driven transmission to semantic-driven transmission; when the monitoring scenario is under normal security, the system automatically reduces resource consumption and protects privacy; and in the moment of danger, the system can concentrate the resources of the entire network to ensure the clarity and smoothness of key images, thereby effectively solving the value gap and ensuring that guardians can obtain effective information at critical moments.

[0015] Example 2: The specific process by which the encoding control unit performs real-time analysis of the raw video stream data is as follows: Extract the sequence of key points of human skeleton in video frames and identify temporal action patterns; The timing action pattern is evaluated according to preset rules. If it matches the preset high-risk behavior pattern library, it is judged as a high-risk behavior. Otherwise, it is considered a safe behavior.

[0016] This embodiment details the specific process by which the encoding control unit performs real-time analysis of the raw video stream data; this process is implemented based on a lightweight AI model on the edge. Extract the sequence of key points of human skeleton in video frames and identify temporal action patterns; In this embodiment, lightweight pose estimation models such as PoseNet or a pruned version of OpenPose are used to estimate the pose of the current video frame. Process and extract the contents A set of key points This set consists of coordinates and confidence level composition; A keypoint sequence within a continuous time period is extracted using a sliding window mechanism and input into a temporal action network such as TSN or ST-GCN to generate a temporal action pattern vector. ; This refers to the probability distribution vector output by the fully connected layer of the temporal action network and normalized using Softmax, denoted as... ,in, For the predefined total number of action categories, each component This indicates that the current action belongs to the first... Confidence level of class behavior; The timing action pattern is evaluated according to preset rules; To quantify the degree of danger of a behavior, this embodiment introduces a risk entropy coefficient. The calculation logic is as follows:

[0017] in, : Represents the risk entropy coefficient at the current moment, a dimensionless value, with a range of values ​​ranging from 0 to 1. ; : Represents the set of indices of predefined high-risk behaviors within the entire set of action categories; let the entire set of action categories be . Define a subset of high-risk behaviors as If such as falling, climbing, etc., then This indicates the index of each action category in the subset. : Represents a temporal action pattern vector The Middle The normalized probability values ​​of each component, and satisfying The first term in the formula represents the cumulative confidence level that the current action belongs to all high-risk categories. : Represents the spatial distance attenuation factor; its calculation process is as follows: using a pre-calibrated homography matrix The center point of the human body in the image coordinate system Mapped to physical coordinates in the world coordinate system ,Right now ;calculate The Euclidean distance between the center point of the pre-defined danger zone and the center point of the danger zone, in centimeters; : Represents the Gaussian kernel bandwidth and standard deviation of distance attenuation, used to control the sensitive range of space risk; its value must satisfy... Ideally, it should be set to half the average shoulder width of a person, for example. ; : Represents the weighting coefficient of the behavioral probability term and the spatial location term, and must satisfy the normalization constraint. as well as For example, take ; Decision logic: If calculated Risk exceeding the preset threshold ,For example, The value range is set to If the system matches a pre-defined library of high-risk behavior patterns, it determines that the current state is a high-risk behavior. Otherwise, that is When this occurs, the system determines it as a safe behavior; By introducing an analysis mechanism based on skeletal key points and temporal actions, the system can quickly and accurately identify semantic events with security risks on the edge without relying on high-performance cloud servers. Compared with traditional motion detection, this method greatly reduces the false alarm rate caused by changes in light and shadow or pet activities, and achieves accurate understanding of the behavioral intent of the monitored object.

[0018] Example 3: The semantically aware high-fidelity encoding and priority transmission process is as follows: The human body region is determined based on the sequence of key points of the human skeleton, and this region is marked as the foreground region of interest. Assign a first quantization parameter to the foreground region of interest that is lower than that to the background region, and increase the encoding frame rate; Insert instantaneous decoding refresh frames into the current encoding sequence and add forward error correction redundancy.

[0019] This embodiment details the specific implementation of semantically aware high-fidelity coding and priority transmission process; this process aims to ensure the integrity of core information through non-uniform coding when network resources are limited. The human body region is determined based on the sequence of key points of the human skeleton, and this region is marked as the foreground region of interest. In this embodiment, the data obtained based on the aforementioned steps... Calculate the minimum bounding rectangle (BoundingBox) that covers all key points, and expand it outward by a preset number of pixels. The region of interest (ROI) is defined as the foreground region; the portion of a video frame other than the ROI is defined as the background region. Assign a first quantization parameter to the foreground region of interest that is lower than that to the background region, and increase the encoding frame rate; To optimize bitrate allocation, this embodiment employs a differentiated quantization strategy; Defined here The basic quantization parameter is dynamically generated by the encoder based on the current network bandwidth detection results, and its value is usually an integer. For the H.264 / H.265 standard, the lower the bandwidth... The larger; Let the basic quantization parameter be The quantization parameters of the ROI region Quantization parameters of the background region The settings are as follows:

[0020]

[0021] in, The quantization parameters for the region of interest in the foreground must meet the following requirements. ; The quantization parameters for the background region must meet the following requirements. ; The basic quantization parameters are dynamically generated by the encoder based on the current network bandwidth. Image enhancement step size indicates the region of interest ( Reduce the quantization step size; to ensure foreground sharpness, it is recommended to use integer values. ; Background compression step size: This indicates the magnitude by which the quantization step size is increased in the background region; to maximize bandwidth savings, it is recommended to use an integer value. ; At the same time, the encoder will increase the video encoding frame rate. Adjust to maximum value For example, 60fps, to capture rapidly changing motion details; Insert instantaneous decoding refresh frames into the current encoding sequence and add forward error correction redundancy; The encoder immediately inserts an Instant Decoding Refresh (IDR) frame the instant a high-risk behavior is detected. The purpose of the IDR frame is to cut off the reference frame chain and ensure that even if there is packet loss in the previous transmission, the decoder can immediately recover the image from the current moment to prevent the spread of screen tearing. In addition, the transport layer enables Forward Error Correction Redundancy (FEC), increasing the proportion of redundant check packets. Upgraded to ; This process sacrifices background image quality for high definition in the core action area, and significantly enhances the video stream's resilience to packet loss in weak network environments by adding redundancy and a forced refresh mechanism. This ensures that in the critical few seconds when a child falls or has an accident, guardians can see clear and continuous action details, rather than pixelated or stuttering images.

[0022] Example 4: Privacy-enhanced low-bandwidth coding combined with conventional transmission processes, specifically any of the following: A dynamic skeletal graphic is reconstructed based on the sequence of key points in the human skeleton, and then used to replace the original video footage for encoding and transmission; or, The raw video stream data is globally blurred before being encoded and transmitted. At the same time, the video encoding frame rate is adjusted to a preset low threshold.

[0023] This embodiment details the specific implementation of privacy-preserving low-bandwidth coding and conventional transmission processes; this process is applicable to normal security scenarios. This process specifically includes one of the following two options: Option 1: Reconstruct a dynamic skeleton graphic based on the sequence of key points of the human skeleton; In this mode, the system no longer transmits any raw pixel data; the encoding control unit will... Key point coordinates are mapped onto a pre-set 2D or 3D virtual puppet model. Specifically, a motion redirection algorithm is used to convert the key point coordinates into rotation angle vectors of the virtual puppet's skeletal joints, and inverse dynamic constraints are applied to limit the range of motion of the joints in order to drive the pre-bound skeletal mesh to render movements that conform to human kinematics, generating a composite video image containing only a solid color background and a line puppet, or directly transmitting key point coordinate data for rendering by the receiving end. Option 2: Perform global blurring on the original video stream data; In this mode, Gaussian blur or mosaic filters are used to... Full-frame processing is performed, making faces and specific environmental details in the image unrecognizable, retaining only the general outlines and changes in light and shadow; At the same time, the video encoding frame rate is adjusted to a preset low threshold; Regardless of which scheme is used, the encoder will output the video encoding frame rate to the transmission execution unit. Reduced to a low threshold ,For example or Meanwhile, the front-end video capture unit continues to operate at a high frame rate for real-time behavior analysis. In safe scenarios such as when children are sleeping or playing quietly, this process reduces the video bitrate by more than 90%, greatly saving network bandwidth and storage costs. More importantly, by transmitting skeleton diagrams or blurred images, the system fully respects the privacy of the monitored individuals while maintaining a sense of confirmation that the monitoring is in place, avoiding the psychological pressure brought about by 24 / 7 high-definition monitoring.

[0024] Example 5: The specific process by which the transmission execution unit executes the network service quality policy is as follows: When the video encoded stream corresponds to the first control command, the data packet is marked as the highest quality of service priority and the maximum available bandwidth is requested; When the video encoded stream corresponds to the second control command, the data packet is marked as a best-effort quality of service level, and redundant bandwidth is released.

[0025] This embodiment, based on embodiment 3 or 4, further elaborates on the process by which the transmission execution unit executes the network service quality (QoS) policy; When the video encoded stream corresponds to the first control command, the data packet is marked as the highest quality of service priority and the maximum available bandwidth is requested; In this embodiment, the system writes a highest priority flag, such as CS6 or EF class, into the Differential Service Code Point (DSCP) field of the IP packet header. Simultaneously, the system sends a bandwidth reservation request to the router or gateway, for example, by calling a gateway interface supporting the TR-069 protocol or UPnPQoS standard to add the MAC address of the monitoring device to a high-priority queue, or by actively probing and occupying the maximum available bandwidth in the current network environment using WebRTC's congestion control mechanism. This ensures that high-risk video streams have absolute transmission priority. When the video encoded stream corresponds to the second control command, the data packet is marked as a best-effort quality of service level, and redundant bandwidth is released; At this point, the system marks the DSCP field as BestEffort, i.e., 0x00, indicating that the data stream can be preferentially dropped by network devices when congested; the system actively releases the bandwidth resources previously occupied, making them available for other network applications within the home. Through deep integration with the network layer, this system achieves direct control of network layer behavior by application layer semantics. This mechanism ensures that, in situations where home network bandwidth is limited, such as when multiple people are online at the same time, the monitoring system can intelligently schedule resources, avoiding normal monitoring from crowding out the normal network experience, while ensuring uninterrupted transmission channels in emergency situations.

[0026] Example 6: The encoding control unit is also configured to perform audio-assisted analysis: Synchronously analyze the collected audio data; When a behavior is determined to be safe, but the audio characteristics exceed a preset threshold or match a preset dangerous audio template, the analysis result will be forcibly corrected to a high-risk behavior, and a first control command will be generated.

[0027] This embodiment is a further optimization of embodiment 2, adding an audio-assisted analysis function to solve the problems of visual blind spots or visual misjudgment; Synchronously analyze the collected audio data; The system strictly aligns the acquired audio and video frames in time based on a unified timestamp, ensuring that audio features and visual actions occur within the same time window; the system also acquires ambient sound signals while acquiring video. And extract audio feature vectors such as Mel frequency cepstral coefficients (MFCC) or sound pressure level energy values. ; When a behavior is determined to be safe, but the audio characteristics exceed a preset threshold or match a preset dangerous audio template, the analysis result will be forcibly corrected to a high-risk behavior, and a first control command will be generated. This embodiment sets up audio correction logic; assuming the video analysis results... However, audio analysis satisfies any of the following conditions: Instantaneous energy exceedance: The instantaneous decibel value of the audio signal. ,in, A preset threshold, such as the sound of a scream or a heavy object falling. ; Template matching: With preset danger sound templates such as help, crying, or breaking glass. similarity Among them, similarity The cosine similarity algorithm is used for calculation, and the specific formula is as follows:

[0028] in, : Refers to the feature vector of the currently acquired audio. ; : Refers to the feature vector of a preset dangerous audio template ; : Represent vectors respectively sum vector In the Numerical components in each feature dimension; : Represents the total dimension of the eigenvectors. If Mel-frequency cepstral coefficients are used, typically... ; The preset similarity threshold has a range of values. This embodiment is preferably... ; The system will ignore the security conclusions of the video analysis and force the final judgment result. Revised to high-risk behavior; This multimodal fusion mechanism effectively compensates for the limitations of single visual monitoring. For example, when a child falls in a blind spot or is covered by a blanket or other obstruction, posing a risk of suffocation, the visual algorithm may misjudge the child as still or safe. However, audio analysis, such as crying or struggling sounds, can promptly trigger alarms and high-definition transmission, significantly improving the system's safety redundancy.

[0029] Example 7: It also includes a terminal presentation unit, configured as follows: Receive and decode the video encoded stream sent by the transmission execution unit; When the decoded result is a skeleton graphic or a blurred image processed by a privacy-enhanced low-bandwidth encoding process, it is displayed normally. When the decoded video is processed through a high-fidelity encoding process and carries a high-risk behavior indicator, a pop-up alarm will be triggered and high-definition real-time footage and cached clips will be played.

[0030] This embodiment relates to a terminal presentation unit and describes in detail the interaction logic of user terminals such as guardian's mobile phone or PC; Receive and decode the video encoded stream sent by the transmission execution unit; The terminal device receives the bitstream from the cloud or the edge and identifies the current encoding mode based on the metadata in the bitstream; When the decoded result is a skeleton graphic or a blurred image processed by a privacy-enhanced low-bandwidth encoding process, it is displayed normally. If the received data is a skeleton image or a fuzzy stream, the terminal is in normal monitoring mode; the interface operates in low power mode, displaying only abstract dynamic graphics or low-resolution overviews, without triggering strong alerts; this allows the caregiver to check on the child's status at any time without feeling anxious or disturbed. When the decoded video is processed by a high-fidelity encoding process and carries a high-risk behavior indicator, a pop-up alarm will be triggered and high-definition real-time footage and cached clips will be played. Upon receiving a high-definition stream with high-risk metadata tags, the terminal immediately switches to emergency response mode: Strong alert: Triggers device vibration, plays a high-frequency alarm sound, and pops up a full-screen notification; High-definition playback: The player not only displays the current real-time high-definition picture, but also automatically splices and plays the pre-buffered clips stored before the event occurred, ensuring that the guardian can see the whole process of the danger, such as the cause of the fall. This tiered presentation mechanism aligns with users' psychological models; it avoids ineffective, "boy who cried wolf" type of interference, ensuring that every major alarm is a genuinely concerning event; simultaneously, combined with real-time pop-ups and retrospective functionality featuring high-definition visuals, it enables guardians to quickly understand the details of the emergency and take appropriate rescue measures, demonstrating the system's ultimate delivery of its core value.

[0031] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention.

Claims

1. A behavior monitoring system based on real-time video image transmission, characterized in that, It includes a video acquisition unit, an encoding and control unit, and a transmission execution unit: The video acquisition unit is configured to acquire and output raw video stream data of the monitored area. The encoding control unit is configured to perform real-time analysis of the raw video stream data and dynamically generate video encoding control strategies based on the analysis results. Video encoding control strategies include the following two types: In response to the analysis results indicating high-risk behavior, a first control command is generated to trigger a semantically aware high-fidelity encoding and priority transmission process on the raw video stream data; In response to the analysis results indicating security behavior, a second control command is generated to trigger the execution of privacy-preserving low-bandwidth encoding and regular transmission processes on the raw video stream data; The transmission execution unit is configured to receive the video encoded stream processed by the encoding control unit and execute the network service quality policy corresponding to the first control instruction or the second control instruction to complete the video image transmission.

2. The behavior monitoring system based on real-time video image transmission according to claim 1, characterized in that, The specific process by which the encoding control unit performs real-time analysis of the raw video stream data is as follows: Extract the sequence of key points of human skeleton in video frames and identify temporal action patterns; The timing action pattern is evaluated according to preset rules. If it matches the preset high-risk behavior pattern library, it is judged as a high-risk behavior. Otherwise, it is considered a safe behavior.

3. The behavior monitoring system based on real-time video image transmission according to claim 2, characterized in that, The semantically aware high-fidelity encoding and priority transmission process is as follows: The human body region is determined based on the sequence of key points of the human skeleton, and this region is marked as the foreground region of interest. Assign a first quantization parameter to the foreground region of interest that is lower than that to the background region, and increase the encoding frame rate; Insert instantaneous decoding refresh frames into the current encoding sequence and add forward error correction redundancy.

4. The behavior monitoring system based on real-time video image transmission according to claim 2, characterized in that, The privacy-preserving low-bandwidth encoding and the regular transmission process are specifically one of the following: A dynamic skeletal graphic is reconstructed based on the sequence of key points in the human skeleton, and then used to replace the original video footage for encoding and transmission; or, The raw video stream data is globally blurred before being encoded and transmitted. At the same time, the video encoding frame rate is adjusted to a preset low threshold.

5. The behavior monitoring system based on real-time video image transmission according to claim 3 or 4, characterized in that, The specific process by which the transmission execution unit executes the network service quality policy is as follows: When the video encoded stream corresponds to the first control command, the data packet is marked as the highest quality of service priority and the maximum available bandwidth is requested; When the video encoded stream corresponds to the second control command, the data packet is marked as a best-effort quality of service level, and redundant bandwidth is released.

6. The behavior monitoring system based on real-time video image transmission according to claim 2, characterized in that, The encoding control unit is also configured to perform audio-assisted analysis: Synchronously analyze the collected audio data; When a behavior is determined to be safe, but the audio characteristics exceed a preset threshold or match a preset dangerous audio template, the analysis result will be forcibly corrected to a high-risk behavior, and a first control command will be generated.

7. The behavior monitoring system based on real-time video image transmission according to claim 1, characterized in that, It also includes a terminal presentation unit, configured as follows: Receive and decode the video encoded stream sent by the transmission execution unit; When the decoded result is a skeleton graphic or a blurred image processed by a privacy-enhanced low-bandwidth encoding process, it is displayed normally. When the decoded video is processed through a high-fidelity encoding process and carries a high-risk behavior indicator, a pop-up alarm will be triggered and high-definition real-time footage and cached clips will be played.

Citation Information

Patent Citations

  • Video transmission quality evaluation and optimization method and system based on user behavior indexes

    CN115278354A

  • Tunnel electromechanical construction visual image processing method and system

    CN118135480A

  • Video coding method and device for privacy protection encryption, equipment and storage medium

    CN119232940A

  • Mobile terminal video monitoring optimized transmission method and system

    CN119893054A

  • Intelligent mine supervision system based on 5G communication and supervision method thereof

    CN120181506A