Hot-pluggable AI perception extension system suitable for multiple types of display terminals
Patent Information
- Application Number
- CN202511128422.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-13
- Publication Date
- 2025-11-21
AI Technical Summary
传统显示终端缺乏可扩展的AI能力,硬件升级困难,无法适应快速迭代的AI算法,端云协同效率低下,隐私保护难以保障,且现有方案存在隐私泄露风险和资源浪费。
设计了一种可热插拔的AI感知扩展系统,包含显示终端模块、AI感知模块和端云协同学习模块,通过磁吸式机械接口实现模块的灵活连接,结合差分隐私联邦学习和跨终端渲染引擎,实现本地推理与云端计算的动态协同,保障用户数据安全。
实现了显示终端的灵活AI扩展,降低了硬件升级成本,提升了资源利用率,满足了实时交互需求,并通过差分隐私技术保障了用户数据安全,适应了快速迭代的AI算法要求。
Smart Images

Figure CN120994343A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of artificial intelligence technology, specifically relating to a hot-swappable AI perception extension system applicable to multiple types of display terminals. Background Technology
[0002] As AI models are gradually moved from the cloud to terminal devices, edge computing power has been significantly improved, requiring ordinary display terminals to expand their AI functions to provide real-time response and offline capabilities. Traditional LCD TVs only receive HDMI / DP video signals and lack cameras, microphones, and AI computing units, making real-time emotional interaction with users impossible. Existing smart TVs with cameras (CN110748305A) have the camera fixed to the top of the unit, making it impossible to upgrade or replace; once damaged, the entire unit must be returned for repair. CN214067565U only implements mechanical quick-release and does not solve the integrated upgrade of "perception-computing-display"; CN113429276A only proposes virtual human display and does not solve cross-terminal hot-swapping and edge-cloud dual-mode inference.
[0003] Traditional displays lack scalable AI capabilities, requiring users to purchase complete devices with integrated AI functions (such as smart meeting tablets), making it impossible to reuse existing display equipment. Fixed integration solutions lead to difficulties in hardware upgrades and an inability to adapt to rapidly iterating AI algorithms. On-device deployment of multiple sensors such as RGB / ToF / infrared requires high computing power, but mobile devices have limited NPU resources. While cloud-based centralized inference has sufficient computing power, high round-trip latency cannot meet the needs of real-time gesture interaction and low-latency dialogue. At the same time, the idle computing power of terminals is not fully utilized, resulting in overall resource waste.
[0004] GDPR, personal information protection laws, and other regulations require "minimum, revocable, and forgettable data collection." Most existing edge-cloud collaboration solutions directly upload raw video / audio to the cloud, posing a risk of privacy leaks; if traditional federated learning is used, it faces problems such as gradient inversion attacks, high communication overhead, and slow model convergence. Summary of the Invention
[0005] To address the aforementioned problems in the existing technology, this invention provides a hot-swappable AI sensing extension system applicable to multiple types of display terminals. The objective of this invention can be achieved through the following technical solutions: The display terminal module integrates HDMI, DP, USB-C video input interfaces and a power supply interface. It is used as a normal display device before the AI sensing module is inserted and reverts to a normal display device after the AI sensing module is removed. The AI perception module is hot-swappably connected to the display terminal module via a magnetic mechanical interface. It has a built-in microphone array, RGB camera, ToF depth camera, infrared camera, edge AI computing unit, wireless communication unit and security chip, and is used to collect multimodal user data in real time and perform local inference calculations. The edge-cloud collaborative learning module is communicatively coupled with the AI perception module, including an edge model for offline dialogue and basic scene recognition, and a cloud model for complex reasoning and global knowledge base invocation; based on the local data shared by the edge model, the scene recognition capability is incrementally optimized through a differential privacy federated continuous learning mechanism; The cross-terminal rendering engine automatically obtains extended display recognition data from the display terminal and matches the target terminal's resolution and refresh rate in real time.
[0006] Specifically, the magnetic mechanical interface consists of a permanent magnet array, a Hall sensor, and a probe connector. The permanent magnet array is used for automatic alignment and maintaining the adsorption force. The Hall sensor is used to output a recognition signal at the moment of contact, triggering the display terminal module to pop up an interactive prompt asking "Enable AI?". The probe connector simultaneously carries the power supply, data channel, and video return channel, realizing zero-delay signal switching during hot-plugging.
[0007] Specifically, the edge model is deployed within the edge AI computing unit, and its specific structure and operation mechanism include: a basic perception and interaction model layer, a collaborative reasoning and task dynamic allocation engine, and a continuous learning module integrating differential privacy; the basic perception and interaction model layer consists of a lightweight pre-trained neural network for processing local real-time tasks; the collaborative reasoning and task dynamic allocation engine is used to balance edge computing power and AI tasks, and dynamically schedule the execution path of AI tasks; the continuous learning module incrementally optimizes scene recognition capabilities.
[0008] Specifically, the basic perception and interaction model layer consists of a video perception submodule, an audio perception submodule, and a local dialogue management and execution submodule. The video perception submodule performs real-time face detection, identity recognition, gesture and posture recognition, and user departure detection and distance judgment based on the RGB camera, ToF depth camera, and infrared camera. The audio perception submodule uses microphone array data to perform sound source localization, speaker separation, keyword wake-up, and offline speech recognition for preset command sets. The local dialogue management and execution submodule parses the commands recognized from offline speech recognition and executes the corresponding local operations.
[0009] Specifically, the collaborative reasoning and task dynamic allocation engine evaluates the resource requirements and computational complexity of the tasks received by the AI perception module and divides the received tasks into local processing tasks and cloud candidate tasks through adaptive hierarchical offloading.
[0010] Specifically, the continuous learning module retains sampling points of the multimodal raw data based on the perceptual contribution on the edge side, performs sparsification on the gradient tensor and then performs nonlinear quantization, and dynamically adjusts the retention rate and quantization step size; the compression strategy is driven by the real-time bandwidth detector; the compression and decompression process is completed by the coprocessor built into the edge NPU; A lightweight Transformer network is deployed, with each encoder containing a local multi-head attention mechanism and a dynamic mask generation layer. The local multi-head attention mechanism divides the input sequence into segments and performs self-attention computation within each segment, reducing computational complexity. The dynamic mask generation layer outputs an importance score matrix based on attention weights and generates gradient-updated masks.
[0011] Specifically, the cloud model includes a gradient decompression and calibration engine, a privacy-enhanced federated aggregation unit, a privacy budget scheduler, and a secure collaboration interface; The gradient decompression and calibration engine receives the compressed gradient data stream uploaded from the receiving end and performs Huffman decoding and sparse matrix reconstruction; it compensates for the information loss caused by gradient sparsity through an adaptive calibration algorithm; the privacy-enhanced federated aggregation unit performs gradient fusion with three levels of noise injection. The privacy budget scheduler tracks the cumulative privacy consumption of edge devices in real time. When it exceeds a preset threshold, it automatically switches to local incremental training mode and pauses gradient upload from the edge device. It also triggers the knowledge distillation channel to compress cloud knowledge into a miniature model and send it to the edge device. The secure collaboration interface interacts with the edge security chip through a hardware encryption channel.
[0012] Specifically, after performing gradient decompression in the cloud, the gradient decompression and calibration engine adopts a dynamic calibration mechanism based on KL divergence to transform information loss compensation into KL divergence minimization. Using the previous iteration parameters of the global model as a priori, it constructs the loss function of the validation set and solves the calibration factor in real time through online gradient descent.
[0013] Specifically, the three-level noise injection includes local noise by adding Gaussian noise to the end-side compressed gradient, aggregated noise by adding calibration noise to the global gradient in the cloud before federated averaging, and sparse masking noise by performing before parameter updates. By generating a shared key, each round of gradients is homomorphically encrypted before being uploaded. The cloud only decrypts the aggregated results, not the gradients of individual devices. The cloud performs weighted aggregation based on timestamps and version numbers.
[0014] Specifically, the secure collaboration interface establishes a two-way encrypted connection, and the terminal certificate is stored in the EEPROM of the edge security chip; the cloud verifies the firmware integrity of the edge security chip through a data center proof method, and refuses access to devices that fail the verification; the edge generates a one-time proof token through the security chip, and all keys, random seeds and noise parameters are synchronized with the sealed storage area of the edge security chip through the cloud hardware security module.
[0015] This invention, through an innovative modular design, enables flexible expansion of display terminals and AI sensing capabilities, solving the problems of difficult hardware upgrades and low efficiency in edge-cloud collaboration in traditional devices. Hot-swappable functionality allows users to obtain real-time AI interaction capabilities without replacing the entire device, significantly reducing upgrade costs. Simultaneously, the dynamic collaborative mechanism between local inference and cloud computing balances real-time performance with the demands of complex task processing, effectively improving resource utilization. Furthermore, the system employs differential privacy and federated learning technologies to ensure user data security while achieving continuous model optimization, meeting privacy protection regulations. The overall solution features high compatibility, easy scalability, and strong privacy protection, providing an intelligent upgrade path for various types of display terminals. Attached Figure Description
[0016] To facilitate understanding by those skilled in the art, the present invention will be further described below with reference to the accompanying drawings.
[0017] Figure 1 This is a schematic diagram of the structure of a hot-swappable AI sensing extension system applicable to multiple types of display terminals according to the present invention; Figure 2 This describes the application scenarios of the projector and television terminals of the present invention. Detailed Implementation
[0018] To further illustrate the technical means and effects of the present invention in achieving its intended purpose, the following detailed description of the specific implementation methods, structures, features, and effects of the present invention, in conjunction with the accompanying drawings and preferred embodiments, is provided.
[0019] Please see Figure 1-2 A hot-swappable AI sensing extension system suitable for multiple types of display terminals, comprising: The display terminal module integrates HDMI, DP, USB-C video input interfaces and a power supply interface. It is used as a normal display device before the AI sensing module is inserted and reverts to a normal display device after the AI sensing module is removed. The AI perception module is hot-swappably connected to the display terminal module via a magnetic mechanical interface. It has a built-in microphone array, RGB camera, ToF depth camera, infrared camera, edge AI computing unit, wireless communication unit and security chip, and is used to collect multimodal user data in real time and perform local inference calculations. The edge-cloud collaborative learning module is communicatively coupled with the AI perception module, including an edge model for offline dialogue and basic scene recognition, and a cloud model for complex reasoning and global knowledge base invocation; based on the local data shared by the edge model, the scene recognition capability is incrementally optimized through a differential privacy federated continuous learning mechanism; The cross-terminal rendering engine automatically obtains extended display recognition data from the display terminal and matches the target terminal's resolution and refresh rate in real time.
[0020] In this embodiment, the display terminal includes, but is not limited to, projectors, LCD TVs, conference screens, and vehicle central control screens, and has at least one video input interface (HDMI / DP / USB-C); the hot-swappable AI sensing module integrates within its housing: a microphone array (≥4 MEMS), an RGB camera (≥8 MP), a ToF depth camera, a 940nm infrared camera, an edge AI computing unit (SoC+NPU, peak computing power 6-15 TOPS), a wireless communication unit (Wi-Fi 6 / BT 5.3 / 4G / 5G eSIM), and a magnetic mechanical interface + 16-pin pogo pin (power supply + USB 3.2 + DP Alt-mode); Dual-mode large model for edge-cloud: Edge model: 7B-13B quantized LLM, supporting offline emotional dialogue and digital human-driven; Cloud model: 70B+ full-precision LLM, supporting long memory and complex inference; Adaptive switching algorithm: dynamically selects the inference path based on network RTT, bandwidth, user privacy switch, and remaining terminal battery power.
[0021] Cross-terminal digital human rendering engine: Receives the display terminal's EDID and automatically matches the resolution and refresh rate; supports multi-channel output for 2D and 3D holographic fans and AR glasses; The hot-swap workflow is as follows: S1: When the user brings the AI sensing module close to the groove on the top of the display terminal, the magnet will automatically align. S2: Hall sensor identifies module type → Terminal prompts "Enable AI Companion?" S3: The terminal transmits EDID back via HDMI-CEC or USB-C DP Alt-mode; S4: The AI perception module starts the local model and projects a 1:1 virtual human onto the TV screen; S5: User voice / gesture interaction → End-to-cloud dual-mode inference → Real-time facial expression and motion feedback for digital humans; S6: When the user unplugs the AI sensing module, the terminal reverts to normal TV functions with zero residue.
[0022] Specifically, the magnetic mechanical interface consists of a permanent magnet array, a Hall sensor, and a probe connector. The permanent magnet array is used for automatic alignment and maintaining the adsorption force. The Hall sensor is used to output a recognition signal at the moment of contact, triggering the display terminal module to pop up an interactive prompt asking "Enable AI?". The probe connector simultaneously carries the power supply, data channel, and video return channel, realizing zero-delay signal switching during hot-plugging.
[0023] Specifically, the edge model is deployed within the edge AI computing unit, and its specific structure and operation mechanism include: a basic perception and interaction model layer, a collaborative reasoning and task dynamic allocation engine, and a continuous learning module integrating differential privacy; the basic perception and interaction model layer consists of a lightweight pre-trained neural network for processing local real-time tasks; the collaborative reasoning and task dynamic allocation engine is used to balance edge computing power and AI tasks, and dynamically schedule the execution path of AI tasks; the continuous learning module incrementally optimizes scene recognition capabilities.
[0024] Specifically, the basic perception and interaction model layer consists of a video perception submodule, an audio perception submodule, and a local dialogue management and execution submodule. The video perception submodule performs real-time face detection, identity recognition, gesture and posture recognition, and user departure detection and distance judgment based on the RGB camera, ToF depth camera, and infrared camera. The audio perception submodule uses microphone array data to perform sound source localization, speaker separation, keyword wake-up, and offline speech recognition for preset command sets. The local dialogue management and execution submodule parses the commands recognized from offline speech recognition and executes the corresponding local operations.
[0025] Specifically, the collaborative reasoning and task dynamic allocation engine evaluates the resource requirements and computational complexity of the tasks received by the AI perception module and divides the received tasks into local processing tasks and cloud candidate tasks through adaptive hierarchical offloading.
[0026] Specifically, the continuous learning module retains sampling points of the multimodal raw data based on the perceptual contribution on the edge side, performs sparsification on the gradient tensor and then performs nonlinear quantization, and dynamically adjusts the retention rate and quantization step size; the compression strategy is driven by the real-time bandwidth detector; the compression and decompression process is completed by the coprocessor built into the edge NPU; A lightweight Transformer network is deployed, with each encoder containing a local multi-head attention mechanism and a dynamic mask generation layer. The local multi-head attention mechanism divides the input sequence into segments and performs self-attention computation within each segment, reducing computational complexity. The dynamic mask generation layer outputs an importance score matrix based on attention weights and generates gradient-updated masks.
[0027] In this embodiment, as Figure 2 As shown, the edge model is deployed in the edge AI computing unit of the AI perception module (adopting a heterogeneous SoC architecture, integrating an NPU computing unit, with a peak computing power of 6-15 TOPS). The video perception submodule synchronously calls an RGB camera (8MP, 30fps), a ToF depth camera (VGA resolution, 10m ranging accuracy ±1%), and an infrared camera (940nm wavelength, supports low-light imaging); the audio perception submodule connects to a 4-channel MEMS microphone array. When the collaborative reasoning and task dynamic allocation engine receives a task, it monitors the NPU utilization, memory usage and battery level in real time. It calculates the task computation cost through a complexity estimator (based on the FLOPs prediction model). If the local resources meet the latency constraints, the task is allocated to the NPU; otherwise, it is packaged into an encrypted task packet and uploaded to the cloud via the wireless unit. When the user triggers the "gesture to adjust volume" action: The video perception submodule captures hand depth data through a ToF camera and combines it with RGB images to recognize the "thumb swipe up" gesture in real time. The task allocation engine determined it to be a low-complexity local task; The local dialogue management submodule generates volume+ commands to directly control the audio output of the display terminal. The continuous learning module records gradient data for this scenario, which is then compressed and uploaded to the cloud via the privacy budget scheduler.
[0028] Specifically, the cloud model includes a gradient decompression and calibration engine, a privacy-enhanced federated aggregation unit, a privacy budget scheduler, and a secure collaboration interface; The gradient decompression and calibration engine receives the compressed gradient data stream uploaded from the receiving end and performs Huffman decoding and sparse matrix reconstruction; it compensates for the information loss caused by gradient sparsity through an adaptive calibration algorithm; the privacy-enhanced federated aggregation unit performs gradient fusion with three levels of noise injection. The privacy budget scheduler tracks the cumulative privacy consumption of edge devices in real time. When it exceeds a preset threshold, it automatically switches to local incremental training mode and pauses gradient upload from the edge device. It also triggers the knowledge distillation channel to compress cloud knowledge into a miniature model and send it to the edge device. The secure collaboration interface interacts with the edge security chip through a hardware encryption channel.
[0029] Specifically, after performing gradient decompression in the cloud, the gradient decompression and calibration engine adopts a dynamic calibration mechanism based on KL divergence to transform information loss compensation into KL divergence minimization. Using the previous iteration parameters of the global model as a priori, it constructs the loss function of the validation set and solves the calibration factor in real time through online gradient descent.
[0030] In this embodiment, the gradient decompression and calibration engine decodes the binary compressed stream (containing sparse matrix metadata in the header) uploaded from the receiving end into a sparse coordinate index set according to the Huffman coding table (preset dictionary ID matching the cloud codebook library) through Huffman decoding and reconstruction; and reconstructs the sparse gradient tensor. The sparse tensor reconstructed after gradient sparsification contains the coordinate indices and values of non-zero elements. The objective function of the KL divergence dynamic calibration mechanism is set as follows: , Where, θ t The parameters are the global model parameters from the previous round, α is the calibration factor to be solved, which is dynamically adjusted to compensate for information loss, and G... sparse The sparse gradient is the result of decompression; P is the model parameter distribution function; KL divergence measures the difference in distribution before and after calibration; G... true To achieve optimal compensation for information loss by minimizing the KL divergence, we can find the true gradient distribution.
[0031] With verification set D val Loss reduction serves as a monitoring signal: , in, ( ) represents the cross-entropy loss, and f() represents the current global model with calibrated weights. Forward propagation on the validation set, compute L calib As a supervisory signal for online gradient descent; Updated via online gradient descent: , Where η is the learning rate adaptively adjusted using the Adam optimizer; The calibration factor is updated in real time using the first-order gradient to minimize the KL objective. Output calibrated gradient: ; This yields the gradient used for the final federated aggregation.
[0032] Specifically, the three-level noise injection includes local noise by adding Gaussian noise to the end-side compressed gradient, aggregated noise by adding calibration noise to the global gradient in the cloud before federated averaging, and sparse masking noise by performing before parameter updates. By generating a shared key, each round of gradients is homomorphically encrypted before being uploaded. The cloud only decrypts the aggregated results, not the gradients of individual devices. The cloud performs weighted aggregation based on timestamps and version numbers.
[0033] Specifically, the secure collaboration interface establishes a two-way encrypted connection, and the terminal certificate is stored in the EEPROM of the edge security chip; the cloud verifies the firmware integrity of the edge security chip through a data center proof method, and refuses access to devices that fail the verification; the edge generates a one-time proof token through the security chip, and all keys, random seeds and noise parameters are synchronized with the sealed storage area of the edge security chip through the cloud hardware security module.
[0034] In this embodiment, the gradient is encrypted using a shared homomorphic public key on the device side, and the cloud only decrypts the aggregation result to obtain the global gradient, ensuring that the gradient is not visible on a single device.
[0035] In terms of security, the system ensures the reliability of edge-cloud data interaction through a multi-layered protection mechanism. The edge-side security chip employs hardware-level isolation technology to completely separate sensitive operations from the ordinary computing environment. The cloud-based hardware security module is equipped with a quantum random number generator to dynamically generate encryption keys for each session, effectively resisting brute-force attacks and replay attacks. Simultaneously, the system implements an identity authentication protocol based on zero-knowledge proofs, ensuring the authenticity of device identities while preventing the leakage of sensitive information. In actual operation, all data transmissions use dynamically adjusted encryption strength, automatically adapting to the network environment and task importance to ensure an optimal balance between performance and security.
[0036] The above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention in any way. Although the present invention has been disclosed above with reference to preferred embodiments, it is not intended to limit the present invention. Any person skilled in the art can make some modifications or alterations to the above-disclosed technical content to create equivalent embodiments without departing from the scope of the present invention. Any simple modifications, equivalent changes and alterations made to the above embodiments based on the technical essence of the present invention without departing from the scope of the present invention shall still fall within the scope of the present invention.
Claims
1. A hot-pluggable AI-aware extension system suitable for multi-type display terminals, characterized in that, The application relates to a display terminal module, an AI sensing module and an end-cloud collaborative learning module. The display terminal module is integrated with HDMI, DP, USB-C video input interfaces and power supply interfaces, and is used as a common display device before the AI sensing module is inserted and is restored to a common display device after the AI sensing module is pulled out. The AI sensing module is electrically connected to the display terminal module through a magnetic mechanical interface, is internally provided with a microphone array, an RGB camera, a ToF depth camera, an infrared camera, an end-side AI operation unit, a wireless communication unit and a security chip, and is used for collecting multi-modal user data in real time and performing local inference operation. The end-cloud collaborative learning module is communicatively coupled with the AI sensing module, comprises an end-side model for offline dialogue and basic scene recognition and a cloud model for complex inference and global knowledge base calling, and performs incremental optimization on scene recognition ability based on local data shared by the end-side model through a differential privacy federated continuous learning mechanism. The cross-terminal rendering engine automatically acquires extended display recognition data of the display terminal and matches target terminal resolution and refresh rate in real time.
2. The method of claim 1, wherein, The magnetic mechanical interface is composed of a permanent magnet array, a Hall sensor and a probe connector, wherein the permanent magnet array is used for automatic alignment and maintenance of adsorption force; the Hall sensor is used for outputting an identification signal at the moment of adhesion to trigger the display terminal module to pop up an "whether to enable AI" interactive prompt; and the probe connector simultaneously bears power supply, data channel and video return channel, and realizes zero-delay signal switching in the process of hot plugging.
3. The method of claim 1, wherein, The end-side model is arranged in the end-side AI operation unit, and specific constitution and operation mechanism thereof comprises a basic perception and interaction model layer, a collaborative inference and task dynamic distribution engine and a continuous learning module integrated with differential privacy; the basic perception and interaction model layer is composed of a lightweight pre-trained neural network and is used for processing local real-time tasks; the collaborative inference and task dynamic distribution engine is used for balancing end-side computing power and AI tasks and dynamically scheduling an execution path of the AI tasks; and the continuous learning module performs incremental optimization on scene recognition ability.
4. The method of claim 3, wherein, The basic perception and interaction model layer is composed of a video perception sub-module, an audio perception sub-module and a local dialogue management and execution sub-module; The video perception sub-module performs real-time face detection, identity recognition, gesture and posture recognition based on the RGB camera, the ToF depth camera and the infrared camera, and performs user off-seat detection and distance judgment based on ToF and infrared data; The audio perception sub-module performs sound source positioning, speaker separation, keyword wake-up and offline speech recognition on a preset instruction set by using microphone array data; and the local dialogue management and execution sub-module analyzes the instruction recognized by the offline speech recognition and executes corresponding local operation.
5. The method of claim 3, wherein, The collaborative inference and task dynamic distribution engine evaluates resource demand and calculation complexity of the tasks received by the AI sensing module, divides the received tasks into local processing tasks and cloud candidate tasks through adaptive hierarchical unloading.
6. The method of claim 3, wherein, The continuous learning module retains sampling points of the multimodal raw data based on the perceived contribution on the edge side, performs sparsification on the gradient tensor and then performs nonlinear quantization, and dynamically adjusts the retention rate and quantization step size; the compression strategy is driven by the real-time bandwidth detector; the compression and decompression process is completed by the coprocessor built into the edge NPU; A lightweight Transformer network is deployed, with each encoder containing a local multi-head attention mechanism and a dynamic mask generation layer. The local multi-head attention mechanism divides the input sequence into segments and performs self-attention computation within each segment, reducing computational complexity. The dynamic mask generation layer outputs an importance score matrix based on attention weights and generates gradient-updated masks.
7. The method of claim 1, wherein, The cloud model includes a gradient decompression and calibration engine, a privacy-enhanced federated aggregation unit, a privacy budget scheduler, and a secure collaboration interface; The gradient decompression and calibration engine receives the compressed gradient data stream uploaded from the receiving end and performs Huffman decoding and sparse matrix reconstruction. The information loss caused by gradient sparsity is compensated by an adaptive calibration algorithm; The privacy-enhanced federated aggregation unit performs gradient fusion with three levels of noise injection; The privacy budget scheduler tracks the cumulative privacy consumption of edge devices in real time. When the cumulative privacy consumption exceeds a preset threshold, it automatically switches to local incremental training mode and pauses gradient uploading from the edge devices. Trigger the knowledge distillation channel to compress cloud knowledge into a miniature model and send it to the edge; The secure collaboration interface interacts with the edge security chip through a hardware encryption channel.
8. The method of claim 7, wherein, After performing gradient decompression in the cloud, the gradient decompression and calibration engine adopts a dynamic calibration mechanism based on KL divergence to transform information loss compensation into KL divergence minimization. Using the previous iteration parameters of the global model as a priori, it constructs the loss function of the validation set and solves the calibration factor in real time through online gradient descent.
9. The method of claim 7, wherein, The three-level noise injection includes local noise by adding Gaussian noise to the end-side compressed gradient, aggregated noise by adding calibration noise to the global gradient in the cloud before federated averaging, and sparse masking noise by performing before parameter updates. By generating a shared key, each round of gradients is homomorphically encrypted before being uploaded. The cloud only decrypts the aggregated results, not the gradients of individual devices. The cloud performs weighted aggregation based on timestamps and version numbers.
10. The method of claim 7, wherein, The secure collaboration interface establishes a two-way encrypted connection, with the terminal certificate stored in the EEPROM of the edge security chip. The cloud verifies the firmware integrity of the edge security chip through a data center verification method, and refuses access to devices that fail verification. The edge generates a one-time verification token through the security chip, and all keys, random seeds, and noise parameters are synchronized with the sealed storage area of the edge security chip through the cloud hardware security module.
Citation Information
Patent Citations
Deepwater riser pipe with protective shell
CN110748305A
Green process synthesis method of 3, 4, 5-trimethoxybenzaldehyde
CN113429276A
Three-dimensional interactive projector for AI voice interaction
CN214067565U
Cited By
Desktop type intelligent digital assistant terminal
CN121900861A