Robot privacy protection communication method based on multi-modal semantics and adaptive compression

CN122845259APending Publication Date: 2026-09-29WUHAN UNIV OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611102177.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-23
Publication Date
2026-09-29

AI Technical Summary

Technical Problem

[0004]本发明需要解决的技术问题是:现有的端侧全量采集后,在云端集中处理的通信方案存在以下明显缺陷:首先,该方案需要将包含儿童敏感生物特征(如面部高清图像、真实声纹)以及家庭物理环境的高清原始音视频流,直接上传至公网服务器,存在极大的隐私数据泄露风险,难以满足儿童产品严苛的数据安全要求

Benefits of technology

本发明提出了一种兼顾高隐私保护、低延迟通信与极高运转安全性的机器人多模态协同方案。针对传统方案全量上传原始音视频导致的隐私泄露与带宽拥塞缺陷,本发明在端侧完成语义张量提取后,利用硬件中断直接覆写底层内存,从介质源头彻底销毁多媒体裸数据;同时根据跨模态情绪激烈度与感知置信度执行自适应量化降维,大幅降低上行网络开销,保障了情感特征的极速无损传输。针对复杂交互环境,系统不仅在局部遮挡导致置信度偏低时自适应提取未压缩的中间层张量以支撑云端的高精度意图重构,更通过引入姿态与触觉差分算子建立了物理防卫优先路由,实现了对暴力推搡等破坏性交互的瞬态预警。此外,配合端侧底层看门狗双轨兜底逻辑,系统能够在遭遇深度拥塞或断网时自主切断网络写入权限并强制执行机械复位,全面保障了机器人在极端物理与通信环境下的综合安全性与交互连贯性。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122845259A_ABST
    Figure CN122845259A_ABST
Patent Text Reader

Abstract

The application discloses a kind of robot privacy protection communication methods based on multimodal semantics and adaptive compression, comprising: step one, acquisition multimodal physical signal, and safely isolated and stored in video ring buffer area, audio ping-pong cache queue, kinematics state cache area and tactile topology cache area;Step two, multidimensional feature extraction and data dimension reduction are carried out to multimodal physical signal, obtain facial semantic feature tensor, voiceprint semantic feature tensor, motion feature matrix and tactile topology tensor;Step three, continuously extract multidimensional feature tensor, and execute fluctuation feature extraction in space-time domain, obtain four-dimensional difference operator, calculate comprehensive emotional intensity score;Step four, construct three-dimensional joint scheduling state machine, and the feature load after being processed differently is packaged as application layer communication data frame, and is transmitted to cloud;Step five, based on cloud collaborative intent analysis and end-side double-track safe drivingThe application can realize high privacy, low delay robot embodied interaction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of robot communication and information security technology, and in particular to a robot privacy-preserving communication method based on multimodal semantics and adaptive compression. Background Technology

[0002] With the rapid development of artificial intelligence and the Internet of Things (IoT) technologies, intelligent interactive robots designed for children's education and emotional companionship are becoming increasingly popular and integrated into home settings. To achieve a more natural and human-like interactive experience, modern companion robots are typically equipped with multimodal sensing hardware such as cameras and microphone arrays to comprehensively acquire interactive information such as children's facial expressions, body movements, and voice intonation. Meanwhile, because children's language expression is often fragmented and their emotional needs are complex, the industry generally adopts an edge-cloud collaborative computing architecture, leveraging large-scale language models (LLMs) deployed in the cloud and abundant cloud computing power for deep semantic understanding and interactive content generation.

[0003] Currently, the closest existing technical solution to this invention typically employs a method of full data acquisition at the edge and centralized processing in the cloud. Specifically, the robot continuously acquires the user's high-definition video and audio streams in real time using multimodal sensors on the local edge. Subsequently, the edge's main control system uses standard audio and video codec protocols (such as H.264 / H.265 video compression and AAC audio compression format) to compress and package these raw multimedia physical data. Finally, the entire audio and video data stream is uploaded to a cloud server via a home Wi-Fi network. Upon receiving the data stream, the cloud server performs centralized decoding, multimodal feature extraction, and intent recognition. Relying on a large cloud model, it derives the user's current emotion and corresponding reassurance strategies, and transmits the generated voice text and action control commands back to the robot's edge for physical execution via a downlink network.

[0004] The technical problem this invention aims to solve is that existing communication solutions that collect full data on the device side and process it centrally in the cloud have the following significant drawbacks: First, these solutions require directly uploading high-definition raw audio and video streams containing children's sensitive biometric features (such as high-definition facial images and real voiceprints) and the home's physical environment to a public network server. This poses a significant risk of privacy data leakage and fails to meet the stringent data security requirements of children's products. Second, transmitting full multimedia data continuously consumes a large amount of uplink network bandwidth. In complex home Wi-Fi environments, this can easily lead to data congestion and packet loss, resulting in excessively high latency in device-to-cloud interaction and failing to meet the real-time requirements for rapidly calming children's transient emotions under multimodal concurrent conditions. Furthermore, existing traditional video compression protocols are often driven solely by network channel conditions and cannot perceive the actual semantic value of the data. When faced with sudden intense emotions in children accompanied by network fluctuations, this can easily lead to the loss of key facial or vocal micro-expression features, making it impossible to accurately judge their emotions.

[0005] To address the issues that existing technologies can only upload original images and audio in their entirety and cannot simultaneously ensure privacy protection and extremely low-latency interaction, there is an urgent need to propose a privacy-preserving communication method for robots based on multimodal semantics and adaptive compression. Summary of the Invention

[0006] The technical problem this invention aims to solve is to address the shortcomings of existing technologies by providing a privacy-preserving robot communication method based on multimodal semantics and adaptive compression. This invention employs a method based on multimodal semantic extraction and adaptive emotion compression. By extracting multimodal features on the local device side and directly truncating and destroying the original physical media data, the risk of privacy leakage can be completely eliminated from the underlying physical link of the communication. Simultaneously, adaptive quantization compression is performed based on the intensity of the extracted emotional features, eliminating the need to upload massive amounts of original video streams. This significantly reduces network bandwidth consumption and ensures ultra-fast, lossless transmission of key emotional features, thereby achieving highly private and low-latency embodied robot interaction.

[0007] The technical solution adopted by this invention to solve its technical problem is: This invention provides a robot privacy-preserving communication method based on multimodal semantics and adaptive compression, the method comprising the following steps: Step 1: During the power-on initialization phase of the robot system, the end-side main control unit divides four consecutive address blocks in the local physical memory, which are respectively configured as a video circular buffer, an audio ping-pong buffer queue, a kinematic state buffer, and a tactile topology buffer. These blocks are used to securely isolate and store the collected multimodal physical signals, which include visual signals, auditory signals, posture kinematic signals, and tactile interaction signals. Step 2: The terminal main control unit performs multi-dimensional feature extraction and data dimensionality reduction on the multimodal physical signals to obtain facial semantic feature tensor, voiceprint semantic feature tensor, motion feature matrix and tactile topology tensor; after the multi-dimensional feature tensors are extracted and written into the secure working area of ​​static random access memory, the multimodal physical signal data stored in the local physical memory is erased. Step 3: Within a set sliding evaluation time window of a certain length, continuously extract multidimensional feature tensors and perform spatiotemporal fluctuation feature extraction on them to obtain facial deformation difference operators, auditory fluctuation difference operators, posture change difference operators, and tactile violence difference operators respectively; construct a cross-modal semantic evaluation function based on these four difference operators and calculate the comprehensive emotion intensity score. Step 4: Obtain the probability of a four-dimensional difference operator belonging to a certain predicted category through a neural network, calculate the dispersion of the four-dimensional multimodal fusion probability distribution in real time, and output the feature extraction confidence score; calculate the physical impact extreme value based on the posture mutation difference operator and the tactile violence difference operator; combine the comprehensive emotional intensity score, confidence score and physical impact extreme value to construct a three-dimensional joint scheduling state machine. The state machine is set with three main routing branches and multiple thresholds. The feature payloads processed by different branches are encapsulated into application layer communication data frames, encrypted and transmitted uplink to the cloud. Step 5: After verifying the timestamp, the cloud performs forward inference based on the application layer communication data frame by reconstructing or retaining the multi-dimensional feature tensor in the cloud-based large language model. It outputs a text byte stream representing the soothing semantics and a timing control matrix representing the physical action, and sends them down via the downlink network link. The terminal main control unit drives the motor to perform corresponding operations and performs voice broadcast according to the sent control signals.

[0008] Furthermore, the method for acquiring multimodal physical signals in step one of the present invention specifically includes: The robot system is equipped with a visual perception array, a microphone array, a six-axis inertial measurement unit, and a flexible piezoresistive sensor array. The system acquires ambient optical signals through a visual sensing array and outputs raw video frame data. Spatial sound field signals are acquired through a microphone array, and after analog-to-digital conversion, the original audio stream in pulse code modulation format is output. The built-in six-axis inertial measurement unit collects the dynamics of the fuselage space and outputs a kinematic digital stream containing three-axis acceleration and three-axis angular velocity. Physical contact pressure is collected by a flexible piezoresistive sensor array on the surface of the device, and after multiplexing and analog-to-digital conversion, a digital matrix of tactile pressure is output.

[0009] Furthermore, the specific method for signal security isolation in step one of the present invention includes: Raw video frame data from visual signals is stored in a video circular buffer; raw audio streams from auditory signals are stored in an audio ping-pong buffer queue; kinematic digital streams from posture kinematic signals are stored in a kinematic state buffer; and tactile pressure digital matrices from tactile interaction signals are stored in a tactile topology buffer. The physical address ranges containing the four independent buffers are set to be invisible to the network communication protocol stack, and the network socket layer on the end side is forcibly stripped of its read permissions for these physical address ranges.

[0010] Furthermore, the specific method of step two of the present invention includes: For visual signals, a facial landmark detection algorithm is executed to calculate the two-dimensional pixel coordinates of preset facial feature points and output the coordinates of the points in dimension 1. Facial semantic feature tensor ,in N This represents the total number of feature points identified by the algorithm. For auditory signals, pre-emphasis, framing, windowing, and discrete Fourier transform are performed on the original audio stream sequence to extract cepstral coefficient features in the frequency domain, and the output dimension is... Voiceprint semantic feature tensor ,in T This represents the number of audio frames within the time window. F Represents the dimension of the frequency band features; For the attitude kinematics signal, moving average and Kalman filtering are performed on the kinematic digital stream to remove high-frequency mechanical vibration noise, and transient components of triaxial acceleration and angular velocity are extracted to output a motion feature matrix representing the continuous changes in the terminal's spatial attitude. ; For tactile interaction signals, the operator is invoked to perform spatial domain feature pooling on the tactile pressure digital matrix in the tactile topology buffer, and outputs a tactile topology tensor representing the spatial distribution of local physical contact intensity and pressure. .

[0011] Furthermore, the specific method of step three of the present invention includes: Facial deformation difference operator Extract facial semantic feature tensors from two adjacent frames within a time window. and For the tensor contained in N For each facial feature point, calculate the Euclidean distance offset of each feature point in the two-dimensional image coordinate system; the formula is:

[0012] In the formula, and Representing the The horizontal and vertical coordinate components of a facial feature point in the current image frame. and This represents the corresponding coordinate component of the feature point in the previous image frame; Auditory wave difference operator Extracting the semantic feature tensor of voiceprint The energy matrix, which includes time and frequency dimensions, is used to calculate the statistical variance of the high-frequency band energy sequence of each audio segment within a time window. The formula is as follows:

[0013] In the formula, To evaluate the total number of valid audio frames contained within a time window, For the first The total frequency band energy of each audio frame. It is the arithmetic mean of the total energy of all audio frames within this time window; Attitude change differential operator According to the motion feature matrix ,extract The triaxial acceleration vector and triaxial angular velocity vector of each consecutive sampling period are used to extract the acceleration representing the physical impact force and the angular acceleration representing the rotational torque by calculating the first-order kinematic difference between adjacent sampling periods. These are then linearly weighted and fused, as shown in the following formula:

[0014] In the formula, and Representing the first time window The and the first During each sampling clock cycle, the fuselage is The quantized value of the acceleration in the axial direction via analog-to-digital conversion; and These represent the corresponding periods of winding. The angular velocity of the axis rotation is converted from analog to digital and quantized. and These are the linear acceleration abrupt change penalty coefficient and the angular acceleration abrupt change penalty coefficient, respectively, which are fixed in the underlying system. Tactile Violence Differential Operator According to the tactile topological tensor Real-time polling covering the fuselage Flexible piezoresistive sensor array, extracting within a time window The underlying pressure space topology matrix for each sampling period is used, and a static physical contact determination threshold is introduced. The arithmetic logic unit is invoked to perform the superthreshold energy integral of the spatial dimension, and its formula is:

[0015] In the formula, Representing the In the sampling period, the first sampling period in the physical coordinate system u Line 1 v The piezoresistive sensor unit outputs the underlying absolute pressure digital quantization value. It is a nonlinear activation algorithm; Construct a cross-modal semantic evaluation function to calculate the comprehensive emotion intensity score. Its formula is:

[0016] In the formula, , , and These are the non-negative weighting coefficients for visual, auditory, postural kinematic, and tactile representations, respectively. .

[0017] Furthermore, in step four of this invention, the output feature extraction confidence score is... The formula is:

[0018] In the formula, The total number of reserved categories set for the output layer of a neural network classification layer. The feature tensor belongs to the first The true floating-point probabilities of each predicted category; Physical impact extreme value The calculation formula is:

[0019] In the formula, and These are independent defense decision weighting coefficients assigned to the underlying firmware for kinematic shock and topological stress overload, respectively.

[0020] Furthermore, the three backbone routing branches in step four of the present invention specifically include: Set confidence level safety threshold First emotional threshold Second emotional threshold ,satisfy and physical impact red line threshold ; In the first backbone routing branch, regardless of and What state is it in when the comparator determines... At this time, the system confirms that it is experiencing destructive physical interaction or that the target object is in an extremely violent and out-of-control state; the state machine forcibly short-circuits all dimensionality reduction quantization operators, attaches a physical pass-through mode, and includes the current... and The four-dimensional physical tensor is pushed directly into the send buffer as the highest priority data stream; In the second backbone routing branch, when determining and At that time, the system confirms that there is no serious physical interference in the sensing link, and the state machine switches to the emotion-driven quantization channel; if This triggers a high-speed compression subroutine; the system attaches a tensor dimensionality reduction operator, mapping the original 32-bit single-precision floating-point feature tensor data to 8-bit unsigned integer data; the mapping equation is... Where R is the floating-point truth value, S is the dynamic scaling factor, and Z is the zero-point offset constant, these are used to generate the ultimate compressibility characteristic load; if This triggers the equalization subroutine, mapping the tensor data structure to a 16.5-bit floating-point type; if This triggers the bypass pass-through subroutine, which forces all dimensionality reduction operators to be short-circuited by the underlying logic gates, and directly outputs the full 32-bit single-precision floating-point original feature tensor. In the third backbone routing branch, when determining and When the system confirms that the current interactive environment has signal attenuation or local occlusion, resulting in insufficient confidence in the extraction of multimodal features on the edge, the state machine triggers an adaptive tensor compensation mechanism, pausing the subsequent pooling and fully connected layer forward inference calculations of the neural network, and calling the direct memory access controller to extract the intermediate layer spatial feature matrix before the pooling layer from the static random access memory. This matrix retains the uncompressed spatial topology and fine-grained frequency domain information, and its mathematical dimension is defined as... The system will compare this matrix with the current... , , and Perform underlying tensor splicing.

[0021] Furthermore, the method for verifying the timestamp in step five of the present invention includes: The cloud server's border communication gateway receives uplink application layer communication data frames and creates an independent verification sandbox in memory; the system reads the timestamp sequence from the data frame header. Get the current local clock of the cloud server. And calculate the one-way communication time overhead. If and only if If the data frame is less than the preset anti-replay attack lifecycle threshold, it is allowed to enter the dequantization parsing queue; otherwise, the cloud server directly discards the frame in the verification sandbox to block network latency tampering attacks.

[0022] Furthermore, the specific method of step five of the present invention includes: For application-layer communication data frames that pass verification, the cloud server reads the quantization precision check bit and the routing branch identifier bit. If the identifier bit indicates an ultra-fast compression state, the system starts the asymmetric inverse quantization operator and uses a pre-trained high-order interpolation algorithm to fill the matrix sparse gaps caused by low-precision quantization, reconstructing a tensor with approximately 32-bit precision. If the identifier bit indicates an adaptive tensor compensation state, the intermediate feature matrix in the high-dimensional space is directly extracted. The cloud-based large language model performs forward inference on the reconstructed or retained multimodal feature tensor, outputs a text byte stream representing soothing semantics and a temporal control matrix representing physical actions, and sends them down via the downlink network link.

[0023] The present invention provides a computer-readable storage medium, characterized in that it stores a computer program for implementing the robot privacy-preserving communication method based on multimodal semantics and adaptive compression when executed by a processor.

[0024] The beneficial effects of this invention are: This invention proposes a multimodal robot collaboration scheme that balances high privacy protection, low-latency communication, and extremely high operational safety. Addressing the privacy leaks and bandwidth congestion issues caused by the full uploading of raw audio and video in traditional solutions, this invention extracts semantic tensors on the edge and then directly overwrites the underlying memory using hardware interrupts, completely destroying raw multimedia data at the source. Simultaneously, it performs adaptive quantization and dimensionality reduction based on cross-modal emotional intensity and perceptual confidence, significantly reducing uplink network overhead and ensuring ultra-fast, lossless transmission of emotional features. For complex interactive environments, the system not only adaptively extracts uncompressed intermediate-layer tensors to support high-precision intent reconstruction in the cloud when local occlusion leads to low confidence, but also establishes a physical defense priority route by introducing gesture and tactile differential operators, enabling transient early warning of destructive interactions such as violent pushing. Furthermore, with the edge-side watchdog dual-track fallback logic, the system can autonomously sever network write permissions and force a mechanical reset when encountering deep congestion or network outages, comprehensively ensuring the robot's overall safety and interactive continuity in extreme physical and communication environments. Attached Figure Description

[0025] The present invention will be further described below with reference to the accompanying drawings and embodiments. In the accompanying drawings: Figure 1 This is a flowchart of the privacy-preserving communication and adaptive scheduling method according to an embodiment of the present invention; Figure 2This is a schematic diagram of the end-to-cloud collaborative multimodal communication system architecture according to an embodiment of the present invention. Detailed Implementation

[0026] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention.

[0027] like Figure 1 and Figure 2 As shown, the robot privacy-preserving communication method based on multimodal semantics and adaptive compression according to an embodiment of the present invention specifically includes the following steps: Step 1: Acquisition and Secure Isolation of Multimodal Physical Signals During the system power-on initialization phase, the memory management unit of the terminal main control unit statically divides four independent contiguous address blocks in the local physical memory, which are respectively configured as a video circular buffer, an audio ping-pong buffer queue, a kinematic state buffer, and a haptic topology buffer.

[0028] The visual sensing array acquires environmental optical signals and outputs raw video frame data via the mobile industrial processor interface protocol. The microphone array acquires spatial sound field signals, which, after analog-to-digital conversion, are output as a raw audio stream in pulse code modulation format via the integrated circuit's built-in audio bus. Simultaneously, the built-in six-axis inertial measurement unit acquires the dynamics of the fuselage space and outputs a kinematic digital stream containing three-axis acceleration and three-axis angular velocity via the integrated circuit interconnect bus; the flexible piezoresistive sensor array on the fuselage surface acquires physical contact pressure, which, after multiplexing and analog-to-digital conversion, is output as a tactile pressure digital matrix.

[0029] To ensure extremely low latency interaction requirements, the main control unit can be configured with a direct memory access controller. This controller responds to request signals from underlying peripherals, directly hardware-transferring the raw video frame data to the video circular buffer, and simultaneously and precisely transferring the raw audio stream, underlying kinematic digital stream, and underlying haptic pressure digital matrix to the audio ping-pong buffer queue, kinematic state buffer, and haptic topology buffer, respectively.

[0030] In the system kernel permission configuration policy, the physical address ranges where the above four independent caches are located are all set to be invisible to the network communication protocol stack. The network socket layer of the end system is forcibly stripped of read permissions for this physical address range, thereby establishing a mandatory isolation wall at the operating system level between raw media and physically sensed data and wireless radio frequency transmission links, providing a security foundation for subsequent privacy data truncation.

[0031] Step 2: Multimodal semantic tensor extraction and physical memory overwriting The edge-side main control unit calls the embedded neural network acceleration hardware and digital signal processing module to perform multi-dimensional feature extraction and data dimensionality reduction on the raw video frame data in the video ring buffer, the raw audio stream in the audio ping-pong buffer queue, the kinematic digital stream in the kinematic state buffer, and the tactile pressure digital matrix in the tactile topology buffer.

[0032] For visual signals, neural networks accelerate hardware execution of facial landmark detection algorithms, calculate the two-dimensional pixel coordinates of preset facial feature points, and output the dimension of... Facial semantic feature tensor ,in N This represents the total number of feature points identified by the algorithm.

[0033] For auditory signals, the main control unit performs pre-emphasis, framing, windowing, and discrete Fourier transform on the original audio stream sequence, extracts the cepstral coefficient features in the frequency domain, and outputs a dimension of Voiceprint semantic feature tensor ,in T This represents the number of audio frames within the time window. F This represents the dimension of the frequency band characteristics.

[0034] For the attitude kinematics signal, the main control unit performs moving average and Kalman filtering on the kinematic digital stream in the kinematic state buffer to remove high-frequency mechanical vibration noise, and extracts the transient components of triaxial acceleration and angular velocity, outputting a motion feature matrix representing the continuous change of the terminal's spatial attitude. .

[0035] For tactile interaction signals, the system calls an operator to perform spatial domain feature pooling on the tactile pressure digital matrix in the tactile topology buffer, and outputs a tactile topology tensor representing the spatial distribution of local physical contact intensity and pressure. .

[0036] A data destruction security mechanism established at the hardware level. When facial semantic feature tensors... Voiceprint semantic feature tensor Motion feature matrix With tactile topological tensor After all data has been fully extracted and written to the secure working area of ​​the system's static random access memory, the system kernel immediately triggers a high-priority hardware interrupt. This interrupt service routine forcibly mounts the direct memory access transfer channel, continuously hard-copying the pre-set all-zero bitstream from read-only memory to the physical address segments occupied by the video circular buffer, audio ping-pong buffer queue, kinematic state buffer, and haptic topology buffer in the current processing cycle. This overwrite operation completely bypasses the operating system's file management layer, achieving irreversible physical erasure of the initial multimedia and sensor raw data containing sensitive biometric features and real physical interaction behaviors, severing the privacy leakage link at the physical level of the storage medium.

[0037] Step 3: Construct a semantic value evaluation function and emotion quantification calculation based on multimodal spatiotemporal difference. The system sets a fixed-length sliding evaluation time window. Within this time window, the system continuously acquires the serialized facial semantic feature tensor, voiceprint semantic feature tensor, motion feature matrix, and tactile topological tensor, and performs spatiotemporal domain fluctuation feature extraction on them.

[0038] For the visual modality, extract the facial semantic feature tensor of two adjacent frames within the time window. and Regarding the tensor containing N For each facial feature point, calculate the Euclidean distance offset of each feature point in the two-dimensional image coordinate system. Sum the offsets of all facial feature points and calculate the arithmetic mean, which is defined as the facial deformation difference operator. Its mathematical expression is:

[0039] In the formula, and Representing the The horizontal and vertical coordinate components of a facial feature point in the current image frame. and This represents the corresponding coordinate component of the feature point in the previous image frame. This facial deformation difference operator physically characterizes the collective displacement amplitude of the facial muscles of the target object within a specific time window.

[0040] For auditory modalities, extract the voiceprint semantic feature tensor. The energy matrix includes time and frequency dimensions. The statistical variance of the high-frequency band energy sequence of each audio segment within the time window is calculated and defined as the auditory wave difference operator. Its mathematical expression is:

[0041] In the formula, To evaluate the total number of valid audio frames contained within a time window, For the first The total frequency band energy of each audio frame. This is the arithmetic mean of the total energy of all audio frames within the time window. This auditory fluctuation difference operator acoustically characterizes the degree of abrupt changes in the intensity and pitch of the target object's sound.

[0042] For attitude kinematic modes, the end-side master control unit reads the low-level discrete sequence of the built-in six-axis inertial measurement unit via the integrated circuit interconnect bus at a preset sampling rate. Within a set sliding evaluation time window, it extracts... The system obtains the triaxial acceleration vector and triaxial angular velocity vector for each consecutive sampling period. It extracts the acceleration representing physical impact force and the angular acceleration representing rotational torque by calculating the first-order kinematic difference between adjacent sampling periods, and then linearly weights and fuses them into an attitude change differential operator. Its discrete mathematical expression is:

[0043] In the formula, and Representing the first time window The and the first During each sampling clock cycle, the fuselage is The quantized value of the acceleration in the axial direction via analog-to-digital conversion; and These represent the corresponding periods of winding. The angular velocity of the axis rotation is converted to a quantized value using analog-to-digital conversion. and These are the linear acceleration abrupt change penalty coefficient and the angular acceleration abrupt change penalty coefficient, respectively, which are fixed in the underlying system. This operator accurately quantifies the transient kinetic energy spillover when the terminal is subjected to physical pushing or violent shaking by capturing high-frequency mechanical jumps.

[0044] For haptic interaction modes, the main control unit polls the area covered by the terminal body in real time. Flexible piezoresistive sensor array. The system extracts data within a time window. The underlying pressure space topology matrix for each sampling period is used, and a static physical contact determination threshold is introduced. The system calls the arithmetic logic unit to perform the super-threshold energy integral in the spatial dimension, defined as the tactile violence differential operator. Its mathematical expression is:

[0045] In the formula, Representing the In the sampling period, the first sampling period in the physical coordinate system u Line 1 vThe piezoresistive sensor unit outputs the digitally quantized value of the underlying absolute pressure. Nonlinear activation operator. All normal safety stroking signals below a set threshold are forcibly filtered out, and mean square summation is performed only on overload pressures exceeding a destructive threshold. This operator accurately characterizes the energy density of a local fuselage subjected to violent gripping or sharp physical pressure from a target object at the physical topology level.

[0046] After obtaining the difference operators for the aforementioned four-dimensional heterogeneous modalities, the system constructs a cross-modal semantic evaluation function to calculate the comprehensive emotion intensity score. The algebraic expression is as follows:

[0047] In the formula, , , and These are the non-negative weight coefficients for visual, auditory, postural kinematics, and tactile representations, respectively. The system kernel enforces this at the logic layer. Mathematical constraints.

[0048] The calculated score of emotional intensity Declared as a global floating-point variable and written to a register, it constitutes a mathematical criterion for measuring the comprehensive semantic value of the multimodal information flow within the current time window, providing a unique trigger condition for the gateway to execute differentiated quantization compression strategies in subsequent communication processing.

[0049] Step 4: Communication strategy scheduling and tensor encapsulation based on fusion of physical and semantic criteria The edge-side main control unit acquires the facial semantic feature tensor. Voiceprint semantic feature tensor Motion feature matrix With tactile topological tensor Within the synchronous clock cycle, the normalized exponential function probability distribution sequence of the neural network acceleration engine output layer is read via the system bus. The system calls the internal floating-point coprocessor to calculate the discreteness of the four-dimensional multimodal fusion probability distribution in the current interaction cycle in real time based on Shannon information entropy theory, and outputs the feature extraction confidence score. Its discrete mathematical expression is:

[0050] In the formula, The total number of reserved categories set for the output layer of the neural network classification layer. The feature tensor belongs to the first The true floating-point probabilities of each predicted category. The floating-point coprocessor will calculate the resulting... Write it into the system status register as a hardware criterion for evaluating physical environment distortion.

[0051] The communication processing gateway integrates independent hardware logic circuits to construct a three-dimensional joint scheduling state machine. The system statically stores the confidence level security threshold in non-volatile memory. First emotional threshold Second emotional threshold (and satisfy) ), and the physical impact red line threshold. The state machine's comparator continuously polls the emotional intensity score in the global register. Confidence score in the system status register Simultaneously, the system's arithmetic logic unit calls an independent hardware multiply-accumulator to extract the attitude change differential operator in real time. Differential operator for tactile violence And calculate the physical impact extreme value. Its algebraic expression is:

[0052] In the formula, and These are independent defense decision weights assigned to the underlying firmware for kinematic shock and topological stress overload, respectively. The state machine is based on... , and These three global variables form the three main routing branches.

[0053] In the first main trunk routing branch (physical defense priority branch), regardless of and What state is it in when the comparator determines... At this time, the system confirms that the terminal is undergoing destructive physical interaction or the target object is in an extremely violent and out-of-control state. The state machine forcibly short-circuits all dimensionality reduction quantization operators, attaches a physical pass-through mode, and includes the current... and The four-dimensional physical tensor is pushed directly into the sending buffer as the highest priority data stream, ensuring millisecond-level perception of extreme violent interactions in the cloud.

[0054] In the second backbone routing branch, when determining and At this point, the system confirms that no serious physical interference has occurred in the sensing link, and the state machine switches to the emotion-driven quantization channel. If This triggers the high-speed compression subroutine. The system attaches a tensor dimensionality reduction operator, mapping the original 32-bit single-precision floating-point data to 8-bit unsigned integer data. The mapping equation is as follows: ,in R For floating-point truth values, S For dynamic scaling factor, ZThis is a zero-point offset constant, used to generate the ultimate compressive characteristic load. If... This triggers the equalization subroutine, mapping the tensor data structure to a 16.5-bit floating-point type. If This triggers the bypass pass-through subroutine, which forces all dimensionality reduction operators to be short-circuited by the underlying logic gates, and directly outputs the full 32-bit single-precision floating-point original feature tensor.

[0055] In the third backbone routing branch, when determining and At this point, the system confirms that the current interactive environment suffers from signal attenuation or partial occlusion, resulting in insufficient confidence in the extraction of multimodal features on the edge. The state machine then triggers an adaptive tensor compensation mechanism, pausing subsequent pooling and fully connected layer forward inference computations in the neural network, and invoking the direct memory access controller to extract the intermediate layer spatial feature matrix before the pooling layer from static random access memory. This matrix preserves the uncompressed spatial topology and fine-grained frequency domain information, and its mathematical dimension is defined as... The system will compare this matrix with the current... , , and The underlying tensor concatenation is performed. The concatenated composite data packet, without containing the original multimedia pixels and raw audio data (i.e. maintaining physical privacy isolation), is sent to the cloud as a compensation payload rich in underlying information to support the cloud server in performing high-precision joint feature inference and intent reconstruction.

[0056] The feature payloads that have completed routing branch processing are then pushed to the application layer protocol encapsulation queue. The serialization engine sequentially pushes in a frame synchronization guide flag of a preset width, a tensor dimension descriptor matching the system standard word length, a three-dimensional quantization precision check bit indicating the current routing branch, and a real-time timestamp sequence, and appends a multi-byte cyclic redundancy check code based on generator polynomial calculation to the end. The merged application layer communication data frames are processed by a symmetric encryption algorithm and then handed over to the wireless network radio frequency front-end baseband chip for wideband carrier modulation and uplink spatial radiation.

[0057] Step 5: Cloud-based Intent Parsing and Dual-Track Security Driven by the Endpoint The cloud server's border communication gateway receives uplink application layer communication data frames and allocates an independent verification sandbox in memory. The system reads the timestamp sequence from the data frame header. Obtain the current local clock of the cloud system hardware. And calculate the one-way communication time overhead. If and only if If the data frame is less than the preset anti-replay attack lifecycle threshold, it is allowed to enter the dequantization parsing queue; otherwise, the hardware directly discards the frame in the sandbox to block network latency tampering attacks.

[0058] For data frames that pass verification, the cloud server reads the quantization precision check bit and the routing branch identifier bit. If the identifier bit indicates a high-speed compression state, the system initiates an asymmetric inverse quantization operator, using a pre-trained high-order interpolation algorithm to fill the matrix sparsity gaps caused by low-precision quantization, reconstructing a tensor with approximately 32-bit precision. If the identifier bit indicates an adaptive tensor compensation state, the intermediate feature matrix in the high-dimensional space is directly extracted. The cloud-based large language model performs forward inference on the reconstructed or retained multimodal tensors, outputting a text byte stream representing soothing semantics and a temporal control matrix representing physical actions, which is then transmitted via the downlink network link.

[0059] The master control unit of the end-side system receives downlink load data through its internal high-speed backbone bus and establishes a dual-track response logic based on a hardware watchdog timer. Upon the initial clock cycle of the characteristic load being pushed to the uplink from the end-side, the master control unit immediately starts an independent hardware watchdog timer.

[0060] In the normal response track, if the main control unit successfully parses the downlink text byte stream and timing control matrix before the watchdog timer overflows, the system immediately resets the timer. The main control unit calls the underlying incremental proportional-integral-differential algorithm module to parse the timing control matrix into the target angular displacement of each physical joint servo motor, and drives the motor to perform smooth, anthropomorphic movements by adjusting the duty cycle of the pulse width modulation signal. Simultaneously, the main control unit passes the text byte stream through to the local speech synthesis coprocessor to generate analog audio waveforms for real-time playback, completing low-latency emotional interaction.

[0061] In the fallback mechanism, if deep network congestion causes the watchdog timer to overflow, the main control unit's underlying hardware immediately mounts an exception interrupt service routine, forcibly blocking write permissions to all external network downlink ports. The system retrieves pre-set security standby action codes and network-buffered voice baseband data from the local read-only memory's security action space pool via the direct memory access controller. The main control unit forcibly drives all servo motors to reset to a mechanical zero-point locked state and drives the audio-to-analog converter module to play the locally buffered voice, maintaining the system's mechanical safety and interactive continuity in harsh communication environments.

[0062] It should be understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.

[0063] It should be understood that those skilled in the art can make improvements or modifications based on the above description, and all such improvements and modifications should fall within the protection scope of the appended claims.

Claims

1. A robot privacy-preserving communication method based on multimodal semantics and adaptive compression, characterized in that, The method includes the following steps: Step 1: During the power-on initialization phase of the robot system, the end-side main control unit divides four consecutive address blocks in the local physical memory, which are respectively configured as a video circular buffer, an audio ping-pong buffer queue, a kinematic state buffer, and a tactile topology buffer. These blocks are used to securely isolate and store the collected multimodal physical signals, including visual signals, auditory signals, posture kinematic signals, and tactile interaction signals. Step 2: The terminal main control unit performs multi-dimensional feature extraction and data dimensionality reduction on the multimodal physical signals to obtain facial semantic feature tensor, voiceprint semantic feature tensor, motion feature matrix and tactile topology tensor; after the multi-dimensional feature tensors are extracted and written into the secure working area of ​​static random access memory, the multimodal physical signal data stored in the local physical memory is erased. Step 3: Within a set sliding evaluation time window of a certain length, continuously extract multidimensional feature tensors and perform spatiotemporal fluctuation feature extraction on them to obtain facial deformation difference operators, auditory fluctuation difference operators, posture change difference operators, and tactile violence difference operators respectively; construct a cross-modal semantic evaluation function based on these four difference operators and calculate the comprehensive emotion intensity score. Step 4: Obtain the probability of a four-dimensional difference operator belonging to a certain predicted category through a neural network, calculate the dispersion of the four-dimensional multimodal fusion probability distribution in real time, and output the feature extraction confidence score; calculate the physical impact extreme value based on the posture mutation difference operator and the tactile violence difference operator; combine the comprehensive emotional intensity score, confidence score and physical impact extreme value to construct a three-dimensional joint scheduling state machine. The state machine is set with three main routing branches and multiple thresholds. The feature payloads processed by different branches are encapsulated into application layer communication data frames, encrypted and transmitted uplink to the cloud. Step 5: After verifying the timestamp, the cloud performs forward inference based on the application layer communication data frame by reconstructing or retaining the multi-dimensional feature tensor in the cloud-based large language model. It outputs a text byte stream representing the soothing semantics and a timing control matrix representing the physical action, and sends it down via the downlink network link. The terminal main control unit drives the motor to perform corresponding operations and performs voice broadcast according to the sent control signals.

2. The robot privacy-preserving communication method based on multimodal semantics and adaptive compression according to claim 1, characterized in that, The method for acquiring multimodal physical signals in step one specifically includes: The robot system is equipped with a visual perception array, a microphone array, a six-axis inertial measurement unit, and a flexible piezoresistive sensor array. The system acquires ambient optical signals through a visual sensing array and outputs raw video frame data. Spatial sound field signals are acquired through a microphone array, and after analog-to-digital conversion, the original audio stream in pulse code modulation format is output. The built-in six-axis inertial measurement unit collects the dynamics of the fuselage space and outputs a kinematic digital stream containing three-axis acceleration and three-axis angular velocity. Physical contact pressure is collected by a flexible piezoresistive sensor array on the surface of the device, and after multiplexing and analog-to-digital conversion, a digital matrix of tactile pressure is output.

3. The robot privacy-preserving communication method based on multimodal semantics and adaptive compression according to claim 2, characterized in that, The specific methods for signal security isolation in step one include: Raw video frame data from visual signals is stored in a video circular buffer; raw audio streams from auditory signals are stored in an audio ping-pong buffer queue; kinematic digital streams from posture kinematic signals are stored in a kinematic state buffer; and tactile pressure digital matrices from tactile interaction signals are stored in a tactile topology buffer. The physical address ranges containing the four independent buffers are set to be invisible to the network communication protocol stack, and the network socket layer on the end side is forcibly stripped of its read permissions for these physical address ranges.

4. The robot privacy-preserving communication method based on multimodal semantics and adaptive compression according to claim 2, characterized in that, The specific method for step two includes: For visual signals, a facial landmark detection algorithm is executed to calculate the two-dimensional pixel coordinates of preset facial feature points and output the coordinates of the points in dimension 1. Facial semantic feature tensor ,in N This represents the total number of feature points identified by the algorithm. For auditory signals, pre-emphasis, framing, windowing, and discrete Fourier transform are performed on the original audio stream sequence to extract cepstral coefficient features in the frequency domain, and the output dimension is... Voiceprint semantic feature tensor ,in T This represents the number of audio frames within a time window. F Represents the dimension of the frequency band features; For the attitude kinematics signal, moving average and Kalman filtering are performed on the kinematic digital stream to remove high-frequency mechanical vibration noise, and transient components of triaxial acceleration and angular velocity are extracted to output a motion feature matrix representing the continuous changes in the terminal's spatial attitude. ; For tactile interaction signals, the operator is invoked to perform spatial domain feature pooling on the tactile pressure digital matrix in the tactile topology buffer, and outputs a tactile topology tensor representing the spatial distribution of local physical contact intensity and pressure. .

5. The robot privacy-preserving communication method based on multimodal semantics and adaptive compression according to claim 4, characterized in that, The specific methods for step three include: Facial deformation difference operator Extract facial semantic feature tensors from two adjacent frames within a time window. and For the tensor contained in N For each facial feature point, calculate the Euclidean distance offset of each feature point in the two-dimensional image coordinate system; the formula is: In the formula, and Representing the The horizontal and vertical coordinate components of a facial feature point in the current image frame. and This represents the corresponding coordinate component of the feature point in the previous image frame; Auditory wave difference operator Extracting the semantic feature tensor of voiceprint The energy matrix, which includes time and frequency dimensions, is used to calculate the statistical variance of the high-frequency band energy sequence of each audio segment within a time window. The formula is as follows: In the formula, To evaluate the total number of valid audio frames contained within a time window, For the first The total frequency band energy of each audio frame. It is the arithmetic mean of the total energy of all audio frames within this time window; Attitude change differential operator According to the motion feature matrix ,extract The triaxial acceleration vector and triaxial angular velocity vector of each consecutive sampling period are used to extract the acceleration representing the physical impact force and the angular acceleration representing the rotational torque by calculating the first-order kinematic difference between adjacent sampling periods. These are then linearly weighted and fused, as shown in the following formula: In the formula, and Representing the first time window The and the first During each sampling clock cycle, the fuselage is The quantized value of the acceleration in the axial direction via analog-to-digital conversion; and These represent the corresponding periods of winding. The angular velocity of the axis rotation is converted from analog to digital and quantized. and These are the linear acceleration abrupt change penalty coefficient and the angular acceleration abrupt change penalty coefficient, respectively, which are fixed in the underlying system. Tactile Violence Differential Operator According to the tactile topological tensor Real-time polling covering the fuselage Flexible piezoresistive sensor array, extracting within a time window The underlying pressure space topology matrix for each sampling period is used, and a static physical contact determination threshold is introduced. The arithmetic logic unit is invoked to perform the superthreshold energy integral of the spatial dimension, and its formula is: In the formula, Representing the In the sampling period, the first sampling period in the physical coordinate system u Line number v The piezoresistive sensor unit outputs the underlying absolute pressure digital quantization value. It is a nonlinear activation algorithm; Construct a cross-modal semantic evaluation function to calculate the comprehensive emotion intensity score. Its formula is: In the formula, , , and These are the non-negative weighting coefficients for visual, auditory, postural kinematic, and tactile representations, respectively. .

6. The robot privacy-preserving communication method based on multimodal semantics and adaptive compression according to claim 5, characterized in that, In step four, the confidence score of feature extraction is output. The formula is: In the formula, The total number of reserved categories set for the output layer of the neural network classification layer. The feature tensor belongs to the first The true floating-point probabilities of each predicted category; Physical impact extreme value The calculation formula is: In the formula, and These are independent defense decision weighting coefficients assigned to the underlying firmware for kinematic shock and topological stress overload, respectively.

7. The robot privacy-preserving communication method based on multimodal semantics and adaptive compression according to claim 6, characterized in that, The three main routing branches in step four specifically include: Set confidence level safety threshold First emotional threshold Second emotional threshold ,satisfy and physical impact red line threshold ; In the first backbone routing branch, regardless of and What state is it in when the comparator determines... At this time, the system confirms that it is experiencing destructive physical interaction or that the target object is in an extremely violent and out-of-control state; the state machine forcibly short-circuits all dimensionality reduction quantization operators, attaches a physical pass-through mode, and includes the current... and The four-dimensional physical tensor is pushed directly into the send buffer as the highest priority data stream; In the second backbone routing branch, when determining and At that time, the system confirms that there is no serious physical interference in the sensing link, and the state machine switches to the emotion-driven quantization channel; if This triggers a high-speed compression subroutine; the system attaches a tensor dimensionality reduction operator, mapping the original 32-bit single-precision floating-point feature tensor data to 8-bit unsigned integer data; the mapping equation is... Where R is the floating-point truth value, S is the dynamic scaling factor, and Z is the zero-point offset constant, these are used to generate the ultimate compressibility characteristic load; if This triggers the equalization subroutine, mapping the tensor data structure to a 16.5-bit floating-point type; if This triggers the bypass pass-through subroutine, which forces all dimensionality reduction operators to be short-circuited by the underlying logic gates, and directly outputs the full 32-bit single-precision floating-point original feature tensor. In the third backbone routing branch, when determining and When the system confirms that there is signal attenuation or partial occlusion in the current interaction environment, resulting in insufficient confidence in the extraction of multimodal features on the edge, the state machine triggers an adaptive tensor compensation mechanism, pausing the subsequent pooling and fully connected layer forward inference calculations of the neural network, and calling the direct memory access controller to extract the intermediate layer spatial feature matrix before the pooling layer from the static random access memory. This matrix retains the uncompressed spatial topology and fine-grained frequency domain information, and its mathematical dimension is defined as... The system will compare this matrix with the current... , , and Perform underlying tensor splicing.

8. The robot privacy-preserving communication method based on multimodal semantics and adaptive compression according to claim 7, characterized in that, The method for verifying the timestamp in step five includes: The cloud server's border communication gateway receives uplink application layer communication data frames and creates an independent verification sandbox in memory; the system reads the timestamp sequence from the data frame header. Get the current local clock of the cloud server. And calculate the one-way communication time overhead. If and only if If the data frame is less than the preset anti-replay attack lifecycle threshold, it is allowed to enter the dequantization parsing queue; otherwise, the cloud server directly discards the frame in the verification sandbox to block network latency tampering attacks.

9. The robot privacy-preserving communication method based on multimodal semantics and adaptive compression according to claim 8, characterized in that, The specific methods for step five include: For application-layer communication data frames that pass verification, the cloud server reads the quantization precision check bit and the routing branch identifier bit. If the identifier bit indicates an ultra-fast compression state, the system starts the asymmetric inverse quantization operator and uses a pre-trained high-order interpolation algorithm to fill the matrix sparse gaps caused by low-precision quantization, reconstructing a tensor with approximately 32-bit precision. If the identifier bit indicates an adaptive tensor compensation state, the intermediate feature matrix in the high-dimensional space is directly extracted. The cloud-based large language model performs forward inference on the reconstructed or retained multimodal feature tensor, outputs a text byte stream representing soothing semantics and a temporal control matrix representing physical actions, and sends them down via the downlink network link.

10. A computer-readable storage medium, characterized in that, The device contains a computer program that, when executed by a processor, implements the robot privacy-preserving communication method based on multimodal semantics and adaptive compression as described in any one of claims 1 to 9.