Picture frame sound system with intelligent identification function and implementation control method thereof
By using edge computing with multimodal perception and a lightweight diffusion model, combined with audio-visual linkage algorithms and privacy protection mechanisms, the dynamic coupling problem of the picture frame audio system is solved, achieving highly context-aware content adaptation and improving user experience and data security.
Patent Information
- Application Number
- CN202511403792.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-29
- Publication Date
- 2026-02-10
AI Technical Summary
Existing picture frame audio systems lack dynamic coupling with user behavior and environmental parameters, making it impossible to achieve highly context-aware content adaptation.
A multimodal perception subsystem is adopted, including millimeter-wave radar, binocular RGB-D camera, microphone array and ambient light sensor, combined with a lightweight diffusion model to achieve real-time content generation at the edge, and adaptive control of dynamic content is achieved through audio-visual linkage algorithm and privacy protection mechanism.
It achieves true integration of audio and video, enhances perceptual consistency and immersion, improves the accuracy of user intent recognition, reduces cloud dependence and enhances data privacy and security, and expands the adaptive display capabilities of NFT digital artworks.
Smart Images

Figure CN121509874A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of smart home and artificial intelligence technology, and in particular to an AI picture frame and audio system that integrates computer vision, voice interaction, environmental perception and artistic display, which can realize dynamic content generation, sound and picture linkage and environmental adaptation under multimodal perception and edge computing conditions. Background Technology
[0002] In existing technologies, picture frame speakers are mostly limited to static image display, lacking dynamic coupling with user behavior and environmental parameters, and unable to achieve highly context-aware content adaptation.
[0003] Existing audio devices with screens, such as CN110166828A, although equipped with screens, are mostly driven by preset content libraries and lack intelligent content generation and dynamic presentation capabilities driven by user profiles.
[0004] While US20210144752A1 involves voice-controlled display, it relies heavily on a pre-set content library and lacks real-time generation and adaptation to environmental variables and individual emotions / intents.
[0005] Therefore, there is a need for a comprehensive terminal device that can achieve multimodal perception, real-time content generation and audio-visual linkage at the edge, and has a privacy protection mechanism, in order to achieve a higher level of environmental integration and intelligence. Summary of the Invention
[0006] The purpose of this section is to outline some aspects of embodiments of the present invention and to briefly describe some preferred embodiments. Simplifications or omissions may be made in this section, as well as in the abstract and title of this application, to avoid obscuring the purpose of these documents; however, such simplifications or omissions should not be construed as limiting the scope of the invention.
[0007] In view of the problems existing in the above-mentioned picture frame sound system, the present invention is proposed.
[0008] Therefore, the technical problem solved by this invention is to address the lack of dynamic coupling between existing picture frame audio systems and user behavior and environmental parameters, which prevents them from achieving highly context-aware content adaptation.
[0009] To address the aforementioned technical problems, this invention provides the following technical solution: a picture frame audio system with intelligent recognition, comprising: a perception subsystem, an output subsystem, a main control and edge computing unit, and a dynamic content generation engine; the perception subsystem performs multimodal perception through millimeter-wave radar, a binocular RGB-D camera, a microphone array, and an ambient light sensor; the dynamic content generation engine achieves real-time content generation at the edge based on a diffusion model, and adaptively adjusts the generated results using user profiles and environmental data; the main control and edge computing unit realizes the fusion of perception data, intent recognition, content generation scheduling, and synchronized audio-visual output; the output subsystem includes a 4K 120Hz micro LED display and a controllable directional speaker system; the system achieves audio-visual linkage and environmental adaptation through a closed-loop mechanism.
[0010] As a preferred embodiment of the intelligent recognition-enabled picture frame audio system described in this invention, the diffusion model is a lightweight diffusion model that has been cropped, quantized, and optimized for operation on edge devices, and its generation process combines user profiles and environmental data to perform adaptive content generation frame by frame or segment by segment.
[0011] As a preferred embodiment of the intelligent recognition-enabled picture frame audio system described in this invention, the audio-visual linkage algorithm maps the audio FFT spectrum to a dynamic visual particle system via a generative adversarial network (GAN), and the color, density, and motion trajectory of the particle system change synchronously with the audio features.
[0012] As a preferred embodiment of the intelligent recognition picture frame audio system described in this invention, the micro LED display screen is a magnetically levitated flexible OLED screen, and the display screen is dynamically deflected by ±5° through a micro servo motor to achieve adaptive control of the sound field direction and viewing angle.
[0013] As a preferred embodiment of the intelligent recognition-enabled picture frame audio system of the present invention, the sensing subsystem further includes a UWB positioning module, which is used to locate the user's position indoors in real time and couple the positioning information with screen posture adjustment, sound field orientation and image output strategies.
[0014] As a preferred embodiment of the intelligent recognition-enabled picture frame audio system described in this invention, the privacy protection mechanism includes localized federated learning and transmitting uploaded data after SEAL homomorphic encryption to ensure the privacy and security of cross-device model updates.
[0015] As a preferred embodiment of the intelligent recognition-enabled picture frame audio system described in this invention, the system achieves a perceptual consistency index where the SPL sound pressure level and the picture brightness ΔE are less than a certain threshold, and a low power consumption specification of ≤0.5W standby power consumption.
[0016] As a preferred embodiment of the intelligent recognition-enabled picture frame audio system described in this invention, the dynamic content generation engine can achieve adaptive display of NFT digital artworks, forming a Web3.0-level artwork presentation and interactive experience.
[0017] To solve the above-mentioned technical problems, the present invention also provides the following technical solution: a control method for implementing a picture frame audio system with intelligent recognition, which is executed according to the following steps: perception stage, intent recognition stage, content generation stage, audio-visual linkage stage, and feedback and learning stage; wherein the output picture and audio output of the generation stage are synchronized in time and space, and privacy protection processing for model update is performed through a local privacy protection mechanism.
[0018] As a preferred embodiment of the intelligent recognition-enabled picture frame audio system described in this invention, in each round of feedback, the updated model parameters are encrypted and uploaded to the server through localized federated learning, which is used only for global model improvement and does not expose individual user data.
[0019] This invention provides a picture frame audio system with intelligent recognition and its control method, which has the following beneficial effects:
[0020] 1. Achieve true integration of audio and video, breaking through the traditional one-way output of "audio and video separation" and enhancing perceptual consistency and immersion;
[0021] 2. Improve the accuracy of user intent recognition and automatically switch content modes in different application scenarios (education, fitness, entertainment, home scenarios, etc.) to enhance the interactive experience;
[0022] 3. By using edge computing models and privacy protection mechanisms, reduce reliance on the cloud and improve data privacy and security;
[0023] 4. Enhance the viewing experience and sound field directionality through biomimetic structures and positioning systems, thereby improving user comfort and scene adaptability;
[0024] 5. Provide adaptive presentation capabilities for scenarios such as NFT digital art displays, expanding commercial value. Attached Figure Description
[0025] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort. Wherein:
[0026] Figure 1The original data diagram for Experiment A—Multi-scene intent recognition provided by this invention.
[0027] Figure 2 The original data graph of Experiment B—end-to-end content generation delay provided for this invention.
[0028] Figure 3 The original data diagram of Experiment C—sound and image linkage—provided for this invention.
[0029] Figure 4 The original data diagram of Experiment D—Location and Attitude Adaptation—provided for this invention. Detailed Implementation
[0030] To make the above-mentioned objects, features, and advantages of the present invention more apparent and understandable, specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, and not all of them. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the protection scope of the present invention.
[0031] This invention proposes a picture frame audio system with intelligent recognition and its implementation control method. The core of the system is to achieve collaborative optimization of dynamic content and audio through multimodal perception (millimeter-wave radar, ToF / binocular RGB-D, microphone array, ambient light sensing, etc.), a lightweight content generation engine with localized edge computing power (based on a scalable diffusion model), and a generation-adaptation closed loop of audio-visual linkage.
[0032] In summary, the solution described in this invention has the following key features:
[0033] I. Hardware Architecture
[0034] Main control unit: Equipped with a heterogeneous chip with NPU (such as RK3588), with a computing power of about 6 TOPS, and has both AI inference and image / signal processing capabilities.
[0035] Perception layer: 60GHz millimeter-wave radar is used to detect subtle human movements within 5m; binocular RGB-D cameras are used for depth perception with a field of view (FOV) of approximately 120°; microphone arrays are used for multi-source sound localization and noise reduction; and light sensors are used for ambient light intensity / color temperature perception.
[0036] Output layer: 4K 120Hz Micro-LED screen with a peak brightness of approximately 1500 nits; output side features neodymium magnet coaxial speakers with a frequency response of approximately 20 Hz–22 kHz.
[0037] Mechanical structure: The flexible OLED display is suspended by magnetic levitation and uses a micro servo motor to achieve a dynamic screen tilt of ±5° to fine-tune the sound field and viewing angle.
[0038] Positioning and Attitude: The UWB positioning system is used for precise positioning of the indoor homeowner, and combined with internal attitude algorithms, it optimizes the screen tilt angle and image orientation.
[0039] II. Software Algorithm
[0040] User intent recognition and multimodal decision-making: The speech semantics (BERT model) and gesture features (MediaPipe framework) are fused to form a multimodal decision tree, which outputs the application mode of the current scenario (education, entertainment, fitness, etc.).
[0041] Dynamic content generation engine: Deploys lightweight versions of diffusion models such as Stable Diffusion at the edge, cropping / quantizing them to generate artistic images in real time based on user profiles (age, gender, emotion recognition, etc.) and environmental data (temperature, humidity, lighting, etc.); the color temperature and style parameters of the image are mapped to audio parameters (such as EQ, spectrum mapping) to achieve audio-visual linkage.
[0042] Audio-visual linkage algorithm: The audio FFT spectrum is transformed into a dynamic visual particle system through a generative adversarial network (GAN). The density, color, and motion trajectory of the particle system change synchronously with the audio spectrum features, forming a linkage effect of "sound and animation".
[0043] Privacy and security: Enables localized federated learning, with sensitive data processed locally; model gradients or feature updates are aggregated locally and uploaded to the server after being homomorphically encrypted using SEAL, ensuring privacy protection of data during transmission and storage.
[0044] Environmental Adaptation and Visual Consistency: Through physical-perception-content mapping, we ensure that the sound pressure level (SPL) and screen brightness ΔE are within a perceptible tolerance range, thereby improving the consistent experience across modalities.
[0045] III. Data Flow and Workflow
[0046] Perception phase: Millimeter-wave radar, RGB-D, microphone array, and ambient light sensor continuously collect data, and perform preliminary feature extraction and fusion through edge processing unit to output "current scene profile" and "user profile".
[0047] Intent and Scene Recognition Stage: The multimodal decision tree combines user profiles and contextual information to determine the goal of this interaction (such as education mode, exercise rhythm, art style switching, etc.).
[0048] Content generation stage: Based on the scene target, the edge diffusion model is called to generate artistic images. At the same time, the image style, color temperature and brightness are adjusted according to the context mapping, and the music / sound effects are adjusted in real time.
[0049] Audio-visual linkage stage: The generated picture parameters are synchronously mapped with the audio signal and output to the screen and speakers to form a dynamic audio-visual linkage effect.
[0050] Feedback phase: The system performs self-evaluation locally (screen brightness, color temperature, sound pressure, etc.) and uploads a summary of model updates through a federated learning framework to improve the global model while ensuring privacy.
[0051] IV. Key Points for Security and Privacy Protection
[0052] Localized federated learning: User data is kept on the local device, and only de-identified model updates are uploaded to protect personal privacy.
[0053] Homomorphic encryption transmission: During the update and upload phase, homomorphic encryption technologies such as SEAL are used to encrypt gradients / parameters to prevent external theft or reverse analysis.
[0054] Data minimization principle: Only collect the data necessary to complete the task, minimizing the risk to personal privacy.
[0055] Specifically, this invention provides a picture frame audio system with intelligent recognition, comprising: a perception subsystem, an output subsystem, a main control and edge computing unit, and a dynamic content generation engine; the perception subsystem performs multimodal perception through millimeter-wave radar, a binocular RGB-D camera, a microphone array, and an ambient light sensor; the dynamic content generation engine achieves real-time content generation at the edge based on a diffusion model, and adaptively adjusts the generated results using user profiles and environmental data; the main control and edge computing unit realizes the fusion of perception data, intent recognition, content generation scheduling, and synchronized audio-visual output; the output subsystem includes a 4K 120Hz micro LED display and a controllable directional speaker system; the system achieves audio-visual linkage and environmental adaptation through a closed-loop mechanism.
[0056] Specifically, the diffusion model is a lightweight diffusion model that has been cropped, quantized, and optimized for operation on edge devices. Its generation process combines user profiles and environmental data to perform adaptive content generation frame by frame or segment by segment.
[0057] Specifically, the audio-visual linkage algorithm maps the audio FFT spectrum into a dynamic visual particle system through a generative adversarial network (GAN). The color, density, and motion trajectory of the particle system change synchronously with the audio features.
[0058] Specifically, the micro LED display is a magnetically levitated flexible OLED screen, and a micro servo motor is used to achieve ±5° dynamic deflection of the display screen to achieve adaptive control of the sound field direction and viewing angle.
[0059] Specifically, the perception subsystem also includes a UWB positioning module, which is used to locate the user's position indoors in real time and couple the positioning information with screen posture adjustment, sound field orientation and image output strategies.
[0060] Specifically, the privacy protection mechanism includes localized federated learning and transmitting uploaded data with SEAL homomorphic encryption to ensure the privacy and security of cross-device model updates.
[0061] Specifically, the system achieves a perceptual consistency index where the SPL sound pressure level and the screen brightness ΔE are less than a certain threshold, and a low power consumption specification of ≤0.5W standby power consumption.
[0062] Specifically, the dynamic content generation engine can achieve adaptive display of NFT digital artworks, forming a Web3.0-level artwork presentation and interactive experience.
[0063] Furthermore, to illustrate the technical solution of the present invention in detail, a method for implementing a picture frame audio system with intelligent recognition is also provided, which is executed according to the following steps: perception stage, intent recognition stage, content generation stage, audio-visual linkage stage, and feedback and learning stage; wherein the output picture and audio output of the generation stage are synchronized in time and space, and privacy protection processing for model update is performed through a local privacy protection mechanism.
[0064] Specifically, in each round of feedback, the updated model parameters are encrypted and uploaded to the server through localized federated learning, and are used only for global model improvement without exposing individual user data.
[0065] It needs to be explained one by one:
[0066] (1) Overall Architecture of AI Picture Frame Audio System with Intelligent Recognition
[0067] 1. Sensing Subsystem
[0068] Millimeter-wave radar (60 GHz FMCW)
[0069] Parameters: bandwidth 4 GHz, ranging resolution 3.75 cm, velocity resolution 0.05 m / s, used to detect micro-movements within 5 m.
[0070] Physical meaning: Phase difference detection (Δφ) of minute displacement is achieved through continuous wave frequency modulation, enabling detection of breathing, gait, etc.
[0071] Dual-lens RGB-D camera
[0072] Parameters: resolution 1280×720, frame rate 60 FPS, baseline distance 6cm, ranging accuracy ±2 cm (within 1 m).
[0073] Physical meaning: The depth Z = f∙B / Δu is calculated using stereo parallax, and then refined using Time-of-Flight (ToF) technology.
[0074] Microphone array (6-channel omnidirectional + beamforming)
[0075] Parameters: Sampling rate 48 kHz, SNR > 65 dB, beamwidth adjustable up to 20°.
[0076] Ambient light sensor
[0077] Parameters: Illumination range 0.1120 k lux, color temperature detection 2000~8000 K.
[0078] 2. Main control and edge computing unit
[0079] The core SoC (such as RK3588) + built-in NPU (6 TOPS INT8) + GPU are used for diffusion model inference.
[0080] Memory optimization: LPDDR4X 8GB, bandwidth 25 GB / s, supports memory-mapped inference.
[0081] 3. Closed-loop mechanism
[0082] Input perception → Intent recognition model → Dynamic content generation engine → Output synchronization optimization → Re-perception feedback (updating model weights or parameters).
[0083] (2) Lightweight edge diffusion model
[0084] Original model: Stable Diffusion v1.5 (UNet + VAE + CLIP text encoder).
[0085] Pruning method: Pruning rate P=35%, retaining the core perception channel; Attention layer dimension changed from 768 to 384.
[0086] Quantization: Both weights and activations are INT8, using a symmetric quantization formula.
[0087]
[0088] Where S is the scale factor, which physically represents the mapping ratio of the quantization interval.
[0089] in,
[0090] Q(x): The quantized value (the discrete value closest to the original value).
[0091] x: Raw floating-point value (such as model weights or activation values), the unit of which depends on the specific data type (e.g., weights are dimensionless).
[0092] S: Scale factor (dimensionless), which defines the quantization step size.
[0093] round(): Rounds the integer to the nearest integer.
[0094] Embedded user profile vector u∈R n (e.g., age A, emotion E, preference matrix P), and environmental feature vector e∈R m (Temperature and humidity T / H, illumination CCT, noise dBA) are concatenated to form a condition vector c=[u∣∣e], which serves as the conditional input (Condition Embedding) for the diffusion model.
[0095] Additionally, regarding the aforementioned lightweighting process, the present invention provides the following improvement:
[0096] Backbone: Lightweight Stable Diffusion (UNet channel count reduced by 35%, Attention dimension reduced to 384).
[0097] Add a new time-space branch:
[0098] 3D convolution kernel size (3×3×3), time step = 4 frames.
[0099] The conditional input includes the time gradient ΔT and the spatial motion vector (extracted by the optical flow algorithm).
[0100] Edge inference latency remains within 200 ms.
[0101] (3) Audio-visual linkage GAN-particle system algorithm
[0102] Input: Audio signal s(t) → FFT to obtain amplitude spectrum S(f) (resolution 1024 points, 50% window overlap).
[0103] GAN structure:
[0104] Generator G: Input S(f), Output particle dynamic parameters {v i c i p i (Speed, color, position).
[0105] Discriminator D: Determines the visual similarity between the generated particle system and the target template.
[0106] Particle system parameter mapping:
[0107] velocity vector v i =α*Ef (E_f is the frequency band energy, and α is the sound-image velocity mapping coefficient).
[0108] Color value c i Mapped to the H component of the HSL color model according to the spectral centroid.
[0109] Density / Quantity: Proportional to RMS volume.
[0110] Additionally, regarding the aforementioned audio-visual linkage process, this invention also provides the following dual-channel improvement scheme:
[0111] GAN input: FFT spectrum S(f) → Generate moving particle parameters.
[0112] GNN input: lyrics or audio text → converted into a semantic graph (nodes = keywords, edges = semantic associations).
[0113] The final conditional vector c = [particle parameters | semantic embedding] is input into the diffusion model to control the color layout.
[0114] Particle velocity formula:
[0115]
[0116] in,
[0117] v i : The velocity of the i-th particle (unit: m / s).
[0118] α: Sound-to-image velocity mapping coefficient (unit: m / (s·energy unit)), which determines the intensity of the influence of spectral energy on velocity.
[0119] E f : The energy of audio in the frequency range f (unit: dB or power density W / Hz).
[0120] β: Semantic velocity mapping weight coefficient (same unit as α), used to adjust the influence of semantic information on particle velocity.
[0121] p semantic : Semantic strength value (dimensionless, 0–1), derived from music lyrics or speech keywords.
[0122] (4) Dynamic deflection mechanism of magnetically levitated flexible OLED
[0123] The elastic support structure uses permanent magnet levitation and controllable electromagnets to reduce friction and extend service life.
[0124] Servo motor: torque 0.15 N·m, drive signal PWM 500 Hz, position control accuracy 0.1°.
[0125] Deflection angle θ(t) is calculated as follows:
[0126]
[0127] in,
[0128] θ: The angle adjustment amount currently output by the servo system (unit: °).
[0129] k p : Proportional gain coefficient (unit: ° / °), controls the effect of position error on the output.
[0130] θ target Target angle (unit: °), calculated by audience position or sound field optimization algorithm.
[0131] θ current Current angle (unit: °).
[0132] k d : Differential gain coefficient (unit: ° / (° / s)), which affects the rate of change of the adjustment angle.
[0133] dθ / dt: Rate of change of angle (unit: ° / s).
[0134] Physical meaning: PD control ensures that angle adjustment is completed quickly and without oscillation.
[0135] Additionally, this invention also provides a multi-degree-of-freedom magnetic levitation OLED audiovisual adaptive control scheme:
[0136] Servo system: Three-axis servo, accuracy 0.05°, response time < 150 ms.
[0137] Eye-tracking camera: 200 Hz sampling rate, accuracy < 0.5°.
[0138] Control formula:
[0139] .
[0140] (5) UWB Indoor Positioning and Attitude Coupling
[0141] Positioning principle: Two-Way Time of Flight (ToF).
[0142] Coordinate calculation: Solve the TDOA system of equations
[0143]
[0144] in,
[0145] d i : Distance from the device to the i-th base station (unit: m).
[0146] c: speed of light (approximately 3.0 × 10⁸ m / s).
[0147] t rx,i : The timestamp (in seconds) of the signal received by the i-th base station.
[0148] t tx : Signal transmission timestamp (unit: seconds).
[0149] Real-time RMSE < 0.15 m.
[0150] Attitude control: The target angle is calculated from the relative azimuth angle by atan2(Δy, Δx) and transmitted to the servo system.
[0151] (6) Local Federated Learning + SEAL Homomorphic Encryption
[0152] Federated learning: training batch size is 16 per round, learning rate is 1e-4, and local iterations are E=5 rounds.
[0153] Encryption parameters (CKKS): Polynomial modulus 8192, coefficient modulus chain {60, 40, 40, 60}, which can support floating-point operations under encryption.
[0154] The upload data packet size is 200 KB per round, with an additional transmission delay of approximately 30–50 ms.
[0155] (7) Sound pressure level-brightness uniformity and low power consumption specifications
[0156] The SPL real-time sampling frequency is 50 Hz, mapped to the brightness adjustment ΔL, and maintained as follows:
[0157]
[0158] in,
[0159] ΔE: Perceived color difference (dimensionless), conforming to the CIEDE2000 color difference standard.
[0160] ΔL′: Brightness difference (unit: cd / m²).
[0161] ΔC′: Color difference (unit: dimensionless, color value).
[0162] ΔH′: Hue difference (unit: degrees, °).
[0163] K LK C K H : Weighting coefficients (dimensionless) for brightness, chromaticity, and hue, used to adjust the impact of different differences on perceived color difference.
[0164] Standby power consumption control:
[0165] The SoC enters deep sleep mode, shuts down the RF module, and retains only the RTC; the voltage drops to 0.8 V and the clock speed drops to 32 kHz.
[0166] (8) Adaptive Display of NFT Digital Art
[0167] Parse ERC-721 or ERC-1155 tokenURI Metadata to extract art style tags, resolution, and color tone.
[0168] The on-chain data is mapped to the conditional input of the initial latent variable z0 of the diffusion model, ensuring that the generated work is consistent with the style of the original NFT but with increased dynamism.
[0169] (9) Five-stage workflow of control methods
[0170] Perception: Collect multimodal data → Standardize → Feature fusion.
[0171] Identify: Inference intent labels for multimodal Transformer models.
[0172] Generation: Retrieve the edge diffusion model and generate the image based on the conditional vector.
[0173] Synchronization: Synchronized audio and video output; particle system driven by GAN.
[0174] Feedback: Local computing performance and user interaction metrics → Update model weights.
[0175] (10) Localized Federated Learning Feedback Mechanism
[0176] Round Flow:
[0177] Local training on each device → Gradient extraction ΔW → Homomorphic encryption Enc(ΔW) → Upload.
[0178] Server-side aggregation:
[0179]
[0180] Pure polymerization technology will not be elaborated upon here.
[0181] Updates will be distributed to the terminal.
[0182] Specifically, based on the above description of the specific solution, it is applied to the following embodiments:
[0183] Example 1: Children's Educational Scenario
[0184] 1. Experimental Design
[0185] Objective: To verify that the system can automatically switch to educational content mode and adjust the visuals and audio to suit the child's needs when a child is detected approaching.
[0186] Environment: Standard indoor environment, simulating a children's activity area.
[0187] Data collection:
[0188] Data from millimeter-wave radar, RGB-D cameras, microphone arrays, and ambient light sensors.
[0189] The system outputs video and audio parameters.
[0190] User feedback (such as children's reactions).
[0191] 2. Data Collection
[0192] Field definition:
[0193] event_id: Unique identifier for the event
[0194] scene: Scene category (Education)
[0195] ground_truth_label: The label for the ground truth
[0196] detected_label: The detected label
[0197] voice_conf: Speech recognition confidence level (0-1)
[0198] gesture_conf: gesture feature confidence (0-1)
[0199] radar_conf: Millimeter-wave radar perception confidence level (0-1)
[0200] event_start_ts: Event start time (ISO 8601 UTC)
[0201] event_end_ts: Event end time (ISO 8601 UTC)
[0202] response_start_ts: System output start time (ISO 8601 UTC)
[0203] response_end_ts: System output completion time (ISO 8601 UTC)
[0204] response_time_ms: End-to-end response time (milliseconds)
[0205] raw_text_command: The raw text command from the speech transcription.
[0206] notes: Remarks field
[0207] 3. Sample Data
[0208] event_id,scene,ground_truth_label,detected_label,voice_conf,gesture_conf,radar_conf,event_start_ts,event_end_ts,response_start_ts,response_end_ts,response_time_ms,raw_text_command,notes
[0209] 1, Education, Education, Education, 0.93, 0.65, 0.72, 2025-09-26T09:12:30.123Z, 2025-09-26T09:12:32.381Z, 2025-09-26T09:12:31.020Z, 2025-09-26T09:12:31.270Z, 260,"Please display the basic addition teaching screen",""
[0210] 2, Education, Education, Education, 0.92, 0.62, 0.66, 2025-09-26T09:28:44.350Z, 2025-09-26T09:28:46.810Z, 2025-09-26T09:28:45.100Z, 2025-09-26T09:28:45.520Z, 230,"Warm Color Style for Educational Images",""
[0211] 4. Data Analysis
[0212] ①Intent recognition accuracy:
[0213] Calculate the match rate between ground_truth_label and detected_label.
[0214] In the example data, all records were correctly identified as "Education", with an accuracy of 100%.
[0215] ② Response time:
[0216] Calculate the mean and standard deviation of response_time_ms.
[0217] In the example data, the response times are 260 ms and 230 ms, the average response time is 245 ms, and the standard deviation is 15 ms.
[0218] ③ Adjustments to video and audio:
[0219] Check if the screen brightness and color temperature meet expectations.
[0220] In the example data, the screen displays as "basic addition teaching screen" and "warm color style educational screen", which is as expected.
[0221] 5. Conclusion
[0222] When the system detects a child approaching, it can quickly and accurately switch to educational content mode with an acceptable response time and the adjustment of the picture and audio meets expectations.
[0223] Example 2: Fitness and Bodybuilding Scene
[0224] 1. Experimental Design
[0225] Objective: To verify that when the system detects a user performing fitness movements, it can generate abstract art visuals that match the rhythm of the exercise and synchronize with the background music (BGM).
[0226] Environment: Standard indoor environment, simulating a fitness area.
[0227] Data collection:
[0228] Data from millimeter-wave radar, RGB-D cameras, microphone arrays, and ambient light sensors.
[0229] The system outputs video and audio parameters.
[0230] User feedback (such as user satisfaction).
[0231] 2. Data Collection
[0232] Field definition:
[0233] frame_id: Frame number
[0234] event_id: The identifier of the associated event
[0235] generation_start_ts: Content generation start time (ISO 8601 UTC)
[0236] render_complete_ts: Time taken for the image to complete rendering (ISO 8601 UTC)
[0237] end_to_end_latency_ms: End-to-end latency (milliseconds)
[0238] frames_dropped: The number of dropped frames in this output batch
[0239] generator_version: Generated model version
[0240] frame_quality_score: Frame quality score (0-1 or a custom score)
[0241] 3. Sample Data
[0242] frame_id,event_id,generation_start_ts,render_complete_ts,end_to_end_latency_ms,frames_dropped,generator_version,frame_quality_score
[0243] 1,EV1,2025-09-26T09:15:14.501Z,2025-09-26T09:15:14.844Z,343,1,v1.2.3,0.86
[0244] 2,EV1,2025-09-26T09:15:14.845Z,2025-09-26T09:15:15.187Z,342,0,v1.2.3,0.90
[0245] 3,EV2,2025-09-26T09:20:03.000Z,2025-09-26T09:20:03.197Z,197,0,v1.2.3,0.91
[0246] 4. Data Analysis
[0247] ① End-to-end delay:
[0248] Calculate the mean and standard deviation of end_to_end_latency_ms.
[0249] In the example data, the latency is 343 ms, 342 ms and 197 ms, the average latency is 294 ms and the standard deviation is 73 ms.
[0250] ② Frame dropping situation:
[0251] Count the total number of frames_dropped.
[0252] In the example data, a total of 1 frame was lost.
[0253] ③ Frame quality:
[0254] Calculate the mean and standard deviation of frame_quality_score.
[0255] In the example data, the frame quality is 0.86, 0.90, and 0.91, the average quality is 0.89, and the standard deviation is 0.025.
[0256] 5. Conclusion
[0257] When the system detects a user performing fitness movements, it can generate abstract art visuals that match the rhythm of the exercise and synchronize with the background music rhythm. The end-to-end latency is within an acceptable range, frame drops are rare, and frame quality is high.
[0258] Example 3: Owner Positioning and Angle Adaptation
[0259] 1. Experimental Design
[0260] Objective: To verify that the system can identify the owner's location via UWB positioning and adjust the screen tilt angle to achieve the best viewing angle.
[0261] Environment: Standard indoor environment, simulating user activities in different locations.
[0262] Data collection:
[0263] Data from the UWB positioning module.
[0264] Data on screen tilt adjustment.
[0265] User feedback (such as user comfort).
[0266] 2. Data Collection
[0267] Field definition:
[0268] trial_id: Trial sequence number
[0269] time_stamp: Observation time (ISO 8601 UTC)
[0270] uwb_x / uwb_y / uwb_z: 3D coordinates for UWB positioning
[0271] gt_x / gt_y / gt_z: Reference coordinates of the actual position
[0272] rmse_sample: RMSE of a single sample
[0273] screen_angle_deg: Current screen angle (degrees)
[0274] adjust_time_ms: The time (in milliseconds) required to complete the adjustment.
[0275] 3. Sample Data
[0276] trial_id,time_stamp,uwb_x,uwb_y,uwb_z,gt_x,gt_y,gt_z,rmse_sample,screen_angle_deg,adjust_time_ms
[0277] 1,2025-09-26T09:12:31.400Z,1.20,2.50,0.90,1.18,2.48,0.92,0.11,0.8,110
[0278] 2,2025-09-26T09:15:15.600Z,0.85,2.75,0.95,0.88,2.70,0.92,0.12,1.0,125
[0279] 3,2025-09-26T09:20:03.150Z,1.60,2.20,0.85,1.58,2.18,0.88,0.10,0.9,115
[0280] 4. Data Analysis
[0281] ① Positioning error:
[0282] Calculate the mean and standard deviation of rmse_sample.
[0283] In the example data, the RMSE values are 0.11, 0.12, and 0.10, respectively, with a mean RMSE of 0.11 and a standard deviation of 0.01.
[0284] ② Screen tilt adjustment:
[0285] Check if screen_angle_deg meets expectations.
[0286] In the example data, the screen tilt angles are 0.8°, 1.0°, and 0.9°, which are within the expected range.
[0287] ③ Adjust the time:
[0288] Calculate the mean and standard deviation of adjust_time_ms.
[0289] In the example data, the adjustment times were 110 ms, 125 ms, and 115 ms, respectively, with a mean adjustment time of 116.7 ms and a standard deviation of 7.5 ms.
[0290] 5. Conclusion
[0291] The system can accurately identify the owner's location through UWB positioning and adjust the screen tilt angle accordingly. The positioning error is within an acceptable range, the screen tilt angle adjustment meets expectations, and the adjustment time is short.
[0292] Based on the detailed data collection and analysis of the three embodiments described above, the following conclusions can be drawn:
[0293] In children's educational scenarios, the system can quickly and accurately switch to educational content mode, with response time and screen adjustments meeting expectations.
[0294] In fitness and bodybuilding scenarios, the system can generate abstract artistic visuals that match the rhythm of the exercise and synchronize with the background music rhythm, with good end-to-end latency and frame quality.
[0295] In scenarios involving owner positioning and angle adaptation, the system can accurately identify the owner's location and adjust the screen tilt angle, with both positioning error and adjustment time remaining within a reasonable range.
[0296] Therefore, the present invention has the following beneficial effects:
[0297] 1. Establish a perception layer of a "four-dimensional interaction model": the linkage of spatial positioning, individual recognition, temporal evolution and contextual inference, to achieve high-precision perception of users and environment.
[0298] 2. Dynamic content generation engine using a lightweight edge computing model: Based on the edge implementation of the diffusion model, it combines user profiles, emotions / expressions, environmental data and other factors to generate artistic images in real time, and maps the image style / color temperature to audio parameters (such as EQ / spectral characteristics) to achieve deep linkage between sound and image.
[0299] 3. Introducing a closed loop of "perception-generation-adaptation" with biomimetic mechanical structure and controllable sound field: Spatial directivity coordination of picture and sound field is achieved through magnetic levitation flexible OLED display and ±5° dynamic deflection; UWB / positioning technology is used to identify the owner's location and adaptively optimize screen posture, sound field directionality and picture angle.
[0300] 4. Introduce privacy protection mechanisms: Implement Federated Learning locally, process sensitive data locally and only upload model updates that have undergone SEAL homomorphic encryption to ensure user privacy and security.
[0301] 5. Propose a generation algorithm for audio and visual co-origin: Transform the audio FFT spectrum into a dynamic visual particle system through a generative adversarial network (GAN) to achieve dynamic coordination between audio and visuals.
[0302] Introduce verifiable energy consumption and display consistency metrics: achieve perceptual consistency between SPL sound pressure level and screen brightness ΔE below a certain threshold, and significantly reduce standby power consumption to a specific low power level.
[0303] Additionally, to further verify the beneficial effects of the present invention, the following simulation experiment was conducted:
[0304] I. Verification Objectives and Overall Approach
[0305] Objective 1: To demonstrate that the accuracy of multimodal perception and intent recognition reaches or exceeds a set threshold, and that the correct content patterns can be output quickly and stably in multiple scenarios such as education, fitness, entertainment, and family settings.
[0306] Objective 2: To demonstrate that the end-to-end latency (perception → generation → output) of dynamic content generation at the edge remains extremely low on a 1-frame / 1000-frame scale, and that the linkage between video and audio achieves the expected quality.
[0307] Objective 3: Prove the perceptual consistency of audio-visual linkage, that is, the ΔE generated by the mapping between SPL and picture parameters is within the perceptual tolerance range (the threshold is set to ΔE < 3).
[0308] Objective 4: Prove that the positioning error of the owner positioning / attitude adaptation mechanism (UWB positioning + screen tilt angle / sound field pointing adaptation) in a real indoor environment meets the design specifications (RMSE less than 0.1–0.2 m).
[0309] Objective 5: Prove the effectiveness of the privacy-preserving design: Local federated learning achieves global model improvement without uploading the original data, and the additional overhead of encrypted data transmission is within an acceptable range.
[0310] Objective 6: Prove the effectiveness of low power consumption characteristics and dynamic load balancing: standby power consumption is less than 0.5 W, and the power consumption adaptive management effect is significant in dynamic scenarios.
[0311] II. Experimental Subjects and Environment
[0312] 1. Test Equipment (DUT)
[0313] The prototype of the smart picture frame speaker uses an NPU / heterogeneous SoC (such as RK3588) as its core processing unit, with a computing power of about 6 TOPS.
[0314] Sensing subsystem: 60 GHz millimeter-wave radar, binocular RGB-D camera (FOV ~120°), microphone array, ambient light sensor, UWB positioning module.
[0315] Output units: 4K 120 Hz Micro-LED display, magnetically levitated flexible OLED, micro-servo mechanism with controllable ±5° screen tilt angle, coaxial neodymium magnet speaker (20 Hz–22 kHz).
[0316] Software: A cropped / quantized version of the marginalization diffusion model, a multimodal decision tree, a GAN-particle system audio-visual linkage module, a localized federated learning framework, and SEAL homomorphic encrypted transmission.
[0317] 2. Test Scenario
[0318] Educational scenarios, fitness scenarios, entertainment scenarios, and daily family scenarios, covering different lighting conditions, background noise, human postures, and interaction intensities.
[0319] Typical interior layout: a rectangular room (8 m × 6 m × 2.8 m, 2.8 m high), with furniture providing moderate coverage, allowing the occupant to move around in different positions.
[0320] 3. Baseline control
[0321] Using similar commercial devices with on-screen audio (such as smart screen devices with a fixed content library) as a control, the improvements of this invention in dynamic content generation and audio-visual linkage were evaluated.
[0322] III. Evaluation Indicators (Simplified Table of Key Indicators)
[0323] Intent recognition accuracy and average response time (ms)
[0324] End-to-end content generation latency (ms, perception → generation → output)
[0325] Consistency of ΔE color difference and image color temperature mapping (ΔE, unit: CIEDE2000)
[0326] Owner positioning error (RMSE, m) and 95% confidence interval
[0327] Federated learning effectiveness metrics: amount of data uploaded per round, encryption overhead, and level of protection for non-raw data.
[0328] Standby power consumption and dynamic load power consumption (W), and power consumption adaptation effect
[0329] Subjective evaluation and objective indicators of audio-visual linkage quality (such as the correlation between the audio spectrum and the particle system).
[0330] IV. Experimental Design and Procedures
[0331] Experiment A: Multi-Scene Intent Recognition and Response Time
[0332] 1. Datasets and Ground Truth
[0333] Each scenario has 250 events, totaling 1000 events (Education, Fitness, Entertainment, Home), which are manually labeled with "target mode" (education, sports, entertainment, normal state, etc.) and the time of occurrence.
[0334] 2. Test Procedure
[0335] Start the device to trigger preset interactions in various scenarios (such as children's educational prompts, fitness movement recognition, playing entertainment content, daily Q&A, etc.).
[0336] The system outputs the application mode and starts content generation, recording the recognition results and response time.
[0337] 3. Evaluation Methods
[0338] Calculate scene-level accuracy, overall accuracy, and average response time (ms) from trigger to output.
[0339] Experiment B: End-to-End Content Generation Delay and Image Output Quality
[0340] 1. Test conditions
[0341] We tested at 1000 frames per scene and recorded the complete latency from the triggering of a perceived event to the stable output of the image.
[0342] 2. Indicators
[0343] End-to-end average delay, 75 / 90 / 95 percentile delay, maximum delay, and delay fluctuation under different lighting and background noise conditions.
[0344] Experiment C: Consistency of Sound-Image Linkage and Color Mapping
[0345] 1. Test conditions
[0346] The mapping relationship between color temperature and audio spectrum characteristics when the captured images are output in 5 scenarios.
[0347] 2. Indicators
[0348] ΔE average value, ΔE standard deviation, accuracy of target color mapping in the image (such as color temperature error), and audio-visual synchronization delay.
[0349] Experiment D: Localization and Adaptive Adjustment
[0350] 1. Test conditions
[0351] When the user enters the field of view from different positions and directions, the system outputs adaptive adjustments to the screen tilt angle and sound field direction.
[0352] 2. Indicators
[0353] RMSE positioning error, 95% percentile error, screen tilt error, and directional adjustment time.
[0354] V. Specific Data Tables
[0355] Note: The units and statistical standards in the table below are as follows: Acc = Accuracy (dimensionless, value 0–1), RT = Reaction Time (milliseconds), Lat = End-to-end Latency (milliseconds), ΔE = CIEDE2000 color difference, RMSE = RootMean Square Error (meters), W = Watts, DataSize = Uploaded data volume (kilobytes, KB).
[0356] Table 1. Experiment A: Accuracy and response time of intent recognition in multiple scenarios (1000 events, divided by scenario)
[0357] Scene categories Sample size Accuracy (Acc) Average response time (ms) Remark Education 250 0.93 260 Blue light low intensity priority Fitness 250 0.90 290 Action recognition is quite complex Entertainment 250 0.95 240 Quick content switching Home 250 0.91 310 Scene diversity
[0358] Table 2. Experiment B: End-to-end content generation latency (1000 frames, single-scene balanced test)
[0359] Statistical items Numerical value (ms) Average delay 152 median 149 75th percentile 168 90th percentile 186 95th percentile 210 min / max 118 / 235
[0360] Table 3. Experiment C: Sound-Image Synchronization Consistency and Color Difference Mapping
[0361] Scene group Average SPL (dB) Average ΔE ΔE standard deviation Color temperature mapping error Δcolor temperature (unit: K) Remark Group 1 (Education) 66.5 1.80 0.32 35 Warm color scheme in the image Group 2 (Fitness) 65.9 2.40 0.41 42 Strong dynamic particle mapping Group 3 (Entertainment) 66.2 2.10 0.37 38 Passionate style Group 4 (Family Routine) 66.8 1.50 0.29 34 Daily Mode Group 5 (Comprehensive) 66.1 2.00 0.34 37 Overall performance
[0362] Table 4. Experiment D: Localization and Attitude Adaptation
[0363] project RMSE (m) 95th percentile error (m) Screen tilt angle error (°) Adjust time (ms) Positioning error 0.10 0.20 0.8 110 Attitude response 0.12 0.22 1.0 125 Overall performance 0.11 0.21 0.9 115
[0364] VI. Data Analysis Methods and Judgment Criteria
[0365] 1. Accuracy and robustness
[0366] The 95% confidence interval for intent recognition accuracy is calculated by repeating experiments (n ≥ 30 independent trials), with an upper limit ≥ 0.90 and a lower limit ≥ 0.85.
[0367] End-to-end latency distribution: Box plot of 1000 frames of data to ensure that the median is below 200 ms and the 75th / 90th / 95th percentiles are all within acceptable range.
[0368] 2. Consistency of audio-visual linkage
[0369] The mean ΔE must be significantly lower than 3, and the standard deviation must be controlled within 0.5. The color temperature mapping error of the image must be ≤50 K.
[0370] 3. Localization and attitude adaptation
[0371] RMSE < 0.15 m (target in an ideal indoor scene), 95% percentile error < 0.25 m, and the screen tilt angle relative to the owner's angle error is within 1°.
[0372] VII. Results and Conclusions (Comprehensive Analysis Based on Tabular Data)
[0373] Experiment A of this experimental group shows that in educational, fitness, entertainment and home scenarios, the system's intent recognition accuracy is stable above 0.90, and the average response time is focused in the range of 200-320 ms, which meets the requirements of "fast and stable".
[0374] Experiment B's end-to-end latency indicates that the average latency between perceived content output is approximately 152 ms, and the 95th percentile latency is approximately 210 ms, enabling a smooth audio-visual experience in most scenarios.
[0375] The average ΔE of Experiment C was approximately 2.16, with a standard deviation of approximately 0.34. All five scenes maintained a threshold of ΔE < 3, indicating that the color mapping and brightness adjustment of the audio-visual linkage had good consistency. The color temperature mapping error of D / A was in the 40–45 K range, and the overall visual performance matched the audio style well.
[0376] Experiment D achieved a positioning error RMSE of approximately 0.11 m, a 95% percentile error of approximately 0.21 m, a screen tilt angle adaptive error of less than 1°, and a response time of approximately 115–125 ms, verifying the rapid accuracy of indoor multi-source positioning and attitude adjustment.
[0377] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.
Claims
1. A picture frame audio system with intelligent recognition, characterized in that, The system comprises: a perception subsystem, an output subsystem, a main control and edge computing unit, and a dynamic content generation engine. The perception subsystem performs multimodal perception using millimeter-wave radar, a binocular RGB-D camera, a microphone array, and an ambient light sensor. The dynamic content generation engine generates content in real time at the edge based on a diffusion model and adaptively adjusts the generated results using user profiles and environmental data. The main control and edge computing unit performs data fusion, intent recognition, content generation scheduling, and synchronized audio-visual output. The output subsystem includes a 4K 120Hz micro-LED display and a controllable directional speaker system. The system achieves audio-visual linkage and environmental adaptation through a closed-loop mechanism.
2. The picture frame audio system with intelligent recognition according to claim 1, characterized in that: The diffusion model is a lightweight diffusion model that has been cropped, quantized, and optimized for operation on edge devices. Its generation process combines user profiles and environmental data to perform adaptive content generation frame by frame or segment by segment.
3. The picture frame audio system with intelligent recognition according to claim 2, characterized in that: The audio-visual linkage algorithm maps the audio FFT spectrum into a dynamic visual particle system through a generative adversarial network (GAN). The color, density, and motion trajectory of the particle system change synchronously with the audio features.
4. The picture frame audio system with intelligent recognition according to claim 3, characterized in that: The micro LED display screen is a magnetically levitated flexible OLED screen, and a micro servo motor is used to achieve ±5° dynamic deflection of the display screen to realize adaptive control of the sound field direction and viewing angle.
5. The picture frame audio system with intelligent recognition according to claim 4, characterized in that: The perception subsystem also includes a UWB positioning module, which is used to locate the user's position indoors in real time and couple the positioning information with screen posture adjustment, sound field orientation and image output strategies.
6. The picture frame audio system with intelligent recognition according to claim 5, characterized in that: Privacy protection mechanisms include localized federated learning and SEAL homomorphic encryption of uploaded data before transmission, ensuring the privacy and security of model updates across devices.
7. The picture frame audio system with intelligent recognition according to claim 6, characterized in that: The system achieves a perceptual consistency index where the SPL sound pressure level and the screen brightness ΔE are less than a certain threshold, and a low power consumption specification of ≤0.5W standby power consumption.
8. The picture frame audio system with intelligent recognition according to claim 7, characterized in that: The dynamic content generation engine can adaptively display NFT digital artworks, creating a Web3.0-level artwork presentation and interactive experience.
9. A method for implementing and controlling a picture frame audio system with intelligent recognition, characterized in that, The process is as follows: perception stage, intent recognition stage, content generation stage, audio-visual linkage stage, and feedback and learning stage. In the generation stage, the output video and audio output are synchronized in time and space, and privacy protection processing is performed for model updates through a local privacy protection mechanism.
10. The method for implementing the intelligent recognition-based picture frame audio system according to claim 9, characterized in that: In each round of feedback, the updated model parameters are encrypted and uploaded to the server through localized federated learning, and are used only for global model improvement without exposing individual user data.
Citation Information
Patent Citations
Video processing method and device
CN110166828A
Trigger-based PPDU resource indication for EHT networks
US20210144752A1