A low-delay intelligent white noise generation and playing method and system

CN122547993APending Publication Date: 2026-08-11SHENZHEN YINYEWANG TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-05-14
Publication Date
2026-08-11

AI Technical Summary

Technical Problem

[0003]当前白噪声播放与生成主要分为两类技术路线,二者均存在低延迟与内容多样性无法兼顾的核心缺陷:

Benefits of technology

[0017] Compared with existing technologies, the beneficial effects of this invention are as follows: This invention provides a low-latency intelligent white noise generation and playback method, which effectively solves the problem that existing white noise playback and generation methods cannot meet the dual requirements of low-latency response and content diversity. By building a two-level linked audio and video storage architecture and adopting a fixed hierarchical retrieval logic that prioritizes the first-level finished audio and video library, content matching and playback can be completed quickly. This fundamentally solves the problem of high latency and inability to respond instantly when generating inference data in large models, thus achieving low-latency playback. At the same time, a dual-track computing power scheduling mechanism is adopted, which uses lightweight front-end retrieval and asynchronous batch generation during back-end computing power troughs. The front-end ensures a fast response to user requests, while the back-end continuously generates and expands the audio and video material library during idle periods. This breaks the limitation of static and singular content in a fixed audio library, achieving dynamic improvement in content diversity. This effectively balances the core contradiction between low-latency response and content richness, better adapts to the instant playback and diverse sleep aid needs in sleep scenarios, and significantly improves the user's audiovisual interaction and sleep aid experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122547993A_ABST
    Figure CN122547993A_ABST
Patent Text Reader

Abstract

This invention provides a low-latency intelligent white noise generation and playback method and system. The method includes constructing a two-level linked audio-visual storage architecture, comprising a primary finished audio-visual library and a secondary fragmented audio-visual material library; performing hierarchical retrieval on user requests according to a fixed order where the primary finished audio-visual library takes precedence over the secondary fragmented audio-visual material library; employing a dual-track computing power scheduling mechanism of lightweight front-end retrieval response and asynchronous batch generation of audio-visual content during background computing power off-peak periods; dynamically expanding the two-level storage library through asynchronous background content generation, and achieving synchronous audio-visual retrieval and playback, thus balancing low-latency playback with content diversity. Beneficial effects: The two-level linked audio-visual storage and the primary finished library with priority hierarchical retrieval achieve low-latency playback; the dual-track scheduling of lightweight front-end retrieval and asynchronous generation during background computing power off-peak periods dynamically expands the content library, balancing the contradiction between low latency and diversity.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of artificial intelligence technology, specifically to a low-latency intelligent white noise generation and playback method and system. Background Technology

[0002] With the rapid development of smart home and sleep aid health fields, white noise playback has become a core function of smart speakers, sleep aids, and sleep assist devices. Users have put forward dual requirements for the response speed and content richness of sleep-aid white noise. They need a low-latency interactive experience of "speak and play" and diverse and personalized audio content to avoid auditory fatigue. At the same time, with the rise of audiovisual fusion sleep aid demand, the smoothness and consistency of audio and video synchronous playback have also become key experience indicators.

[0003] Currently, white noise playback and generation mainly fall into two technical categories, both of which suffer from the core drawback of being unable to simultaneously achieve low latency and diverse content: Fixed audio library retrieval and playback technology uses a pre-set audio file library and keyword matching to directly retrieve pre-stored audio files locally or in the cloud. Its advantages lie in its simple retrieval and playback process, low computational consumption, and millisecond-level low-latency response, meeting basic requirements for fast playback. However, this technology uses a statically stored audio library with fixed and limited content that cannot be updated independently. Long-term use can lead to issues such as repetitive playback and content homogenization, resulting in a severe lack of content diversity and failing to meet users' personalized and diverse sleep aid needs.

[0004] The large-scale model real-time generation and playback technology uses an AI audio large model to infer and generate content from user commands in real time. It can output diverse and personalized white noise content, with a much richer content than fixed-library solutions. However, the large-scale model inference computation is computationally intensive and time-consuming, and the real-time generation process has significant delays, making it impossible to achieve the interactive requirement of "speak and play". At the same time, the high latency of audio generation will directly lead to the misalignment of the generation sequence of the accompanying visual content, causing audio-visual asynchrony problems, poor audio-visual fusion experience, and difficulty in adapting to immersive sleep aid usage scenarios.

[0005] In summary, existing white noise playback and generation technologies suffer from an irreconcilable core contradiction: fixed audio libraries can achieve low latency but the content is static and monotonous, while large models can generate rich content in real time but have high inference latency. Neither can simultaneously meet the dual requirements of low latency response and content diversity. Moreover, the high latency problem directly leads to derivative defects such as audio-visual asynchrony and poor audiovisual experience, which has become a technical problem that urgently needs to be solved in this field. Summary of the Invention

[0006] To address the problems in the prior art, this invention provides a low-latency intelligent white noise generation and playback method and system.

[0007] This invention provides a low-latency intelligent white noise generation and playback method, comprising: A two-level interconnected audio and video storage architecture is constructed, which includes a first-level finished audio and video library and a second-level fragmented audio and video material library; The system performs a tiered search on user requests, prioritizing the primary finished audio and video library over the secondary fragmented audio and video material library. A dual-track computing power scheduling mechanism is adopted, which involves lightweight front-end retrieval response and asynchronous batch generation of audio and video content during back-end computing power off-peak periods. By dynamically expanding the two-level repository through asynchronous content generation in the background, and synchronously retrieving and playing audio and video, both low-latency playback and content diversity are achieved.

[0008] The present invention is further improved by receiving a user's playback request, parsing the request into an acoustic-visual fusion feature vector, and performing hierarchical retrieval and matching based on the fusion feature vector.

[0009] The present invention is further improved by adopting a dual matching degree threshold judgment result for hierarchical retrieval: if the matching degree of the first-level finished audio and video library meets the standard, it is played directly; if it does not meet the standard, the materials in the second-level fragmented audio and video material library are retrieved and quickly spliced ​​together before playback.

[0010] The present invention is further improved by using a dynamic adaptive threshold that is adjusted based on the user's historical retrieval data for the dual matching degree threshold.

[0011] The present invention is further improved by dynamically expanding the asynchronous content generation in the background, which includes: real-time collection of user playback behavior and feedback data, and automatic identification of audio and video content gaps; automatic generation of multimodal generation prompts based on the content gaps; targeted generation of audio and video content based on the multimodal generation prompts; and updating the generated audio and video content in the database, forming a fully automated content evolution closed loop without human intervention.

[0012] This invention is further improved by identifying audio and video content gaps through cluster analysis and association rule mining, accurately identifying content gaps, low-satisfaction content, and long-tail demand gaps.

[0013] The invention is further improved by dynamically expanding the asynchronously generated content in the background, which also includes performing fully automated multi-dimensional quality inspection on the generated and spliced ​​audio and video. The quality inspection dimensions include acoustic indicators, visual indicators, scene emotional adaptability and audiovisual fusion consistency, and unqualified content is automatically removed.

[0014] The present invention also provides a low-latency intelligent white noise generation and playback system, comprising: A hierarchical storage module is used to construct a two-level linked audio and video storage architecture, which includes a first-level finished audio and video library and a second-level fragmented audio and video material library. The hierarchical retrieval module performs hierarchical retrieval on user playback requests in a fixed order, prioritizing the first-level finished audio and video library over the second-level fragmented audio and video material library. The dual-track computing power scheduling module performs dual-track computing power scheduling, which involves front-end lightweight retrieval response and back-end asynchronous batch generation of audio and video content during computing power off-peak periods. The multimodal generation module continuously generates and pre-assembles content asynchronously, supplementing high-quality content for multi-level libraries; The dual-track computing power scheduling module dynamically expands the two-level storage repository by asynchronously generating content in the background, and the hierarchical retrieval module combines with the front-end retrieval to achieve low-latency playback.

[0015] The invention is further improved by including an audio and video splicing module and a sleep scene optimization module. The audio and video splicing module is used to complete the seamless splicing of secondary fragmented materials through acoustic feature matching and smooth fusion of transition frames. The sleep scene optimization module is used to optimize long audio playback without repetition and generate low-stimulation visual content specifically for sleep.

[0016] The invention is further improved by including a data closed-loop iteration module and an automated quality inspection and control module. The data closed-loop iteration module is used to construct a fully algorithm-automated content evolution closed loop to achieve autonomous updates of the library. The automated quality inspection and control module is used to perform full-dimensional automated detection of audio and video content and remove unqualified content.

[0017] Compared with existing technologies, the beneficial effects of this invention are as follows: This invention provides a low-latency intelligent white noise generation and playback method, which effectively solves the problem that existing white noise playback and generation methods cannot meet the dual requirements of low-latency response and content diversity. By building a two-level linked audio and video storage architecture and adopting a fixed hierarchical retrieval logic that prioritizes the first-level finished audio and video library, content matching and playback can be completed quickly. This fundamentally solves the problem of high latency and inability to respond instantly when generating inference data in large models, thus achieving low-latency playback. At the same time, a dual-track computing power scheduling mechanism is adopted, which uses lightweight front-end retrieval and asynchronous batch generation during back-end computing power troughs. The front-end ensures a fast response to user requests, while the back-end continuously generates and expands the audio and video material library during idle periods. This breaks the limitation of static and singular content in a fixed audio library, achieving dynamic improvement in content diversity. This effectively balances the core contradiction between low-latency response and content richness, better adapts to the instant playback and diverse sleep aid needs in sleep scenarios, and significantly improves the user's audiovisual interaction and sleep aid experience. Attached Figure Description

[0018] To more clearly illustrate the solutions in this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0019] Figure 1 This is a system block diagram of a low-latency intelligent white noise generation and playback system according to the present invention; Figure 2 This is a flowchart of a low-latency intelligent white noise generation and playback method according to the present invention. Detailed Implementation

[0020] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs; the terminology used herein in the specification of the application is for the purpose of describing particular embodiments only and is not intended to limit the application; the terms “comprising” and “having” and any variations thereof in the specification, claims and foregoing description of the drawings are intended to cover non-exclusive inclusion.

[0021] In this document, the term "embodiment" means that a particular feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of this application. The appearance of this phrase in various places throughout the specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment mutually exclusive with other embodiments. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.

[0022] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings.

[0023] like Figure 1 As shown, this invention provides a low-latency intelligent white noise generation and playback method system. This system is suitable for collaborative deployment between sleep aid terminals and cloud servers in sleep scenarios. The hardware environment includes: edge-side intelligent sleep aids, smart speakers, mobile phones, tablets, and other low-computing-power interactive devices, with functions such as voice acquisition, audio and video playback, local caching, and network communication; the cloud server has multimodal large model inference, computing power scheduling, massive data storage, and data analysis capabilities. The software environment includes: a large language model engine, a multimodal encoder, an audio feature extraction model, a visual generation model, clustering analysis algorithms, cosine similarity retrieval algorithms, etc., supporting fully automated operation of the system.

[0024] The core of this system includes a hierarchical storage module, a graded retrieval module, a dual-track computing power scheduling module, a multimodal generation module, an audio and video splicing module, a sleep scene optimization module, a data closed-loop iteration module, and an automated quality inspection and control module. Each module has its own independent division of labor and works together to cover all technical functions such as low-latency playback, content diversity expansion, sleep adaptation, autonomous evolution, and full-dimensional quality inspection.

[0025] For tiered storage modules: This module serves as the core data foundation of the system, constructing a two-tiered, interconnected audio and video storage architecture, comprising a primary finished audio and video library and a secondary fragmented audio and video material library: Level 1 finished audio and video library: Stores long audio and video files that can be played directly for long sleep scenarios, categorized and labeled by scenario, emotion, and sleep aid intensity, and supports direct retrieval and playback on the device. Secondary fragmented audio and video material library: Stores short-term, high-quality fragmented audio and video materials, binds and associates them according to acoustic features and visual elements, and reserves splicing rules and feature indexes.

[0026] It adopts a distributed deployment architecture of local caching on the device side and full storage in the cloud: the device side caches high-frequency hot content to achieve zero network latency retrieval; the cloud stores the full audio and video content to support backend generation and iteration.

[0027] Implement intelligent management rules for prioritizing the caching of hot content and automatically archiving cold content: content with a retrieval hit rate of ≥50% in the past 30 days is considered hot content and is prioritized for storage on the client side; cold content with a retrieval hit rate of <10% is automatically archived to cloud cold storage to improve storage efficiency and retrieval speed.

[0028] For the hierarchical retrieval module: This module is the central decision-making unit of the system. It receives user voice / text playback requests collected from the receiving end and parses the requests into acoustic-visual fusion feature vectors through a multimodal encoder. The acoustic-visual fusion feature vectors contain acoustic and visual information and extract core features such as scene, emotion, and sleep needs.

[0029] Hierarchical retrieval is performed in a fixed order, prioritizing the primary finished audio and video library over the secondary fragmented audio and video material library, and the cosine similarity algorithm is used to complete the matching calculation between the fused feature vector and the audio and video content.

[0030] Configure dual matching thresholds to determine search results. The thresholds are dynamically adaptive thresholds adjusted based on the user's historical search data: if the matching degree of the first-level finished product library meets the standard, it will be directly retrieved and played; if it does not meet the standard, it will trigger the splicing and playback of the second-level material library; the system response latency is ≤30ms, achieving millisecond-level low-latency interaction.

[0031] For the dual-track computing power scheduling module: This module is the core unit ensuring low latency and content diversity of the system, and specifically performs the following functions: Front-end scheduling: Executes lightweight search logic, relies on hierarchical search modules to complete fast matching, does not consume high computing resources, and ensures immediate response to user requests.

[0032] Background scheduling: During periods of idle computing power or low computing power usage at night, asynchronously batch generate fragmented audio and video materials and finished long audio and video products, dynamically expand the primary finished audio and video library and the secondary fragmented audio and video material library, continuously enrich the diversity of content, without consuming front-end interactive computing power.

[0033] Front-end and back-end collaboration: Low-latency playback on the front end and content generation on the back end run synchronously, achieving a dual unity of low-latency response and unlimited content diversity.

[0034] For the multimodal generation module: This module, serving as the core of backend content production, continuously generates and pre-assembles content asynchronously, supplementing multi-level libraries with high-quality content. The core achieves a high degree of audio-visual matching through a dual audio-user intent mapping module: first, it analyzes the acoustic details (frequency, rhythm, emotional atmosphere) of the generated audio, then integrates the user's original intent features to generate an audio-visual matching feature vector. Based on this, it optimizes the visual generation parameters, ensuring a high degree of consistency between the visual content and the audio in terms of scene, emotion, and detail. The generated short-duration audio-visual fragments undergo initial inspection before being placed into a secondary fragmented audio-visual material library. The long audio and accompanying video, automatically assembled in the backend, undergo full-process quality inspection before being placed into a primary finished audio-visual library, achieving an automated flow of "generation-assembly-entry into the library."

[0035] For the "audio-user intent dual mapping module" setup, the concept of "Human Feedback Reinforcement Learning (RLHF)" is introduced to achieve multi-module co-evolution while retaining the core logic of multi-level storage hierarchical retrieval and full algorithm evolution. The implicit behavioral feedback of users (such as completing audio and video playback = positive reward, skipping = negative reward, high audio-visual matching degree = positive reward, high first-level library retrieval hit rate = positive reward) is converted into reward signals to train a multimodal quality reward model; The automated quality inspection and control module uses the score of the multimodal quality reward model as the core quality inspection indicator, and feeds back the reward signal to the multimodal generation module, the audio-user intent dual mapping module and the audio-video splicing module. The three models improve the quality of subsequent generated / spliced ​​audio and video, sleep scene adaptability, audio-visual matching degree and splicing success rate based on the reward signal self-optimization generation strategy, matching strategy and splicing strategy. The collaborative evolution of the quality inspection, generation, matching, and splicing modules is achieved, further improving audio and video quality, user intent matching, audiovisual fusion effect, and retrieval hit rate of the multi-level storage system. The core multi-level storage hierarchical retrieval, low-latency dual-track scheduling, full algorithm closed loop, and sleep scenario adaptation remain unchanged.

[0036] For the audio / video splicing module: This module is a fragmented material processing unit. It receives secondary fragmented audio and video materials from the hierarchical retrieval module and completes seamless and rapid splicing through acoustic feature matching and smooth transition frame fusion technology.

[0037] The splicing process is seamless with no audio breaks or abrupt loops, and the splicing time is ≤50ms. After splicing, a long audio video specifically for sleep is generated and pushed to the playback device.

[0038] The spliced ​​audio and video are synchronized with the visual content to ensure smooth playback.

[0039] For the audio / video synchronization unit: This module is an audiovisual consistency guarantee unit, enabling the system to support full-process audio and video synchronization. It maintains timing alignment throughout the entire chain of generation, retrieval, splicing, and playback, achieving synchronized audio and video playback. This completely solves the problems of audiovisual timing deviation and audio-visual asynchrony, ensuring audiovisual fusion consistency.

[0040] For the sleep scene optimization module: This module is a dedicated adaptation unit for sleep mode, and it performs the following operations: Audio optimization: The spliced ​​long audio is optimized for non-repetitive playback. The low-frequency ratio is adjusted and the loudness fluctuation is reduced through a sleep-specific auditory model to adapt to the long-term playback needs during sleep.

[0041] Visual optimization: Generate low brightness ≤50cd / m 2 Low-stimulation visual content with low color saturation ≤30% and slow dynamic frame rate ≤10fps, without abrupt images or strong light flicker, and image switching interval ≥30 seconds, reduces visual disturbances during sleep.

[0042] For the data closed-loop iteration module: This module is the system's autonomous evolution unit, constructing a fully automated content evolution closed loop with no human intervention throughout the process. Collect playback feedback data such as user playback duration, favorites, skipped videos, completed videos, and subjective ratings; Cluster analysis and association rule mining are used to identify content gaps, low-satisfaction content, and long-tail demand. Automatically generate multimodal generation prompts based on content gaps to drive audio and video generation in a targeted manner; Qualified content is categorized and updated in the database to enable the audio and video library to evolve autonomously and sustainably.

[0043] For automated quality inspection and control modules: This module serves as the system's content quality control unit, performing fully automated, multi-dimensional quality checks on the generated and spliced ​​audio and video. Unqualified content is automatically removed, and only qualified content is added to the database. Acoustic performance quality inspection: Spectrum integrity, signal-to-noise ratio ≥40dB, no popping sounds, no background noise, and no audio abrupt changes; Visual quality control: Detection of sharpness, brightness, saturation, and dynamic frame rate to ensure compliance with low-stimulation sleep standards; Scene-based emotional adaptation quality inspection: Determine whether the semantic and emotional fit between the audio / video and the user's request meets the standards; Audiovisual fusion consistency quality inspection: verify the matching degree of audio and video in terms of scene, emotion, rhythm and details to ensure audiovisual unity.

[0044] The overall collaborative workflow of this low-latency intelligent white noise generation and playback system is as follows: (1) Low-latency playback process in front end The terminal-side interactive playback module of the terminal-side interactive device collects the voice / text commands issued by the user and uploads them to the hierarchical retrieval module; The hierarchical retrieval module is based on a large language model and a multimodal encoder. It deeply analyzes the scene description, emotional tendency, and sleep needs (such as "low stimulation" and "no repetition") in the instructions and transforms them into high-dimensional acoustic-visual fusion feature vectors. Based on the cosine similarity algorithm, it performs hierarchical retrieval in the order of first-level finished audio and video library → second-level fragmented audio and video material library, and links dynamic feature library for millisecond-level retrieval to determine whether the content matching degree meets the standard. If the matching score of the first-level finished audio and video library is ≥90 (upper limit of double threshold), the finished audio and video will be directly retrieved; if the matching score is 70-89 (range of double threshold), highly relevant materials will be extracted from the second-level fragmented audio and video material library for quick splicing; if the matching score is <70 (lower limit of double threshold), alternative content will be played and the demand gap will be marked. The terminal-side interactive playback module of the terminal-side interactive device: completes audio-visual timing calibration, pushes audio and video synchronously to the terminal, realizes "speak and play, audio and video synchronization", and controls the response latency to the millisecond level; After the sleep scene optimization module completes the audio and video adaptation optimization, it is pushed to the terminal synchronously, and the synchronous playback unit ensures audio and video synchronization. Meanwhile, the system collects user playback feedback (playback duration, skip / favorite / complete playback status, subjective satisfaction rating) in real time and uploads it to the user intent log database, which is then uploaded to the data closed-loop iteration module to provide accurate data support for backend evolution.

[0045] (2) Backend autonomous evolution process Construct a fully automated closed loop of algorithms for "data acquisition - gap identification - targeted generation - quality inspection and warehousing".

[0046] The data closed-loop iteration module periodically scans the user intent log database, analyzes user feedback, and identifies content gaps in the multi-level storage system (such as high-frequency unmatched needs, low-satisfaction content categories, and missing long-tail scenario materials) through cluster analysis and association rule mining. Automatically generate multimodal prompts, transform content gap requirements into standardized multimodal prompts (clearly defining acoustic parameters, visual features, and sleep adaptation requirements), extract audio, user intent, and basic visual generation parameters, and send them to the dual-track computing power scheduling module; The multimodal generation module asynchronously generates audio and video clips in batches during periods of low computing power (when equipment is idle or during low usage at night), and automatically calls splicing algorithms to convert the materials into finished audio and video products; All generated / assembled content undergoes comprehensive quality inspection by the automated quality control module and sleep-specific optimization by the sleep scene optimization module. Qualified content is then stored by type in the secondary fragmented audio and video material library (segments) and the primary finished audio and video library (finished products) of the hierarchical storage module. This enables the autonomous and sustainable expansion of the multi-level storage system. The two-level storage libraries are dynamically updated, allowing the system to become "richer in content and more accurate in matching the more it is used," thus continuously evolving the system.

[0047] (3) Full-process quality inspection process After audio and video generation and splicing are completed, the automated quality inspection and control module sequentially performs acoustic index testing (Mel spectrum / spectral integrity / signal-to-noise ratio ≥40dB / no background noise / pops / abruptness / continuity, etc.), visual index testing (clarity / brightness / saturation / dynamic frame rate / low stimulation, etc. meet the standards), scene emotion adaptation testing (audio and video semantic / emotional fit with Prompt ≥80 points), and audiovisual fusion consistency testing (matching degree of four aspects: scene / emotion / rhythm / detail ≥90 points). Unqualified content is directly rejected, and qualified content enters the storage and playback chain, with no manual spot checks throughout the entire process.

[0048] This embodiment solves the core problem of the trade-off between low latency and content diversity: front-end retrieval response latency ≤30ms, splicing time ≤50ms, achieving millisecond-level playback; the back-end asynchronously generates and continuously expands the content library, with infinitely diverse content. It achieves full-process synchronization of audio and video, completely eliminating the problem of audio-visual asynchrony and providing a superior audio-visual experience; For sleep scenarios, a dedicated optimization system is built to address core sleep-aid needs, combining audio and visual elements throughout the entire process of front-end playback and back-end generation / stitching. On the audio side: a high-quality segment intelligent selection algorithm extracts high-quality materials from a secondary fragmented audio and video material library. Through acoustic feature matching and smooth transition frame fusion technology, seamless and non-repetitive long audio stitching is achieved, avoiding auditory fatigue caused by traditional loop playback. On the visual side: low-brightness (≤50cd / m²) audio is generated. 2 It features low color saturation (≤30%), slow dynamics (frame rate ≤10fps), and low-stimulation content without abrupt visuals, supporting slow-paced image switching (interval ≥30 seconds) and long-duration slow-motion video playback; at the same time, through a sleep-specific auditory model (optimizing low-frequency proportion and reducing loudness fluctuations) and a visual model (filtering strong light and weakening color contrast), it further enhances the sleep-aiding effect, fully adapting to the low-interference and high-immersion needs of sleep scenarios, and the low-stimulation audio and video significantly improve the sleep-aiding effect; A closed-loop system for autonomous algorithm evolution is constructed without human intervention, and the library becomes richer and more accurate with use. Fully automated, multi-dimensional quality inspection ensures audio and video quality and consistency with audiovisual content, eliminating substandard content.

[0049] (4) Hierarchical retrieval scheduling logic As a core guarantee for low latency, an intelligent scheduling mechanism of "dual thresholds + priority sorting" is adopted. Based on the fused feature vector of user intent, the system prioritizes searching the primary finished audio and video library (response latency < 30ms). If the matching degree meets the standard, it plays directly; if not, it searches the secondary fragmented audio and video material library (splicing time < 50ms), and quickly splices the materials after sorting them by material relevance and quality score; if neither library meets the standard, it triggers asynchronous generation in the background, and the closest alternative content is played in the foreground, with the user experiencing no waiting time. This logic ensures that more than 90% of user requests achieve millisecond-level response, balancing low latency and content coverage.

[0050] (5) Equivalent implementation method The modular splitting, deployment methods, threshold parameters, algorithm replacements, and other conventional variations of this system are all equivalent implementations of this invention. For example, expanding two-level storage to three-level storage, replacing cosine similarity with Euclidean distance algorithm, and changing edge-cloud deployment to local centralized deployment all fall within the protection scope of this invention.

[0051] like Figure 2As shown, this invention also provides a low-latency intelligent white noise generation and playback method. This method is used in the aforementioned system and fully realizes two-level storage construction, hierarchical retrieval, dual-track computing power scheduling, low-latency playback, and dynamic content expansion. It also achieves all technical effects such as feature matching, threshold determination, seamless splicing, audio-visual synchronization, sleep adaptation, content evolution, automated quality inspection, and intelligent storage management. The specific implementation steps are as follows: 1. Construct a two-level interconnected audio and video storage architecture A two-tiered, interconnected audio and video storage architecture is established, consisting of a primary finished audio and video library and a secondary fragmented audio and video material library. This architecture employs a distributed deployment architecture that combines local caching on the client side with full storage in the cloud. (1) Level 1 finished audio and video library: Stores long audio and video that can be played directly in the long sleep scene, classified by scene and emotional tags, and caches high-frequency hot content locally on the device to achieve zero network latency retrieval; (2) Secondary fragmented audio and video material library: Stores short-term high-quality audio and video fragmented materials, which are associated and bound according to acoustic features and visual elements, and have reserved splicing rules. The full amount of materials are stored in the cloud and the terminal is incrementally synchronized as needed; The design approach of "distributed deployment architecture of local caching on the device side + full storage in the cloud" can also adopt "centralized storage in the cloud".

[0052] Building upon this foundation, further optimizations can be made by adding a hot / cold tiered management mechanism and dynamic migration rules. This will automatically identify content popularity through algorithms, enabling intelligent transfer between different storage levels while retaining the core logic of tiered retrieval. Level 1 finished audio and video library (hot library): Only stores "high-hot finished products" with a search hit rate of ≥50% in the past 30 days, with full coverage of local cache on the device side to ensure extremely low latency; Secondary fragmented audio and video material library (warm library): stores "medium-hot material" generated / stitched in the past 90 days, sorted by popularity, with high-frequency material being synchronized to the end first; The third-level backend generation scheduling layer (cold storage association layer) associates with an independent "cold storage repository" to store "cold finished products / cold materials" with a retrieval hit rate of <10%, without occupying the core storage resources of the first two levels of the repository; Dynamic migration rules: The content popularity is monitored in real time by the automated data analysis module. High-popularity content is upgraded from cold storage to warm storage to hot storage, while low-population content is downgraded from hot storage to warm storage to cold storage. The migration process does not affect the front-end search. Hierarchical retrieval logic: Prioritize retrieval of the first-level hot storage → second-level warm storage → third-level cold storage related layer. If no cold storage is found, trigger the background generation.

[0053] This approach can significantly improve the retrieval efficiency and storage utilization of the first two levels of the database, reduce the pressure on the client-side cache, and prevent cold content from occupying core resources. The core multi-level storage hierarchical retrieval and the closed-loop logic of the whole algorithm evolution remain unchanged.

[0054] 2. Receive user requests and parse them into acoustic-visual fusion feature vectors. By collecting user voice / text playback requests through edge devices, such as "play low-stimulation deep-sea white noise to help you sleep," the requests are deeply analyzed using a large language model and a multimodal encoder to extract features such as scene, emotion, and sleep needs. These features are then encoded into an acoustic-visual fusion feature vector, which serves as the core basis for subsequent retrieval and matching.

[0055] Users can express their needs in various ways, such as voice / text, uploading sleep scene pictures (e.g., "deep sea pictures"), humming melodies, selecting sleep mood tags (e.g., "relaxed, quiet"), and describing visual needs (e.g., "low brightness, slow motion"). The multimodal encoder maps multimodal data such as images, audio clips, text, tags, and visual descriptions to the same acoustic-visual fusion feature vector. The subsequent front-end multi-level storage hierarchical retrieval, back-end multimodal generation, audio-visual matching, quality inspection, and sleep scene optimization logic are completely consistent with the present invention.

[0056] By setting up a multimodal encoder, the ways users can express their intentions are expanded, improving the richness of interaction and the accuracy of intent parsing. The core features, such as multi-level storage hierarchical retrieval, low-latency dual-track scheduling, full algorithm closed loop, sleep scene adaptation, and audio-video high-matching generation, remain unchanged.

[0057] 3. Perform hierarchical retrieval and feature matching according to fixed priorities. Following a fixed order where the primary finished audio and video library takes precedence over the secondary fragmented audio and video material library, a cosine similarity algorithm is used to perform matching calculations between the fused feature vectors and the audio and video content. The search results are then determined by combining a dual-matching degree threshold, which can be dynamically adjusted based on the user's historical search data. (1) If the matching degree of the first-level finished product library is ≥90 points, it is judged as a matching standard and the finished product audio and video can be directly retrieved and played. (2) If the matching degree of the first-level finished product library is 70-89 points, it is judged as unqualified, and materials from the second-level material library are retrieved for rapid splicing; (3) If the matching score is less than 70, play the alternative content and mark the content gap; (4) The front-end hierarchical retrieval response latency is ≤30ms to ensure low-latency interaction.

[0058] For the "Dynamic Adaptive Threshold" based on historical retrieval data from the user intent log database, the matching threshold of the primary finished audio and video library / secondary fragmented audio and video material library is dynamically adjusted in real time: for example, if the primary finished audio and video library has a high hit rate for a certain type of content, the threshold of the primary finished audio and video library is appropriately increased to ensure playback quality; if the primary finished audio and video library has a low hit rate for a certain type of content, the threshold of the secondary fragmented audio and video material library is appropriately decreased to improve the coverage of spliced ​​playback.

[0059] The design of "dynamic adaptive threshold" is not limited to this; it can also be "fixed double matching degree threshold".

[0060] 4. Seamless and rapid splicing of secondary fragmented materials The system rapidly splices fragmented audio and video materials at level 2 using acoustic feature matching and smooth transition frame fusion technology. The splicing time is ≤50ms. After splicing, the audio is seamless and without abrupt loops, generating a unique long audio file for sleep without repetition. Corresponding visual content is generated simultaneously to ensure smooth playback.

[0061] Synchronous playback (device-side interactive playback module): Completes audio-visual timing calibration, synchronously pushes audio and video to the terminal, realizes "instant playback, audio and video synchronization", and controls the response latency to the millisecond level.

[0062] 5. Dual-track computing power scheduling and synchronized audio and video playback A dual-track computing power scheduling mechanism is adopted, which combines front-end lightweight retrieval with back-end asynchronous generation of computing power during off-peak periods. (1) Front end: Low latency response is achieved through hierarchical retrieval, audio and video are generated, retrieved and played synchronously, eliminating the problem of audio-visual asynchrony and ensuring the consistency of audio-visual fusion; (2) Backend: During periods of idle computing power / low usage at night, asynchronously generate audio and video materials and long audio and video finished products in batches, dynamically expand the primary finished audio and video library and the secondary fragmented audio and video material library, and continuously enrich the diversity of content.

[0063] 6. Optimized for specific sleep scenarios Specifically adapted for sleep scenarios: (1) Audio end: The spliced ​​long audio is optimized without repetition. The low frequency ratio is optimized and the loudness fluctuation is reduced through the sleep-specific auditory model to adapt to long-term sleep aid playback; (2) Visual end: Generate low brightness ≤50cd / m 2 Low-stimulation visual content with low color saturation ≤30% and slow dynamic frame rate ≤10fps, without abrupt changes in the image, and slow image switching interval ≥30 seconds to reduce visual interference.

[0064] 7. Fully automated multi-dimensional quality inspection The generated and spliced ​​audio and video are subjected to fully automated multi-dimensional quality inspection, which includes the following dimensions: (1) Acoustic indicators: Detection of spectrum integrity, signal-to-noise ratio ≥40dB, no popping / background noise / abrupt changes, etc., and any unqualified items will be directly rejected; (2) Visual indicators: Parameters such as clarity, brightness, saturation, and dynamic frame rate are detected. Those that do not meet the requirements for low stimulation during sleep are eliminated. (3) Scene and emotional fit: The scene, semantics of the user request, and emotional fit of the audio and video with the prompt words are judged to be ≥80 points, and qualified content is selected by combining the naturalness score; (4) Audiovisual integration consistency: The matching degree of the four aspects of scene, emotion, rhythm and detail is ≥90 points. If the standard is not met, it will be backtracked for optimization or directly eliminated. (5) Perform final adaptability verification for sleep aid scenarios. Unqualified content is sent to the generation module for optimization. If it still does not meet the standards, it is discarded. (6) Only the full inspection content is used to enter the optimized storage stage, and the inspection data is fed back to generate and quality inspection models to continuously improve accuracy.

[0065] 8. Fully Algorithm-Automated Content Evolution Closed Loop Building a fully algorithm-automated content evolution closed loop without human intervention: (1) Collect user playback duration, favorites, skips, subjective evaluations, and imperfect matches, and write them into the user intent log database to provide a data source for backend iteration; (2) Identify content gaps, low-satisfaction content and long-tail demand, and low-matching categories through cluster analysis and association rule mining, and generate a structured gap analysis report; (3) Transform the demand gap into standardized multimodal generation prompts containing acoustic, visual, and sleep adaptation parameters; (4) Receive standardized multimodal generated prompts from the backend and extract audio, user intent, and visual basic generation parameters; (5) Extract acoustic features from the initially generated audio segments, analyze frequency, rhythm, emotion, and scene details, and generate acoustic feature vectors; (6) The audio acoustic features are fused with the user intent features to generate an audiovisual matching feature vector, which serves as the core basis for visual generation; (7) Based on the fusion feature vector, the visual generation parameters are corrected in real time to ensure that the visual style, dynamics and audio and user intent are highly consistent; (8) Based on the optimization parameters, generate matching visual segments to achieve synchronous generation of audio and video from the same source; (9) Calculate the audiovisual matching score. If the score is not met, backtrack and optimize until it meets the preset standard. (10) Conduct full-process quality control on generated content to screen out inferior and non-compliant content; (11) Optimize qualified content in an integrated manner with low interference, low frequency and slow dynamics to enhance the sleep aid effect; (12) Extract audio and video feature vectors and scene tags, establish audio-visual associations, and update them to the two-level storage repository and dynamic feature repository; (13) New content is integrated into the front-end search link to form a fully automatic evolution loop of "demand-analysis-generation-playback-feedback" and continuously improve content adaptability.

[0066] Qualified content is categorized and updated in the database, enabling the audio and video library to evolve autonomously and sustainably. The more the system is used, the richer the content becomes and the more accurate the matching becomes.

[0067] Through the above steps, the core contradictions of limited content in a fixed audio library and high latency in generating large models are perfectly resolved, achieving a dual unity of millisecond-level low-latency response and unlimited content diversity. At the same time, it completes dedicated adaptation for sleep scenarios, fully automated content evolution, multi-dimensional quality inspection, and synchronized audio-visual playback, providing users with an immersive sleep aid experience. The front-end retrieval response latency is ≤30ms, the material splicing time is ≤50ms, and more than 90% of user requests can be played instantly. The audio and video library continues to expand autonomously without human intervention, fully meeting the sleep aid needs of sleep scenarios.

[0068] Key term definitions: Acoustic-visual fusion feature vector: A high-dimensional vector that is extracted from user natural language commands through multimodal encoding and has semantic, scene, emotion, acoustic and visual attributes. It is the core matching basis for retrieval and generation.

[0069] Dual-track computing power scheduling mechanism: The architecture of front-end millisecond-level lightweight retrieval + back-end asynchronous generation during computing power troughs ensures low-latency response in the front-end and no front-end computing power consumption in the back-end, solving the pain point of high latency in large model generation.

[0070] Multimodal generation prompts: These synchronously cover audio and visual generation requirements, including standardized instructions for scene, emotion, parameters, sleep adaptation, and audiovisual matching rules, and are the core basis for targeted generation.

[0071] A fully algorithm-driven, automated content evolution closed loop: a process without human intervention, from playback feedback to gap analysis, prompt generation, audio and video generation, audiovisual matching, automated quality inspection, and database updates, enabling the content library to evolve autonomously and sustainably.

[0072] Long audio splicing for sleep scenarios: Employing segment selection, acoustic matching, and smooth transition technology, it achieves seamless and non-repeating splicing of fragmented audio, adapting to the needs of long-term playback during sleep.

[0073] Low-stimulation visual content for sleep scenarios: Visual materials that meet the characteristics of low brightness, low saturation, slow motion, and no abrupt visuals, and are suitable for the need for low-interference sleep aid.

[0074] Audio-User Intent Dual Mapping: A technology that integrates audio acoustic features and user intent features to generate audiovisual matching vectors and guide visual generation, ensuring that audio and video are highly aligned with user needs.

[0075] Fully automated multi-dimensional quality inspection: A fully automated quality control system that integrates acoustic, visual, scene emotion, and audiovisual fusion detection, with no manual sampling throughout the process, strictly controlling content quality and adaptability.

[0076] Computing power trough: A period of low computing power utilization on the terminal / cloud. During this period, asynchronous generation does not affect the front-end interactive experience.

[0077] Audiovisual integration consistency: The degree to which audio and visual elements match in terms of scene, emotion, rhythm, and detail is a core indicator of an immersive sleep aid experience.

[0078] Long-tail demand: Low-frequency, personalized audio-visual sleep aid needs in sleep scenarios can be accurately covered through a closed-loop algorithm.

[0079] Compared with the prior art, the present invention has the following innovative features: (1) Relying on multi-level storage and hierarchical retrieval, true millisecond-level ultra-low latency is achieved. This invention innovatively designs a multi-level audio and video storage architecture and hierarchical retrieval and scheduling logic to solve the pain point of long real-time processing of complex user requests from the bottom layer: After a user request is sent, the first-level finished product library (pre-assembled long audio and video) is retrieved first, and if a match is found, it is played directly; if the matching degree of the finished product library is low, the second-level segment library (high-quality fragmented material) is retrieved, and it is quickly and intelligently spliced ​​and played; background processing is only triggered when there is no suitable material in the segment library, and the whole process does not require the user to wait, truly realizing "speak and play", which is different from the high latency defects of existing single-level storage.

[0080] (2) The combination of low-latency dual-track scheduling and multi-level storage achieves a dual unity of unlimited content diversity and ultimate response. This invention deeply integrates a dual-track computing power scheduling architecture—combining millisecond-level lightweight feature retrieval in the front end with asynchronous batch generation during off-peak computing power periods in the back end—with multi-level storage: the front end achieves hierarchical retrieval based on multi-level storage, and completes synchronized audio and video playback with low computing power and low latency; the back end utilizes idle device periods / off-peak computing power periods to continuously and asynchronously generate fragmented audio and video segments in batches, and automatically splices them into finished audio and video, continuously expanding the first-level finished product library and the second-level segment library, realizing infinite diversity of content, and making the system "richer in content the more it is used, higher in retrieval hit rate the more it is used, and lower in latency the more it is used."

[0081] (3) Construct a fully algorithmic and automated audio and video library evolution loop to completely eliminate manual intervention. This invention replaces traditional manual operation with a dynamic feedback and prompt generation network, forming a closed-loop algorithm of "playback feedback - gap analysis - automatic prompt generation - targeted generation - fully automated quality inspection - dual storage of segments / finished products": The algorithm automatically analyzes user behavior feedback and audio-visual quality gaps, transforms gap requirements into standardized prompt parameters, and simultaneously drives the targeted generation of audio and visual content. After quality inspection, the generated segments are stored in a secondary segment library, and the background automatically splices the segments into finished products and stores them in a primary finished product library. The entire process is done without human intervention, which not only solves the problems of low creativity and easy recognition as conventional technical means by manual intervention, but also realizes the autonomous, efficient, and sustainable evolution of the audio and video library.

[0082] (4) Fully automated, multi-dimensional, and robust quality inspection ensures consistent audio-visual quality and sound quality. This invention constructs a fully automated, multi-dimensional quality inspection system encompassing acoustic indicators, visual indicators, semantic / scene / emotional adaptation, and audiovisual fusion consistency, eliminating all manual sampling steps. It utilizes Fourier transform, acoustic fingerprinting, and Mel-spectrum analysis to detect core acoustic indicators such as audio signal-to-noise ratio and spectral integrity; and image / video clarity, color saturation, and frame rate to detect core visual indicators. A multimodal large-scale model is combined to detect the semantic, scene, and emotional adaptability of audio / video content to user commands. Furthermore, an audiovisual fusion consistency detection system is added to ensure a high degree of matching between visual content and audio details. This fully automated system controls audio / video quality and audiovisual fusion effects, ensuring that only content that passes quality inspection enters the multi-level storage system, guaranteeing the quality of content within the library.

[0083] (5) Intelligent audio and video optimization specifically for sleep scenarios, completely solving the problem of auditory and visual fatigue. Addressing industry pain points such as fragmented AI audio, stiff traditional loop filling, and highly stimulating visual content, this invention features a dual optimization technology specifically designed for sleep scenarios: intelligent long audio splicing and low-stimulation visual content generation. This optimization capability extends throughout the entire multi-level storage process. On the audio side: a high-quality segment intelligent selection algorithm is used to dynamically extract high-quality segments from the secondary segment library. Through acoustic feature matching and smooth transition frame fusion technology, seamless and natural long audio splicing is achieved. Each playback is a brand new combination with no repeated melodies. The spliced ​​product is stored in the primary product library. Visual aspect: Based on the needs of sleep scenarios, generate low-brightness, low-color saturation, slow-motion, and non-jarring low-stimulation images and videos to meet the core needs of long sleep duration, low interference, and high immersion. Simultaneously, through automated quality inspection and optimization using a sleep-specific auditory and visual model, it completely eliminates the cyclical and harsh feeling of traditional white noise and the high stimulation of visual content, significantly improving the sleep-aiding effect.

[0084] (6) Audio and video are highly matched to generate an immersive audio-visual fusion sleep aid experience. This invention breaks through the limitations of existing shallow scene matching, achieving highly accurate generation of visual content based on the dual mapping of audio acoustic features and user intent. Audio-video matching is integrated throughout the entire process of generation, splicing, and storage: A multimodal large model analyzes the acoustic details, rhythm, emotional atmosphere, and deep intent of the audio, generating images and videos that highly match the audio in terms of scene, emotion, rhythm, and detail. For example, if the audio is "gentle bubbles in the deep sea + slow waves," a video of "low-brightness deep-sea scenes, slowly swaying seaweed, and faint bubble light and shadow" is generated simultaneously, achieving an immersive audio-visual fusion experience with synchronized audio and video, detailed matching, and emotional consistency. Furthermore, the matched audio-visual segments / finished products are stored in their respective storage levels, ensuring audio-visual consistency for subsequent retrieval.

[0085] (7) Low computing power consumption and high compatibility with end-side devices The multi-level storage hierarchical retrieval architecture and lightweight front-end retrieval and multimodal generation modules of this invention significantly reduce the computing power consumption of edge devices: the first-level finished product library can be cached locally on the edge to achieve zero network latency playback, the second-level segment library can be incrementally synchronized to the edge as needed, and only a few complex requests require cloud collaboration, which can be directly adapted to low-computing edge devices such as smart speakers, sleep aids, and mobile phones; the asynchronous generation in the background and the fully automated algorithm process do not require additional computing power support on the edge, taking into account both device compatibility and user experience.

[0086] The specific embodiments described above are preferred embodiments of the present invention and are not intended to limit the specific scope of the present invention. The scope of the present invention includes, but is not limited to, these specific embodiments. All equivalent changes made in accordance with the present invention are within the protection scope of the present invention.

Claims

1. A low latency intelligent white noise generation and playback method, characterized by: include: A two-level interconnected audio and video storage architecture is constructed, which includes a first-level finished audio and video library and a second-level fragmented audio and video material library; The system performs a tiered search on user requests, prioritizing the primary finished audio and video library over the secondary fragmented audio and video material library. A dual-track computing power scheduling mechanism is adopted, which involves lightweight front-end retrieval response and asynchronous batch generation of audio and video content during back-end computing power off-peak periods. By dynamically expanding the two-level repository through asynchronous content generation in the background, and synchronously retrieving and playing audio and video, both low-latency playback and content diversity are achieved.

2. The low-latency intelligent white noise generation and playback method of claim 1, wherein: After receiving a user's playback request, the request is parsed into an acoustic-visual fusion feature vector, and a hierarchical retrieval and matching is performed based on the fusion feature vector.

3. The low-latency intelligent white noise generation and playback method of claim 2, wherein: The hierarchical retrieval uses a dual matching degree threshold to determine the result: if the matching degree of the first-level finished audio and video library meets the standard, it is played directly; if it does not meet the standard, the materials in the second-level fragmented audio and video material library are retrieved, quickly spliced ​​together, and then played.

4. The low-latency intelligent white noise generation and playback method of claim 1, wherein: The dual matching threshold is a dynamically adaptive threshold adjusted based on the user's historical search data.

5. The low-latency intelligent white noise generation and playback method of claim 1, wherein: The background asynchronous content generation and dynamic expansion includes: real-time collection of user playback behavior and feedback data to automatically identify gaps in audio and video content; automatic generation of multimodal generation prompts based on the content gaps; targeted generation of audio and video content based on the multimodal generation prompts; and updating the generated audio and video content in the database, forming a fully automated content evolution closed loop without human intervention.

6. The low-latency intelligent white noise generation and playback method of claim 5, wherein: The identification of audio and video content gaps is achieved through cluster analysis and association rule mining, which accurately identifies content gaps, low-satisfaction content, and long-tail demand gaps.

7. The low latency intelligent white noise generation and playback method of claim 5, wherein: The asynchronous content generation and dynamic expansion in the background also includes fully automated multi-dimensional quality inspection of the generated and spliced ​​audio and video. The quality inspection dimensions include acoustic indicators, visual indicators, scene emotional adaptability and audiovisual fusion consistency, and unqualified content is automatically removed.

8. A low-latency intelligent white noise generation and playback system for implementing the low-latency intelligent white noise generation and playback method of any one of claims 1-7, characterized by: include: A hierarchical storage module is used to construct a two-level linked audio and video storage architecture, which includes a first-level finished audio and video library and a second-level fragmented audio and video material library. The hierarchical retrieval module performs hierarchical retrieval on user playback requests in a fixed order, prioritizing the first-level finished audio and video library over the second-level fragmented audio and video material library. The dual-track computing power scheduling module performs dual-track computing power scheduling, which involves front-end lightweight retrieval response and back-end asynchronous batch generation of audio and video content during computing power off-peak periods. The multimodal generation module continuously generates and pre-assembles content asynchronously, supplementing high-quality content for multi-level libraries; The dual-track computing power scheduling module dynamically expands the two-level storage repository by asynchronously generating content in the background, and the hierarchical retrieval module combines with the front-end retrieval to achieve low-latency playback.

9. The low latency intelligent white noise generation and playback system of claim 8, wherein: It also includes an audio and video splicing module and a sleep scene optimization module. The audio and video splicing module is used to seamlessly splice secondary fragmented materials through acoustic feature matching and smooth fusion of transition frames. The sleep scene optimization module is used to optimize long audio playback without repetition and generate low-stimulation visual content specifically for sleep.

10. The low latency intelligent white noise generation and playback system of claim 9, wherein: It also includes a data closed-loop iteration module and an automated quality inspection and control module. The data closed-loop iteration module is used to build a fully algorithm-automated content evolution closed loop to achieve autonomous updates of the library. The automated quality inspection and control module is used to perform full-dimensional automated detection of audio and video content and remove unqualified content.