White noise seamless loop playing method and system based on audio frame pre-caching
Through the white noise classification model, multi-dimensional audio tags are identified and similar audio frames are pre-cachedated, the problem of interruption of white noise playback is solved, seamless loop playback is achieved, and fluency and system efficiency are improved.
Patent Information
- Application Number
- CN202510766373.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-10
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2045-06-10
AI Technical Summary
In the prior art, there is a sense of interruption when playing white noise during switching, which affects the user's sleep aid and hearing experience, and has a high resource occupancy rate and low system operation efficiency.
Multidimensional audio tags are identified through the white noise classification model, and the next audio frame with the same tag is pre-cachedated to achieve seamless loop playback and reduce resource usage.
Eliminates the interruption of white noise playback, optimizes playback fluency, improves system operation efficiency, and enhances user experience and system stability.
Smart Images

Figure CN120295601A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data processing, and in particular to a method and system for seamless loop playback of white noise based on audio frame pre-buffering. Background Art
[0002] White noise can help people relax, block other noise interferences in the surrounding environment, create a relatively quiet and uniform sound environment, contribute to improving sleep quality, and help those with difficulty falling asleep enter the sleep state faster.
[0003] In related technologies, since it is necessary to decode to a sleep aid device after playback to play, it involves a complex process of transmitting audio data from a playback device to a sleep aid device and then decoding. Factors such as unstable Bluetooth connection and Wi-Fi network congestion will slow down the transmission. The performance of the decoding chip and the complexity of the algorithm in the sleep aid device also affect the decoding speed, which breaks the coherence of audio playback. For users who rely on continuous white noise to aid sleep, the waiting time after the end of a piece of audio will disrupt the sleep aid atmosphere and affect sleep quality. Users who are sensitive to sound changes may even be awakened and have difficulty falling asleep again.
[0004] In addition, in related technologies, in terms of audio connection, there is a time gap between single-piece audio data, which seriously affects the auditory experience. White noise is supposed to block interferences and help users relax with a continuous and stable sound. However, interruptions will distract attention, reduce the sleep aid and relaxation effects, and easily cause irritable emotions, reducing the product satisfaction.
[0005] Therefore, there is an urgent need for a brand-new solution to improve the quality of credit assessment, assist in improving the efficiency of credit assessment, and reduce the risk of credit assessment. Summary of the Invention
[0006] In view of the technical problems existing in the prior art, the present invention provides a method and system for seamless loop playback of white noise based on audio frame pre-buffering, which is used to obtain multi-dimensional audio tags through a white noise classification model, pre-buffer the white noise data frames to be played based on the multi-dimensional audio tags, thereby eliminating the sense of interruption in white noise playback, optimizing playback fluency, reducing the cache resource occupancy rate at the same time, and improving the system operation efficiency.
[0007] In a first aspect, an embodiment of the present application provides a method for seamless loop playback of white noise based on audio frame pre-buffering, including: For a first white noise data frame currently being played, identify the multi-dimensional audio tags corresponding to the first white noise data frame through a white noise classification model; the multi-dimensional audio tags at least include: white noise type, sound intensity level, audio rhythm type, timbre type, audio scene; Based on the multi-dimensional audio tag, select a second white noise data frame to be played from a pre-established candidate audio library, and pre-cache the second white noise data frame; there is at least one identical audio tag between the second white noise data frame and the first white noise data frame; Set the playback mode of the second white noise data frame, and directly play the second white noise data frame after the first white noise data frame finishes playing.
[0008] In a second aspect, an embodiment of the present application provides a white noise seamless loop playback system based on audio frame pre-caching. The system includes the following units: An identification unit, configured to, for a first white noise data frame currently being played, identify the multi-dimensional audio tag corresponding to the first white noise data frame through a white noise classification model; the multi-dimensional audio tag at least includes: white noise type, sound intensity level, audio rhythm type, timbre type, audio scene; A selection unit, configured to select a second white noise data frame to be played from a pre-established candidate audio library based on the multi-dimensional audio tag; there is at least one identical audio tag between the second white noise data frame and the first white noise data frame; A playback unit, configured to set the playback mode of the second white noise data frame, and directly play the second white noise data frame after the first white noise data frame finishes playing.
[0009] In a third aspect, an embodiment of the present application provides an electronic device, and the electronic device includes: At least one processor, a memory, and an input / output unit; Wherein, the memory is used to store a computer program, and the processor is used to call the computer program stored in the memory to execute the white noise seamless loop playback method based on audio frame pre-caching in the first aspect.
[0010] In a fourth aspect, a computer-readable storage medium is provided, which includes instructions. When the instructions are run on a computer, the computer is caused to execute the white noise seamless loop playback method based on audio frame pre-caching in the first aspect.
[0011] The beneficial effects of the present invention are as follows: A white noise seamless loop playback method and system based on audio frame pre-caching are provided. In this technical solution, for the first white noise data frame currently being played, a multi-dimensional audio label corresponding to the first white noise data frame is identified through a white noise classification model; the multi-dimensional audio label at least includes: white noise type, sound intensity level, audio rhythm type, timbre type, audio scene. Furthermore, based on the multi-dimensional audio label, a second white noise data frame to be played is selected from a pre-established candidate audio library; there is at least one identical audio label between the second white noise data frame and the first white noise data frame. Finally, the playback mode of the second white noise data frame is set, and the second white noise data frame is directly played after the first white noise data frame is played.
[0012] In the embodiment of the present application, by directly switching to the second frame after the first frame is played and combining with the pre-caching technology, the sense of playback interruption is eliminated, and the playback fluency is optimized. In scenarios such as sleep aid, meditation, and focused learning, a smooth auditory experience is provided for users, and the user's immersion is maintained. By means of a white noise classification model to identify multi-dimensional audio labels, and based on the labels, the next frame with at least one identical label is selected from the candidate audio library, ensuring that the content styles are coherent and similar, enhancing the sense of scene reality and immersion, while meeting the diverse needs of users and improving satisfaction. And the pre-caching technology only caches relevant audio frames, reducing resource occupancy, reducing the memory and power consumption of mobile devices, improving the system operation efficiency, and the accurate matching and smooth playback reduce errors and exceptions, enhancing the system stability and ensuring the user experience. In short, through the white noise classification model, multi-dimensional audio labels are obtained, and based on the multi-dimensional audio labels, the white noise data frames to be played are pre-cached, thereby eliminating the sense of interruption in white noise playback, optimizing the playback fluency, while reducing the cache resource occupancy rate and improving the system operation efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0013] Figure 1 is a schematic flowchart of a white noise seamless loop playback method based on audio frame pre-caching according to an embodiment of the present application; Figure 2 is a schematic structural diagram of a white noise seamless loop playback system based on audio frame pre-caching according to an embodiment of the present application; Figure 3 is a schematic structural diagram of an electronic device according to an embodiment of the present application; Figure 4 is a schematic structural diagram of a medium device according to an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0014] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative efforts belong to the scope of protection of the present application.
[0015] The embodiment of the present application provides a method and system for seamless loop playback of white noise based on audio frame pre-buffering. In this technical solution, for the first white noise data frame currently being played, a multi-dimensional audio label corresponding to the first white noise data frame is identified through a white noise classification model; the multi-dimensional audio label at least includes: white noise type, sound intensity level, audio rhythm type, timbre type, audio scene. Furthermore, based on the multi-dimensional audio label, a second white noise data frame to be played is selected from a pre-established candidate audio library, and the second white noise data frame is pre-buffered; there is at least one same audio label between the second white noise data frame and the first white noise data frame. Finally, the playback mode of the second white noise data frame is set, and the second white noise data frame is directly played after the first white noise data frame is played.
[0016] In the embodiment of the present application, a multi-dimensional audio label is obtained through a white noise classification model, and the white noise data frame to be played is pre-buffered based on the multi-dimensional audio label, thereby eliminating the sense of interruption in white noise playback, optimizing playback smoothness, reducing the cache resource occupancy rate at the same time, and improving the system operation efficiency.
[0017] It should be particularly emphasized that the embodiment of the present application is mainly used for cached playback of the same song. For example, in the case of continuously playing the same song, the first white noise data frame and the second white noise data frame belong to the same song. In this case, a multi-dimensional audio label is obtained through a white noise classification model, and the white noise data frames belonging to the same song to be played are pre-buffered based on the multi-dimensional audio label, which can effectively avoid the interruption of white noise playback, eliminate the sense of interruption in white noise playback, and optimize playback smoothness.
[0018] The seamless loop playback solution of white noise based on audio frame pre-buffering provided by the embodiment of the present application can also be executed by an electronic device, which can be a server, a server cluster, or a cloud server. The electronic device can also be a terminal device such as a mobile phone, a computer, a tablet computer, a wearable device, or a dedicated device (such as a dedicated terminal device with a seamless loop playback system of white noise based on audio frame pre-buffering). These electronic devices can also be equipped with the chips introduced in the above embodiments. Or, these electronic devices can also install a service program for executing the seamless loop playback solution of white noise based on audio frame pre-buffering.
[0019] Figure 1 Schematic diagram of a white noise seamless loop playback method based on audio frame pre-buffering provided by an embodiment of the present application. As Figure 1 shown, the method includes the following steps: 101. For the first white noise data frame currently being played, identify the multi-dimensional audio tags corresponding to the first white noise data frame through a white noise classification model; 102. Based on the multi-dimensional audio tags, select the second white noise data frame to be played from a pre-established candidate audio library, and pre-buffer the second white noise data frame; 103. Set the playback mode of the second white noise data frame, and directly play the second white noise data frame after the first white noise data frame finishes playing.
[0020] In the embodiment of the present application, the first white noise data frame refers to the basic unit in the white noise audio data currently being played. The first white noise data frame is the starting point of the entire audio processing flow, carrying key information for subsequent analysis and matching. Through analysis by the white noise classification model, the corresponding multi-dimensional audio tags are identified. These tags include white noise types (such as ocean wave sound, rain sound, bird chirping sound, etc.), sound intensity levels (low, medium, high), audio rhythm types (linear change, periodic change, random change, etc.), timbre types (sharp, soft, etc.), and audio scenes (forest, seaside, city, etc.). Based on these tags, the system selects the second white noise data frame having at least one same audio tag from the pre-established candidate audio library, so as to achieve seamless switching playback after the first white noise data frame finishes playing. The feature recognition and analysis of the first white noise data frame play a crucial role in realizing the seamless loop playback of white noise, improving audio matching degree, and optimizing system performance.
[0021] It can be understood that for the case of continuously playing the same song, the first white noise data frame and the second white noise data frame belong to the same song. In this case, the first white noise data frame and the second white noise data frame correspond to the multi-dimensional audio tags of the same song.
[0022] In the embodiment of the present application, the audio tag is an identifier or metadata used to describe the characteristics and attributes of audio data. It can quantify and classify information in various dimensions of the audio, so as to facilitate the computer system or user to identify, manage, retrieve, and process the audio. In the above technical solution of white noise seamless loop playback based on audio frame pre-buffering, the audio tags specifically include the following categories: The white noise type is used to distinguish white noises generated by different natural or artificial environments, such as wind sound, rain sound, ocean wave sound, campfire burning sound, etc. It clarifies the basic sound source characteristics of the audio, helping the system and users quickly identify the main content of the audio.
[0023] The sound intensity level reflects the strength of the audio signal, which can generally be measured by indicators such as volume and sound pressure level, and is often expressed as low, medium, high or a specific decibel value range. This tag helps the system select the audio with an appropriate volume for playback according to user needs or scenario requirements, avoiding the situation where the sound is too loud or too soft and affecting the user experience.
[0024] The audio rhythm type describes the variation law of the audio signal over time, such as whether it has periodicity, the speed of the rhythm, and whether the rhythm change is stable or sudden. For example, the sound of a heartbeat may have a stable periodic rhythm, while the rhythm of the rustling of leaves may be relatively random. This tag enables the system to select audio frames with a coherent and coordinated rhythm, ensuring the smoothness and naturalness of the playback.
[0025] The timbre type reflects the unique sound quality characteristics of the audio, which are determined by factors such as the harmonic components and waveform envelopes of the audio, such as soft, sharp, bright, deep, etc. Different timbres can bring different auditory experiences to users. Through this tag, the system can select audio frames that match in timbre, enhancing the overall texture and coherence of the audio.
[0026] The audio scene is used to summarize the overall environment or situation simulated or represented by the audio, such as a forest, a seaside, an office, a bedroom, etc. It classifies the audio from a more macroscopic perspective, enabling the system to accurately select the appropriate audio according to different scenarios or needs of the user, creating an auditory atmosphere that is more in line with the actual needs for the user.
[0027] It can be understood that the classification system of audio tags can be adjusted according to different needs. From the perspective of application scenarios, for example, in the sleep assistance scenario, if it is mainly used for sleep assistance products, it may focus more on classifying audio tags according to relevant dimensions of the sleep assistance effect. In addition to the basic types mentioned above, tags such as "depth of sleep assistance" can be added. According to professional data such as the impact of audio on the human brain waves, the audio can be divided into different depth of sleep assistance levels; or a tag of "proportion of calming ingredients" can be set to analyze the proportion of elements such as gentle melodies and low-frequency rhythms in the audio that are helpful for calming the nerves.
[0028] In the music creation scenario, in the field of music creation, the classification of audio tags will pay more attention to music theory and creative elements. For example, a tag of "chord progression type" can be added to mark whether the chord progression contained in the audio is the classic Canon progression or other special chord connections; a tag of "melody development mode", such as whether it is a sequence, inversion, etc. in melody development techniques, which is convenient for creators to quickly locate the audio that conforms to the creative idea when looking for materials.
[0029] From the perspective of user groups, for professional audio engineers, the classification requirements for audio tags will be more professional and detailed. It may be necessary to add tags related to audio technical parameters such as "audio sampling rate" and "bit depth" to facilitate their accurate screening of audio materials that meet technical requirements when performing audio editing, mixing, etc. There may also be a need for a "dynamic range" tag to measure the difference between the strongest and weakest sounds in an audio signal, which is very important for post-processing and effect adjustment of audio.
[0030] For ordinary users, they may be more concerned about the emotional attributes of audio and the convenience of usage scenarios. Tags such as "mood matching degree" can be added to classify audio into different emotional categories such as soothing, pleasant, and exciting; a "suitable activity" tag, such as sports, party, study, etc., to facilitate users to quickly find suitable audio according to the activity they are engaged in.
[0031] From the perspective of technological development, deep learning algorithms can be used to perform more complex feature extraction and classification on audio, thus adding new types of tags. For example, "audio semantic tags", by understanding the semantics of audio content, label specific events or scenarios contained in the audio, such as "crowd cheering" and "car driving"; "emotional tendency tags", using sentiment analysis technology to more accurately judge the positive, negative, or neutral emotional tendency conveyed by the audio.
[0032] New audio formats and coding methods are constantly emerging, and it may be necessary to add tags related to these new technologies. For example, for some audio files that support spatial audio, a "spatial audio mode" tag needs to be added to label specific spatial audio types such as Dolby Atmos and surround sound; an "audio coding type" tag to distinguish whether it is MP3, FLAC, or other new coding formats, to facilitate devices and software to correctly decode and play the audio.
[0033] It is worth noting that traditional audio classification methods may only simply divide audio into large categories such as music and speech, which are difficult to meet the more in-depth and detailed management requirements for audio content. Multidimensional audio tags can establish a more accurate and comprehensive classification system for audio by covering multiple dimensions such as white noise type, sound intensity level, audio rhythm type, timbre type, and audio scene, enabling each audio to find an accurate position in a complex audio library. When users need to search for specific audio, multidimensional audio tags allow users to screen and retrieve from multiple dimensions according to their specific needs. For example, if a user wants to find a natural environment white noise audio for sleep aid with a low sound intensity, through the accurate labeling of multidimensional audio tags, the system can quickly and accurately match the audio resources that meet the requirements for the user, greatly improving the retrieval efficiency and accuracy.
[0034] By analyzing users' listening history and behavioral data and combining multi-dimensional audio tags, audio platforms can gain an in-depth understanding of users' preferences in terms of white noise type, sound intensity, rhythm, timbre, scene, etc. For example, if a user often listens to white noise such as rain sounds and prefers audio with a low sound intensity and a uniform rhythm, the platform can accurately recommend similar audio content based on these tag features.
[0035] Different users have different audio needs in different scenarios. Multi-dimensional audio tags can help the platform better meet these diverse needs and provide personalized audio recommendations for users. For example, for users who like to listen to audio with a fast rhythm and a moderate sound intensity during exercise, the platform can recommend suitable music or audio for exercise based on tags such as rhythm type and sound intensity level.
[0036] In different life scenarios, users' audio needs vary. The audio scene tags in multi-dimensional audio tags can help users quickly find audio suitable for the current scene. For example, when a user is in a meditation and relaxation scene, by selecting the audio scene tag related to "meditation", they can quickly obtain a series of audio suitable for meditation, such as natural wind sounds and running water sounds, creating a more scene-appropriate atmosphere and enhancing the user's experience in this scene.
[0037] Users can customize their own auditory experiences using multi-dimensional audio tags according to their preferences in dimensions such as sound intensity and timbre. For example, users who like soft timbres can filter out audio with soft timbres and adjust the sound intensity to their comfortable range to obtain a more satisfactory auditory enjoyment.
[0038] For audio creators, multi-dimensional audio tags can serve as a reference basis for creation. By analyzing the tag features of different types of audio, creators can understand the characteristics of popular audio in the market and draw on the design ideas of excellent works in terms of white noise type, rhythm, timbre, etc. to create audio works that better meet users' needs. For example, if a creator wants to create a sleep aid audio, they can refer to the common white noise types, appropriate sound intensities, and rhythms in existing sleep aid audios and conduct targeted creation.
[0039] In the field of audio research, multi-dimensional audio tags provide rich data dimensions for researchers. Researchers can study the feature distributions of audio in different dimensions and the preference trends of users for different audio features by analyzing a large amount of audio tag data. For example, by analyzing the preference data of users of different age groups for timbre types, they can understand the relationship between audio perception and age, providing strong support for research in fields such as audio psychology and acoustics.
[0040] As an alternative embodiment, in 101, identifying the multi-dimensional audio tag corresponding to the first white noise data frame through the white noise classification model can be implemented as the following steps: Extract the first white noise data frame feature corresponding to the first white noise data frame; Obtain the candidate audio classification information corresponding to the first white noise data frame feature through a multi-branch perception layer; each branch perception layer in the multi-branch perception layer is used to extract the corresponding key feature information in the first white noise data frame feature, and obtain the corresponding candidate audio classification information based on the extracted key feature information; each branch perception layer is set based on the corresponding classification logic; Convert the candidate audio classification information corresponding to the first white noise data frame feature into the corresponding multi-dimensional audio tag.
[0041] In the embodiment of the present application, there is at least one same audio tag between the second white noise data frame and the first white noise data frame. That is, when performing audio analysis and annotation on these two white noise data frames, the tag contents in one or more dimensions of the multi-dimensional audio tags are the same.
[0042] It should be noted that the embodiment of the present application mainly aims at the situation of continuously playing the same song. In this case, the first white noise data frame and the second white noise data frame correspond to the multi-dimensional audio tags of the same song.
[0043] For example, the multi-dimensional audio tag includes multiple dimensions such as white noise type, sound intensity level, audio rhythm type, timbre type, audio scene, etc. For example, the multi-dimensional audio tag of the first white noise data frame is {white noise type: rain sound, sound intensity level: medium intensity, audio rhythm type: uniform type, timbre type: soft type, audio scene: indoor}, and the multi-dimensional audio tag of the second white noise data frame is {white noise type: rain sound, sound intensity level: high intensity, audio rhythm type: variable speed type, timbre type: crisp type, audio scene: outdoor}. In this case, they have the same audio tag "rain sound" in the dimension of "white noise type".
[0044] In this way, these two white noise data frames have similar audio characteristics in a certain aspect. For example, they are the same in terms of the type of white noise, which means they have the same categorical attribution in the essential attributes of the sound, and they may both be the same sound phenomenon in the simulated natural environment. In audio processing and analysis, the same audio label can be used as a basis for judging whether two audio data frames are related or similar. For example, in an audio retrieval system, if a user searches for white noise of "rain sound", then the first and second white noise data frames with the same audio label of "rain sound" may both be retrieved as relevant results. For audio data frames that are uncertain or difficult to accurately label, by comparing them with other data frames with known audio labels, if at least one same audio label is found, the other label information of the known data frames can be used to assist in the analysis and improve the annotation of the current data frame. For example, if the audio scene annotation of the second white noise data frame is more accurate, while the audio scene annotation of the first white noise data frame is not very certain, but they have the same white noise type label, then the audio scene of the second white noise data frame can be referred to for further judgment of the audio scene of the first white noise data frame.
[0045] In practical applications, when constructing an audio database, by identifying the same audio labels between different audio data frames, the audio data can be classified, stored, and managed, which is convenient for subsequent query and invocation. According to the user's preference for audio data frames with certain audio labels, other audio data frames with the same or similar audio labels can be recommended. For example, if the user likes the "rain sound" white noise type of the first white noise data frame, the system can recommend the second white noise data frame with the label of "rain sound" or other similar audio data. When developing and evaluating audio processing algorithms, different data frames with the same audio label are used to test the accuracy and stability of the algorithms. If the algorithm can obtain accurate processing results for different data frames with the same audio label, it indicates that the algorithm has good generalization ability.
[0046] Specifically, in the above optional embodiment of 101, for the first white noise data frame, various signal processing and feature extraction methods can be adopted, such as the short-time Fourier transform (STFT), which converts the white noise data in the time domain to the frequency domain to obtain its spectral features, including frequency distribution, energy distribution, etc. The Mel-frequency cepstral coefficients (MFCC) can also be calculated, which can simulate the human ear's perception characteristics of sound frequencies and extract features that are more in line with human auditory perception. At the same time, time-domain features such as zero-crossing rate and energy entropy can also be extracted to comprehensively describe the characteristics of the first white noise data frame.
[0047] Set up a multi-branch perception layer according to the key feature information to be extracted and the classification logic. For example, if you want to extract white noise type features, you can set up a branch perception layer, and its classification logic can be based on the spectral feature differences of different white noises. For example, the spectrum of ocean wave sound has a specific energy distribution in certain frequency bands, which is different from other white noises such as rain sound. For the sound intensity level, another branch perception layer can be set up, and its classification logic is based on the energy size of the audio to divide different intensity levels.
[0048] For the audio rhythm type, the branch perception layer can determine the rhythm type by analyzing features such as the periodicity and beats of the audio signal. For example, a uniform rhythm type has relatively stable periodic features, while a gradually changing rhythm type has the characteristic of gradually changing period. For the timbre type, the branch perception layer can judge based on features such as the harmonic structure and formants of the spectrum. For example, a soft timbre type has relatively fewer high-frequency components and a smoother harmonic structure. For the audio scene, the classification logic of the branch perception layer can be set according to the specific environmental characteristic sounds contained in the audio. For example, an indoor scene may contain some sound characteristics of electrical appliances running, while an outdoor forest scene has rich bird calls, insect sounds, etc.
[0049] After each branch perception layer receives the first white noise data frame features, it processes them according to their respective classification logics. Taking the white noise type branch as an example, it will focus on analyzing spectral features, etc., and compare and match them with the feature templates of various established white noise types to determine which white noise type the first white noise data frame most likely belongs to, and obtain the corresponding candidate audio classification information. For example, it is judged as rain sound type white noise.
[0050] The sound intensity level branch perception layer calculates the energy value of the audio, compares it with the preset energy thresholds of different intensity levels, and determines the candidate information of its sound intensity level, such as low intensity, medium intensity or high intensity. The audio rhythm type branch perception layer determines the candidate information of its rhythm type, such as uniform type, gradually changing type or random type, etc., through the periodic analysis of the audio signal. The timbre type branch perception layer judges the candidate information of the timbre, such as soft type, sharp type or mellow type, etc., according to features such as the harmonics of the spectrum. The audio scene branch perception layer obtains the candidate information of whether it is an indoor scene, an outdoor scene or a special scene based on the environmental characteristic sounds.
[0051] Integrate the candidate audio classification information obtained by each branch perception layer and convert it into a multi-dimensional audio label. For example, if the white noise type branch judges it as ocean wave sound, the sound intensity level branch determines it as medium intensity, the audio rhythm type branch obtains uniform type, the timbre type branch judges it as soft type, and the audio scene branch determines it as outdoor scene, then the finally converted multi-dimensional audio label can be expressed as: {white noise type: ocean wave sound, sound intensity level: medium intensity, audio rhythm type: uniform type, timbre type: soft type, audio scene: outdoor scene}.
[0052] Extract the features of the second white noise data frame. Using the same feature extraction method as the first white noise data frame, obtain the features of the second white noise data frame. Then input the features of the second white noise data frame into the same multi-branch perception layer to obtain its candidate audio classification information.
[0053] Compare the candidate audio classification information obtained from the first white noise data frame and the second white noise data frame in each branch perception layer. Since they have at least one same audio label, the classification results on some branches should be consistent. For example, if it is known that they are the same in terms of white noise type, the results obtained from the white noise type branch perception layer should be the same type judgment such as the sound of ocean waves.
[0054] Based on the setting of the same audio label, the multi-dimensional audio labels of the first white noise data frame can be optimized and calibrated. If the classification result of the second white noise data frame on a certain branch is more accurate or reliable (for example, because the audio quality of the second white noise data frame is better, etc.), consider using the result of the second white noise data frame on this branch to correct the candidate audio classification information of the corresponding branch of the first white noise data frame.
[0055] For example, if the second white noise data frame is determined to be "weak medium intensity" in terms of sound intensity level through more precise analysis, while the first white noise data frame was originally judged as simply "medium intensity", then according to the relationship of their same audio labels, the candidate information of the sound intensity level of the first white noise data frame can be optimized to "weak medium intensity", so that the multi-dimensional audio labels of the first white noise data frame are more accurate and perfect.
[0056] Further optionally, in 101, obtaining the candidate audio classification information corresponding to the features of the first white noise data frame through the multi-branch perception layer includes: Input the features of the first white noise data frame into the classification target layer to identify the white noise type in the features of the first white noise data frame; Based on the identified white noise type, adaptively configure the types of branch perception layers involved in the corresponding multi-branch perception layer; Call the corresponding type of multi-branch perception layer to identify the audio change features corresponding to different frequency bands, different time sequence intervals, and different change waveforms in the features of the first white noise data frame, so as to obtain the candidate audio classification information corresponding to the features of the first white noise data frame.
[0057] Specifically, the white noise type is a key feature of audio. Different types of white noise, such as the sound of wind, rain, ocean waves, etc., have unique acoustic characteristics. Accurately identifying the white noise type provides a basic direction for subsequent audio processing and analysis because different types of white noise may also exhibit different patterns in other feature dimensions.
[0058] The classification target layer is usually constructed based on machine learning or deep learning models. For example, a convolutional neural network (CNN) can be used. Its convolutional layers can automatically extract local patterns in audio features, and the pooling layers reduce the dimensionality of the features while retaining key information. By training on a large amount of audio data with known white noise types, the model can learn the feature patterns of different types of white noise. Inputting the features of the first white noise data frame into the trained model, the model outputs the prediction result of the white noise type for this data frame. Taking a simple CNN model as an example, assume that the features of the first white noise data frame input are the preprocessed spectrogram. The CNN model extracts features such as the frequency distribution and energy concentration regions in the spectrogram through convolutional operations. After multiple layers of convolution and pooling, the fully connected layer maps these features to different white noise type categories, and the softmax function outputs the probabilities of each category (such as the sound of wind, rain, etc.). The category with the highest probability is the identified white noise type.
[0059] Different types of white noise have different significant features in different frequency bands, time series intervals, and changing waveforms. For example, rain sounds may have more detailed information in the high-frequency band because the high-frequency sounds generated by raindrops hitting different object surfaces are relatively rich; while ocean wave sounds may have stronger energy in the low-frequency band, manifested as low-frequency roars. Therefore, adaptively adjusting the branch types of the multi-branch perception layer according to the identified white noise type can extract key features more targeted.
[0060] If the identified white noise type is rain sound, due to the wide frequency range and rich high-frequency components of rain sound, more branch perception layers that focus on the high-frequency band may be configured. These branch perception layers can use specific filters or convolutional kernels to highlight high-frequency features. At the same time, considering that rain sounds may be intermittent in time series, branch perception layers that specifically analyze the time series interval features will be configured to capture information such as the intervals and rhythms of raindrops falling. For the changing waveform, the sound waveform of raindrops falling may have a certain periodicity or randomness, and the corresponding configured branch perception layer can analyze this waveform change feature. If the identified sound is ocean wave sound, due to its low-frequency characteristics, branch perception layers that are sensitive to the low-frequency band will be emphasized, as well as time series analysis branch perception layers that can capture the periodic changes of ocean wave sounds on a long time scale (such as the undulation rhythm of ocean waves).
[0061] When identifying the characteristics of different frequency bands, each branch perception layer analyzes a specific frequency band. For example, the high-frequency branch perception layer focuses on the audio change characteristics in the high-frequency band through a band-pass filter or a specific convolution kernel. It calculates features such as the energy distribution and frequency peaks within the high-frequency band. Assuming that the audio data is transformed into the frequency domain through the fast Fourier transform (FFT), the high-frequency branch perception layer performs a convolution operation on the spectrum of the high-frequency part to extract the high-frequency feature patterns. These feature patterns can reflect the detailed information of white noise in the high-frequency band, such as the high-frequency sound characteristics generated by raindrops splashing. Through the analysis and learning of these features, the branch perception layer can determine which type the characteristics of the white noise in the high-frequency band belong to. For example, if the high-frequency components are rich and sharp, it may correspond to the sound characteristics of raindrops hitting a hard object surface.
[0062] When identifying the characteristics of different time-sequence intervals, the time-sequence analysis branch perception layer focuses on the changes in audio within different time intervals. It can divide the audio data into multiple intervals in chronological order and analyze the audio characteristics within each interval, such as energy changes and frequency changes. For example, for the sound of rain, analyze the audio energy changes within the interval between each raindrop falling and the change in the frequency of raindrops falling over a period of time. Through the extraction and analysis of these time-sequence characteristics, it is possible to determine the rhythm type of the audio, such as whether it is a uniform raindrop rhythm or a rhythm with fast and slow changes.
[0063] When identifying the characteristics of different changing waveforms, the branch perception layer responsible for analyzing the changing waveforms can use time-frequency analysis methods such as wavelet transform to convert the audio signal into a time-frequency diagram and observe the changes in the waveform at different times and frequencies. For the sound of ocean waves, its waveform may show periodic undulations. This branch perception layer identifies the characteristic patterns of the ocean wave sound by analyzing features such as the period and amplitude changes of the waveform. For example, calculate the autocorrelation function of the waveform to determine its periodicity, and analyze the peaks and valleys of the waveform to judge the intensity changes of the ocean waves.
[0064] Each branch perception layer integrates the characteristics of different frequency bands, time-sequence intervals, and changing waveforms extracted to form candidate audio classification information. These information comprehensively reflect the characteristics of white noise in multiple dimensions. For example, the characteristics in the high-frequency band indicate the sharpness of the sound, the time-sequence characteristics reflect the regularity of the rhythm, and the changing waveform characteristics reflect the dynamic change pattern of the sound. These information can be further used to determine other multi-dimensional audio labels such as the sound intensity level, timbre type, and audio scene of the white noise, or directly used as a classification basis to provide a reference for selecting the second white noise data frame from the candidate audio library later.
[0065] Further optionally, after extracting the first white noise data frame features corresponding to the first white noise data frame in 101, the interference noise feature matrix in the first white noise data frame can also be extracted; the interference noise elements in the interference noise feature matrix are expressed by the following formula: ; wherein, represents the first white noise data frame feature corresponding interference noise element at the t-th time step, and is the dynamic expansion rate in the interference noise feature matrix, and is the dynamic step size in the interference noise feature matrix, and is the interference noise analysis dimension index, and are the first white noise data frame features output coordinate values, and is in the interference noise element feature components in the dimension. Furthermore, in 101, the interference noise feature matrix is fused with the first white noise data frame features, and the fusion result is input into the multi-branch perception layer to increase the robustness of the first white noise data frame feature processing process.
[0066] In the above steps, fusing the interference noise feature matrix with the first white noise data frame features is to enable the model to simultaneously consider the interference noise information existing in the white noise data frame features when processing the white noise data frame features. In actual audio data, interference noise is inevitable. By explicitly extracting the interference noise features and fusing them with the white noise features, the model can better learn the true features of the white noise and avoid being misled by the interference noise. In this way, in the face of various complex audio environments and different degrees of interference, the model can more accurately identify and process the white noise data, improving the stability and accuracy of the model, that is, increasing the robustness of the processing process.
[0067] Exemplarily, taking the identification of white noise in rain sounds as an example, if there is interference noise such as wind noise, without considering the interference noise feature matrix, the model may misjudge some features of the wind noise as the features of the rain sound white noise, resulting in misidentification. By fusing the interference noise feature matrix, the model can accurately distinguish the wind noise as interference noise and more accurately extract the features of the rain sound white noise, thereby improving the accuracy of white noise type identification.
[0068] In different audio environments, the intensity and characteristics of interfering noise may change. For example, in indoor and outdoor environments, the interfering noise received by white noise is different. By fusing the interfering noise feature matrix, the model can better adapt to these changes, and will not cause a large deviation in the extraction of white noise features due to the change of interfering noise, enabling the model to maintain relatively stable performance in various audio environments and enhancing the stability of the model.
[0069] After the model is trained by fusing the interfering noise feature matrix, it can also perform better in the extraction and processing of white noise features in unseen new audio data. For example, only some types of interfering noise situations are included in the training data, and a new type of interfering noise is encountered in actual applications. Since the model has learned how to handle the relationship between interfering noise and white noise, it can more effectively extract white noise features from new data, improving the generalization ability of the model and enabling it to better apply to audio processing tasks in various actual scenarios.
[0070] Further optionally, after extracting the first white noise data frame feature corresponding to the first white noise data frame in 101, the real-time delay information, real-time network fluctuation information, and real-time network environment state corresponding to the playback of the first white noise data frame can also be obtained. Furthermore, based on the real-time delay information, real-time network fluctuation information, and real-time network environment state, the predicted network environment information when the first white noise data frame is played to a preset period is predicted; where the preset period at least includes: a preset duration before the end of the first white noise data frame. Then, according to the predicted network environment information, the adaptability between the predicted network environment information and the memory resources corresponding to the pre-loaded area in the playback device is evaluated. Finally, based on the adaptability, the network transition type corresponding to the first white noise data frame is determined; the network transition type is used to indicate the network pre-loading state before transitioning to the second white noise data frame.
[0071] Exemplarily, assume that an online white noise playback application is being used, and the first white noise data frame currently being played is an audio segment simulating a forest environment. Through the network monitoring module within the application, the real-time delay information during the playback of the current first white noise data frame is obtained as 50 milliseconds. The real-time network fluctuation information shows that the network bandwidth fluctuates within the range of 1 - 1.2 Mbps in the past 10 seconds. The real-time network environment status indicates that the current is in a 4G network environment with a signal strength of 3 bars. Using a prediction model based on historical data and machine learning algorithms, the predicted network environment information 10 seconds (preset duration) before the end of the first white noise data frame is predicted according to the above real-time information. For example, the prediction result shows that 10 seconds before the end of the first white noise data frame, the network delay will increase to 80 milliseconds, the network bandwidth may drop to 0.8 Mbps, the network environment remains a 4G network, but the signal strength may drop to 2 bars.
[0072] Furthermore, the memory resources in the preloading area of the playback device are set to be able to cache the audio data for 10 seconds at a bandwidth of 1 Mbps to ensure smooth playback. According to the predicted network environment information, the amount of audio data for 10 seconds at a bandwidth of 0.8 Mbps is less than the capacity that the preloading area memory resources can accommodate, and after calculation, the adaptability is relatively high. Assume that the adaptability evaluation is 80% (full score is 100%).
[0073] Finally, according to the adaptability of 80%, the network transition type is determined as "smooth transition", which means that before transitioning to the second white noise data frame, the network preloading state is relatively stable, and there are sufficient memory resources to preload the second white noise data frame to ensure the continuity of playback.
[0074] As can be seen from the above example, the real-time delay information can be obtained by measuring the time interval for audio data to be transmitted from the server to the playback device; the real-time network fluctuation information can be obtained by monitoring the change of network bandwidth within a certain time window; the real-time network environment status can be determined by the network connection status detection module of the device, such as judging the network types such as Wi-Fi, 4G, 5G, etc. and the corresponding signal strength.
[0075] Here, the predicted network environment information can adopt time series analysis or machine learning algorithms. For example, a prediction model is constructed using historical network data (including network delay, bandwidth, network type, and signal strength at different time periods). Common algorithms include the autoregressive integrated moving average model (ARIMA) for predicting delays and bandwidth in time series. For the network environment status (such as network type and signal strength), rule-based prediction or simple machine learning classification algorithms can be used to predict according to the current and recent network environment change trends. These models estimate the network environment for a preset future period by learning the patterns in historical data.
[0076] Furthermore, the memory resources corresponding to the preloading area have certain capacity limitations, which are related to the transmission rate and duration of the audio data. According to the predicted network environment information, the amount of audio data to be transmitted within a preset time period can be calculated. For example, given the predicted network bandwidth, the amount of data that can be transmitted within a preset duration can be calculated and compared with the amount of data that can be accommodated by the memory resources in the preloading area to obtain the adaptation degree. The adaptation degree evaluation can use simple ratio calculation (such as predicted transmission data volume / preloading area capacity), or consider more complex factors such as data transmission stability and error range for weighted calculation.
[0077] In this way, different thresholds can be set according to the adaptation degree to determine the network transition type. For example, when the adaptation degree is higher than 70%, the network transition type is considered "smooth transition", indicating that the network preloading state is good and the second white noise data frame can be preloaded smoothly; when the adaptation degree is between 40% and 70%, it may be determined as "transition requiring optimization", meaning that some optimization measures need to be taken, such as adjusting the preloading strategy and reducing the audio quality; when the adaptation degree is lower than 40%, it is determined as "unstable transition", and it may be necessary to pause the playback or prompt the user that the network condition is poor.
[0078] The above steps can, by predicting the network environment information in advance and evaluating its adaptation degree with the memory resources, know in advance whether the network preloading state is good when switching to the second white noise data frame during the playback of the first white noise data frame. If it is predicted that the network environment may deteriorate but is still within the optimizable range, the preloading strategy can be adjusted in advance, such as starting to preload the second white noise data frame in advance, or caching more data when the network condition is good, so as to reduce the stuttering or buffering phenomenon caused by network problems during the switch and improve the playback fluency. It avoids the sudden playback interruption or long-time buffering caused by network problems during the playback for the user, making the entire white noise playback process more coherent and comfortable. Especially for some users who need to play white noise for a long time to assist in sleeping, meditation or focused work, this smooth playback experience can effectively reduce external interference and improve the user's satisfaction with the application. According to the predicted network environment information and adaptation degree, the memory resources of the playback device are reasonably allocated and utilized. It avoids the waste of memory resources caused by excessive preloading when the network environment is poor, or the insufficient preloading when the network environment is good, which affects the playback fluency, realizes the effective coordination of memory resources and network resources, and improves the resource utilization efficiency.
[0079] As an optional embodiment, in 102, selecting the second white noise data frame to be played from the pre-established candidate audio library based on the multi-dimensional audio label includes: Construct a virtual playback scenario corresponding to the first white noise data frame based on the multi-dimensional audio tags; select candidate white noise data frames with at least one identical node from the candidate audio library according to the nodes corresponding to the respective audio tags in the virtual playback scenario; the arrangement of the nodes corresponding to the respective audio tags in the virtual playback scenario is determined based on the attributes of the white noise data frame; determine the second white noise data frame from the candidate white noise data frames based on user preference data.
[0080] Exemplarily, assume that the multi-dimensional audio tags of the first white noise data frame are: {white noise type: ocean wave sound, sound intensity level: medium intensity, audio rhythm type: uniform speed, timbre type: soft, audio scenario: seaside}. Construct a virtual playback scenario based on these audio tags. For example, with the "seaside" audio scenario as the core, the "ocean wave sound" white noise type unfolds around this scenario, and the "medium intensity" sound intensity level, "uniform speed" audio rhythm type, and "soft" timbre type serve as the attribute nodes of the ocean wave sound in this scenario. One can imagine a virtual seaside scenario where the ocean waves rise and fall with a medium intensity, uniform speed, and soft sound.
[0081] The candidate audio library stores a large number of white noise data frames with various audio tags. According to the nodes corresponding to the respective audio tags in the virtual playback scenario, search for candidate white noise data frames with at least one identical node in the candidate audio library. For example, there is a data frame in the candidate audio library with the label {white noise type: ocean wave sound, sound intensity level: high intensity, audio rhythm type: variable speed, timbre type: crisp, audio scenario: seaside in a storm}. Since its white noise type is the same as that of the first white noise data frame (both are ocean wave sounds), this data frame is selected as a candidate white noise data frame. Assume there are other data frames, such as one with the label {white noise type: wind sound, sound intensity level: medium intensity, audio rhythm type: uniform speed, timbre type: sharp, audio scenario: mountaintop}. Because its sound intensity level and audio rhythm type have the same nodes as the first white noise data frame, it is also selected as a candidate white noise data frame.
[0082] Assume that through user preference data, it is known that the user prefers white noise with a moderate sound intensity and a stable rhythm. Among the above candidate white noise data frames, the second data frame (sound intensity level: medium intensity, audio rhythm type: uniform speed) better meets the user's preference, so it is determined as the second white noise data frame.
[0083] As can be seen from the above example, during the construction of the virtual playback scenario, the multi-dimensional audio tags provide rich audio feature description information. The construction of the virtual playback scenario by combining these tags is based on the correlation between audio tags. For example, a specific audio scenario is usually closely related to a specific type of white noise, sound intensity, etc. Taking the "seaside" scenario as an example, it will naturally be associated with the white noise type of ocean waves, and in this scenario, there will also be corresponding common attributes for sound intensity, rhythm, timbre, etc. This construction method transforms the abstract audio tags into a more intuitive and relevant virtual scenario, facilitating the subsequent screening of data frames based on the scenario.
[0084] Furthermore, each white noise data frame in the candidate audio library carries its own audio tags. By comparing the nodes corresponding to the audio tags in the virtual playback scenario with the audio tags of the data frames in the candidate audio library, the data frames with at least one identical node are selected as candidates. This is based on the assumption that data frames with the same audio tag nodes are similar in some aspects. Even if other tags are different, they are at least consistent in one key feature, which provides a relatively large range of options for subsequent selection of the appropriate second white noise data frame.
[0085] The user preference data reflects the user's preference degree for different audio tags. Among the numerous candidate white noise data frames, by analyzing the user preference data, the data frame that best matches the user's preference is selected as the second white noise data frame. For example, if a user often selects white noise with medium-intensity sound, then among the candidate data frames, the data frames with the label of medium-intensity sound are more likely to be selected. This can ensure that the selected second white noise data frame better meets the user's personalized needs.
[0086] Therefore, by constructing a virtual playback scenario and selecting candidate white noise data frames based on the audio tag nodes in the scenario, the feature similarity between audio data frames can be considered more comprehensively. Compared with simply selecting based on individual audio tags, this method can perform comprehensive screening from multiple dimensions, improving the matching degree of the selected white noise data frame with the first white noise data frame in terms of overall features, thereby providing a more coherent and expected audio experience for users. Determining the second white noise data frame from the candidate white noise data frames based on the user preference data fully considers the user's personalized needs. Different users have different preferences for sound intensity, rhythm, timbre, etc. Through this method, the user's personalized choices can be satisfied, enhancing the user's satisfaction and loyalty to the audio playback service. The above steps can make full use of the data resources in the candidate audio library. Through the screening mechanism based on the virtual playback scenario and user preferences, each data frame in the candidate audio library has the opportunity to be selected according to its audio tag characteristics, avoiding the situation that some data frames are left idle for a long time due to a single-dimensional screening criterion, and improving the utilization rate of the entire audio library resources.
[0087] Further optionally, the optimization of the virtual playback scene can be performed from multiple dimensions such as rich scene details, enhanced user interaction, and intelligent dynamic adjustment. In addition to the existing white noise type, sound intensity level and other tags, more dimensional audio tags such as audio spatial location information (from which direction the sound comes from the left front, front, etc.), audio reverberation effect (simulating reverberation in different spaces such as caves and rooms) can also be introduced to make the virtual playback scene more three-dimensional and realistic. It can be considered to combine with visual elements to match the corresponding static pictures or dynamic video images for the virtual playback scene. For example, when playing the white noise of the sound of the waves, a video of the scenery at the seaside is displayed at the same time. It can also be further combined with olfactory elements, and aromatherapy equipment can be connected through smart devices to release smells that match the scene, such as releasing a faint smell of the sea in the seaside scene, which can enhance the user's immersion in all aspects. Provide users with more custom permissions, so that users can freely combine audio tags according to their preferences to build a unique virtual playback scene. For example, users can set the intensity and rhythm of the sound of the waves, and add other sound effects such as the cry of seagulls to create their own virtual seaside scene. Develop real-time interactive functions to enable users to interact with virtual playback scenes in real time. For example, users can use voice commands to make the sound of waves in the virtual scene louder or softer, or switch the sound effects in the scene from waves to rain.
[0088] Optionally, a machine learning algorithm is used to analyze the behavioral data of users using virtual playback scenes, such as the scene types and dwell time selected by users at different times and in different environments, and automatically adjust the parameters of the virtual playback scenes. For example, if it is found that users are more inclined to choose white noise scenes with lower sound intensity at night, the system can automatically reduce the sound intensity of the scene at night. Combined with sensor data from smart devices, such as light sensors and temperature sensors, the virtual playback scene can be adaptively adjusted according to the actual environment in which the user is located. For example, when the light becomes brighter, the audio rhythm in the virtual playback scene can be appropriately accelerated to create a more vibrant atmosphere; when the temperature drops, the sound effects can appropriately add some elements such as wind sounds to make the scene more in line with the actual feeling.
[0089] When switching between different virtual playback scenarios, smooth transition technology is adopted to avoid discomfort to users caused by sudden changes in elements such as audio and vision. For example, when switching from a seaside scene to a forest scene, the sound of the waves can gradually fade, while the sound effects such as the chirping of birds in the forest gradually increase to achieve a natural transition. Analyze the relevance between different virtual playback scenarios, and optimize the order and method of scene switching according to users' usage habits and scene logic. For example, if a user often switches from a seaside scene to a beach bonfire scene, the system can make the switching between these two scenes more convenient and natural, and even add some transition sound effects and pictures during the switching process, such as the sound of the waves gradually changing to the sound of the bonfire burning, and the seaside picture gradually transitioning to the bonfire picture.
[0090] In 102, pre-cache the second white noise data frame. In 103, set the playback mode of the second white noise data frame.
[0091] Specifically, pre-caching is to fetch and store the upcoming second white noise data frame from a storage device or network into a cache space such as memory during the playback of the current first white noise data frame. In this way, when the second white noise data frame needs to be played, the data can be quickly read directly from the cache, without having to fetch it temporarily from a slower storage medium or network, thus greatly reducing the waiting time during playback and avoiding stuttering.
[0092] Specifically, in 102, the timing and amount of pre-caching can be determined according to pre-set playback rules or playback configuration methods. For example, based on factors such as the playback progress of the first white noise data frame, network conditions, and device performance, the appropriate pre-caching timing and pre-caching amount can be calculated in advance. Generally, when the first white noise data frame is played to a certain proportion (such as 80%), the second white noise data frame will start to be pre-cached to ensure that the second frame is ready to be played immediately when the first frame finishes playing.
[0093] For example, when continuously playing the same song, pre-caching the second white noise data frame is mainly to ensure smooth playback and avoid stuttering. Further optionally, a buffer (such as a queue) can be used to store the pre-cached data frames. In this way, the data frames can be managed according to the first-in, first-out (FIFO) principle to ensure that the data frames are played in order. When playing the first white noise data frame, calculate the position of the next second white noise data frame to be cached based on the total number of frames of the song and the current playback position. For example, dynamically configure the length of the data frames to be cached in the future period based on the network environment and device cache situation. For example, a fixed-length caching strategy can be adopted, such as caching the next 10 data frames.
[0094] It should be noted that in order not to affect the playback of the current data frame, the second white noise data frame can be read from the candidate audio library in an asynchronous loading manner and added to the buffer.
[0095] For example, assume that the user is using a white noise playback application to listen to the rain white noise (the first white noise data frame). According to the preset rules, when 20% of the rain sound is left to play, the application starts to obtain the next bird song white noise (the second white noise data frame) from the server in the background through the network and stores it in the device's cache. If the network condition is good, before the rain white noise finishes playing, the bird song white noise has been completely pre-cached in the cache space. Then, when the bird song needs to be played, the data can be directly read from the cache for playback to achieve seamless connection.
[0096] In this way, through pre-caching, it can effectively avoid playback interruption or stuttering caused by data acquisition delay, providing a smoother audio playback experience for users. Users do not need to wait for data loading and can more naturally transition from one white noise to the next, making the entire audio playback process more coherent and enhancing users' satisfaction and stickiness with the application.
[0097] In the above steps, different users have different preferences and requirements for the playback of white noise. Setting the playback method is to enable users to customize the playback behavior of the second white noise data frame according to their own preferences. For example, various parameters of the audio can be set, such as volume size, playback speed, channel balance, etc. By adjusting these parameters, the playback effect of the audio can be changed to adapt to different scenarios and user needs. Multiple playback modes are provided for users to choose, such as single loop, sequential playback, random playback, etc. Single loop can allow users to focus on a certain white noise to achieve a specific relaxation or sleep aid effect; sequential playback is suitable for scenarios with a certain audio content sequence arrangement; random playback can bring more freshness and diversity to users.
[0098] For example, assume that the user uses the white noise application to aid sleep before going to bed. For the upcoming second white noise data frame (such as the sound of running water), the user can set the volume to 30% to create a gentle sleep atmosphere. At the same time, the user selects the single loop mode to let the sound of running water play continuously to help themselves relax and fall asleep better. In addition, the user can also adjust the channel balance according to the headphone or speaker device they use to make the sound more stereo and comfortable.
[0099] In this way, users can personalize the playback of the second white noise data frame according to their preferences and actual needs, better meeting the usage requirements of different users in different scenarios. By providing a rich set of playback mode setting options, users can more actively participate in the audio playback process, increasing the interactivity between the user and the application and enhancing the user's interest in and frequency of using the application.
[0100] As an alternative embodiment, in 103, obtain the real-time playback environment status, and configure a preloading strategy based on the real-time playback environment status, the timbre type of the second white noise data frame, and the audio scenario; based on the preloading strategy, set the playback mode of the second white noise data frame.
[0101] In the embodiments of the present application, the playback mode at least includes: the number of preloaded frames of the second white noise data frame, the preloading sound intensity level, the preloading audio rhythm type, the preloading timbre type, and the audio transition effect between the second white noise data frame and the first white noise data frame.
[0102] When continuously playing the same song, setting the playback mode of the second white noise data frame is mainly to achieve seamless loop playback.
[0103] In an alternative embodiment of 103, set the playback parameters of the second white noise data frame, such as volume, channel, etc., to ensure consistency with the playback parameters of the first white noise data frame to achieve seamless transition. After the first white noise data frame finishes playing, immediately retrieve the second white noise data frame from the buffer for playback to ensure playback continuity. When playing to the last data frame of the song, reset the playback position to the first data frame of the song to achieve loop playback.
[0104] Specifically, in 103, obtaining the real-time playback environment status includes factors such as ambient noise level, device battery level, and network stability. For example, in a noisy environment, it may be necessary to increase the sound intensity level of the white noise to ensure that the user can hear clearly; if the device battery level is low, to save power, the number of preloaded frames may be appropriately reduced or the preloading sound intensity level may be adjusted. Network stability affects the preloading method and timing. When the network is unstable, it may be necessary to increase the number of preloaded frames to prevent playback stuttering.
[0105] Different timbre types have different requirements for preloading and playback. For example, a gentle wind white noise may require a more delicate audio transition effect to create a natural and smooth feeling; while a rushing water white noise may require a higher preloading sound intensity level to highlight its characteristics. Configuring the preloading strategy and setting the playback mode according to the characteristics of the timbre type can better display the effects of each timbre.
[0106] The audio scenario determines the overall atmosphere and requirements. If it is an audio scenario for sleep aid, the pre-loaded audio rhythm type may choose a slow and steady rhythm, and the pre-loaded timbre type will tend to be soft and soothing sounds, such as the sound of waves, rain, etc. Moreover, the audio transition effect will be set very smoothly to avoid abrupt sound changes that may affect the user's sleep. If it is an audio scenario for work concentration, relatively monotonous but regular white noise may be selected, such as the sound of a fan, air conditioner, etc., and the pre-loading strategy will pay more attention to ensuring the continuity and stability of the sound.
[0107] Scenario 1: Office environment Suppose the office environment has a certain background noise, stable network, and sufficient device battery. Suppose the timbre type of the second white noise data frame is selected as the sound of keyboard typing to simulate the natural sound in the office, helping the user better integrate into the working environment and improve concentration. Suppose the user is in a work concentration scenario. Based on the foregoing assumptions, the pre-loaded frame number is set to be moderate to ensure the continuity of the sound; the pre-loaded sound intensity level is slightly higher than the office background noise so that the sound of keyboard typing can be clearly heard but not overly abrupt; the pre-loaded audio rhythm type is similar to the normal keyboard typing rhythm, with a certain regularity; the pre-loaded timbre type is the sound of keyboard typing itself; the audio transition effect is set to be a gradual change, smoothly transitioning from the first white noise data frame to the sound of keyboard typing, making the user feel natural and smooth.
[0108] Scenario 2: Outdoor camping environment Suppose the real-time playback environment status: The outdoor environmental noise is relatively complex, with the sounds of insects, wind, etc., the network signal is weak, and the device battery is average. Suppose the timbre type of the second white noise data frame: The crackling sound of a bonfire burning is selected as the timbre type of the second white noise data frame to enhance the camping atmosphere. Suppose the audio scenario: Outdoor leisure scenario. Based on the above assumptions, due to the weak network signal, the pre-loaded frame number is appropriately increased to prevent playback interruption; the pre-loaded sound intensity level is adjusted according to the surrounding environmental noise, which should not only highlight the characteristics of the bonfire sound but also not overpower other natural sounds; the pre-loaded audio rhythm type simulates the sound rhythm of a real bonfire burning, with a certain degree of randomness and variation; the pre-loaded timbre type is the sound of a bonfire burning; the audio transition effect is set to natural integration, allowing the bonfire sound to better blend with the first white noise data frame and the surrounding natural sounds, creating a realistic outdoor camping atmosphere.
[0109] Therefore, by configuring the preloading strategy and setting the playback mode according to the real-time playback environment status, timbre type, and audio scene, it can make users feel more immersive, better integrate into the scene created by the audio, and enhance the user's immersion and experience. By reasonably setting the number of preloaded frames, sound intensity level, audio rhythm type, timbre type, and audio transition effect, it can ensure that the white noise is played more smoothly and naturally, avoid problems such as stuttering and abruptness, and improve the audio playback quality. This method can be dynamically adjusted according to different environments and user needs, making the white noise playback system more adaptable and flexible, capable of meeting the user's usage requirements in various scenarios, and improving user satisfaction.
[0110] In the embodiment of this application, on the first aspect, after the first white noise data frame is played, the second white noise data frame is directly played, ensuring the continuity of audio playback and avoiding problems such as pauses and stutters that may occur in traditional playback methods, providing users with a smooth auditory experience. For example, in a sleep aid scenario, users can continuously immerse themselves in a stable white noise environment without being awakened or having their sleep disturbed due to playback interruption. Through the pre-caching technology, the second white noise data frame to be played is prepared in advance, making the playback switching process rapid and without delay. In some scenarios with high real-time requirements, such as meditation and focused learning, it can ensure that the user's concentration is not disturbed and the user always remains in the specific atmosphere created by the white noise.
[0111] On the second aspect, a white noise classification model is used to identify the multi-dimensional audio tags corresponding to the first white noise data frame, including white noise type, sound intensity level, audio rhythm type, timbre type, audio scene, etc., and based on these tags, the second white noise data frame to be played is selected from the candidate audio library, and at least one of them has the same audio tag. This ensures that the selected audio frames are coherent and similar in content and style. For example, in the playback of white noise simulating a forest environment, it can ensure that the subsequent played audio frames also contain similar elements such as bird calls and the rustling of leaves, enhancing the realism and immersion of the scene. The multi-dimensional audio tags cover multiple dimensions and can adapt to different users' preferences for different types of white noise. For example, for users who like gentle ocean wave sounds to aid sleep, the system will accurately match similar ocean wave audio frames according to tags such as sound intensity level and white noise type, meeting the user's personalized needs and improving the user's satisfaction with the audio content.
[0112] In a third aspect, the pre-caching technology only caches the second white noise data frames related to the currently playing audio frames, avoiding a large amount of unnecessary audio data storage and loading, reducing the occupation of system resources, and improving the system operation efficiency. For example, on mobile devices, it can reduce memory and power consumption and extend the device's battery life. Through precise audio frame matching and a smooth playback process, errors and abnormal situations that may occur during playback are reduced, enhancing the stability and reliability of the system. In scenarios where white noise is played for a long time, the system can operate continuously and stably without frequent error correction or restart, ensuring the user experience.
[0113] In summary, by directly switching to the second frame after the first frame is played and combining the pre-caching technology, the sense of playback interruption is eliminated and the playback smoothness is optimized, providing a smooth auditory experience for users in scenarios such as sleep aid, meditation, and focused learning, and maintaining the user's immersion. By using the white noise classification model to identify multi-dimensional audio tags and selecting the next frame with at least one identical tag from the candidate audio library based on the tags, the content style is ensured to be coherent and similar, enhancing the scene realism and immersion, while meeting the diverse needs of users and improving satisfaction. Moreover, the pre-caching technology only caches relevant audio frames, reducing resource occupation, lowering memory and power consumption of mobile devices, improving system operation efficiency, and reducing errors and anomalies through precise matching and smooth playback, enhancing system stability and ensuring the user experience.
[0114] Figure 2 The following is a schematic structural diagram of a white noise seamless loop playback system based on audio frame pre-caching provided by an embodiment of the present application, as Figure 2 shown. The system includes the following steps: An identification unit, configured to identify, for the first white noise data frame currently being played, the multi-dimensional audio tags corresponding to the first white noise data frame through a white noise classification model; the multi-dimensional audio tags at least include: white noise type, sound intensity level, audio rhythm type, timbre type, audio scene; A selection unit, configured to select, based on the multi-dimensional audio tags, a second white noise data frame to be played from a pre-established candidate audio library; there is at least one identical audio tag between the second white noise data frame and the first white noise data frame; A playback unit, configured to set the playback mode of the second white noise data frame and directly play the second white noise data frame after the first white noise data frame is played.
[0115] Further optionally, the identification unit identifies the multi-dimensional audio tags corresponding to the first white noise data frame through a white noise classification model, specifically: Extract the first white noise data frame features corresponding to the first white noise data frame; Obtain candidate audio classification information corresponding to the first white noise data frame feature through a multi-branch perception layer; each branch perception layer in the multi-branch perception layer is used to extract key feature information corresponding to the first white noise data frame feature, and obtain corresponding candidate audio classification information based on the extracted key feature information; each branch perception layer is set based on the corresponding classification logic; Convert the candidate audio classification information corresponding to the first white noise data frame feature into a corresponding multi-dimensional audio label.
[0116] Further optionally, the recognition unit obtains candidate audio classification information corresponding to the first white noise data frame feature through a multi-branch perception layer, specifically for: Input the first white noise data frame feature into a classification target layer to identify the white noise type in the first white noise data frame feature; Based on the identified white noise type, adaptively configure the type of branch perception layer involved in the corresponding multi-branch perception layer; Call the corresponding type of multi-branch perception layer to identify audio change features corresponding to different frequency bands, different time sequence intervals, and different change waveforms in the first white noise data frame feature, so as to obtain candidate audio classification information corresponding to the first white noise data frame feature.
[0117] Further optionally, after the recognition unit extracts the first white noise data frame feature corresponding to the first white noise data frame, it is also used for: Extract the interference noise feature matrix in the first white noise data frame; the interference noise elements in the interference noise feature matrix are represented by the following formula: ; where, represents the interference noise element corresponding to the first white noise data frame feature at the t-th time step, and is the dynamic expansion rate in the interference noise feature matrix, and is the dynamic step size in the interference noise feature matrix, and is the interference noise analysis dimension index, and are the output coordinate values of the first white noise data frame feature and is in the interference noise element feature components in the dimension; Fuse the interference noise feature matrix with the first white noise data frame feature, and input the fusion result into a multi-branch perception layer to increase the robustness of the first white noise data frame feature processing process.
[0118] Further optionally, after the identification unit extracts the first white noise data frame feature corresponding to the first white noise data frame, it is further used for: Obtain the real-time delay information, real-time network fluctuation information, and real-time network environment state corresponding to the playback of the first white noise data frame; Based on the real-time delay information, real-time network fluctuation information, and real-time network environment state, predict the predicted network environment information when the first white noise data frame is played to a preset time period; wherein, the preset time period at least includes: a preset duration before the end of the first white noise data frame; According to the predicted network environment information, evaluate the adaptability between the predicted network environment information and the memory resources corresponding to the pre-loaded area in the playback device; Based on the adaptability, determine the network transition type corresponding to the first white noise data frame; the network transition type is used to indicate the network pre-loading state before transitioning to the second white noise data frame.
[0119] Further optionally, the selection unit selects the second white noise data frame to be played from a pre-established candidate audio library based on the multi-dimensional audio label, specifically used for: Construct a virtual playback scene corresponding to the first white noise data frame based on the multi-dimensional audio label; According to the nodes corresponding to each audio label in the virtual playback scene, select candidate white noise data frames with at least one same node from the candidate audio library; the arrangement of the nodes corresponding to each audio label in the virtual playback scene is determined based on the attributes of the white noise data frame; Determine the second white noise data frame from the candidate white noise data frames based on the user preference data.
[0120] Further optionally, the selection unit sets the playback mode of the second white noise data frame, specifically used for: obtaining the real-time playback environment state, and configuring a pre-loading strategy based on the real-time playback environment state, the timbre type of the second white noise data frame, and the audio scene; Based on the pre-loading strategy, set the playback mode of the second white noise data frame; the playback mode at least includes: the number of pre-loaded frames of the second white noise data frame, the pre-loaded sound intensity level, the pre-loaded audio rhythm type, the pre-loaded timbre type, and the audio transition effect between the second white noise data frame and the first white noise data frame.
[0121] Please refer to Figure 3 ,Figure 3 This is a schematic diagram of an embodiment of an electronic device provided by an embodiment of the present application. As Figure 3 shown, an embodiment of the present application provides an electronic device 500, including a memory 510, a processor 520, and a software program 511 stored on the memory 510 and executable on the processor 520. When the processor 520 executes the software program 511, the foregoing embodiments are implemented.
[0122] Please refer to Figure 4 , Figure 4 This is a schematic diagram of an embodiment of a computer-readable storage medium provided by an embodiment of the present application. As Figure 4 shown, this embodiment provides a computer-readable storage medium 600, on which a computer program 611 is stored. When the computer program 611 is executed by a processor, the foregoing embodiments are implemented.
[0123] It should be noted that in the above embodiments, the descriptions of the respective embodiments have their own emphases. For parts not described in detail in a certain embodiment, reference may be made to the relevant descriptions of other embodiments.
[0124] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) containing computer-usable program code.
[0125] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to the embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram, and the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded computer, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the functions specified in Figure 1 one or more of the flows or multiple flows and / or blocks Figure 1 one or more of the blocks or multiple blocks. Figure 1 These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device, and the instruction device implements the functions specified in Figure 1The functions specified in one or more blocks. These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process. Thus, the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one or more processes and / or blocks Figure 1 One or more processes and / or blocks Figure 1 The steps of the functions specified in one or more blocks. Although the preferred embodiments of the present invention have been described, those skilled in the art can make additional changes and modifications to these embodiments once they know the basic creative concept. Therefore, the appended claims are intended to be construed to include the preferred embodiments as well as all changes and modifications falling within the scope of the present invention. Obviously, those skilled in the art can make various changes and variations to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalent technologies, the present invention is also intended to include these modifications and variations
Claims
1. A method for seamless loop playback of white noise based on audio frame pre-buffering, characterized in that, The method at least includes: For the first white noise data frame being currently played, identify the multi-dimensional audio tags corresponding to the first white noise data frame through a white noise classification model; the multi-dimensional audio tags at least include: white noise type, sound intensity level, audio rhythm type, timbre type, audio scene; Based on the multi-dimensional audio tags, select a second white noise data frame to be played from a pre-established candidate audio library; there is at least one identical audio tag between the second white noise data frame and the first white noise data frame; Set the playing mode of the second white noise data frame and directly play the second white noise data frame after the first white noise data frame is played.
2. The method for seamless loop playback of white noise based on audio frame pre-buffering according to claim 1, wherein The step of identifying the multi-dimensional audio tags corresponding to the first white noise data frame through a white noise classification model includes: Extract the first white noise data frame features corresponding to the first white noise data frame; Obtain candidate audio classification information corresponding to the first white noise data frame features through a multi-branch perception layer; each branch perception layer in the multi-branch perception layer is used to extract key feature information corresponding to the first white noise data frame features and obtain corresponding candidate audio classification information based on the extracted key feature information; each branch perception layer is set based on the corresponding classification logic; Convert the candidate audio classification information corresponding to the first white noise data frame features into corresponding multi-dimensional audio tags.
3. The method for seamless loop playback of white noise based on audio frame pre-buffering according to claim 2, wherein The step of obtaining candidate audio classification information corresponding to the first white noise data frame features through a multi-branch perception layer includes: Input the first white noise data frame features into a classification target layer to identify the white noise type in the first white noise data frame features; Based on the identified white noise type, adaptively configure the types of branch perception layers involved in the corresponding multi-branch perception layer; Call the corresponding type of multi-branch perception layer to identify the audio change features corresponding to different frequency bands, different time sequence intervals, and different change waveforms in the first white noise data frame features, so as to obtain candidate audio classification information corresponding to the first white noise data frame features.
4. The white noise seamless loop playback method based on audio frame pre-buffering according to claim 2, wherein After extracting the first white noise data frame features corresponding to the first white noise data frame, it further includes: Extract the interference noise feature matrix in the first white noise data frame; the interference noise elements in the interference noise feature matrix are expressed by the following formula: ; Among them, represents the first white noise data frame feature corresponding interference noise element at the t - th time step, and is the dynamic expansion rate in the interference noise feature matrix, and is the dynamic step size in the interference noise feature matrix, and is the interference noise analysis dimension index, and is the output coordinate value of the first white noise data frame feature and is in the interference noise element feature component under the dimension; Perform fusion processing on the interference noise feature matrix and the first white noise data frame features, and input the fusion result into the multi-branch perception layer to increase the robustness of the first white noise data frame feature processing process.
5. The method for seamless loop playback of white noise based on audio frame pre-buffering according to claim 2, wherein After extracting the first white noise data frame features corresponding to the first white noise data frame, it further includes: Obtain the real-time delay information, real-time network fluctuation information, and real-time network environment status corresponding to the playing of the first white noise data frame; Based on the real-time delay information, real-time network fluctuation information, and real-time network environment status, predict the predicted network environment information when the first white noise data frame is played to a preset time period; where the preset time period at least includes: a preset duration before the end of the first white noise data frame; Evaluate the adaptability between the predicted network environment information and the memory resources corresponding to the preloading area in the playback device according to the predicted network environment information; Determine the network transition type corresponding to the first white noise data frame based on the adaptability; the network transition type is used to indicate the network preloading state before transitioning to the second white noise data frame.
6. The method for seamless loop playback of white noise based on audio frame pre-buffering according to claim 1, wherein The selecting the second white noise data frame to be played from the pre-established candidate audio library based on the multi-dimensional audio tag includes: Construct a virtual playback scene corresponding to the first white noise data frame based on the multi-dimensional audio tag; Select candidate white noise data frames having at least one identical node from the candidate audio library according to the nodes corresponding to the respective audio tags in the virtual playback scene; the arrangement manner of the nodes corresponding to the respective audio tags in the virtual playback scene is determined based on the attributes of the white noise data frame; Determine the second white noise data frame from the candidate white noise data frames based on user preference data.
7. The method for seamless loop playback of white noise based on audio frame pre-buffering according to claim 1, wherein The setting the playback mode of the second white noise data frame includes: Obtain the real-time playback environment state, and configure a preloading strategy based on the real-time playback environment state, the timbre type and audio scene of the second white noise data frame; Set the playback mode of the second white noise data frame based on the preloading strategy; the playback mode at least includes: the number of preloaded frames of the second white noise data frame, the preloading sound intensity level, the preloading audio rhythm type, the preloading timbre type, and the audio transition effect between the second white noise data frame and the first white noise data frame.
8. A white noise seamless loop playback system based on audio frame pre-buffering, characterized in that, The system at least includes the following units: An identification unit, configured to identify the multi-dimensional audio tag corresponding to the first white noise data frame being currently played through a white noise classification model for the first white noise data frame being currently played; The multi-dimensional audio tag at least includes: white noise type, sound intensity level, audio rhythm type, timbre type, audio scene; A selection unit, configured to select the second white noise data frame to be played from the pre-established candidate audio library based on the multi-dimensional audio tag; there is at least one identical audio tag between the second white noise data frame and the first white noise data frame; A playback unit, configured to set the playback mode of the second white noise data frame and directly play the second white noise data frame after the first white noise data frame is played.
9. An electronic device, characterized in that, Including a memory for storing computer software programs; A processor, configured to read and execute the computer software program, thereby implementing the white noise seamless loop playback method based on audio frame pre-caching according to any one of claims 1-7.
10. A non-transitory computer-readable storage medium, characterized in that, The computer software program is stored in the storage medium, and when the computer software program is executed by the processor, it implements the white noise seamless loop playback method based on audio frame pre-caching according to any one of claims 1-7.
Citation Information
Patent Citations
White noise playing system and portable sound equipment
CN106792297A
Vehicle-mounted audio playing method and device, equipment and storage medium
CN117827141A
Mode switching optimization method, system and equipment of audio input player and storage medium
CN119201030A
System and method for dynamically adjusting settings of audio output devices to reduce noise in adjacent spaces
US20210392451A1
Information processing device, information processing system, and computer program
WO2025110207A1
Cited By
White noise adaptive recommendation method and device based on electroencephalogram signal feedback
CN121506428A