Event processing method and device based on sound signal analysis, equipment and medium

By performing segmented processing and feature extraction of sound signals, combined with preset feature database comparison, the problem of inaccurate multi-event type recognition in the prior art is solved, and automatic identification and adaptive response processing of events such as snoring, apnea and breathing is realized, improving identification accuracy and response flexibility.

CN120544570APending Publication Date: 2025-08-26PING AN TECH (SHENZHEN) CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510763022.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-09
Publication Date
2025-08-26

AI Technical Summary

Technical Problem

The prior art cannot accurately identify and distinguish multiple respiratory event types, and lacks an automatic response processing mechanism, resulting in low accuracy of event recognition and inability to achieve flexible and automated event processing.

Method used

By obtaining sound signals, performing segmented processing, extracting sound characteristics, and comparing them with reference sound characteristics in the preset feature library, determining the event type, and performing corresponding processing operations when the preset conditions are met.

Benefits of technology

Accurate recognition and adaptive response processing of multiple event types such as snoring, apnea and gasping are achieved, improving the accuracy of sound event recognition and flexibility in response.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120544570A_ABST
    Figure CN120544570A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of data analysis, can be applied to business scenes of financial science and technology, medical health and the like, and discloses a sound signal analysis-based event processing method, device, equipment and medium, and the method comprises the steps: obtaining a to-be-analyzed sound signal, carrying out the segmentation processing of the to-be-analyzed sound signal, and obtaining a target sound segment; extracting sound features of the target sound fragment, and comparing the sound features with reference sound features in a preset feature library to generate a comparison result; and determining an event type corresponding to the target sound segment based on the comparison result, and when the event type satisfies a preset condition, executing a processing operation corresponding to the event type. According to the method, through sound signal segmentation processing and feature extraction, event types such as snore, apnea and enzootic pneumonia are accurately recognized, the event recognition accuracy is improved by combining feature comparison, corresponding processing is automatically executed based on the event types, and self-adaptive response of multiple event types is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data analysis technology, and in particular to an event processing method, device, equipment and storage medium based on sound signal analysis. Background Art

[0002] In the healthcare sector, the diagnosis and monitoring of obstructive sleep apnea syndrome (OSAHS) has long relied on polysomnography (PSG), the current gold standard diagnostic method. However, PSG requires patients to spend an entire night in the hospital connected to multiple sensors (such as EEG, ECG, and blood oxygen saturation). This is not only complex and expensive, but can also lead to a "first-night effect" due to unfamiliarity with the environment, where sleep data can be distorted by environmental changes. Furthermore, patients must remain still for extended periods while wearing multiple electrodes, significantly reducing patient compliance. Even portable screening devices (such as oximeters or wristwatches) can only monitor blood oxygen saturation or pulse rate, lacking the ability to identify snoring characteristics and apnea patterns, posing a risk of missed diagnosis. While acoustic analysis technology has been tried for OSAHS detection, most acoustic monitoring systems rely on microphones to record snoring and are combined with endoscopes or pressure catheters for auxiliary diagnosis. These complex procedures and reliance on manual interpretation make real-time monitoring and automatic identification difficult.

[0003] In the field of fintech business, smart health insurance products are gradually emerging, and health data monitoring based on sound signal analysis is regarded as a key means to assess the health status of policyholders, calculate premiums and risks. However, existing monitoring methods based on static health data (such as blood oxygen and heart rate) lack the ability to identify multiple event types and cannot accurately distinguish between various respiratory abnormalities such as snoring, apnea, and gasping, resulting in one-sided health assessment data that cannot accurately reflect the policyholder's true health risks. In addition, the data analysis of existing products mostly relies on single signals or static voiceprint analysis, and fails to form a complete recognition of the dynamic sequence of "snoring-apnea-gasping", resulting in insufficient data accuracy. At the same time, existing systems usually require the use of professional equipment for data collection and analysis, and cannot achieve real-time warning and processing based on mobile terminals, and fail to effectively meet the real-time and convenience requirements of insurance companies in health monitoring and risk assessment. Summary of the Invention

[0004] The main purpose of the present invention is to provide an event processing method, device, equipment and storage medium based on sound signal analysis, aiming to solve the technical problems that the existing technology cannot accurately identify and distinguish multiple respiratory event types, and lacks an automatic response processing mechanism based on event type, resulting in low event recognition accuracy and inability to achieve flexible and automated event processing.

[0005] To achieve the above object, the present invention provides an event processing method based on sound signal analysis, comprising:

[0006] Acquiring a sound signal to be analyzed;

[0007] Segmenting the sound signal to be analyzed to obtain target sound segments;

[0008] Extracting sound features of the target sound segment;

[0009] Comparing the sound feature with reference sound features in a preset feature library to generate a comparison result;

[0010] Based on the comparison result, determining the event type corresponding to the target sound segment;

[0011] When the event type meets a preset condition, a processing operation corresponding to the event type is executed.

[0012] Furthermore, to achieve the above-mentioned purpose, the present invention provides an event processing device based on sound signal analysis, comprising:

[0013] An acquisition module, used to obtain the sound signal to be analyzed;

[0014] A segmentation processing module, configured to perform segmentation processing on the sound signal to be analyzed to obtain target sound segments;

[0015] A feature extraction module, configured to extract the sound features of the target sound segment;

[0016] A feature comparison module, configured to compare the sound feature with reference sound features in a preset feature library to generate a comparison result;

[0017] An event determination module, configured to determine an event type corresponding to the target sound segment based on the comparison result;

[0018] The event response module is used to execute a processing operation corresponding to the event type when the event type meets a preset condition.

[0019] Furthermore, to achieve the above-mentioned purpose, the present invention also provides a computer device, which includes a memory, a processor, and an event processing program based on sound signal analysis stored in the memory and executable on the processor. When the event processing program based on sound signal analysis is executed by the processor, the steps of the event processing method based on sound signal analysis as described above are implemented.

[0020] Furthermore, to achieve the above-mentioned purpose, the present invention also provides a computer-readable storage medium, on which an event processing program based on sound signal analysis is stored. When the event processing program based on sound signal analysis is executed by a processor, the steps of the event processing method based on sound signal analysis as described above are implemented.

[0021] Beneficial effects: The present invention relates to the field of data analysis technology and can be applied to business scenarios such as financial technology and medical health. It discloses an event processing method, device, equipment and medium based on sound signal analysis, including: obtaining a sound signal to be analyzed, performing segmentation processing on the sound signal to be analyzed to obtain a target sound segment; extracting sound features of the target sound segment, comparing the sound features with reference sound features in a preset feature library, and generating a comparison result; determining the event type corresponding to the target sound segment based on the comparison result, and executing a processing operation corresponding to the event type when the event type meets the preset conditions. The present invention can accurately extract features of different events such as snoring, apnea and gasping from a variety of sound events through segmentation processing and feature extraction of sound signals; effectively identify event types by comparing the extracted sound features with the reference feature library; and automatically generate and execute corresponding processing operations based on the event type, thereby realizing automatic recognition and adaptive response processing of multiple event types, and improving the accuracy of sound event recognition and the flexibility of response. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] The present invention will be further described below with reference to the accompanying drawings and embodiments, in which:

[0023] Figure 1 Schematic diagram of an application environment of an event processing method based on sound signal analysis in one embodiment of the present invention;

[0024] Figure 2 1 is a flow chart of an embodiment of an event processing method based on sound signal analysis according to the present invention;

[0025] Figure 3 Schematic diagram of functional modules of a preferred embodiment of an event processing device based on sound signal analysis of the present invention;

[0026] Figure 4 A schematic diagram of the structure of a computer device according to an embodiment of the present invention;

[0027] Figure 5 FIG. 2 is another structural diagram of a computer device according to an embodiment of the present invention. DETAILED DESCRIPTION

[0028] It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.

[0029] The event processing method based on sound signal analysis provided by the embodiment of the present invention can be applied in Figure 1 In an application environment, the user terminal communicates with the server terminal through a network. The server terminal can obtain the sound signal to be analyzed through the user terminal, segment the sound signal to be analyzed, and obtain the target sound segment; extract the sound features of the target sound segment, compare the sound features with the reference sound features in the preset feature library, and generate a comparison result; determine the event type corresponding to the target sound segment based on the comparison result, and when the event type meets the preset conditions, execute the processing operation corresponding to the event type. The present invention can accurately extract the features of different events such as snoring, apnea and gasping from a variety of sound events through segmentation processing and feature extraction of sound signals; effectively identify the event type by comparing the extracted sound features with the reference feature library; and automatically generate and execute corresponding processing operations based on the event type, thereby realizing automatic recognition and adaptive response processing of multiple event types, improving the accuracy of sound event recognition and the flexibility of response. The user terminal can be, but is not limited to, various personal computers, laptops, smart phones, tablet computers and portable wearable devices. The server terminal can be implemented using an independent server or a server cluster consisting of multiple servers. The present invention is described in detail below through specific embodiments.

[0030] See also Figure 2 , Figure 2 This is a flow chart of an embodiment of an event processing method based on sound signal analysis provided by the present invention. It should be noted that although a logical order is shown in the flow chart, in some cases, the steps shown or described may be performed in a different order than that shown here.

[0031] like Figure 2 As shown, the event processing method based on sound signal analysis proposed by the present invention includes the following steps:

[0032] S10, obtaining a sound signal to be analyzed;

[0033] In this embodiment, the sound signal is collected by a sound sensing device, which may include a microphone, an acoustic sensor in a wearable device, a built-in microphone in a smart terminal, or other sound wave receiving device. The collection of sound signals should be carried out in an environment without obvious external noise interference to ensure that the collected sound signal contains the acoustic characteristics of the target sound source. The way of collecting sound signals can be diversified, including directly collecting sound signals in the air through the microphone of the device, or collecting bone conduction sound waves through a vibration sensor in contact with the user's body. In specific implementation, a high-sensitivity microphone can be selected to improve the clarity of sound collection, or a wearable device with a noise cancellation design can be selected to reduce the interference of environmental noise.

[0034] The collected sound signal is usually a continuous original acoustic signal. This signal itself contains a variety of noises and background noises, and the data volume is large, which is not conducive to subsequent processing. In order to improve the clarity of the sound signal, the collected original acoustic signal can first enter the noise reduction processing module. Noise reduction processing can adopt a variety of technical means, such as adaptive filtering algorithm, which eliminates environmental noise in real time by analyzing the frequency characteristics of background noise and dynamically adjusting the filter coefficient. Bandpass filtering can also be used to retain only the frequency band that matches the frequency range of the target sound source (such as breathing sounds, snoring). Furthermore, the noise reduction module can be combined with beamforming technology to focus on the direction of the target sound source through a multi-sensor array to reduce the interference of non-target sounds. The sound signal after noise reduction is regarded as the noise-reduced acoustic signal.

[0035] Based on the noise-reduced acoustic signal, the acquisition module samples the signal. The sampling frequency is set according to the application scenario. For example, a sampling frequency of 8kHz to 16kHz is suitable for collecting most human voice signals. Sampling converts the analog signal into a digital signal, generating a sampled acoustic signal. To ensure data continuity and analysis accuracy, sampling can be combined with a segmented time window, set to a fixed sampling interval such as 10 milliseconds or 25 milliseconds. The sampled signal is sent to a mobile terminal or cloud processing platform via a wireless transmission module. Transmission methods can include Bluetooth, Wi-Fi, cellular data networks, etc.

[0036] In one embodiment, sound signals are collected by a microphone on a wearable device, which is held close to the user's throat to capture a clear acoustic signal. The raw acoustic signal collected by the microphone enters a noise reduction module, where an adaptive filtering algorithm removes ambient noise. The noise-reduced acoustic signal is sampled at a 16kHz sampling rate to generate a sampled acoustic signal. The sampled acoustic signal is then transmitted to a mobile terminal via Bluetooth for subsequent processing.

[0037] In another implementation, sound signals are collected via a smartphone's built-in microphone, eliminating the need for the user to wear any other device. A smartphone app automatically identifies the sound signal and initiates noise reduction, employing bandpass filtering to retain only frequency components within the 300Hz to 3kHz range, adapted to the characteristics of human speech and breathing. The sampling frequency is set at 8kHz, and the sampled sound signal is automatically uploaded to the cloud for processing.

[0038] Home health monitoring devices, such as smart speakers or smart health mirrors, can also capture the user's breathing sounds. The device uses a built-in microphone array to localize the sound source, focusing on the user's location, capturing the acoustic signal and performing real-time noise reduction and sampling. The sampled signal is then transmitted via the home local area network to a mobile device or cloud-based analysis platform.

[0039] Example: In the healthcare field, a user wearing smart health headphones automatically collects acoustic signals from the throat while sleeping at night. The device's built-in noise reduction module eliminates ambient noise and collects sound signals at a 16kHz sampling rate. The sampled signals are transmitted to a mobile device via Bluetooth, where the user can view the waveform data of the sound signals in real time through a mobile app. The system further analyzes the sound signals, identifies snoring and apnea events, and generates a health report.

[0040] In the financial sector, after purchasing a smart health insurance product, users wear a smart health device that automatically collects and monitors sound signals. These signals are uploaded to the insurance company's cloud platform in real time via a 5G network. The platform analyzes the user's sound signal results to assess their health risk level and dynamically adjusts premiums accordingly. The insurance company uses long-term sound signal data to analyze the user's health status and automatically generates a health risk report.

[0041] This embodiment acquires sound signals through sound sensing devices and combines them with various noise reduction processing techniques to effectively reduce ambient noise interference and improve the clarity of sound signals. By setting flexible sampling frequencies and multiple wireless transmission methods, the high fidelity and real-time performance of the sampled sound signals are ensured. Compared with existing technologies, this can achieve high-quality sound signal acquisition without relying on specialized equipment and is adaptable to a variety of usage scenarios.

[0042] S20, performing segmentation processing on the sound signal to be analyzed to obtain target sound segments;

[0043] In this embodiment, the core of the sound signal segmentation processing is to divide the continuous sound signal into multiple independent segments according to a fixed time window, and each segment represents the sound data within a period of time. The main purpose of segmentation processing is to convert the continuous sound signal into structured sound segments for subsequent feature extraction and event recognition. The first step of segmentation processing is to set the segmentation time window, which defines the duration of each segment, such as 30 seconds, 1 minute or other durations that adapt to the target application. The length of the time window directly affects the granularity of the sound segment. A time window that is too long may cause the sound features to be diluted, and a time window that is too short may cause the sound segments to be too fragmented and difficult to analyze.

[0044] Each segmented sound fragment requires a start and end timestamp. These timestamps record the fragment's time range, ensuring that each fragment is accurately mapped to the original sound signal's timeline. This timestamp information is crucial for subsequent event location, frequency analysis, and response processing. Timestamps can be automatically generated based on the system time or accumulated based on sampling timing to ensure that each fragment's time range is accurate and continuous.

[0045] To further enhance the accuracy of sound signal analysis, each segment can be processed using overlapping framing. This involves further dividing the segment into multiple overlapping frames based on the segmented time window. Overlapping framing aims to reduce boundary effects, where sudden changes at the beginning or end of a segment can distort the sound signature. The length and overlap ratio of the overlapping frames can be flexibly configured, for example, with a frame length of 25 milliseconds and a frame overlap ratio of 50%. Each frame can be independently subjected to feature extraction and analysis.

[0046] After processing the segmented sound clips, each segment can be stored as a separate audio file and associated with a corresponding timestamp. Audio files can be in a variety of formats, such as WAV, PCM, or MP3, depending on storage space and data processing requirements. File names can automatically be appended with timestamps or serial numbers for subsequent retrieval and analysis.

[0047] In one embodiment, the collected sound signal is segmented using a fixed 30-second time window, with each 30-second segment appended with a start and end timestamp. Each segment is framed with a 50% overlap, with each frame length of 25 milliseconds and a 50% overlap. Each segment is stored as a WAV audio file with a file name of the format "segment_2023_10_12_22_30_00.wav," ensuring that the file name directly reflects the time information.

[0048] In another embodiment, the sound signal segmentation window is set to 1 minute, and the start and end timestamps of each segment are automatically generated based on the system time. The segmented sound segments are processed using overlapping frames with a 25 millisecond frame length and a 75% overlap ratio. Each segment is stored in MP3 format, which is suitable for storage-limited scenarios. The audio files are stored locally, and the timestamps are recorded in both the file name and the file metadata.

[0049] Dynamic segmentation can also be used, where the segmentation window is dynamically adjusted based on the characteristics of the sound signal. For example, when a continuous sound signal (such as snoring or breathing) is detected, the segmentation window is automatically extended; when a long period of silence or an invalid signal is detected, the segmentation window is automatically shortened. Each segmented sound segment is recorded with a dynamic timestamp to record its start and end time and stored as a WAV format audio file.

[0050] Example: In the healthcare field, a user wears smart health headphones for sleep monitoring. The continuous sound signal collected by the device is automatically segmented into independent sound segments every 30 seconds. Each segment is affixed with a start and end timestamp and processed using overlapping frames. Each sound segment is automatically stored as a WAV file with a timestamp in the file name. Subsequent snoring and apnea event recognition is analyzed based on these segmented sound segments.

[0051] In the financial sector, sound signals collected by users through smart health insurance apps are automatically segmented and processed, generating a separate sound clip every minute and storing it in MP3 format. Each clip is uploaded to the insurance company's cloud platform in real time via Bluetooth. The platform analyzes the user's breathing pattern based on the clip's timestamp and sound characteristics, dynamically assessing the user's health risks and automatically generating a health risk report.

[0052] This embodiment uses segmentation to divide a continuous sound signal into independent sound segments, ensuring that each segment is temporally independent and has a clear start and end time. Overlapping framing reduces distortion at segment boundaries and ensures the continuity of sound features. Storing independent audio files facilitates subsequent feature extraction and event recognition. Compared to existing technologies, this method can precisely locate the time range of sound segments, and through segmentation and framing, improves the clarity and accuracy of sound features.

[0053] S30, extracting the sound features of the target sound segment;

[0054] In this embodiment, sound feature extraction involves extracting numerical data describing the sound features from a target sound segment for subsequent sound event recognition. Sound feature extraction typically involves multiple stages: preprocessing, framing, frequency feature extraction, and energy feature extraction.

[0055] During the preprocessing phase, the sound signal first undergoes pre-emphasis. Pre-emphasis is a high-pass filtering technique that enhances the high-frequency components of the sound signal, particularly the sibilance and high-frequency harmonics characteristic of human voice. Pre-emphasis is achieved by adding a negative value proportional to the previous sample value to each sample point in the sound signal, thereby increasing the proportion of high-frequency components. Pre-emphasis effectively improves the resolution of sound features, especially the capture of high-frequency components.

[0056] After pre-emphasis, the sound signal enters the framing stage. Framing divides the continuous sound signal into multiple short time segments, each of which is called a frame. The frame length can be flexibly set according to the application scenario, such as 20 milliseconds, 25 milliseconds, or 30 milliseconds. The frame shift defines the degree of overlap between frames and is usually set to half the frame length. For example, if the frame length is 25 milliseconds, the frame shift can be 12.5 milliseconds. The purpose of framing is to ensure the continuity of the sound signal and reduce the impact of sudden sound changes caused by time segment boundaries on feature extraction. Each frame is processed using a weighted window function, such as a Hamming window or Hanning window, which can smooth frame edges and reduce spectral leakage.

[0057] Based on the framed sound signal, frequency feature extraction begins. One of the most commonly used sound features is the Mel-Frequency Cepstral Coefficient (MFCC). MFCC is achieved by converting the sound signal from the time domain to the frequency domain and further converting it into a characteristic value on the frequency scale (Mel frequency) perceived by the human ear. This feature can effectively describe the frequency distribution of the sound signal, especially the characteristics of speech and breathing sounds. MFCC features are usually composed of characteristic coefficients in multiple frequency bands, such as 12 to 20 eigenvalues, which can effectively capture the spectral structure of the sound.

[0058] In addition to frequency features, energy features can also be extracted from sound signals. The most commonly used feature is short-time energy. Short-time energy reflects the energy distribution of a sound signal over a short period of time and effectively describes changes in sound intensity. Short-time energy is important for identifying changes in respiratory intensity, snoring, and apnea. To extract short-time energy, the square of each frame of the sound signal is summed to obtain the energy value for that frame. Combining short-time energy with frequency features can capture both the frequency and energy characteristics of a sound.

[0059] After obtaining MFCC and short-time energy features, they can be normalized. The purpose of normalization is to eliminate differences in the numerical scales of different features and improve comparability between them. Normalization can use zero-mean normalization or minimum-maximum normalization to unify the numerical values ​​of different features to the same scale. The normalized features can be weighted and fused according to their weights, and the weights of different features can be flexibly adjusted according to the application scenario. For example, in respiratory monitoring, the weight of short-time energy features can be higher than that of frequency features to highlight changes in respiratory intensity.

[0060] Example: In the healthcare field, a user wears smart health headphones for sleep monitoring. The sound signal collected by the device is automatically pre-emphasized, framed, and feature extracted. The system extracts MFCC and short-term energy features from the sound signal and generates the final sound signature through normalization and weighted fusion. The feature data is uploaded to the cloud, where the platform analyzes the sound features, identifies snoring and apnea events, and generates a health report.

[0061] In the financial sector, after purchasing smart health insurance, users use their smartphones to collect daily breathing and voice signals. The app automatically performs pre-emphasis, framing, and feature extraction. The generated sound signatures are uploaded to the insurance company's cloud platform in real time via the 5G network. The platform analyzes the user's health status based on the sound signatures, dynamically adjusts premium rates, and generates health risk reports.

[0062] This embodiment accurately captures the frequency distribution and energy variations of sound from target sound segments through pre-emphasis, framing, and frequency and energy feature extraction. Pre-emphasis enhances high-frequency features, framing reduces boundary effects, and the combination of MFCC and short-time energy features achieves dual descriptions of frequency and energy. Normalization and weighted fusion ensure feature stability and comparability. Compared to existing technologies, this method can extract accurate and stable sound features in complex sound environments, making it particularly suitable for the precise identification of events such as snoring, apnea, and gasping.

[0063] S40, comparing the sound feature with a reference sound feature in a preset feature library to generate a comparison result;

[0064] In this embodiment, the sound feature comparison with the reference sound features in the preset feature library is performed by comparing the extracted sound features with the standard sound features in the feature library to determine whether a specific event (such as snoring, apnea, or gasping) exists in the sound signal to be analyzed. This comparison process is the core of sound event recognition and ensures that specific event types are accurately identified from the sound signal.

[0065] A sound feature library is a structured collection of reference features, each of which is associated with a specific event type, such as snoring, apnea, and gasping. The feature library can be pre-built and trained and optimized based on a large number of sound signal samples. The reference sound features in the feature library can be generated in various ways, such as training with a large sample of sound signals or constructing them from sound datasets annotated by medical experts. The feature library can also be dynamically updated using the user's historical sound data, allowing it to gradually adapt to the user's individual characteristics.

[0066] The first step in sound feature comparison is to retrieve multiple reference sound features from the feature library. Each reference sound feature is associated with a specific event type. Reference sound features can include Mel-Frequency Cepstral Coefficients (MFCCs), short-term energy, spectral band energy, and other features, each reflecting different characteristics of the sound signal. The combination of multiple features ensures the accuracy of the comparison results.

[0067] During the comparison process, the extracted sound features are compared with each reference sound feature for similarity. Similarity is a measure of the similarity between sound features. Common similarity metrics include cosine similarity, Euclidean distance, and Dynamic Time Warping (DTW). A higher similarity indicates a closer match between the extracted sound features and the reference sound features, meaning the sound signal being analyzed is more likely to belong to the reference event type.

[0068] To improve comparison accuracy, a dynamically adjusted similarity threshold can be set for each event type. This threshold is optimized based on the user's historical data and individual differences. For example, the similarity threshold for snoring events can be dynamically updated based on the frequency and intensity of the user's snoring. This dynamically adjusted threshold allows the comparison results to better adapt to the user's individual characteristics, avoiding misjudgments caused by fixed thresholds.

[0069] In addition to similarity calculations, the duration and number of occurrences per unit time of each event type's similarity value can also be recorded. Duration refers to the duration of each occurrence of the sound signature of each event type, while the number of occurrences per unit time counts the frequency of the event within a preset time window. This temporal information helps determine the severity of the event. For example, the duration and frequency of snoring can help distinguish between normal snoring and sleep apnea.

[0070] Based on the comparison results, a preliminary determination result for each event type can be generated. The preliminary determination result is determined based on similarity and time information. When the similarity value of a certain event type exceeds the dynamically adjusted threshold, and the duration and frequency meet the preset conditions (such as a single duration exceeding 10 seconds and a frequency exceeding 5 times per unit time), a preliminary determination result for that event type is generated.

[0071] In one embodiment, the sound feature library includes three event types: snoring, apnea, and gasping. Each event type contains multiple reference sound features, and each reference sound feature is represented by MFCC and short-time energy. The system retrieves the reference sound features of each event type from the feature library, and calculates the cosine similarity of the extracted sound features with each reference sound feature in turn. The system dynamically adjusts the similarity threshold of each event type based on the user's historical data. For example, the similarity threshold of snoring events is set to 0.85, and the similarity threshold of apnea events is set to 0.90. During the comparison process, the system calculates the similarity value of each event type and records the single duration and the number of occurrences per unit time when the similarity exceeds the threshold. For example, when the similarity value of the snoring event continuously exceeds the threshold and the single duration exceeds 10 seconds, and the number of occurrences per unit time exceeds 5 times, a preliminary judgment result of the snoring event is generated.

[0072] In another embodiment, the feature library can be dynamically updated based on the user's historical data. Each sound signal collected by the user is automatically added to the feature library, and the reference sound feature value is adjusted based on the user's sound characteristics. During the comparison, the system dynamically adjusts the similarity threshold based on the user's individual characteristics, such as processing the user's snoring characteristics in a quiet environment at night separately from the breathing characteristics in a daytime environment.

[0073] The diversity of feature comparison can also be expanded through a cloud-based feature library. This library contains voice features from multiple users and dynamically optimizes reference voice features and similarity thresholds through big data analysis. When performing feature comparisons, the user's local device can reference the reference voice features in the cloud-based feature library, enabling personalized and dynamic event recognition.

[0074] Example: In the healthcare field, a user wearing smart health headphones collects sound signals and uploads them to a cloud platform. The platform retrieves reference sound features for snoring, apnea, and gasping from a feature library and compares the extracted sound features with the reference sound features in sequence. The system records the similarity value and duration of the snoring event and determines whether the snoring event meets the judgment criteria based on a dynamic threshold. If the similarity value of the snoring event exceeds the threshold, the duration of a single snoring event exceeds 10 seconds, and the frequency per unit time exceeds 5 times, the system generates a preliminary judgment result for the snoring event.

[0075] In the financial sector, users use smart health insurance apps to collect daily breathing and voice signals. The app automatically extracts features and uploads them to the insurance company's cloud platform. The platform then retrieves reference sound signatures for snoring and apnea from a dynamic feature library and compares the user's voice signatures with the library. Based on the comparison results, the platform automatically determines the user's respiratory event type, records the similarity, duration, and frequency of snoring and apnea events, and automatically generates a health risk assessment report based on the event type.

[0076] This embodiment accurately identifies multiple sound event types by comparing the extracted sound features with reference sound features in the feature library. By dynamically adjusting the similarity threshold, the comparison results can be optimized based on the user's individual characteristics and historical data, reducing the false positive rate. The recording of duration and frequency ensures that event determination is not only based on feature similarity, but also quantifies the severity of the event. Compared with existing technologies, it can dynamically adapt to multiple event types, automatically determine the event type, and generate preliminary determination results.

[0077] S50, determining an event type corresponding to the target sound segment based on the comparison result;

[0078] In this embodiment, determining the event type of a target sound segment based on sound feature comparison results involves parsing and judging the resulting data to determine the specific event type, such as snoring, apnea, or gasping, that the target sound segment belongs to. The core of this step is to select the most representative event type from multiple comparison results and ensure the stability and accuracy of the event type determination.

[0079] First, the similarity value, single duration, and number of occurrences per unit time for each preset event type are extracted from the comparison results. The similarity value measures the similarity between the sound signature and each event type in the reference signature library, with higher values ​​indicating a closer match. The single duration represents the duration of each event type in a single occurrence, such as snoring lasting 10 seconds or apnea lasting 15 seconds. The number of occurrences per unit time counts the cumulative number of occurrences of the event type within a fixed time window, for example, snoring occurring 20 times per hour.

[0080] Next, three ratios are calculated for each event type: the ratio of the similarity value to the dynamically adjusted threshold for the corresponding event type, the ratio of the single duration to the preset single duration threshold, and the ratio of the number of occurrences per unit time to the preset number threshold. These ratios quantify the sound signal's performance in these three aspects, ensuring that event type determination is based not only on similarity but also on time and frequency. These three ratios for each event type reflect its proximity to the reference event type and its actual occurrence characteristics.

[0081] After obtaining the three ratios for each event type, they are weighted and summed according to preset weighting factors to generate a composite score. The weighting factors can be flexibly adjusted based on the application scenario. For example, in snoring recognition, the similarity weight can be set to 0.5, the single duration weight to 0.3, and the number of occurrences per unit time to 0.2. The weighting ensures that the composite score accurately reflects the overall match of the event type.

[0082] All event types are sorted in descending order based on their combined scores, and the event type with the highest score is selected as the candidate event type. This step ensures that the most likely event type is selected from multiple possible event types. To further improve accuracy, the combined score of a candidate event type must exceed a preset score threshold before it is finally confirmed. If the combined score of a candidate event type does not meet the score threshold, the target sound clip is determined to be an "invalid event" or "unrecognized event."

[0083] Once the target event type is determined, the system automatically generates a determination record containing the event type and its occurrence time. The determination record should include the event type name, the start and end timestamps of the target sound clip, and the overall score, ensuring that the determination of the event type is traceable and data-supported.

[0084] In one embodiment, the system extracts the similarity values, single durations, and number of occurrences per unit time of three event types: snoring, apnea, and gasping from the comparison results. The similarity threshold is dynamically adjusted, with the similarity threshold for snoring events being 0.85 and that for apnea events being 0.90. The single duration threshold is set to 10 seconds, and the frequency threshold per unit time is set to 5 times per hour. The system calculates the three ratios for each event type in turn, and weights them up according to weight coefficients of 0.5, 0.3, and 0.2 to generate a comprehensive score. All event types are sorted according to the comprehensive score, and the event type with the highest score is selected as the candidate event type. When the comprehensive score of a candidate event type exceeds the score threshold of 0.7, the system determines it as the event type of the target sound segment.

[0085] In another embodiment, the weight coefficient is dynamically adjusted based on the user's historical data. For example, if the user snores frequently, the weight of the number of occurrences per unit time of the snoring event is increased to 0.4, and the weights of other event types are adjusted accordingly. The similarity threshold and score threshold can also be dynamically set based on the user's individual characteristics. For example, when the user sleeps in a quiet environment, the similarity threshold is set to 0.85, while in an environment with background noise, the similarity threshold is automatically increased to 0.90.

[0086] Machine learning models can also automatically determine event types. The system uses the similarity, duration, and frequency in the comparison results as feature vectors and feeds them into a machine learning model (such as a random forest or support vector machine). The model then automatically determines the event type based on pre-trained sample data. Machine learning models can gradually improve event recognition accuracy through continuous user data updates.

[0087] This embodiment extracts similarity, duration, and frequency from the comparison results, and calculates a comprehensive score based on the numerical values ​​of these three aspects, ensuring that the determination of the event type is not only based on the similarity of the sound features, but also takes into account the persistence and frequency of the sound event. The weighted summation of the comprehensive score provides a flexible event type determination mechanism that can adjust the weight coefficient according to the application scenario and adapt to different types of sound events. By dynamically adjusting the threshold and score threshold, the recognition results of the event type can be dynamically optimized based on the user's individual characteristics and historical data. Compared with the existing technology, it can effectively reduce misjudgments and missed judgments, and improve the accuracy and flexibility of event recognition.

[0088] S60: When the event type meets a preset condition, execute a processing operation corresponding to the event type.

[0089] In this embodiment, the processing operation corresponding to the event type is performed based on the event type (such as snoring, apnea, or gasping) identified in the target sound segment, triggering a corresponding automated response. This response can be a prompt, warning, device control, or cloud data transmission, depending on the event type and the user's preset processing strategy.

[0090] First, the system retrieves the corresponding processing strategy from the early warning strategy library based on the event type. The early warning strategy library is a structured rule set in which each event type (such as snoring, apnea) is associated with a set of processing strategies. Each processing strategy contains three core parameters: execution action type, target device identifier, and execution time constraint. The execution action type defines the operation that should be triggered in response to the event, such as activating a ventilator, sending a mobile phone reminder, or recording an event log. The target device identifier specifies the execution target of the response operation, such as smart health headphones, smartphones, or cloud servers. The execution time constraint limits the triggering time period of the processing operation, for example, it is only valid from 22:00 at night to 6:00 the next day.

[0091] After retrieving the processing strategy, the system first verifies that the start and end timestamps of the target sound segment fall within the preset validity period. This time verification ensures that processing actions are triggered only within the predetermined timeframe, preventing them from being triggered by events during the day or other non-monitoring periods. For example, a snoring event during nighttime monitoring could trigger a smart headset to automatically play a notification tone, while ignoring it during daytime.

[0092] If the time verification passes, the system generates a handling level identifier based on the preset severity level corresponding to the event type. Severity levels quantify the importance and urgency of the event, such as snoring being classified as medium severity (warning level) and apnea being classified as high severity (emergency level). Severity levels can be automatically assigned based on the potential harm of the event type or adjusted based on user-defined configurations. Severity levels directly influence the priority and response method of subsequent handling operations.

[0093] After determining the processing level, the system automatically generates a control instruction containing the action type, the target device identifier, the start and end timestamps of the target sound clip, and the processing level identifier. A control instruction is a structured data message that defines the specific action to take in response to the event, such as sending a notification, activating a device, or uploading data. Control instructions can use various formats, such as JSON, XML, or binary, depending on the requirements of the device and communication protocol.

[0094] Finally, the system transmits control commands to the target device via a wireless communication protocol (such as Bluetooth, Wi-Fi, or cellular networks). The choice of wireless communication protocol can be automatically adapted based on the device type and communication environment, such as prioritizing Wi-Fi in a home environment and Bluetooth for portable devices. Upon receiving the control command, the target device automatically performs the corresponding processing, such as playing a prompt tone, adjusting ventilator parameters, or recording an event log.

[0095] In one embodiment, when a user's smart health headset detects a snoring event, the system retrieves a snoring processing strategy from the early warning strategy library: the execution action type is to play a reminder tone, the target device is the smart health headset, and the execution time constraint is between 10:00 PM and 6:00 AM the next day. The system verifies that the timestamp of the snoring event is within the valid time period and generates a medium severity processing level. The system generates a control instruction in JSON format, which includes a command to play the reminder tone and a target device identifier. The control instruction is sent to the smart health headset via Bluetooth, and the headset automatically plays a reminder tone to remind the user to adjust their sleeping position.

[0096] In another embodiment, when a user's smart health insurance app detects an apnea event, the system retrieves an apnea handling strategy from the early warning strategy library: the execution action type is to activate the smart ventilator, the target device is the user's smart health ventilator, and the execution time constraint is 24 hours a day. The system verifies that the event timestamp meets the valid time period and generates a high severity handling level. The system generates a control instruction in XML format, which includes the command to activate the ventilator and the target device identifier. The control instruction is sent to the smart health ventilator via Wi-Fi, and the ventilator automatically adjusts the ventilation pressure to ensure smooth breathing for the user.

[0097] Multi-device collaborative processing operations can also be achieved through the cloud platform. Users wear smart headphones and smart watches to monitor sound and blood oxygen data at the same time. When a combined event of snoring and decreased blood oxygen saturation is detected, the system calls a joint processing strategy from the early warning strategy library: play a headphone prompt tone, send a mobile phone notification, and upload data to the cloud. The control instructions are divided into three parts, which are sent to the headphones via Bluetooth, sent to the mobile phone via 5G network, and uploaded to the cloud platform via Wi-Fi. The cloud platform automatically generates a health report and pushes it to the user.

[0098] This embodiment enables real-time response to sound events by automatically triggering corresponding processing operations based on the event type. The early warning policy library ensures the flexibility and adaptability of processing operations, and time verification avoids false triggering during non-monitoring time periods. The processing level identifier quantifies the severity of the event, ensuring that critical events are responded to first. The automatic generation and wireless transmission of control instructions ensure that processing operations can quickly reach the target device. Compared with existing technologies, it can achieve multi-device collaboration, automated response, and personalized policy configuration, significantly improving the accuracy and intelligence of event response.

[0099] The present invention relates to the field of data analysis technology and can be applied to business scenarios such as financial technology and medical health. It discloses an event processing method, device, equipment and medium based on sound signal analysis, including: obtaining a sound signal to be analyzed, performing segmentation processing on the sound signal to be analyzed to obtain a target sound segment; extracting sound features of the target sound segment, comparing the sound features with reference sound features in a preset feature library, and generating a comparison result; determining the event type corresponding to the target sound segment based on the comparison result, and when the event type meets the preset conditions, executing the processing operation corresponding to the event type. The present invention can accurately extract features of different events such as snoring, apnea and gasping from a variety of sound events through segmentation processing and feature extraction of sound signals; effectively identify event types by comparing the extracted sound features with the reference feature library; and automatically generate and execute corresponding processing operations based on the event type, thereby realizing automatic recognition and adaptive response processing of multiple event types, and improving the accuracy of sound event recognition and the flexibility of response.

[0100] In one embodiment, the above step S10 includes:

[0101] S101, collecting original acoustic signals from the throat of the target subject through a wearable device;

[0102] S102, filtering the original acoustic signal using a noise reduction module of the wearable device to generate a noise-reduced acoustic signal;

[0103] S103, sampling the noise-reduced acoustic signal according to a preset sampling frequency to generate a sampled acoustic signal;

[0104] S104: Send the sampled acoustic signal as a sound signal to be analyzed to a mobile terminal through the wireless transmission module of the wearable device.

[0105] In this embodiment, acquiring the sound signal to be analyzed is the basis of event analysis, ensuring that the collected sound signal can truly and stably reflect the sound characteristics of the target object. This process generally includes four stages: acquisition, noise reduction, sampling, and wireless transmission.

[0106] During the acquisition phase, sound signals are collected through wearable devices. Wearable devices are devices that are easy for users to wear daily and can continuously collect sound signals. Common forms include smart earphones, neck-mounted devices, headphones, and smart watches. Wearable devices should be equipped with a high-sensitivity microphone and installed in the throat of the target subject. Sound signal collection in the throat area can effectively avoid interference from external environmental noise and accurately capture the user's sound characteristics, such as snoring, breathing, and speech. The choice of the throat area can reduce the attenuation of oral and nasal breathing sounds and enhance the ability to capture high-frequency and low-frequency sound signals.

[0107] During the noise reduction phase, the wearable device's noise reduction module filters the original acoustic signal. This module can employ both active and passive noise reduction techniques. Active noise reduction eliminates ambient noise by picking up ambient noise and generating sound waves with opposite phases to counteract it. Passive noise reduction reduces interference from external noise through physical isolation, such as a windscreen or silicone patch on the microphone. Noise reduction processing can significantly improve the signal-to-noise ratio of the sound signal, ensuring that the collected sound signal is clear and stable. In practice, the noise reduction algorithm can be based on adaptive filtering, spectral subtraction, or frequency domain filtering to ensure high-quality sound signals even in complex acoustic environments.

[0108] During the sampling phase, the system samples the noise-reduced sound signal at a preset acquisition frequency. The acquisition frequency is the number of sound data points collected per second and is usually determined based on the characteristics of the sound signal. The commonly used acquisition frequency for speech signals is 8000 Hz to 48000 Hz. The higher the acquisition frequency, the clearer the details of the sound signal are captured. The sampling bit depth also needs to be set reasonably, such as 16 bits or 24 bits. The higher the bit depth, the greater the dynamic range of the sound signal. During the sampling process, the noise-reduced sound signal is divided into equally spaced sampling points, and each sampling point represents the amplitude of the sound signal at that moment.

[0109] During the wireless transmission stage, the sampled sound signal is sent to the mobile terminal through the wireless transmission module of the wearable device. The wireless transmission module can use Bluetooth, Wi-Fi or cellular network for data transmission. Bluetooth is suitable for short-range real-time transmission, especially for headphones or neck-mounted devices; Wi-Fi is suitable for high-bandwidth, high-speed transmission, such as home health monitoring systems; cellular networks are suitable for remote data upload, such as cloud analysis in health monitoring applications. The wireless transmission process needs to ensure the stability and integrity of the data. Error checking technology (such as CRC check) and end-to-end encryption technology (such as AES encryption) can be used to ensure data security.

[0110] This embodiment collects sound signals from the throat through a wearable device, which can effectively reduce the interference of environmental noise and ensure the clarity and accuracy of the sound signal. The noise reduction module ensures that the target sound (such as snoring and breathing) in the sound signal is clear and stable, while the high acquisition frequency and reasonable bit depth in the sampling stage can accurately preserve the frequency and energy characteristics of the sound signal. The wireless transmission module ensures that the sound signal can be transmitted to the mobile terminal in real time, realizing seamless real-time monitoring and data recording.

[0111] In one embodiment, the above step S20 includes:

[0112] S201, dividing the acoustic signal to be analyzed into multiple segments according to a segmented time window of a preset fixed length;

[0113] S202, adding a segment identifier including a start timestamp and an end timestamp to each segment;

[0114] S203, performing overlapping frame processing on the segments with the segment identifiers to generate target sound segments;

[0115] S204: Store the target sound segment as an independent audio file and associate it with a corresponding segment identifier.

[0116] In this embodiment, the sound signal to be analyzed is segmented to divide the continuous sound signal into multiple independent, structured sound segments to facilitate subsequent feature extraction and event recognition. This processing process includes four core steps: segmentation, identification, overlapping framing, and storage.

[0117] First, the sound signal to be analyzed is divided into multiple segments according to a preset segmented time window of a fixed length. The segmented time window is a periodic time interval for segmenting the sound signal, such as 30 seconds, 1 minute or 5 minutes. The setting of the segmented time window should be flexibly adjusted according to the characteristics of the target event and the monitoring needs. For example, the event recognition of snoring and apnea in sleep monitoring is suitable for a 30-second time window, while speech recognition may require a shorter window. The core of segmented processing is to ensure that the sound signal within each segment remains continuous, and the segments between segments can be processed independently. The segmented time window can be adjusted dynamically, for example, a 30-second window is used for night monitoring and a 1-minute window is used for daytime monitoring.

[0118] After each segment is generated, a segment identifier containing a start timestamp and an end timestamp is added to each segment. The timestamps are the start and end times of each segment in the sound signal, usually expressed with millisecond accuracy. This time identifier ensures the traceability of the timing of each segment and allows the accurate location of each segment in subsequent analysis. Timestamps can be automatically generated based on the system time or the timeline of the sound signal acquisition. The start timestamp indicates the start time of the segment, and the end timestamp indicates the end time of the segment. For example, a 30-second segment starting at 22:00:00 and ending at 22:00:30 can be represented as "22:00:00-22:00:30".

[0119] After adding the timestamp, the segments with segment identifiers are subjected to overlapping framing. Overlapping framing is to further divide each segment into multiple short time windows (frames). Each frame is usually 20 milliseconds to 50 milliseconds. There is partial overlap between frames (such as 50% overlap), which can enhance the continuity of sound features. The purpose of overlapping framing is to avoid the loss of sound signal features due to fixed window segmentation, especially when the sound event crosses the segment boundary. The frame length and frame shift (overlap rate) can be adjusted according to the characteristics of the sound signal and the type of target event. For example, for snoring events, a 25 millisecond frame length and a 50% overlap rate can be used, while for speech signals, a 20 millisecond frame length and a 75% overlap rate can be used. Each framed sound segment retains the temporal characteristics of the original sound signal to ensure the stability of subsequent feature extraction.

[0120] Finally, the segmented target sound segments are stored as independent audio files and associated with corresponding segment identifiers. The storage formats of independent audio files can include WAV, PCM, or MP3, and the choice of format should be flexibly adjusted according to data analysis requirements and storage space limitations. The naming of independent audio files should be associated with the segment identifier, for example, the file name can contain a start timestamp and an end timestamp. The segment identifier and other auxiliary information, such as the acquisition device identifier, sampling frequency, and acquisition location, can also be recorded in the metadata of the audio file. The structured storage of independent audio files facilitates subsequent feature extraction and event analysis, and can ensure the traceability and manageability of data during multiple monitoring processes.

[0121] This embodiment can ensure the structuring and continuity of the sound signal by segmenting the sound signal to be analyzed according to segmented time windows of fixed length, which facilitates subsequent feature extraction and event identification. The dynamic adjustment of the segmented time window ensures that the segmentation can adapt to different monitoring scenarios, such as short-term segmentation at night to improve the sensitivity of event detection, and long-term segmentation during the day to reduce the storage burden. By adding a start timestamp and an end timestamp to each segment, the time positioning and tracing of the sound segment can be achieved. Overlapping frame processing ensures the continuity of the sound signal characteristics and avoids the edge effect caused by fixed segmentation. Storing the segmented sound segments as independent audio files facilitates data management, and the segment identifiers recorded in the file can ensure the traceability of the data.

[0122] In one embodiment, the above step S30 includes:

[0123] S301, performing pre-emphasis filtering on the target sound segment to generate a pre-emphasis sound signal;

[0124] S302, performing frame processing on the pre-emphasized sound signal using a Hamming window to generate a framed sound signal;

[0125] S303, extracting Mel-frequency cepstral coefficients from the framed sound signal as a first sound feature;

[0126] S304, extracting short-time energy from the framed sound signal as a second sound feature;

[0127] S305: Perform normalization fusion processing on the first sound feature and the second sound feature to generate the sound feature of the target sound segment.

[0128] In this embodiment, extracting the sound features of the target sound segment is the core process of analyzing the sound signal. Using various feature extraction methods, the sound signal is converted into a numerical feature vector, facilitating subsequent event recognition and classification. This process includes four stages: pre-emphasis filtering, framing, feature extraction, and feature fusion.

[0129] First, the target sound segment is pre-emphasized and filtered to generate a pre-emphasized sound signal. Pre-emphasis filtering is a high-frequency enhancement technology. Its core principle is to enhance high-frequency components through filters to offset the attenuation of high-frequency energy during the acquisition and transmission of the voice signal. The pre-emphasis filter is a first-order high-pass filter. Its formula is usually expressed as follows: the output signal is the input signal minus its previous sampling point multiplied by the pre-emphasis coefficient. The pre-emphasis coefficient is usually between 0.95 and 0.98. The larger the value, the more obvious the high-frequency enhancement effect. In specific implementation, the pre-emphasis filter can be applied in real-time signal processing through software algorithms or implemented at the microphone input through hardware circuits. Pre-emphasis processing can improve the signal-to-noise ratio of high-frequency components and is particularly suitable for capturing high-frequency features in snoring and breathing signals.

[0130] After generating the pre-emphasized sound signal, the system uses a Hamming window to perform frame processing on the pre-emphasized sound signal, generating a framed sound signal. Framing involves segmenting the continuous sound signal into multiple short-duration windows (frames). Each frame is typically 20 to 50 milliseconds long, with some overlap between frames (e.g., 50%). Framing aims to break the sound signal into multiple stable, short-duration segments, each of which can be considered a nearly stable signal. The Hamming window is a weighted window function that gradually decays to zero at both ends of each frame, effectively reducing boundary effects between frames and improving the stability of feature extraction. The Hamming window's weight coefficient is largest at the center of the window function and gradually decreases to zero toward the ends, ensuring a smooth framed signal. The frame length and frame rate can be flexibly adjusted based on the characteristics of the sound signal. For example, a 25 millisecond frame length and a 50% overlap rate can be used for snoring and breathing signals.

[0131] Mel-frequency cepstral coefficients (MFCCs) are extracted from the framed sound signal as the first sound feature. Mel-frequency cepstral coefficients are a feature extraction method widely used in speech recognition and sound signal analysis. Its core principle is to convert the sound signal from the time domain to the frequency domain and perform nonlinear mapping of the frequency based on the Mel scale. The specific implementation includes the following steps: first, a fast Fourier transform (FFT) is performed on each framed sound signal to obtain a frequency domain representation; then, the frequency domain signal is mapped to the Mel frequency scale through a Mel filter bank. This frequency scale is more consistent with the human ear's perception of sound frequency; then, the Mel spectrum is logarithmically transformed and converted into a logarithmic Mel spectrum; finally, a discrete cosine transform (DCT) is performed on the logarithmic Mel spectrum to generate a set of Mel-frequency cepstral coefficients. Usually 12 to 13 cepstral coefficients are taken as sound features, which can effectively characterize the spectral structure of the sound signal.

[0132] In addition to MFCC, short-time energy is extracted from the framed sound signal as the second sound feature. Short-time energy is a feature that measures the energy changes of the sound signal and can effectively distinguish the changes in the strength of the sound signal, such as snoring, breathing and background noise. Short-time energy is calculated by squaring the amplitude of the sound signal in each frame and summing them. Frames with high short-time energy indicate high sound signal intensity, such as snoring or gasping, while frames with low short-time energy indicate background noise or silence. Short-time energy can effectively reflect the energy distribution of the sound signal and plays an important role in identifying respiratory events, especially in identifying respiratory events (such as apnea and gasping).

[0133] Finally, the extracted first sound feature (MFCC) and the second sound feature (short-time energy) are normalized and fused to generate the sound feature of the target sound segment. Normalization is to convert the feature value into a fixed range (such as 0 to 1) to ensure that each feature has the same numerical scale. Normalization methods can include minimum-maximum normalization (linearly mapping the feature value to the range of 0-1) or Z-score normalization (adjusting the feature value to a distribution with a mean of 0 and a standard deviation of 1). The normalized features can be generated by feature splicing (directly combining MFCC and short-time energy into a vector) or weighted fusion (setting weights according to feature importance) to generate the final sound features. Feature fusion can ensure that the sound features are fully characterized in terms of both frequency and energy, thereby improving the accuracy of subsequent event recognition.

[0134] This embodiment enhances the high-frequency components of the sound signal through pre-emphasis filtering to ensure that the high-frequency details of the sound features are preserved. Hamming window framing can reduce the boundary effect between frames and improve the stability of feature extraction. Mel-frequency cepstral coefficients (MFCC) ensure that the frequency characteristics of the sound signal can be accurately characterized, while short-time energy provides the energy characteristics of the sound signal. The combination of the two can simultaneously characterize the frequency and energy distribution of the sound signal. Normalization fusion ensures that the feature values ​​are on a unified scale to avoid the difference in feature values ​​affecting the subsequent recognition accuracy.

[0135] In one embodiment, the above step S40 includes:

[0136] S401, retrieving multiple reference sound features from a voiceprint feature library, each reference sound feature being associated with a preset respiratory event type;

[0137] S402, respectively determining a similarity value between the sound feature and each reference sound feature;

[0138] S403, recording the single duration and the number of occurrences per unit time of each respiratory event type whose similarity value exceeds the current matching similarity threshold;

[0139] S404, determining whether the single duration of the same respiratory event type exceeds a preset single duration threshold, and whether the number of occurrences per unit time exceeds a preset number threshold;

[0140] S405, when the single duration of the same respiratory event type exceeds a preset single duration threshold, and the number of occurrences per unit time exceeds a preset number threshold, generating a preliminary determination result of the respiratory event type;

[0141] S406 , summarizing preliminary determination results of all respiratory event types to generate a final comparison result including event type, occurrence time, and severity level.

[0142] In this embodiment, the core process of identifying sound events is to compare sound features with reference sound features in a preset feature library. Accurately identifying respiratory events is achieved through feature matching and dynamic threshold determination. This process includes four stages: calling reference sound features, calculating similarity, determining thresholds, and generating results.

[0143] First, multiple reference sound features are retrieved from the voiceprint feature library, and each reference sound feature is associated with a preset respiratory event type. The voiceprint feature library is a structured database that contains multiple types of reference sound features, such as snoring, apnea, and gasping. Each reference sound feature represents a typical respiratory event pattern, and these reference features can be extracted through a pre-trained sound model or from a large amount of sample data. For example, a snoring reference feature may contain a sound feature with enhanced low-frequency intensity, while an apnea reference feature is a sound pattern that recovers quickly after a short period of silence. The reference sound feature library can be updated regularly and adaptively adjusted according to the user's individual sound characteristics to ensure the adaptability and accuracy of sound event recognition.

[0144] After retrieving the reference sound features, the system determines the similarity value between the sound features and each reference sound feature respectively. Similarity is a measure of the similarity between the sound feature and the reference sound feature in the feature space, and is usually calculated using cosine similarity, Euclidean distance or dynamic time warping (DTW) algorithm. Cosine similarity can measure the angle between two feature vectors. The closer the value is to 1, the higher the similarity. Euclidean distance measures the absolute distance between two feature vectors. The smaller the value, the higher the similarity. Dynamic time warping (DTW) is applicable to time series signals, and the optimal alignment path of the feature sequence is calculated through a dynamic programming algorithm. During the similarity calculation process, each sound feature is compared with all reference sound features to generate a set of similarity values, and each similarity value corresponds to a reference sound feature.

[0145] After calculating the similarity, the system records the single duration and number of occurrences per unit time of the similarity value of each respiratory event type exceeding the current matching similarity threshold. The single duration refers to the length of time that the continuous similarity value exceeds the threshold, and the number of occurrences per unit time refers to the number of times the similarity value exceeds the threshold within the preset time window. The similarity threshold is a dynamically adjusted threshold that can be adaptively updated based on the user's voice characteristics and historical data. For example, the similarity threshold for snoring can be set to 0.8, while the similarity threshold for apnea can be set to 0.9. The system automatically records the similarity duration and number of occurrences for each event type to ensure accurate statistics of events.

[0146] After recording the similarity information, the system determines whether the single duration of the same respiratory event type exceeds the preset single duration threshold, and whether the number of occurrences per unit time exceeds the preset number threshold. The single duration threshold is the minimum duration of each respiratory event type. For example, snoring can be set to 2 seconds, and apnea can be set to 10 seconds. The number of occurrences per unit time threshold is the minimum number of times each event occurs within a fixed time window (such as 1 hour), for example, snoring is 5 times per hour, and apnea is 10 times per hour. This dual condition ensures that event recognition is not only based on the single duration, but also takes into account the frequency of event occurrence, avoiding single short-term similarity peaks being misjudged as valid events.

[0147] When the duration of a single respiratory event of the same type exceeds the preset single duration threshold, and the number of occurrences per unit time exceeds the preset number threshold, the system generates a preliminary determination result for that respiratory event type. The preliminary determination result effectively identifies the event type in the current sound signal and includes the event type, similarity value, duration, and number of occurrences. This preliminary determination result forms the basis for subsequent event identification and analysis, ensuring that each event type is fully validated.

[0148] Finally, the system summarizes the preliminary determination results of all respiratory event types and generates a final comparison result that includes the event type, occurrence time, and severity level. The severity level indicates the potential harmfulness or urgency of the event. For example, snoring can be moderately severe (warning level) and apnea can be highly severe (emergency level). The severity level can be automatically assessed based on the event type, single duration, and number of occurrences, or it can be adjusted according to the user's customized strategy. The final comparison result is the final output of event recognition, which facilitates subsequent processing operations or user notifications.

[0149] This embodiment automatically identifies various respiratory events (such as snoring, apnea, and gasping) by comparing sound signatures with reference sound signatures in a preset signature library. A dynamic similarity threshold and dual conditions (single duration and number of occurrences) ensure accurate and flexible event recognition. By summarizing multiple event types based on preliminary determination results, the severity of each event can be accurately assessed and structured comparison results generated.

[0150] In one embodiment, the above step S50 includes:

[0151] S501, extracting the similarity value, single duration, and number of occurrences per unit time of each preset respiratory event type from the comparison result;

[0152] S502, respectively determining a first ratio of a similarity value of each preset respiratory event type to a corresponding matching similarity threshold;

[0153] S503, respectively determining a second ratio of a single duration of each preset respiratory event type to a preset single duration threshold;

[0154] S504, respectively determining a third ratio of the number of occurrences per unit time of each preset respiratory event type to a preset number threshold;

[0155] S505, for each preset respiratory event type, performing a weighted summation on the first ratio, the second ratio, and the third ratio according to a preset weight coefficient to generate a comprehensive score;

[0156] S506, sorting all preset respiratory event types in descending order according to the comprehensive scores, and selecting the type with the highest score as the candidate event type;

[0157] S507, when the comprehensive score of the candidate event type exceeds a preset score threshold, determining the event type of the target sound clip as the candidate event type;

[0158] S508: Generate a determination record including the event type, the start timestamp, and the end timestamp of the target sound segment.

[0159] In this embodiment, the event type corresponding to the target sound clip is determined based on the comparison results. This process selects the most likely event type from multiple candidate events, ensuring an accurate and adaptable screening process. This process calculates the ratio of three features: similarity, duration, and number of occurrences, and then uses a weighted score to achieve a multi-dimensional assessment of event types.

[0160] First, the similarity value, single duration and number of occurrences per unit time of each preset respiratory event type are extracted from the comparison results. The comparison results are the matching results of the sound features with multiple event types in the reference sound feature library. Each event type such as snoring, apnea and gasping has a corresponding similarity value. The similarity value indicates the degree of similarity between the sound feature and the reference sound feature. The higher the value, the higher the similarity. The single duration indicates the length of time that the similarity value exceeds the threshold continuously, for example, the single duration of snoring is 3 seconds and apnea is 15 seconds. The number of occurrences per unit time indicates the cumulative number of occurrences of this event type within a specific time window (such as 1 hour), for example, snoring occurs 8 times per hour and apnea occurs 12 times per hour. Extracting these data is the basis for event type recognition, ensuring that the core features of each event type are captured.

[0161] After extracting the event features, the system determines the first ratio of the similarity value of each preset respiratory event type to the corresponding matching similarity threshold. The similarity threshold is the minimum matching standard for each event type. For example, the similarity threshold for snoring can be set to 0.8, and the similarity threshold for apnea can be set to 0.9. The first ratio is the ratio of the similarity value to the threshold, indicating the degree of deviation of the similarity of the current event from the standard. A ratio greater than 1 indicates that the similarity is higher than the standard, and a ratio less than 1 indicates that the similarity is lower than the standard. The first ratio can quantify the effectiveness of the event type in feature matching.

[0162] Next, the system determines the second ratio of the single duration of each preset respiratory event type to the preset single duration threshold. The single duration threshold is the minimum valid duration of each event type. For example, the single duration threshold for snoring is 2 seconds, and the single duration threshold for apnea is 10 seconds. The second ratio indicates the degree of deviation of the current event duration from the standard. A ratio greater than 1 indicates that the duration exceeds the standard, and a ratio less than 1 indicates that the duration is insufficient. The second ratio can measure the effectiveness of the event duration and ensure that short events (such as 1-second snoring) are not mistakenly judged as valid events.

[0163] Then, the system determines a third ratio of the number of occurrences per unit time for each preset respiratory event type to the preset number threshold. The number of occurrences per unit time represents the cumulative number of occurrences of the event within a specific time window (such as 1 hour), and the number threshold is the minimum frequency standard for each event type, such as 5 times per hour for snoring and 10 times per hour for apnea. The third ratio represents the degree of deviation of the number of occurrences of the event from the standard. A ratio greater than 1 indicates that the number of occurrences exceeds the standard, and a ratio less than 1 indicates that the number of occurrences is insufficient. The third ratio can measure the effectiveness of the frequency of event occurrences.

[0164] After calculating the three ratios, the system weights and sums these three ratios for each preset respiratory event type according to the preset weight coefficient to generate a composite score. The weight coefficient is used to balance the importance of the three features: similarity, duration, and number of occurrences. For example, for snoring events, the similarity weight can be set to 0.5, the duration weight to 0.3, and the number of occurrences weight to 0.2. The weighted sum ensures that different features have an appropriate influence in the overall score, avoiding incorrect judgments due to excessively high values ​​of a single feature. The higher the composite score, the more reliable the recognition of the event type.

[0165] Based on the combined scores, the system sorts all pre-defined respiratory event types in descending order and selects the type with the highest score as the candidate event type. This descending sorting ensures that the highest-scoring event types are selected first, while lower-scoring types are automatically excluded. A candidate event type is the most likely event type for the current sound clip, but its final confirmation requires further judgment.

[0166] When the comprehensive score of the candidate event type exceeds the preset score threshold, the system determines that the event type of the target sound clip is a candidate event type. The score threshold is the final standard for event recognition. For example, if it is set to 0.8, it means that the comprehensive score must reach 80% to confirm the event type. The score threshold can be dynamically adjusted according to the user's risk preference and monitoring scenario. A higher threshold (such as 0.85) can be set for night monitoring, and a lower threshold (such as 0.75) can be set for daytime monitoring. When the comprehensive score does not reach the threshold, the system will not confirm the event type to avoid misidentification.

[0167] Finally, the system generates a judgment record containing the event type and the start and end timestamps of the target sound segment. This judgment record, representing the result of event recognition, accurately records the event type and time of occurrence, facilitating subsequent processing and analysis. This record can be stored locally or in a cloud-based platform in a structured format. The record format can include the event type, timestamp, composite score, and three feature ratios to ensure data traceability.

[0168] This embodiment uses the ratio calculation of three features based on the comparison results to perform a multi-dimensional assessment of each respiratory event type, ensuring accurate event type identification. A weighted summation method allows for flexible configuration of different features to accommodate a variety of monitoring scenarios. Comprehensive score sorting and threshold determination enable precise screening of event types, preventing misidentification.

[0169] In one embodiment, the above step S60 includes:

[0170] S601, retrieve a corresponding processing strategy from a warning strategy library according to the event type, wherein the processing strategy includes an execution action type, a target device identifier, and an execution time constraint;

[0171] S602, verifying whether the start timestamp and the end timestamp of the target sound segment are within the valid period defined by the execution time constraint;

[0172] S603, when the time verification passes, generating a processing level identifier according to a preset severity level corresponding to the event type;

[0173] S604, generating a control instruction including the execution action type, the target device identifier, the start timestamp and the end timestamp of the target sound segment, and a processing level identifier;

[0174] S605: Send the control instruction to the target device corresponding to the target device identifier through a wireless communication protocol, and execute a corresponding processing operation.

[0175] In this embodiment, when an event type meets preset conditions, the corresponding processing operation is executed, aiming to automatically trigger appropriate treatment measures based on the event type's risk level. This process includes five stages: early warning strategy retrieval, time validity verification, processing level generation, control instruction generation, and control instruction delivery.

[0176] First, the corresponding processing strategy is retrieved from the early warning strategy library according to the event type. The early warning strategy library is a structured database that contains preset response strategies for different event types. Each event type (such as snoring, apnea, and gasping) has a corresponding processing strategy, which defines the execution action type, target device identifier, and execution time constraint. The execution action type indicates the specific operation that the system should take when an event occurs, such as starting a ventilator, sending a warning notification, recording event data, or adjusting environmental parameters. The target device identifier is the device identifier that performs the processing operation, such as the device ID of a smart ventilator, the notification ID of a mobile phone application, or the control ID of a smart home device. The execution time constraint is the effective time range of the event processing operation, for example, only monitoring sleep events from 22:00 at night to 6:00 the next day. The early warning strategy library can be dynamically updated through the cloud to ensure the adaptability of the response strategy.

[0177] After calling the processing strategy, the system verifies whether the start timestamp and end timestamp of the target sound segment are within the valid period defined by the execution time constraint. The timestamp is the start and end time of the sound segment, indicating the time when the event occurred. Time verification ensures that the processing operation is triggered only within the specified time period, such as monitoring snoring or apnea at night, but not monitoring during the day. The logic of time verification is to compare the start and end time of the sound segment with the execution time constraint to ensure that the event occurs within the valid period. Events that fail time verification will be automatically ignored and no processing operations will be triggered. This time constraint can avoid false alarms and incorrect processing during unnecessary time periods (such as daytime).

[0178] When the time verification passes, the system generates a processing level identifier based on the preset severity level corresponding to the event type. The preset severity level is the risk level of each event type, such as snoring is low risk (warning level) and apnea is high risk (emergency level). The severity level can be automatically evaluated based on the characteristic parameters of the event type (such as duration and frequency), or it can be customized by the user. The processing level identifier is a level label represented by a number or character, such as "1" for the emergency level, "2" for the warning level, and "3" for the prompt level. The division of processing levels ensures that different risk events can trigger different response measures to avoid all events triggering the same processing operations.

[0179] After generating the processing level identifier, the system generates a control instruction containing the type of action to be performed, the target device identifier, the start and end timestamps of the target sound segment, and the processing level identifier. A control instruction is an operation command sent to the target device and is typically generated in a structured data format (such as JSON or XML). A control instruction contains the following fields:

[0180] Action type: such as "start ventilator", "send warning notification", "adjust ambient lighting";

[0181] Target device identifier: such as "smart ventilator 001", "user mobile phone application", "smart lighting system";

[0182] Timestamp: The start and end time of the sound clip, used for event time tracing;

[0183] Processing level identifier: Indicates the risk level of the event and determines the specific policy executed by the device.

[0184] Control instructions can be generated by a local computing device (such as a smartphone) or a cloud platform. The format of the control instructions should comply with the communication protocol requirements of the target device to ensure that the instructions can be correctly parsed and executed.

[0185] Finally, the control instruction is sent to the target device corresponding to the target device identifier through the wireless communication protocol, and the corresponding processing operation is performed. Wireless communication protocols may include Bluetooth, Wi-Fi, cellular networks (such as 4G / 5G) or Internet of Things protocols (such as MQTT, ZigBee). Control instructions can be sent through point-to-point communication (such as Bluetooth headsets controlling mobile phones) or from the cloud (such as cloud platforms controlling smart home devices). After the control instruction is sent, the target device performs the predetermined operation according to the instruction, such as starting the ventilator, sending health warnings from a mobile phone application, or dimming the lights of a smart lighting system. Wireless communication ensures that control instructions can be transmitted in real time between devices, enabling multi-device collaboration.

[0186] Example: In the healthcare field, a user wears smart health headphones while sleeping. The headphones use an integrated high-sensitivity microphone to collect the user's original acoustic signal. The headphones' noise reduction module automatically filters the original acoustic signal, removing ambient noise and background interference, to generate a noise-reduced acoustic signal. The headphones sample the noise-reduced acoustic signal at a sampling rate of 16,000 times per second to generate a sampled acoustic signal. Through the headphones' wireless transmission module, the sampled acoustic signal is transmitted in real time to the user's smartphone app as the sound signal to be analyzed. After receiving the sound signal to be analyzed, the smartphone app automatically segments it. The app segments the sound signal to be analyzed into fixed 30-second time windows, identifying each segment with a start and end timestamp. To ensure the continuity and feature integrity of the sound segments, the app performs overlapping framing (25 milliseconds frame length, 50% overlap) on the identified sound segments to generate target sound segments. Each target sound segment is stored as a separate audio file and associated with the corresponding segment identifier. The app extracts sound features from each target sound segment. First, the sound clip undergoes pre-emphasis filtering to enhance high-frequency components and generate a pre-emphasized sound signal. The application then frames the pre-emphasized sound signal using a Hamming window to ensure smooth transitions between frames. The application extracts Mel-Frequency Cepstral Coefficients (MFCCs) from the framed signal as the first sound feature, which describes the spectral characteristics of the sound. The application also extracts short-term energy from the framed signal as the second sound feature, which measures the energy variation of the sound signal. These two sound features are then normalized and fused to generate the final sound feature. The application then compares the sound feature with reference sound features in a pre-set feature library. This library includes three reference sound features: snoring, apnea, and gasping. The application calculates the similarity between the sound feature and each reference sound feature. For snoring, the similarity threshold is set to 0.8, and for apnea, to 0.9. The application records the duration of each event type whose similarity value exceeds the threshold and the number of occurrences per unit time. The system identifies snoring events as events that last longer than 2 seconds and occur more than 5 times per hour, and as events that last longer than 10 seconds and occur more than 10 times per hour. Events that meet the conditions generate preliminary judgment results. The application further determines the event type based on the comparison results. The system extracts the similarity, duration and number of occurrences of snoring and apnea from the preliminary judgment results. The similarity value, duration and number of occurrence ratio of each event type are calculated separately, and the weighted sum is calculated according to the preset weights (similarity 0.5, duration 0.3, number of occurrence 0.2) to generate a comprehensive score. The application sorts all event types in descending order according to the comprehensive score and selects the event type with the highest score as the candidate event type. When the comprehensive score of a candidate event type exceeds the score threshold of 0.8, the application determines that type as the final event type.When the event type meets the preset conditions (such as apnea is a high-risk event), the application automatically calls the corresponding processing strategy from the early warning strategy library. For apnea, the processing strategy called by the application is "Start ventilator", the target device identifier is "Smart ventilator 001", and the execution time constraint is 22:00-06:00. The application verifies whether the start timestamp and end timestamp of the sound clip are within the valid time period. When the time verification passes, the application generates a processing level identifier "1" based on the high-risk level of apnea, and generates a control instruction. The control instruction is sent to the smart ventilator via the Bluetooth wireless communication protocol, and the ventilator automatically starts to provide respiratory support. All processing operations and event recognition results are automatically recorded in the user's mobile phone application to generate a daily sleep health report.

[0187] In the financial sector, users use smartwatches to monitor their daily breathing signals through smart health insurance apps. The smartwatch collects acoustic signals from the user's throat and automatically filters ambient noise through a noise reduction module to generate a noise-reduced acoustic signal. The smartwatch samples the noise-reduced acoustic signal at a rate of 8,000 times per second and transmits it in real time to the user's smartphone app via Bluetooth. After receiving the sound signal for analysis, the smartphone app automatically segments it into fixed time windows of one minute, adding start and end timestamps to each segment. The app performs overlapping framing (20 millisecond frame length, 50% overlap) to generate target sound segments and stores them as separate audio files. The app extracts sound features from each target sound segment. It first performs pre-emphasis filtering to enhance high-frequency components. It then applies Hamming window framing to ensure smooth signal transitions. From the framed signal, the app extracts Mel-Frequency Cepstral Coefficients (MFCCs) as the first sound feature and short-term energy as the second sound feature. These two sound features are fused through normalization to generate the final sound feature. The application compares the sound signature with reference sound signatures in the insurance company's cloud-based signature library. The signature library includes three reference signatures: normal breathing, mild breathing abnormalities, and severe breathing abnormalities. The application calculates the similarity between the sound signature and each reference signature and records the duration of each event in which the similarity value exceeds the threshold and the number of occurrences per unit time for each event type. The application automatically determines whether the preset threshold conditions are met, such as the duration of a single severe breathing abnormality exceeding 15 seconds and the number of occurrences exceeding 8 per hour. The application determines the event type based on the comparison results. The application extracts the similarity, duration, and number of occurrences for each event type from the preliminary judgment results. The application calculates the ratio of the similarity, duration, and number of occurrences for each event type and sums them according to the weights (similarity 0.4, duration 0.4, and number of occurrences 0.2). The application selects the event type with the highest score as the candidate event type and determines the final event type when the combined score exceeds 0.75. When the event type meets the preset conditions (such as severe breathing abnormalities), the application automatically calls the processing strategy from the warning strategy library. The app invokes the "Send Warning Notification" policy, targeting the user's phone and valid all day. The app generates a control command and sends it to the user's phone via the mobile communication network. The user receives a warning notification: "Abnormal breathing detected. Please monitor your health." Simultaneously, the app uploads the event log and processing operations to the insurance company's cloud platform. The cloud platform dynamically adjusts health insurance premiums based on the user's respiratory health score. Insurance companies assess users' health risks based on their long-term health monitoring data and generate premium adjustment policies.

[0188] This embodiment dynamically calls the early warning strategy based on the event type, ensuring that each event type triggers the most appropriate processing operation. Time validity verification prevents invalid processing from being triggered during unnecessary time periods (such as daytime). Processing level identifiers ensure that high-risk events (such as apnea) are prioritized for processing, while low-risk events (such as snoring) are only recorded or notified. Control instruction generation and wireless communication ensure that multiple devices work together to respond to events in real time.

[0189] In one embodiment, an event processing device based on sound signal analysis is provided, and the event processing device based on sound signal analysis corresponds one-to-one to the event processing method based on sound signal analysis in the above embodiment. Figure 3 , Figure 3 This is a functional module diagram of a preferred embodiment of an event processing device based on sound signal analysis according to the present invention. It includes an acquisition module 10, a segmentation processing module 20, a feature extraction module 30, a feature comparison module 40, an event determination module 50, and an event response module 60. Each functional module is described in detail below:

[0190] Acquisition module 10, used to acquire the sound signal to be analyzed;

[0191] A segmentation processing module 20 is used to perform segmentation processing on the sound signal to be analyzed to obtain target sound segments;

[0192] A feature extraction module 30 is used to extract the sound features of the target sound segment;

[0193] A feature comparison module 40 is used to compare the sound feature with reference sound features in a preset feature library to generate a comparison result;

[0194] An event determination module 50 is configured to determine an event type corresponding to the target sound segment based on the comparison result;

[0195] The event response module 60 is configured to execute a processing operation corresponding to the event type when the event type meets a preset condition.

[0196] In one embodiment, the acquisition module 10 is specifically configured to:

[0197] Collecting original acoustic signals from the target subject's throat through a wearable device;

[0198] Performing noise filtering on the original acoustic signal using a noise reduction module of the wearable device to generate a noise-reduced acoustic signal;

[0199] Sampling the noise-reduced acoustic signal at a preset sampling frequency to generate a sampled acoustic signal;

[0200] The sampled acoustic signal is sent to a mobile terminal as a sound signal to be analyzed through the wireless transmission module of the wearable device.

[0201] In one embodiment, the segment processing module 20 is specifically configured to:

[0202] Dividing the acoustic signal to be analyzed into multiple segments according to a segmented time window of a preset fixed length;

[0203] Add a segment identifier containing a start timestamp and an end timestamp to each segment;

[0204] Performing overlapping frame processing on the segments with the segment identifiers to generate target sound segments;

[0205] The target sound segment is stored as an independent audio file and associated with a corresponding segment identifier.

[0206] In one embodiment, the feature extraction module 30 is specifically configured to:

[0207] Performing pre-emphasis filtering on the target sound segment to generate a pre-emphasis sound signal;

[0208] Performing frame processing on the pre-emphasized sound signal using a Hamming window to generate a framed sound signal;

[0209] extracting Mel-frequency cepstral coefficients from the framed sound signal as a first sound feature;

[0210] extracting short-time energy from the framed sound signal as a second sound feature;

[0211] Normalization and fusion processing is performed on the first sound feature and the second sound feature to generate the sound feature of the target sound segment.

[0212] In one embodiment, the feature comparison module 40 is specifically configured to:

[0213] Retrieving multiple reference sound features from a voiceprint feature library, each reference sound feature being associated with a preset respiratory event type;

[0214] Determining a similarity value between the sound feature and each reference sound feature respectively;

[0215] Record the single duration and the number of occurrences per unit time of each respiratory event type whose similarity value exceeds the current matching similarity threshold;

[0216] Determine whether the single duration of the same respiratory event type exceeds a preset single duration threshold, and whether the number of occurrences per unit time exceeds a preset number threshold;

[0217] When the single duration of the same respiratory event type exceeds a preset single duration threshold, and the number of occurrences per unit time exceeds a preset number threshold, a preliminary determination result of the respiratory event type is generated;

[0218] The preliminary determination results of all respiratory event types were summarized to generate the final comparison results including event type, occurrence time and severity level.

[0219] In one embodiment, the event determination module 50 is specifically configured to:

[0220] Extracting the similarity value, single duration, and number of occurrences per unit time of each preset respiratory event type from the comparison results;

[0221] Determine a first ratio of a similarity value of each preset respiratory event type to a corresponding matching similarity threshold value;

[0222] respectively determining a second ratio of a single duration of each preset respiratory event type to a preset single duration threshold;

[0223] respectively determining a third ratio of the number of occurrences per unit time of each preset respiratory event type to a preset number threshold;

[0224] For each preset respiratory event type, performing a weighted summation on the first ratio, the second ratio, and the third ratio according to a preset weight coefficient to generate a comprehensive score;

[0225] All preset respiratory event types are sorted in descending order according to the comprehensive scores, and the type with the highest score is selected as the candidate event type;

[0226] When the comprehensive score of the candidate event type exceeds a preset score threshold, determining the event type of the target sound segment as the candidate event type;

[0227] A determination record is generated, which includes the event type, a start timestamp, and an end timestamp of the target sound segment.

[0228] In one embodiment, the event response module 60 is specifically configured to:

[0229] Retrieving a corresponding processing strategy from a warning strategy library according to the event type, the processing strategy including an execution action type, a target device identifier, and an execution time constraint;

[0230] Verifying whether the start timestamp and the end timestamp of the target sound segment are within a valid period defined by the execution time constraint;

[0231] When the time verification passes, a processing level identifier is generated according to a preset severity level corresponding to the event type;

[0232] generating a control instruction including the execution action type, a target device identifier, a start timestamp and an end timestamp of the target sound segment, and a processing level identifier;

[0233] The control instruction is sent to the target device corresponding to the target device identifier through a wireless communication protocol, and a corresponding processing operation is performed.

[0234] In one embodiment, a computer device is provided. The computer device may be a server, and its internal structure diagram may be as follows: Figure 4 As shown. The computer device includes a processor, a memory, a network interface and a database connected via a system bus. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile and / or volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external user terminal via a network connection. When the computer program is executed by the processor, it realizes the functions or steps on the server side of an event processing method based on sound signal analysis.

[0235] In one embodiment, a computer device is provided. The computer device may be a user terminal, and its internal structure diagram may be as follows: Figure 5 As shown. The computer device includes a processor, memory, network interface, display screen and input device connected via a system bus. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external server via a network connection. When the computer program is executed by the processor, it realizes the functions or steps on the user side of an event processing method based on sound signal analysis.

[0236] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the following steps are performed:

[0237] Acquiring a sound signal to be analyzed;

[0238] Segmenting the sound signal to be analyzed to obtain target sound segments;

[0239] Extracting sound features of the target sound segment;

[0240] Comparing the sound feature with reference sound features in a preset feature library to generate a comparison result;

[0241] Based on the comparison result, determining the event type corresponding to the target sound segment;

[0242] When the event type meets a preset condition, a processing operation corresponding to the event type is executed.

[0243] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the following steps are implemented:

[0244] Acquiring a sound signal to be analyzed;

[0245] Segmenting the sound signal to be analyzed to obtain target sound segments;

[0246] Extracting sound features of the target sound segment;

[0247] Comparing the sound feature with reference sound features in a preset feature library to generate a comparison result;

[0248] Based on the comparison result, determining the event type corresponding to the target sound segment;

[0249] When the event type meets a preset condition, a processing operation corresponding to the event type is executed.

[0250] It should be noted that the above functions or steps that can be implemented by the computer-readable storage medium or computer device can be found in the relevant descriptions of the server side and the user side in the aforementioned method embodiment. To avoid repetition, they will not be described one by one here.

[0251] Those skilled in the art will appreciate that all or part of the processes in the above-mentioned embodiments can be implemented by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM).

[0252] Those skilled in the art will clearly understand that for the sake of convenience and brevity of description, only the division of the above-mentioned functional units and modules is used as an example. In actual applications, the above-mentioned functions can be distributed and completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.

[0253] It should be noted that if any software tools or components other than those of the Company appear in the embodiments of this application, they are merely for illustration and do not represent actual use. The above embodiments are intended only to illustrate the technical solutions of the present invention, not to limit them. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the above embodiments, or replace some of the technical features therein with equivalents. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included in the scope of protection of the present invention.

Claims

1. An event processing method based on sound signal analysis, characterized in that: The following steps are involved: Acquiring a sound signal to be analyzed; Segmenting the sound signal to be analyzed to obtain target sound segments; Extracting sound features of the target sound segment; Comparing the sound feature with reference sound features in a preset feature library to generate a comparison result; Based on the comparison result, determining the event type corresponding to the target sound segment; When the event type meets a preset condition, a processing operation corresponding to the event type is executed.

2. The event processing method based on sound signal analysis according to claim 1, characterized in that: Obtain the sound signal to be analyzed, including: Collecting original acoustic signals from the target subject's throat through a wearable device; Performing noise filtering on the original acoustic signal using a noise reduction module of the wearable device to generate a noise-reduced acoustic signal; Sampling the noise-reduced acoustic signal at a preset sampling frequency to generate a sampled acoustic signal; The sampled acoustic signal is sent to a mobile terminal as a sound signal to be analyzed through the wireless transmission module of the wearable device.

3. The event processing method based on sound signal analysis according to claim 1, characterized in that: Segmenting the sound signal to be analyzed to obtain target sound segments includes: Dividing the acoustic signal to be analyzed into multiple segments according to a segmented time window of a preset fixed length; Add a segment identifier containing a start timestamp and an end timestamp to each segment; Performing overlapping frame processing on the segments with the segment identifiers to generate target sound segments; The target sound segment is stored as an independent audio file and associated with a corresponding segment identifier.

4. The event processing method based on sound signal analysis according to claim 1, characterized in that: Extracting the sound features of the target sound segment includes: Performing pre-emphasis filtering on the target sound segment to generate a pre-emphasis sound signal; Performing frame processing on the pre-emphasized sound signal using a Hamming window to generate a framed sound signal; extracting Mel-frequency cepstral coefficients from the framed sound signal as a first sound feature; extracting short-time energy from the framed sound signal as a second sound feature; Normalization and fusion processing is performed on the first sound feature and the second sound feature to generate the sound feature of the target sound segment.

5. The event processing method based on sound signal analysis according to claim 1, characterized in that: Comparing the sound feature with a reference sound feature in a preset feature library to generate a comparison result includes: Retrieving multiple reference sound features from a voiceprint feature library, each reference sound feature being associated with a preset respiratory event type; Determining a similarity value between the sound feature and each reference sound feature respectively; Record the single duration and the number of occurrences per unit time of each respiratory event type whose similarity value exceeds the current matching similarity threshold; Determine whether the single duration of the same respiratory event type exceeds a preset single duration threshold, and whether the number of occurrences per unit time exceeds a preset number threshold; When the single duration of the same respiratory event type exceeds a preset single duration threshold, and the number of occurrences per unit time exceeds a preset number threshold, a preliminary determination result of the respiratory event type is generated; The preliminary determination results of all respiratory event types were summarized to generate the final comparison results including event type, occurrence time and severity level.

6. The event processing method based on sound signal analysis according to claim 1, characterized in that: Determining the event type corresponding to the target sound segment based on the comparison result includes: Extracting the similarity value, single duration, and number of occurrences per unit time of each preset respiratory event type from the comparison results; Determine a first ratio of a similarity value of each preset respiratory event type to a corresponding matching similarity threshold value; respectively determining a second ratio of a single duration of each preset respiratory event type to a preset single duration threshold; respectively determining a third ratio of the number of occurrences per unit time of each preset respiratory event type to a preset number threshold; For each preset respiratory event type, performing a weighted summation on the first ratio, the second ratio, and the third ratio according to a preset weight coefficient to generate a comprehensive score; All preset respiratory event types are sorted in descending order according to the comprehensive scores, and the type with the highest score is selected as the candidate event type; When the comprehensive score of the candidate event type exceeds a preset score threshold, determining the event type of the target sound segment as the candidate event type; A determination record is generated, which includes the event type, a start timestamp, and an end timestamp of the target sound segment.

7. The event processing method based on sound signal analysis according to claim 1, characterized in that: When the event type meets the preset conditions, executing a processing operation corresponding to the event type includes: Retrieving a corresponding processing strategy from a warning strategy library according to the event type, the processing strategy including an execution action type, a target device identifier, and an execution time constraint; Verifying whether the start timestamp and the end timestamp of the target sound segment are within a valid period defined by the execution time constraint; When the time verification passes, a processing level identifier is generated according to a preset severity level corresponding to the event type; generating a control instruction including the execution action type, a target device identifier, a start timestamp and an end timestamp of the target sound segment, and a processing level identifier; The control instruction is sent to the target device corresponding to the target device identifier through a wireless communication protocol, and a corresponding processing operation is performed.

8. An event processing device based on sound signal analysis, characterized in that: The event processing device based on sound signal analysis includes: An acquisition module, used to obtain the sound signal to be analyzed; A segmentation processing module, configured to perform segmentation processing on the sound signal to be analyzed to obtain target sound segments; A feature extraction module, configured to extract the sound features of the target sound segment; A feature comparison module, configured to compare the sound feature with reference sound features in a preset feature library to generate a comparison result; An event determination module, configured to determine an event type corresponding to the target sound segment based on the comparison result; The event response module is used to execute a processing operation corresponding to the event type when the event type meets a preset condition.

9. A computer device, characterized in that: The computer device includes a memory, a processor, and an event processing program based on sound signal analysis stored in the memory and capable of running on the processor. When the event processing program based on sound signal analysis is executed by the processor, the steps of the event processing method based on sound signal analysis as described in any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium, characterized in that The storage medium stores an event processing program based on sound signal analysis, and when the event processing program based on sound signal analysis is executed by the processor, the steps of the event processing method based on sound signal analysis according to any one of claims 1 to 7 are implemented.

Citation Information

Cited By

  • OSA monitoring and early warning method and device integrating snore and blood oxygen characteristics

    CN121101468A