Hospital ward-oriented distributed sound detection method, apparatus, device and system

Through the distributed sound detection method, audio data is collected and feature extraction is performed using microphone arrays to build a multi-modal sound event detection model. Combined with the JS weighted fusion algorithm, the problem of inefficient sound recognition in hospital wards is solved, and efficient and accurate audio monitoring and information management are achieved.

CN120544595APending Publication Date: 2025-08-26NORTH CHINA UNIVERSITY OF TECHNOLOGY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510870520.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-26
Publication Date
2025-08-26

AI Technical Summary

Technical Problem

In the sound detection of hospital wards, the prior art has problems such as low sound recognition efficiency and low accuracy, and it is difficult to adapt to the complex acoustic environment unique to the hospital. The audio data processing efficiency is low and susceptible to environmental noise. The data processing efficiency collected by multiple microphones is low, the sound modal processing is inaccurate, and the model output results are difficult to effectively integrate.

Method used

The distributed sound detection method is adopted to collect audio data through a microphone array, and MFCC and BEATS models are extracted respectively, and TDY, FDY and SKCRNN models are constructed. Combined with the JS weighted fusion algorithm, the output results of different models are integrated to improve the accuracy of sound category prediction.

Benefits of technology

It improves the accuracy of sound detection and the overall performance of the system, enhances the monitoring capabilities of the smart audio platform, ensures the stable operation and information management of the hospital ward environment, and improves patient safety guarantees.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120544595A_ABST
    Figure CN120544595A_ABST
Patent Text Reader

Abstract

The invention provides a distributed sound detection method, device, equipment and system for hospital wards, and relates to the technical field of sound detection. The method comprises the following steps: collecting sound data; performing feature extraction on the sound data to obtain audio data; constructing a multi-modal sound event detection model; inputting the audio data into the sound event detection model for processing to obtain a sound category prediction probability result based on different models; and fusing the sound category prediction probability results obtained by the plurality of models by adopting a JS-based weighted fusion algorithm to obtain a final sound category prediction probability result. The application effect of the intelligent audio platform in the inpatient department of the hospital is greatly improved, more reliable sound monitoring service is provided for medical staff, the safety guarantee of the patient is enhanced, meanwhile, the progress of hospital informatization management is promoted, and the application prospect is huge.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of sound detection technology, and in particular to a distributed sound detection method, device, equipment and system for hospital wards. Background Art

[0002] In practical applications, the acoustic environment of hospital inpatient departments is extremely complex, requiring smart audio platforms to intelligently monitor these departments through sound modalities. The diverse sounds within wards, including conversations between doctors and patients, the operation of medical equipment, and alarms, intertwine and create unique acoustic challenges. Therefore, to ensure the continuity of medical services and patient safety, real-time monitoring and intelligent analysis of the inpatient department's sound environment are crucial. This monitoring not only helps detect anomalies promptly but also provides data support for hospital management.

[0003] However, current technologies have some difficult-to-overcome problems. In terms of the deployment of electronic architectures based on sound event monitoring, traditional audio systems are often unable to adapt to the unique acoustic environment of hospitals, resulting in poor monitoring results. In terms of audio data feature extraction and preprocessing, the data processing efficiency collected by multiple microphones is low and is easily affected by environmental noise, which reduces the accuracy of sound recognition. In terms of sound modal processing, existing technologies have difficulty accurately identifying and analyzing complex sound scenes, limiting their value in practical applications. In terms of model output fusion, the output results of different models are difficult to effectively integrate, resulting in limited performance improvements in the overall monitoring system.

[0004] Therefore, there is an urgent need to provide a sound event detection system for hospital wards to solve the above problems. Summary of the Invention

[0005] The present application provides a distributed sound detection method, device, equipment and system for hospital wards, which are used to solve the problems of low efficiency and low accuracy of sound detection and recognition in existing hospital wards.

[0006] According to a first aspect disclosed in the present application, the present application provides a distributed sound detection method for hospital wards, comprising: Collecting sound data; wherein the sound data includes first audio data, second audio data, and third audio data collected from a bed and a door in a hospital ward, respectively; Performing feature extraction on the sound data to obtain audio data; the feature extraction includes: performing MFCC feature extraction on the first audio data and the second audio data respectively and performing BEATS model feature extraction on the third audio data to obtain first MFCC audio data, second MFCC audio data, and third BEATS audio data; Constructing a multimodal sound event detection model; the multimodal sound event detection model includes: a TDY model for processing the first MFCC audio data, an FDY model for processing the second MFCC audio data, and a SKCRNN model for processing the third BEATS audio data; Inputting the audio data into the sound event detection model for processing to obtain sound category prediction probability results based on different models; The sound category prediction probability results obtained by the multiple models are fused using a JS-based weighted fusion algorithm to obtain the final sound category prediction probability result.

[0007] In a feasible implementation manner, the collected sound data includes first audio data, second audio data, and third audio data collected from a bed and a door in a hospital ward respectively; including: Microphone array 1 and microphone array 2, which are set near the beds in the hospital ward, are used to collect audio data in the hospital ward to obtain first audio data and second audio data, and microphone array 3, which is set at the door of the hospital ward, is used to collect audio data in the hospital ward to obtain third audio data, and the above audio data are transmitted to edge device 1 for processing.

[0008] In a feasible implementation, the feature extraction is performed on the sound data to obtain audio data; the feature extraction includes: performing MFCC feature extraction on the first audio data and the second audio data respectively and performing BEATS model feature extraction on the third audio data to obtain first MFCC audio data, second MFCC audio data and third BEATS audio data; including: Edge device one performs MFCCT feature extraction on the first audio data and the second audio data obtained to obtain first MFCC audio data and second MFCC audio data, and sends them to edge device two and edge device three respectively; Edge device one performs BEATS model feature extraction on the acquired third audio data to obtain third BEATS audio data, and sends it to edge device four.

[0009] In a feasible implementation, the step of inputting the audio data into the sound event detection model for processing to obtain sound category prediction probability results based on different models includes: The TDY model in the second edge device processes the first MFCC audio data to obtain a sound category prediction probability result based on the TDY model; The FDY model in the third edge device processes the second MFCC audio data to obtain a sound category prediction probability result based on the FDY model; The SKCRNN model in edge device four processes the third BEATS audio data to obtain the sound category prediction probability result based on the SKCRNN model.

[0010] In a feasible implementation, the TDY model in the second edge device processes the first MFCC audio data to obtain a sound category prediction probability result based on the TDY model; specifically, the process includes: Convolving the first MFCC audio data with a convolution kernel to obtain N convolution results; The N convolution results are fused through the attention matrix to obtain the sound category prediction probability result based on the TDY model processing.

[0011] In a feasible implementation manner, the first MFCC audio data is convolved with a convolution kernel to obtain N convolution results; the specific method is:

[0012] in, is the convolution kernel, size is , is the bias term; and is the number of channels of model input and output; and are the number of time slices of input and output respectively; and are the number of frequency segments of input and output respectively; It is the time-frequency domain input feature of the model; that is, the first MFCC audio data; is the convolution operation; is the output of the convolution, with a total of indivual; The N convolution results are fused through the attention matrix to obtain the sound category prediction probability result based on the TDY model. The specific method is:

[0013]

[0014] Among them, the attention matrix represent dimensional attention weight, is the output of the temporal dynamic convolution; represents element-wise multiplication; Represents the nonlinear activation function ReLU; Represents the SoftMax function; Represents a one-dimensional convolution operation.

[0015] In a feasible implementation, the FDY model in the edge device three processes the second MFCC audio data to obtain a sound category prediction probability result based on the FDY model; specifically: The second MFCC audio data is processed by average pooling, and then processed by one-dimensional convolution, batch normalization and activation function, one-dimensional convolution and SoftMax function in sequence to obtain attention weight; The attention weights are processed by convolution kernel to obtain an overall weight matrix, which is then subjected to two-dimensional convolution. The output of the convolution is then subjected to batch normalization and activation function to obtain the sound category prediction probability result based on the FDY model.

[0016] In a feasible implementation, the attention weight is subjected to convolution kernel processing to obtain an overall weight matrix, which is then subjected to two-dimensional convolution. The output of the convolution is then subjected to batch normalization processing and an activation function to obtain a sound category prediction probability result based on the FDY model; the specific method is:

[0017] in, is the attention weight; is the convolution kernel; is the bias term; is the convolution operation; is the output of convolution; and is the number of channels of model input and output; and are the number of time slices of input and output respectively; and are the number of frequency segments of input and output respectively; It is the time-frequency domain input feature of the model, that is, the second MFCC audio data; The output of the convolution is then batch normalized and activated to obtain the sound category prediction probability based on the FDY model.

[0018] In a feasible implementation, the SKCRNN model in the fourth edge device processes the third BEATS audio data to obtain a sound category prediction probability result based on the SKCRNN model; specifically: Processing the third BEATS audio data through a convolution module, wherein the convolution module processing includes one-dimensional convolution processing, batch normalization processing, activation function, maximum pooling layer processing, and drop layer processing to obtain an output of the convolution module; The output of the convolution module is input into the spatial attention module and the temporal attention module for processing, respectively, to obtain the spatial attention weight matrix and the temporal attention weight matrix; The spatial attention weight matrix and the temporal attention weight matrix are added element by element and divided by 2, and multiplied with the output of the convolution module. The output result is then passed through the bidirectional GRU layer to obtain the output result based on the SKCRNN model. After batch normalization processing and activation function, the sound category prediction probability result based on the SKCRNN model is obtained.

[0019] In a feasible implementation, the output of the convolution module is input into the spatial attention module and the temporal attention module for processing, respectively, to obtain the spatial attention weight matrix and the temporal attention weight matrix; specifically: The output of the convolution module is input into the spatial attention module for processing, specifically: The output of the convolution module is input into the spatial attention module, and is subjected to average pooling and maximum pooling respectively to obtain two different temporal context information; The above context information is then fed into the two shared dense layers, and the spatial attention weight is generated by element-by-element summation and output of the feature vector, which is then processed by the sigmoid function to obtain the spatial attention weight matrix , the size is ; is the convolution kernel in the spatial attention weight matrix; is the bias in the spatial attention weight matrix; The output of the convolution module is input into the temporal attention module for processing, specifically: The output of the convolution module is input into the temporal attention module, and is processed by average pooling and maximum pooling respectively. Then, the cascaded feature vector is input into the two-dimensional convolution to obtain the temporal attention weight; and then processed by the sigmoid function to obtain the temporal attention weight matrix , the size is ; is the convolution kernel in the spatial attention weight matrix; is the bias in the temporal attention weight matrix; In a feasible implementation, the spatial attention weight matrix and the temporal attention weight matrix are element-wise added and divided by 2, and multiplied with the output of the convolution module. The output result is then passed through a bidirectional GRU layer to obtain the output result based on the SKCRNN model. The specific method is:

[0020] in, It is the time-frequency domain input feature of the model, that is, the third BEATS audio data; and is the number of channels of model input and output; and are the number of time slices of input and output respectively; and are the number of frequency segments of input and output respectively; It is the element-wise addition operation; The output result passes through the bidirectional GRU layer, which consists of a forward GRU layer and a reverse GRU layer. The bidirectional GRU merges information from two opposite directions. The specific method is:

[0021] in, Represents the SoftMax function; represents element-wise multiplication; and They are update gate and reset gate respectively; activation Is the previous activation and candidate activation Linear interpolation between ; is a trainable parameter; is the output of the bidirectional GRU.

[0022] In a feasible implementation, the JS-based weighted fusion algorithm fuses the sound category prediction probability results obtained by the multiple models to obtain a final sound category prediction probability result; specifically, the method includes: The sound category prediction probability result based on the FDY model, the sound category prediction probability result based on the TDY model, and the sound category prediction probability result based on the SKCRNN model are transmitted to the edge device 5, and the above results are weightedly fused using the JS weighted fusion algorithm, specifically including: Calculate the JS distance between evidences; the specific calculation method is:

[0023] in, (k = 1, 2, ... p) represents the different sound categories output by the model, and there are a total of p kinds of sounds; (i=1, 2, ..n) and (j=1, 2, ...n) represents two different models, each model is independent evidence, there are n models, and ; Indicates the The output of the model The probability of a sound; Indicates the The output of the model The probability of a sound; Based on the JS distance between the evidences, the conflict degree between the evidences is calculated; specifically:

[0024] Among them, Sim is The conflict matrix, It is the amount of evidence; It is Evidence and the degree of conflict between the pieces of evidence; is the conflict degree of each piece of evidence itself, set to 0; Calculate the degree of support between evidences; specifically:

[0025] in, , indicating evidence and the degree of similarity between them; Express evidence and The degree of conflict between The similarity is obtained by subtracting the conflict degree from 1. The diagonal elements of the matrix are all 1, indicating that each piece of evidence is 100% similar to itself; the elements in other positions are Indicates the relative support between two pieces of evidence, with a value between 0 and 1; evidence supported by other evidence The confidence level is:

[0026] in, Indicates the Evidence support level; Indicates that except for All other evidence except the first The sum of the support of the evidence; Indicates the Evidence for the The support of the evidence; Calculate the weights of evidence; specifically:

[0027] in, is a specific piece of evidence to be calculated The weight of Indicates the Evidence support level; Represents from the set Take the maximum value among From 1 to ; Revise the original BPA and modify the BPA of each piece of evidence; specifically:

[0028] in, represents the modified basic probability assignment (BPA); is a specific piece of evidence to be calculated The weight of Representing an event ; The Dempster combination rule is used for fusion to obtain the final fusion output data, that is, the final sound category prediction probability result; specifically:

[0029] in, is the final fused basic probability assignment (BPA), which represents the comprehensive trust in all evidence; It is the basic probability assignment of a single piece of evidence after modification, indicating the effect of evidence from different sources on the event. Trust It is the operator of Dempster's combination rule, which is used to combine the basic probability assignments of two or more pieces of evidence.

[0030] According to a second aspect disclosed in the present application, the present application provides a distributed sound detection device for hospital wards, comprising: A data acquisition module for collecting environmental sound data, wherein the sound data includes first audio data, second audio data, and third audio data collected from a bed and a door in a hospital ward, respectively; A sound data feature extraction module is used to perform feature extraction on the sound data to obtain audio data; the feature extraction includes: performing MFCC feature extraction on the first audio data and the second audio data respectively and performing BEATS model feature extraction on the third audio data to obtain first MFCC audio data, second MFCC audio data and third BEATS audio data; A multimodal sound event detection model construction module, used to construct a multimodal sound event detection model; the multimodal sound event detection model includes: a TDY model for processing the first MFCC audio data, an FDY model for processing the second MFCC audio data, and a SKCRNN model for processing the third BEATS audio data; a sound category probability detection module, configured to input the audio data into the sound event detection model for processing, and obtain sound category prediction probability results based on different models; The sound category prediction probability fusion module is used to fuse the sound category prediction probability results obtained by the multiple models using a JS-based weighted fusion algorithm to obtain the final sound category prediction probability result.

[0031] According to a third aspect disclosed in the present application, the present application provides a distributed sound detection device for hospital wards, comprising a processor, and a memory communicatively connected to the processor; The memory stores computer-executable instructions; The processor executes the computer-executable instructions stored in the memory to implement the method according to any one of the first aspects.

[0032] According to a fourth aspect disclosed in the present application, the present application provides a distributed sound detection system for hospital wards, comprising the detection as described in the third aspect, and a server; The server is used to receive the final sound category prediction probability result output by the distributed sound detection device, and perform display and logic control.

[0033] According to the fifth aspect disclosed in the present application, the present application provides a computer-readable storage medium, in which computer-executable instructions are stored. When the computer-executable instructions are executed, they are used to implement any one of the methods in the first aspect.

[0034] According to a sixth aspect disclosed in the present application, the present application provides a computer program product, including a computer program, which, when executed, is used to implement any one of the methods in the first aspect.

[0035] Compared with the existing technology, this application has the following beneficial effects: The present application provides a distributed sound detection method, device, equipment and system for hospital wards. By providing an electronic architecture deployment solution based on sound event monitoring, the solution can better adapt to the acoustic characteristics of the hospital inpatient department and ensure the stable operation of the monitoring system; at the same time, an efficient audio data feature extraction and preprocessing method is adopted to significantly improve the quality and speed of data processing; in terms of sound modal processing, sound event detection technology is used to analyze complex sound scenes, and finally, the outputs of different models are effectively integrated through the model output fusion strategy, thereby greatly improving the overall monitoring performance and accuracy of the smart audio platform; greatly improving the application effect of the smart audio platform in the hospital inpatient department, providing medical staff with more reliable sound monitoring services, enhancing patient safety, and promoting the advancement of hospital information management, and has huge application prospects. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application.

[0037] Figure 1 A schematic diagram of a flow chart of a distributed sound detection method for hospital wards provided in an embodiment of the present application; Figure 2 A block diagram of a distributed sound detection system for hospital wards provided in an embodiment of the present application; Figure 3 A schematic block diagram of an embodiment of the present application providing an embodiment of the invention for extracting features from audio data using MFCC; Figure 4 An overall flow chart for processing audio data collected by a microphone array 1 provided in an embodiment of the present application; Figure 5 An overall flow chart for processing audio data collected by the microphone array 2 provided in an embodiment of the present application; Figure 6 An overall flow chart for processing audio data collected by the microphone array 3 provided in an embodiment of the present application; Figure 7 Flowchart of audio data processing by the TDY model in the second edge device provided in an embodiment of the present application; Figure 8 Flowchart of audio data processing by the FDY model in the edge device three provided in an embodiment of the present application; Figure 9 This is an overall flowchart of audio data processing by the SKCRNN model in edge device 4 provided in an embodiment of the present application.

[0038] Figure 10 A schematic diagram of the structure of a bidirectional GRU provided in an embodiment of the present application; Figure 11 A flowchart of the JS-based weighted fusion algorithm provided in this application embodiment; Figure 12 A schematic diagram of an embodiment of a JS-based weighted fusion algorithm provided in an embodiment of the present application; Figure 13 A schematic structural diagram of a distributed sound detection device for hospital wards provided in an embodiment of the present application; Figure 14 A schematic diagram of the structure of a distributed sound detection device for hospital wards provided in an embodiment of the present application.

[0039] The above drawings illustrate specific embodiments of the present application, which will be described in more detail below. These drawings and the textual description are not intended to limit the scope of the present application in any way, but rather to illustrate the concepts of the present application to those skilled in the art by reference to specific embodiments. DETAILED DESCRIPTION

[0040] Exemplary embodiments will be described in detail herein, with examples illustrated in the accompanying drawings. In the following description, when referring to the drawings, identical numerals in different figures represent identical or similar elements, unless otherwise indicated. The embodiments described in the following exemplary embodiments are not intended to represent all embodiments consistent with the present application. Rather, they are merely examples of apparatus and methods consistent with certain aspects of the present application, as detailed in the appended claims.

[0041] This application proposes a distributed sound detection method, device, equipment and system for hospital wards. This method improves the adaptability and stability of the electronic architecture based on sound event detection, so that it can better adapt to the complex acoustic environment of the hospital inpatient department and ensure the efficient operation of the audio monitoring system; optimizes the feature extraction and preprocessing process of audio data, improves the speed and accuracy of data processing, reduces the impact of environmental noise on sound recognition, and improves the robustness of the system; innovates sound modal processing technology to achieve high-precision recognition and analysis of multiple sound sources in the hospital inpatient department, thereby enhancing the intelligent monitoring capability of the system; develops model output fusion strategy to effectively integrate the monitoring results of different models and improve the overall performance and decision-making accuracy of the smart audio platform.

[0042] The following is a detailed description of the technical solution of the distributed sound detection method for hospital wards provided by this application through specific embodiments. It should be noted that the following embodiments can exist independently or in combination with each other, and the same or similar content may not be repeated in different embodiments.

[0043] Figure 1 A schematic diagram of a flow chart of a distributed sound detection method for hospital wards provided in an embodiment of the present application is provided. Figure 2 A block diagram of a distributed sound detection system for hospital wards provided in an embodiment of the present application; see Figure 1 and Figure 2 In some embodiments, the detection method comprises the following steps: S101, collecting environmental sound data; wherein the environmental sound data includes human voice and environmental sound.

[0044] wherein, collecting sound data; wherein, the sound data includes first audio data, second audio data and third audio data collected from a bed and a door in a hospital ward respectively; Specifically, in this embodiment, the scenario considered is a single ward in the hospital inpatient area, such as Figure 2 As shown, microphone arrays are planned to be placed at the door of the ward and near the bed to collect audio data; microphone array one and microphone array two are respectively placed near the bed in the hospital ward to collect audio data in the hospital ward to obtain first audio data and second audio data, and microphone array three is placed at the door of the hospital ward to collect audio data in the hospital ward to obtain third audio data, and the above audio data are respectively transmitted to edge device one for processing.

[0045] S102, performing feature extraction on the sound data to obtain audio data; the feature extraction includes: performing MFCC feature extraction on the first audio data and the second audio data respectively and performing BEATS model feature extraction on the third audio data to obtain first MFCC audio data, second MFCC audio data, and third BEATS audio data; Specifically, the edge device 1 performs MFCCT feature extraction on the first audio data obtained to obtain the first MFCC audio data, and sends it to the edge device 2; Specific extraction flow chart Figure 3 ; The data is pre-emphasized to increase the amplitude of the high-frequency part and make the spectrum of the signal flat. It is usually used to eliminate the nasal sound of the sound wave. It is achieved through a first-order high-pass filter. The formula is:

[0046] in, is the sample value of the original audio signal at time point n; is the pre-emphasized audio signal at time point The sample value of is the pre-emphasis coefficient, which is usually between 0.95 and 0.99. In this embodiment, the value is 0.97.

[0047] After data pre-emphasis, framing is performed to divide the continuous audio signal into frames of fixed length. The following are the steps for framing: first, determine the frame length. and frame shift , frame length Usually set to 20ms to 40ms, frame shift Usually set to half the frame length or less; then, starting from the start of the audio signal, every Samples extract a frame of length N, the The mathematical representation of a frame is as follows:

[0048] in, is the sample sequence of the kth frame, is the index of the frame, starting from 0; is the frame shift, i.e. the number of samples between adjacent frames; is the frame length, that is, the number of samples in each frame.

[0049] After the audio data is framed, it is windowed to reduce the edge effects caused by the framing operation and make the signal within the frame smoother. The windowing process involves multiplying the window function by the samples of each frame. The following is the mathematical formula for the windowing operation:

[0050] In the above formula, It is The sample sequence of frames has a length of ; The length is also The window function of is the sample index in the window function, ; After adding window Frame No. samples, is the window function, and commonly used window functions include rectangular window, Hamming window and Hanning window. Here we choose Hamming window, and its formula is:

[0051] In the above formula, The length is ,in is the sample index in the window function, .

[0052] Then, a fast Fourier transform (FFT, an efficient algorithm for calculating discrete Fourier transform DFT) is performed on each frame of the windowed signal to convert the time domain signal into a frequency domain signal. The following is the mathematical formula of DFT:

[0053] In the above formula, is a windowed audio frame with a length of , is the sample index in the window function, ; yes The DFT of complex frequency components; is the kernel function of Fourier transform, is an imaginary unit, satisfying ; is the sample index in the frequency domain, ranging from ; For the FFT algorithm in practical applications, the above DFT calculation process will be optimized, usually using a divide-and-conquer strategy to decompose the DFT into multiple smaller DFTs to reduce the computational complexity; The following are the basic steps of the FFT algorithm: First, Divided into two lengths Subsequences of: one containing all even-indexed samples, the other containing all odd-indexed samples; perform FFT on each of these two subsequences; then, combine the FFT results of the two subsequences and calculate the final FFT using the following formula:

[0054] The data obtained after the fast Fourier transform is then passed through a Mel filter bank (Mel filter bank); the Mel filter bank simulates the sensitivity of the human ear to different frequencies and maps the spectrum output by the FFT to the Mel scale; each filter is a triangular filter in the frequency domain, covering a specific frequency range, and the output of the filter bank is the sum of the energy in the area covered by each filter.

[0055] In the above formula, is the FFT result of even-indexed samples; is the FFT result of the odd-indexed samples; is the rotation factor; this process is recursively applied to each subsequence until the simplest DFT is reached (i.e. ), and then gradually merge the results upwards to get the final FFT output.

[0056] After the audio data is processed by Fast Fourier Transform (FFT), a Mel filter bank is used to simulate human sensitivity to different frequencies. The output of the Mel filter bank is Each filter pair FFT output The following is the mathematical formula for the Mel filter bank to process FFT output data:

[0057] In the above formula, is the frequency index, is the number of FFT points, is a frequency domain signal The energy (i.e., the square of the amplitude) of It is The weights of the Mel filters; the Mel filter bank consists of triangular filters, each filter Corresponding to a specific Mel frequency range; Mel filter is defined as follows:

[0058] In the above formula, is the actual frequency corresponding to the frequency index k; are the edge frequencies of the mth Mel filter, which are converted according to the Mel frequency scale.

[0059] After the audio data is processed by the Mel filter bank, a logarithmic operation is usually performed on the output energy of each filter. This is to simulate the nonlinear response of the human ear to sound energy and to compress the dynamic range to reduce the impact of large values ​​on subsequent processing. The mathematical formula for the logarithmic operation is as follows:

[0060] In the above formula, It is The logarithm of the filter output energy; It is The output energy of a Mel filter.

[0061] The output energy is transformed into a set of coefficients by discrete cosine transform (DCT). The coefficients represent the energy distribution of different frequency components. The mathematical formula of DCT is as follows:

[0062] In the above formula, It is Mel-frequency cepstral coefficients (MFCC) coefficients, is the DC component, representing the energy mean of the entire spectrum; It is The logarithm of the filter output energy; is the number of Mel filters; is the number of MFCC coefficients that need to be calculated, .

[0063] After calculating the DCT, These coefficients are MFCC (Mel-Frequency Cepstral Coefficients), which are used as effective features of audio for sound event detection or other audio processing tasks.

[0064] The first audio data collected by the microphone array 1 is converted into audio features, namely the first MFCC audio data, after being extracted by the edge device 1, and sent to the edge device 2. Figure 4 As shown; the second audio data collected by the microphone array 2 is converted into audio features after the edge device 1 performs the above MFCC feature extraction, that is, the second MFCC audio data, and is sent to the edge device 3, as shown in the following example. Figure 5shown.

[0065] The third audio data collected by the microphone array That is, it is input into the BEATS model of edge device 1; the BEATS model, namely Bidirectional Encoder representation from Audio Transformers, is an iterative audio pre-training framework based on the Transformer architecture, which uses self-supervised learning to learn deep representations of audio signals from a large amount of unlabeled audio data; the BEATS model can process raw audio waveforms and learn features by predicting masked parts or future segments in the audio signal, without relying on traditional manual feature extraction; the bidirectional nature of this model enables it to consider the historical and future contextual information of the audio signal at the same time, improving the ability to understand the audio content; after pre-training is completed, the BEATS model can effectively extract the audio signal The extracted advanced features, namely the third BEATS audio data, are sent to the edge device 4, as shown in the following example: Figure 6 shown.

[0066] S103, constructing a multimodal sound event detection model; the multimodal sound event detection model includes: a TDY model for processing the first MFCC audio data, an FDY model for processing the second MFCC audio data, and a SKCRNN model for processing the third BEATS audio data; Specifically, a TDY model is constructed in edge device two, an FDY model is constructed in edge device three, and an SKCRNN model is constructed in edge device four, and the first MFCC audio data, the second MFCC audio data, and the third BEATS audio data are processed respectively; S104, inputting the audio data into the sound event detection model for processing to obtain sound category prediction probability results based on different models; specifically including: The TDY model in the second edge device processes the first MFCC audio data to obtain a sound category prediction probability result based on the TDY model; The FDY model in the third edge device processes the second MFCC audio data to obtain a sound category prediction probability result based on the FDY model; The SKCRNN model in edge device four processes the third BEATS audio data to obtain the sound category prediction probability result based on the SKCRNN model.

[0067] The TDY model in the second edge device processes the first MFCC audio data to obtain a sound category prediction probability result based on the TDY model; specifically, the following steps are performed: Convolving the first MFCC audio data with a convolution kernel to obtain N convolution results; The N convolution results are fused through the attention matrix to obtain the sound category prediction probability result based on the TDY model processing.

[0068] Specifically, the first MFCC audio data is convolved with a convolution kernel to obtain N convolution results; the specific method is:

[0069] in, is the convolution kernel, size is , is the bias term; and is the number of channels of model input and output; and are the number of time slices of input and output respectively; and are the number of frequency segments of input and output respectively; It is the time-frequency domain input feature of the model; that is, the first MFCC audio data; is the convolution operation; is the output of the convolution, with a total of indivual; The N convolution results are fused through the attention matrix to obtain the sound category prediction probability result based on the TDY model. The specific method is:

[0070]

[0071] Among them, the attention matrix represent dimensional attention weight, is the output of the temporal dynamic convolution; represents element-wise multiplication; Represents the nonlinear activation function ReLU; Represents the SoftMax function; Represents a one-dimensional convolution operation.

[0072] In this embodiment, the TDY (Temporal dynamic convolution neural networks) algorithm is deployed in the edge device 2. The specific model structure is as follows: Figure 7 shown in the figure; and is the number of channels of model input and output, and are the number of time slices of input and output, and are the number of frequency segments of input and output respectively; It is the time-frequency domain input feature of the model. Here, the MFCC is obtained by feature extraction of the audio data collected by the microphone array 1, which is obtained by convolution with the convolution kernel. ,capturing local features in time series data; is the convolution operation; is the output of the convolution, with a total of indivual; is the convolution kernel, size is ; is the bias term; the specific calculation formula is as follows:

[0073] Together Convolution results , which is achieved through the attention matrix Fusion, attention matrix represent dimensional attention weights, which are based on the fragments of temporal features But it is different; is the output of the temporal dynamic convolution; represents element-wise multiplication; Represents the nonlinear activation function ReLU; the specific aggregation formula is as follows:

[0074] Among them, the attention matrix MFCC is averaged along the channel and frequency axes to reduce the dimension of the data and obtain the main features. and ; Then use two one-dimensional convolutions and ReLU to connect the features To facilitate subsequent data processing, the output signal is converted into a probability distribution through the SoftMax function; the specific formula is as follows:

[0075] Represents the SoftMax function; Represents a one-dimensional convolution operation; Represents the nonlinear activation function ReLU.

[0076] In summary, the aggregated output result of the audio feature MFCC after the TDY model is obtained After batch normalization and activation function, the predicted probability of multiple sound categories such as shouting, talking, and instruments appearing in each time segment is obtained, that is, the sound category prediction probability result based on TDY model processing.

[0077] Specifically, the FDY model in the third edge device processes the second MFCC audio data to obtain a sound category prediction probability result based on the FDY model; specifically, the process includes: The second MFCC audio data is processed by average pooling, and then processed by one-dimensional convolution, batch normalization and activation function, one-dimensional convolution and SoftMax function in sequence to obtain attention weight; The attention weights are processed by convolution kernel to obtain the overall weight matrix, which is then subjected to two-dimensional convolution. The output of the convolution is then subjected to batch normalization and activation function to obtain the sound category prediction probability result based on the FDY model; specifically:

[0078] in, is the attention weight; is the convolution kernel; is the bias term; is the convolution operation; is the output of convolution; and is the number of channels of model input and output; and are the number of time slices of input and output respectively; and are the number of frequency segments of input and output respectively; It is the time-frequency domain input feature of the model, that is, the second MFCC audio data; The output of the convolution is then batch normalized and activated to obtain the sound category prediction probability based on the FDY model.

[0079] In this embodiment, the FDY (Frequency dynamic convolution neural networks) algorithm is deployed in the edge device 3. The specific model structure is as follows: Figure 8 shown in the figure; and is the number of channels of model input and output; and are the number of time slices of input and output respectively; and are the number of frequency segments of input and output respectively; It is the time-frequency domain input feature of the model. Here, the MFCC obtained after feature extraction of the audio data collected by the microphone array 2 is used, that is, the second MFCC audio data; it uses the average pooling layer along the channel axis to reduce the computational complexity and obtain a fixed-size feature vector; then press Figure 8 The order in the example is passed through two one-dimensional convolutions conv1D to capture the local features in the frequency sequence data. Batch normalization BN and activation function ReLU are also used between the two one-dimensional convolutions. The former can accelerate the training process and improve the stability of the model, and the latter can introduce nonlinearity, enabling the model to learn more complex features. Finally, the output is converted into a probability distribution through the SoftMax function, and the attention weight is output. .

[0080] Set a group size to be The convolution kernel and bias , to capture different frequency components and combine them with attention weights Multiply and combine all weighted base kernels into an overall weight matrix [W(f), b(f)] by addition, with a size of , which dynamically adjusts the weight of each convolution kernel according to the frequency characteristics of the input signal, and finally undergoes a two-dimensional convolution to further enhance the expressiveness of the features; is the convolution operation; is the output result of convolution; the specific calculation formula is as follows:

[0081] In summary, the convolution output of the audio feature MFCC after passing through the FDY model is ,After batch normalization and activation function, the predicted probability of various sound categories such as shouting, speaking and instruments appearing in each time segment is obtained, which is the sound category prediction probability result based on the FDY model.

[0082] Specifically, the SKCRNN model in the fourth edge device processes the third BEATS audio data to obtain a sound category prediction probability result based on the SKCRNN model; specifically: Processing the third BEATS audio data through a convolution module, wherein the convolution module processing includes one-dimensional convolution processing, batch normalization processing, activation function, maximum pooling layer processing, and drop layer processing to obtain an output of the convolution module; The output of the convolution module is input into the spatial attention module and the temporal attention module for processing, respectively, to obtain the spatial attention weight matrix and the temporal attention weight matrix; specifically: The output of the convolution module is input into the spatial attention module for processing, specifically: The output of the convolution module is input into the spatial attention module, and is subjected to average pooling and maximum pooling respectively to obtain two different temporal context information; The above context information is then fed into the two shared dense layers, and the spatial attention weight is generated by element-by-element summation and output of the feature vector, which is then processed by the sigmoid function to obtain the spatial attention weight matrix , the size is ; The output of the convolution module is input into the temporal attention module for processing, specifically: The output of the convolution module is input into the temporal attention module, and is processed by average pooling and maximum pooling respectively. Then, the cascaded feature vector is input into the two-dimensional convolution to obtain the temporal attention weight; and then processed by the sigmoid function to obtain the temporal attention weight matrix , the size is .

[0083] The spatial attention weight matrix and the temporal attention weight matrix are added and divided by 2, and multiplied with the output of the convolution module. The output result is then passed through the bidirectional GRU layer to obtain the output result based on the SKCRNN model. After batch normalization and activation function, the sound category prediction probability result based on the SKCRNN model is obtained; specifically:

[0084] in, It is the time-frequency domain input feature of the model, that is, the third BEATS audio data; and is the number of channels of model input and output; and are the number of time slices of input and output respectively; and are the number of frequency segments of input and output respectively; It is the element-wise addition operation; The output result passes through the bidirectional GRU layer, which consists of a forward GRU layer and a reverse GRU layer. The bidirectional GRU merges information from two opposite directions. The specific method is:

[0085] in, Represents the SoftMax function; represents element-wise multiplication; and They are update gate and reset gate respectively; activation Is the previous activation and candidate activation Linear interpolation between ; is a trainable parameter; is the output of the bidirectional GRU.

[0086] In this embodiment, the SKCRNN (Space-Temporal Kernel Convolutional Recurrent Neural Network) algorithm is deployed in the edge device 4. The specific model structure is as follows: Figure 9 As shown, it integrates convolutional modules, spatiotemporal attention modules, and bidirectional gated recurrent unit (GRU) layers; and is the number of channels of model input and output; and are the number of time slices of input and output respectively; and are the number of frequency segments of input and output respectively; It is the time-frequency domain input feature of the model. The audio data collected by the microphone array three is used here as the output feature of the BEATS model. It first passes through the convolution module, which includes 2 one-dimensional convolutions for extracting local features, batch normalization (BN) to accelerate the convergence of the model during training, activation function (ReLU) to avoid gradient disappearance, as well as a maximum pooling layer and a dropout layer. Then, in order to fully extract representative local features, the spatiotemporal attention module is introduced, including the spatial attention mechanism and the temporal attention mechanism. The output of the convolution module is input into the spatial attention module and the temporal attention module respectively.

[0087] The spatial attention module assigns weights to the channels of the feature map to focus on the channels that contribute more to the representation of the audio signal. In the spatial attention mechanism, average pooling and maximum pooling aggregate the temporal information of each channel of the input features, respectively, to obtain two different temporal contexts. Average pooling is effective in capturing global changes, while maximum pooling is better at capturing local changes. Both temporal contexts are then fed into two shared dense layers. The output feature vectors are then merged by element-by-element summation to generate spatial attention weights. Finally, the spatial attention weights are compressed to 0-1 using the sigmoid function to obtain the spatial attention weight matrix. , the size is ; The temporal attention module assigns weights to the temporal segments of the feature map to focus on processing the segments with more information in the audio signal. In the temporal attention mechanism, average pooling and maximum pooling respectively aggregate the spatial information (also known as channel information) of the spatial fine features, and the pooling operation along the channel dimension helps to highlight the information segments. Then, the merged features are combined together through the cascade operation. The concatenated feature vector is then input into the two-dimensional convolution to obtain the temporal attention weight. Similarly, the temporal attention weight is compressed to 0-1 through the sigmoid function to obtain the temporal attention weight matrix. , the size is ; Add the spatial attention weight matrix and the temporal attention weight matrix element-wise and divide them by 2, and multiply them with the output of the convolution module; It is the element-wise addition operation; the specific formula is as follows:

[0088] The output result passes through the bidirectional GRU layer, which consists of a forward GRU layer and a reverse GRU layer, and uses the GRU's gated adjustment mechanism to extract global features; at a certain time, the local features of all past time steps are summarized by the forward GRU, and the local features of all future time steps are summarized by the reverse GRU; the bidirectional GRU merges the information from two opposite directions to obtain contextual information, such as Figure 10 The specific calculation formula is as follows:

[0089] in Represents the SoftMax function; represents element-wise multiplication; and They are the update gate and the reset gate, which determine how to update the current activation , and how to forget the previous activation ;activation Is the previous activation and candidate activation Linear interpolation between ; is a trainable parameter; is the output of the bidirectional GRU.

[0090] In summary, the output features of the BEATS model are output after the SKCRNN model After maximum pooling, batch normalization and activation function, the predicted probability of various sound categories such as shouting, talking and instruments appearing in each time segment is obtained, that is, the sound category prediction probability result based on the SKCRNN model.

[0091] S105, using a JS-based weighted fusion algorithm to fuse the sound category prediction probability results obtained by the multiple models to obtain a final sound category prediction probability result; specifically including: The predicted probabilities of the three models FDY, TDY and SKCRNN for various sounds such as shouting, speaking and instruments are transmitted to the edge device 5 through the wired network, and the output results are fused here. The weighted fusion algorithm based on JS (such as Figure 11 shown); represents the set of different sound categories output by the model, Represents a collection of different models, each of which is independent evidence; represents the kth sound category, Indicates the Models, Indicates the The output of the model The probability of a sound; In this embodiment, the specific embodiment of the weighted fusion algorithm of JS is as follows: Figure 12 As shown, the model only outputs the predicted probabilities of three sounds: shouting, speaking, and instruments; On behalf of Cry, stands for Speech, Represents Machine; Represents the TDY model, Represents the FDY model, Represents the SKCRNN model.

[0092] The probability of Cry output by the TDY model is 0.13, The probability of Speech output by the TDY model is 0.27. The probability of Machine representing the output of the TDY model is 0.6.

[0093] The probability of Cry output by the FDY model is 0.3, The probability of Speech output by the FDY model is 0.43. The probability of Machine representing the output of the FDY model is 0.27.

[0094] The probability of Cry output by the SKCRNN model is 0.21. The probability of Speech output by the SKCRNN model is 0.32. The probability of Machine representing the output of the SKCRNN model is 0.47.

[0095] The output of the JS weighted fusion algorithm is the final sound category and its probability; the specific steps of the JS weighted fusion algorithm are as follows Figure 12 shown.

[0096] Step 1: Calculate the JS distance between evidences: Given the probability distribution of different evidences, use the JS distance to represent the similarity between two pieces of evidence. The specific calculation formula is:

[0097] In the above formula, (k = 1, 2, ... p) represents the different sound categories output by the model, and there are a total of p kinds of sounds; (i=1, 2, ..n) and (j=1, 2, ...n) represents two different models, and each model is independent evidence. There are n models in total. ; Indicates the The output of the model The probability of a sound; Indicates the The output of the model The probability of a sound; The value of represents the correlation between the two pieces of evidence, and its value range is [0,1]. When the value gradually tends to 1, the greater the conflict between the two pieces of evidence, the lower the similarity; As the value gradually approaches 0, the conflict between the two pieces of evidence is smaller and the similarity is higher.

[0098] Step 2: Calculate the degree of conflict between evidences: According to the calculation formula in step 1, the degree of conflict between two pieces of evidence can be calculated and formed. The conflict matrix Sim; Since the JS distance satisfies symmetry and non-negativity, The value on the diagonal represents the degree of conflict in the evidence itself, indicating that the data itself is completely consistent and there is no difference, so the value is 0, that is, ;

[0099] In the above formula, Sim is The conflict matrix, It is the amount of evidence; It is Evidence and the degree of conflict between the pieces of evidence; It is the conflict degree of each piece of evidence. It is set to 0 here, which means that a piece of evidence will not conflict with itself.

[0100] Step 3: Calculate the support between evidences: The transformed support matrix can be expressed as:

[0101] In the above formula, , representing evidence and , the degree of similarity between them; Represents evidence and The degree of conflict between The similarity is obtained by subtracting the conflict degree from 1; the diagonal elements of the matrix are all 1, indicating that the similarity between each piece of evidence and itself is 100%; the elements in other positions are It reflects the relative support between two pieces of evidence. Its value ranges from 0 to 1. The closer it is to 1, the more similar the evidence is. The closer it is to 0, the greater the conflict between the evidences. The confidence level is:

[0102] In the above formula, Indicates the Evidence support level; Indicates that except for All evidence other than the first The sum of the support of the evidence; Indicates the Evidence for the The support of evidence.

[0103] Step 4: Calculate the weights between evidences: The weights reflect the degree of trust in the entire fusion system. The calculation process is mainly through batch normalization of the support between evidences. Therefore, the weights of evidences in the fusion system are Weight It can be expressed as:

[0104] In the above formula, is a specific piece of evidence to be calculated The weight of Indicates the Evidence support level; Represents from the set Take the maximum value among From 1 to .

[0105] Step 5: Modify the original BPA: In order to highlight the importance of different evidences and improve the reliability and fault tolerance of the fusion results, the BPA of each evidence is modified in a mutually supporting manner. The modified BPA can be expressed as:

[0106] In the above formula, represents the modified basic probability assignment (BPA); is a specific piece of evidence to be calculated The weight of Representing an event , the object for which a decision needs to be made.

[0107] Step 6: Fusion using the Dempster combination rule: When there is complete distrust between the evidence, the BPA value is 0. To ensure that the value of the least significant bit is not 0, a minimum value of 10-8 is added to this value. This not only does not change the BPA's support for the real data, but also effectively avoids the occurrence of 0 values, further improving the accuracy of the BPA. Finally, the synthesis formula for fusion using the Dempster combination rule is expressed as:

[0108] In the above formula, is the final fused basic probability assignment (BPA), which represents the comprehensive trust in all evidence; is the basic probability assignment of a single piece of evidence after modification, which respectively represents the probability of the evidence from different sources on the event Trust, It is the operator of Dempster's combination rule, which is used to combine the basic probability assignments of two or more pieces of evidence.

[0109] In summary, the final fusion output data is obtained and transmitted to the host computer; the host computer presents a simple and easy-to-use human-computer interface and various specific functions to the end user. At present, two modules, the interface display module and the logic control module, have been preliminarily designed; the interface display module is responsible for drawing and refreshing the application interface, obtaining the final output data from the edge device five, and drawing it in real time to update it to the screen in chronological order; the logic control module is responsible for realizing user interaction operations, allowing users to download results and control the start and end of the entire smart audio platform according to their own needs.

[0110] Figure 13 This is a schematic diagram of the structure of a distributed sound detection device for hospital wards provided in an embodiment of the present application. Figure 13The distributed sound detection device 1300 includes: a data collection module 1301 for collecting environmental sound data; wherein the sound data includes first audio data, second audio data, and third audio data collected from a bed and a door in a hospital ward respectively; The sound data feature extraction module 1302 is configured to perform feature extraction on the sound data to obtain audio data; the feature extraction includes performing MFCC feature extraction on the first audio data and the second audio data, and performing BEATS model feature extraction on the third audio data, to obtain first MFCC audio data, second MFCC audio data, and third BEATS audio data; A multimodal sound event detection model construction module 1303 is used to construct a multimodal sound event detection model; the multimodal sound event detection model includes: a TDY model for processing the first MFCC audio data, an FDY model for processing the second MFCC audio data, and a SKCRNN model for processing the third BEATS audio data; A sound category probability detection module 1304 is configured to input the audio data into the sound event detection model for processing, and obtain sound category prediction probability results based on different models; The sound category prediction probability fusion module 1305 is used to fuse the sound category prediction probability results obtained by the multiple models using a JS-based weighted fusion algorithm to obtain a final sound category prediction probability result.

[0111] The distributed sound detection device 1300 provided in the embodiment of the present application is used to execute the technical solution provided in the aforementioned detection method embodiment. Its implementation principle and technical effects are similar to those in the embodiment of the aforementioned method and will not be repeated here.

[0112] Figure 14 The present invention provides a structural diagram of a distributed sound detection device for hospital wards, see Figure 14 The detection device 1400 includes a processor 1401 and a memory 1402 communicatively connected to the processor 1401; Memory 1402 stores computer-executable instructions; The processor 1401 executes the computer-executable instructions stored in the memory 1402 to implement the technical solution of the aforementioned detection method.

[0113] In the detection device 1400 described above, the memory 1402 and processor 1401 are directly or indirectly electrically connected to each other to enable data transmission or interaction. For example, these components can be electrically connected via one or more communication buses or signal lines, such as a bus. The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus. Buses can be categorized as address buses, data buses, control buses, etc., but this does not mean there is only one bus or only one type of bus. Memory 1402 stores computer-executable instructions for implementing the aforementioned emergency call method, including at least one software functional module stored in memory 1402 in the form of software or firmware. Processor 1401 executes various functional applications and data processing by running the software programs and modules stored in memory 1402.

[0114] The memory 1402 includes at least one type of readable storage medium, including but not limited to random access memory (RAM), read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), and electrically erasable programmable read-only memory (EEPROM). The memory 1402 is used to store programs, and the processor 1401 executes the programs after receiving execution instructions. Furthermore, the software programs and modules in the memory 1402 may also include an operating system, which may include various software components and / or drivers for managing system tasks (e.g., memory management, storage device control, power management, etc.), and may communicate with various hardware or software components to provide an operating environment for other software components.

[0115] The processor 1401 can be an integrated circuit chip with signal processing capabilities. The above-mentioned processor 1401 can be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), a digital signal processor (DSP), an application-specific integrated circuit (ASIC), etc. It can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of this application. The general-purpose processor can be a microprocessor, or the processor 801 can also be any conventional processor, etc.

[0116] The detection device 1400 is used to execute the technical solution provided by the aforementioned detection method embodiment. Its implementation principle and technical effects are similar to those in the aforementioned method embodiment and will not be repeated here.

[0117] An embodiment of the present application further provides a computer-readable storage medium, in which computer-executable instructions are stored. When the computer-executable instructions are executed, they are used to implement the technical solution of the aforementioned method for calling for help.

[0118] The computer-readable storage medium may be implemented by any type of volatile or non-volatile memory device, or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk. The computer-readable storage medium may be any available medium that can be accessed by a general-purpose or special-purpose computer.

[0119] An exemplary readable storage medium is coupled to the processor, so that the processor can read information from the readable storage medium and write information to the readable storage medium. Of course, the readable storage medium can also be an integral part of the processor. The processor and the readable storage medium can be located in an application specific integrated circuit (ASIC). Of course, the processor and the readable storage medium can also be present as discrete components in the control device of the distress call device.

[0120] An embodiment of the present application further provides a computer program product, including a computer program, which, when executed, is used to implement the technical solution of the aforementioned method for calling for help.

[0121] In the above embodiments, those skilled in the art will appreciate that the above method embodiments can be implemented in whole or in part via software, hardware, firmware, or any combination thereof. When implemented using software, they can be implemented in whole or in part in the form of a computer program product. A computer program product comprises one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions according to the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless network, microwave, etc.) means. The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or data center that integrates one or more available media. Available media may be magnetic media (eg, floppy disks, hard disks, magnetic tapes), optical media (eg, DVDs), or semiconductor media (eg, solid-state drives (SSDs)).

[0122] In the above embodiments, the description of each embodiment has its own emphasis. For parts not described in detail in a particular embodiment, please refer to the relevant description of other embodiments. The technical features of the above embodiments can be combined in any way. To keep the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0123] Those skilled in the art will readily appreciate other embodiments of the present invention after considering the specification and practicing the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of the present invention that follow the general principles of the present invention and include common knowledge or customary techniques in the art not disclosed herein. The description and examples are to be considered merely as exemplary, and the true scope and spirit of the present application are indicated by the appended claims.

[0124] It should be understood that the present application is not limited to the exact structure described above and shown in the drawings, and that various modifications and changes may be made without departing from the scope thereof. The scope of the present application is limited only by the appended claims.

Claims

1. A distributed sound detection method for hospital wards, characterized in that: include: Collecting sound data; wherein the sound data includes first audio data, second audio data, and third audio data collected from a bed and a door in a hospital ward, respectively; Performing feature extraction on the sound data to obtain audio data; the feature extraction includes: performing MFCC feature extraction on the first audio data and the second audio data respectively and performing BEATS model feature extraction on the third audio data to obtain first MFCC audio data, second MFCC audio data, and third BEATS audio data; Constructing a multimodal sound event detection model; the multimodal sound event detection model includes: a TDY model for processing the first MFCC audio data, an FDY model for processing the second MFCC audio data, and a SKCRNN model for processing the third BEATS audio data; Inputting the audio data into the sound event detection model for processing to obtain sound category prediction probability results based on different models; The sound category prediction probability results obtained by the multiple models are fused using a JS-based weighted fusion algorithm to obtain the final sound category prediction probability result.

2. The method according to claim 1, characterized in that The collected sound data includes first audio data, second audio data and third audio data collected from the bed and the door in the hospital ward respectively; including: Microphone array 1 and microphone array 2, which are set near the beds in the hospital ward, are used to collect audio data in the hospital ward to obtain first audio data and second audio data, and microphone array 3, which is set at the door of the hospital ward, is used to collect audio data in the hospital ward to obtain third audio data, and the above audio data are transmitted to edge device 1 for processing.

3. The method according to claim 2, characterized in that The feature extraction of the sound data is performed to obtain audio data; the feature extraction comprises: performing MFCC feature extraction on the first audio data and the second audio data respectively and performing BEATS model feature extraction on the third audio data to obtain first MFCC audio data, second MFCC audio data and third BEATS audio data; comprising: Edge device one performs MFCCT feature extraction on the first audio data and the second audio data obtained to obtain first MFCC audio data and second MFCC audio data, and sends them to edge device two and edge device three respectively; Edge device one performs BEATS model feature extraction on the acquired third audio data to obtain third BEATS audio data, and sends it to edge device four.

4. The method according to claim 3, characterized in that The step of inputting the audio data into the sound event detection model for processing to obtain sound category prediction probability results based on different models comprises: The TDY model in the second edge device processes the first MFCC audio data to obtain a sound category prediction probability result based on the TDY model; The FDY model in the third edge device processes the second MFCC audio data to obtain a sound category prediction probability result based on the FDY model; The SKCRNN model in edge device four processes the third BEATS audio data to obtain the sound category prediction probability result based on the SKCRNN model.

5. The method according to claim 4, characterized in that The TDY model in the second edge device processes the first MFCC audio data to obtain a sound category prediction probability result based on the TDY model; specifically, the process includes: Convolving the first MFCC audio data with a convolution kernel to obtain N convolution results; The N convolution results are fused through the attention matrix to obtain the sound category prediction probability result based on the TDY model processing.

6. The method according to claim 5, characterized in that The first MFCC audio data is convolved with a convolution kernel to obtain N convolution results; the specific method is: in, is the convolution kernel, size is , is the bias term; and is the number of channels of model input and output; and are the number of time slices of input and output respectively; and are the number of frequency segments of input and output respectively; It is the time-frequency domain input feature of the model; that is, the first MFCC audio data; is the convolution operation; is the output of the convolution, with a total of indivual; The N convolution results are fused through the attention matrix to obtain the sound category prediction probability result based on the TDY model. The specific method is: Among them, the attention matrix represent dimensional attention weight, is the output of the temporal dynamic convolution; represents element-wise multiplication; Represents the nonlinear activation function ReLU; Represents the SoftMax function; Represents a one-dimensional convolution operation.

7. The method according to claim 6, characterized in that The FDY model in the edge device three processes the second MFCC audio data to obtain a sound category prediction probability result based on the FDY model; specifically: The second MFCC audio data is processed by average pooling, and then processed by one-dimensional convolution, batch normalization and activation function, one-dimensional convolution and SoftMax function in sequence to obtain attention weight; The attention weights are processed by convolution kernel to obtain an overall weight matrix, which is then subjected to two-dimensional convolution. The output of the convolution is then subjected to batch normalization and activation function to obtain the sound category prediction probability result based on the FDY model.

8. The method according to claim 7, characterized in that The attention weight is processed by convolution kernel to obtain the overall weight matrix, and then subjected to two-dimensional convolution. The output of the convolution is subjected to batch normalization and activation function to obtain the sound category prediction probability result based on the FDY model; the specific method is: in, is the attention weight; is the convolution kernel; is the bias term; is the convolution operation; is the output of convolution; and is the number of channels of model input and output; and are the number of time slices of input and output respectively; and are the number of frequency segments of input and output respectively; It is the time-frequency domain input feature of the model, that is, the second MFCC audio data; The output of the convolution is then batch normalized and activated to obtain the sound category prediction probability based on the FDY model.

9. The method according to claim 8, characterized in that The SKCRNN model in the fourth edge device processes the third BEATS audio data to obtain the sound category prediction probability result based on the SKCRNN model; specifically: Processing the third BEATS audio data through a convolution module, wherein the convolution module processing includes one-dimensional convolution processing, batch normalization processing, activation function, maximum pooling layer processing, and drop layer processing to obtain an output of the convolution module; The output of the convolution module is input into the spatial attention module and the temporal attention module for processing, respectively, to obtain the spatial attention weight matrix and the temporal attention weight matrix; The spatial attention weight matrix and the temporal attention weight matrix are added element by element and divided by 2, and multiplied with the output of the convolution module. The output result is then passed through the bidirectional GRU layer to obtain the output result based on the SKCRNN model. After batch normalization processing and activation function, the sound category prediction probability result based on the SKCRNN model is obtained.

10. The method according to claim 9, characterized in that The output of the convolution module is input into the spatial attention module and the temporal attention module for processing, respectively, to obtain the spatial attention weight matrix and the temporal attention weight matrix; specifically: The output of the convolution module is input into the spatial attention module for processing, specifically: The output of the convolution module is input into the spatial attention module, and is subjected to average pooling and maximum pooling respectively to obtain two different temporal context information; The above context information is then fed into the two shared dense layers, and the spatial attention weight is generated by element-by-element summation and output of the feature vector, which is then processed by the sigmoid function to obtain the spatial attention weight matrix , the size is ; The output of the convolution module is input into the temporal attention module for processing, specifically: The output of the convolution module is input into the temporal attention module, and is processed by average pooling and maximum pooling respectively. Then, through cascade processing, the cascaded feature vector is input into the two-dimensional convolution to obtain the temporal attention weight; Then process it through the sigmoid function to get the time attention weight matrix , the size is .

11. The method according to claim 10, characterized in that The spatial attention weight matrix and the temporal attention weight matrix are added element by element and divided by 2, and multiplied by the output of the convolution module. The output result is then passed through the bidirectional GRU layer to obtain the output result based on the SKCRNN model. The specific method is: in, It is the time-frequency domain input feature of the model, that is, the third BEATS audio data; and is the number of channels of model input and output; and are the number of time slices of input and output respectively; and are the number of frequency segments of input and output respectively; It is the element-wise addition operation; The output result passes through the bidirectional GRU layer, which consists of a forward GRU layer and a reverse GRU layer. The bidirectional GRU merges information from two opposite directions. The specific method is: in, Represents the SoftMax function; represents element-wise multiplication; and They are update gate and reset gate respectively; activation Is the previous activation and candidate activation Linear interpolation between ; is a trainable parameter; is the output of the bidirectional GRU.

12. The method according to claim 11, characterized in that The JS-based weighted fusion algorithm fuses the sound category prediction probability results obtained by the multiple models to obtain the final sound category prediction probability result; specifically, it includes: The sound category prediction probability result based on the FDY model, the sound category prediction probability result based on the TDY model, and the sound category prediction probability result based on the SKCRNN model are transmitted to the edge device 5, and the above results are weightedly fused using the JS weighted fusion algorithm, specifically including: Calculate the JS distance between evidences; the specific calculation method is: in, (k = 1, 2, ... p) represents the different sound categories output by the model, and there are a total of p kinds of sounds; (i=1, 2, ..n) and (j=1, 2, ...n) represents two different models, each model is independent evidence, there are n models, and ; Indicates the The output of the model The probability of a sound; Indicates the The output of the model The probability of a sound; Based on the JS distance between the evidences, the conflict degree between the evidences is calculated; specifically: Among them, Sim is The conflict matrix, It is the amount of evidence; It is Evidence and the degree of conflict between the pieces of evidence; is the conflict degree of each piece of evidence itself, set to 0; Calculate the degree of support between evidences; specifically: in, , indicating evidence and the degree of similarity between them; Express evidence and The degree of conflict between The similarity is obtained by subtracting the conflict degree from 1. The diagonal elements of the matrix are all 1, indicating that each piece of evidence is 100% similar to itself; the elements in other positions are Indicates the relative support between two pieces of evidence, and its value is between 0 and 1; other evidence supports The confidence level is: in, Indicates the Evidence support level; Indicates that except for All other evidence except the first The sum of the support of the evidence; Indicates the Evidence for the The support of the evidence; Calculate the weights of evidence; specifically: in, is a specific piece of evidence to be calculated The weight of Indicates the Evidence support level; Represents from the set Take the maximum value among From 1 to ; Revise the original BPA and modify the BPA of each piece of evidence; specifically: in, represents the modified basic probability assignment (BPA); is a specific piece of evidence to be calculated The weight of Representing an event ; The Dempster combination rule is used for fusion to obtain the final fusion output data, that is, the final sound category prediction probability result; specifically: in, It is Sound category fusion The basic probability assignment (BPA) after the model represents the comprehensive trust in all evidence; It is the basic probability assignment of a single piece of evidence after modification, indicating the effect of evidence from different sources on the event. Trust It is the operator of Dempster's combination rule, which is used to combine the basic probability assignments of two or more pieces of evidence.

13. A distributed sound detection device for hospital wards, characterized in that: include: A data acquisition module for collecting environmental sound data, wherein the sound data includes first audio data, second audio data, and third audio data collected from a bed and a door in a hospital ward, respectively; A sound data feature extraction module is used to perform feature extraction on the sound data to obtain audio data; the feature extraction includes: performing MFCC feature extraction on the first audio data and the second audio data respectively and performing BEATS model feature extraction on the third audio data to obtain first MFCC audio data, second MFCC audio data and third BEATS audio data; A multimodal sound event detection model construction module, used to construct a multimodal sound event detection model; the multimodal sound event detection model includes: a TDY model for processing the first MFCC audio data, an FDY model for processing the second MFCC audio data, and a SKCRNN model for processing the third BEATS audio data; a sound category probability detection module, configured to input the audio data into the sound event detection model for processing, and obtain sound category prediction probability results based on different models; The sound category prediction probability fusion module is used to fuse the sound category prediction probability results obtained by the multiple models using a JS-based weighted fusion algorithm to obtain the final sound category prediction probability result.

14. A distributed sound detection device for hospital wards, characterized in that: comprising a processor, and a memory communicatively connected to the processor; The memory stores computer-executable instructions; The processor executes the computer-executable instructions stored in the memory to implement the method according to any one of claims 1 to 12.

15. A distributed sound detection system for hospital wards, characterized in that: comprising the distributed sound detection device as claimed in claim 14, and a server; The server is used to receive the final sound category prediction probability result output by the distributed sound detection device, and perform display and logic control.