Alarm device based on artificial intelligence voice analysis
By laying audio sensors in key monitoring areas, identifying environmental fluctuations and building an abnormal keyword database, and dynamically adjusting the alarm threshold, the existing alarm devices have solved the problem of false alarms and missed reports in complex environments, achieving smarter and timely abnormal warnings, and improving the safety of public places.
Patent Information
- Application Number
- CN202510073594.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-17
- Publication Date
- 2025-08-05
- Estimated Expiration
- 2045-01-17
AI Technical Summary
The existing alarm devices lack accurate means of identifying audio content in complex environments, resulting in frequent false alarms and missed reports, and the inability to accurately identify complex voice abnormalities, affecting the safety of public places.
The audio sensing layout module determines the key monitoring area and arranges audio acquisition sensors, combines the environment fluctuation recognition module to identify environmental feature information, builds an abnormal keyword database, and uses a voice abnormality analyzer to perform real-time abnormality recognition, and dynamically adjusts the alarm threshold to adapt to different environments.
It improves the accuracy and response speed of the alarm device in complex environments, reduces false alarms and missed reports, enhances the accurate recognition of voice content, and improves the level of security guarantee in public places.
Smart Images

Figure CN119811377B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of speech analysis technology, and in particular to an alarm device based on artificial intelligence speech analysis. Background Art
[0002] Safety and security have always been a key concern in public places, such as campuses, parks, and other crowded areas. To promptly detect anomalies and protect personnel, an increasing number of alarm devices are being widely used in these settings. Audio-based alarm devices, thanks to their ability to detect changes in ambient sound, have become a common monitoring method. These devices use audio sensors to detect environmental sound changes and typically employ a fixed threshold-based alarm system. When ambient noise or sound exceeds a preset threshold, an alarm is automatically triggered, providing an early warning. However, fixed thresholds cannot adapt to varying environments and time periods. For example, in crowded environments with high background noise, unnecessary alarms may be triggered due to misidentification of background sounds, resulting in false alarms. Conversely, in quiet or low-noise environments, actual abnormal sounds may not be detected in a timely manner, resulting in missed alarms. Furthermore, existing alarm devices typically rely solely on simple changes in volume or frequency, lacking analysis of specific speech content and unable to identify complex speech anomalies such as arguments or calls for emergency assistance. This results in poor performance in complex speech environments, making them unable to accurately identify potential dangers or emergencies, hindering public safety. Summary of the Invention
[0003] This application provides an alarm device based on artificial intelligence voice analysis, which solves the technical problem that traditional alarm devices lack the means to accurately identify audio content in complex environments, resulting in frequent false alarms and missed alarms, and achieves the technical effect of improving the accuracy and response speed of alarms in public places.
[0004] In view of the above problems, the present application provides an alarm device based on artificial intelligence voice analysis. The device includes: an audio sensing layout module for determining Q key monitoring areas in a target park, arranging audio acquisition sensors based on the Q area scene information of the Q key monitoring areas to obtain Q audio acquisition sensor layout arrays, where each audio acquisition sensor area layout array corresponds to one key monitoring area, and Q is an integer greater than or equal to 1; an environmental fluctuation recognition module for collecting Q area environmental feature information sequences of the Q key monitoring areas within a preset monitoring window, identifying the information fluctuation degree of the Q area environmental feature information sequences to obtain Q area environmental fluctuation coefficients; an alarm threshold configuration module for determining Q area alarm thresholds based on the magnitudes of the Q area environmental fluctuation coefficients, where the area environmental fluctuation coefficient is proportional to the area alarm threshold; a historical voice data collection module for extracting the voice data collected by the Q audio acquisition sensor layout arrays within a preset historical window to obtain Q historical voice data record sequences; an abnormal keyword library module for traversing the Q historical voice data record sequences to construct an abnormal voice data keyword library to obtain Q area abnormal keyword libraries; a voice anomaly analyzer module for constructing Q area intelligent voice anomaly analyzers based on the Q area abnormal keyword libraries and the Q area alarm thresholds; an anomaly recognition module for performing real-time voice data collection on the Q audio acquisition sensor layout arrays to obtain Q real-time voice data, using the Q area intelligent voice anomaly analyzers to perform anomaly recognition on the Q real-time voice data, and obtaining alarm information according to the Q recognition results.
[0005] One or more technical solutions provided in the present application have at least the following technical effects or advantages:
[0006] The key monitoring areas of the target park are determined by the audio sensing layout module, and audio acquisition sensors are arranged based on the regional scene information, providing a targeted layout for subsequent audio data acquisition and ensuring the acquisition of effective audio data in the key monitoring areas. The environmental fluctuation recognition module collects the regional environmental characteristic information and identifies its fluctuation degree, enabling the real-time perception of environmental changes and providing a basis for the reasonable setting of the alarm threshold to avoid false alarms or missed alarms caused by environmental changes. The alarm threshold configuration module determines the alarm threshold according to the environmental fluctuation coefficient, capable of dynamically adjusting the threshold size according to the environmental conditions in different regions, improving the alarm accuracy and environmental adaptability. The historical voice data acquisition module extracts historical voice data, which are the basic materials for constructing the abnormal keyword library and providing a data basis for subsequent accurate identification of abnormal voices. The abnormal keyword library module traverses the historical voice data to construct the abnormal keyword library, capable of establishing an abnormal voice pattern related to the actual situation in the region and providing specific criteria for subsequent abnormal analysis. The voice anomaly analyzer module constructs a regional intelligent voice anomaly analyzer based on the abnormal keyword library and the alarm threshold, capable of performing targeted abnormal voice analysis according to the situations in different regions. The anomaly recognition module collects the real-time voice data and uses the constructed regional intelligent voice anomaly analyzer for anomaly recognition, and finally obtains the alarm information according to the recognition result, realizing a rapid response to abnormal events and ensuring that the alarm device makes a reaction in the first time.
[0007] In summary, through audio acquisition, environmental fluctuation analysis, construction of an abnormal keyword library, and intelligent voice anomaly analysis, by customarily arranging audio sensors in different monitoring regions and dynamically adjusting the alarm threshold in combination with the environmental fluctuation coefficient, this application can adapt to different environmental conditions, significantly reducing false alarms and missed alarms. In addition, by constructing the abnormal keyword library of historical voice data and the intelligent voice anomaly analyzer, it can accurately identify abnormal signals in complex voices, thus improving the alarm accuracy and response speed. Overall, it not only enhances the adaptability to complex environments but also enhances the accurate recognition of voice content, achieving a more intelligent and timely abnormal early warning and effectively improving the safety guarantee level of public places.
[0008] The above description is only an overview of the technical solution of this application. In order to be able to understand the technical means of this application more clearly, it can be implemented according to the content of the specification. And in order to make the above and other purposes, features and advantages of this application more obvious and understandable, the following specifically gives the specific implementation manners of this application. Brief Description of the Drawings
[0009] Figure 1 It is a schematic structural diagram of an alarm device based on artificial intelligence voice analysis provided by an embodiment of this application.
[0010] Figure 2 Schematic diagram of the process for obtaining Q regional abnormal keyword libraries in the alarm device based on artificial intelligence voice analysis provided by an embodiment of the present application.
[0011] Figure 3 Schematic diagram of the process for constructing Q regional intelligent voice anomaly analyzers in the alarm device based on artificial intelligence voice analysis provided by an embodiment of the present application.
[0012] Explanation of reference numerals: Audio sensing layout module 10, environmental fluctuation recognition module 20, alarm threshold configuration module 30, historical voice data acquisition module 40, abnormal keyword library module 50, voice anomaly analyzer module 60, anomaly recognition module 70. Detailed implementation manners
[0013] By providing an alarm device based on artificial intelligence voice analysis in an embodiment of the present application, the technical problem that traditional alarm devices lack accurate recognition means for audio content in complex environments, resulting in frequent false alarms and missed alarms, is solved, and the technical effect of improving the accuracy and response speed of public place alarms is achieved.
[0014] As Figure 1 shown, an embodiment of the present application provides an alarm device based on artificial intelligence voice analysis, and the device includes:
[0015] An audio sensing layout module 10, configured to determine Q key monitoring areas of a target park, perform layout of audio acquisition sensors based on Q regional scene information of the Q key monitoring areas, and obtain Q audio acquisition sensor layout arrays, where each audio acquisition sensor regional layout array corresponds to one key monitoring area, and Q is an integer greater than or equal to 1.
[0016] Specifically, the target park refers to a specific park that needs to be monitored for security, such as a school, an industrial park, etc. The key monitoring areas are areas within the target park that need to focus on security situations, such as stairways, laboratories, etc. in a school. These areas may have higher security risks. The regional scene information is information related to the environment, function, etc. of the key monitoring area. For example, the regional scene information of the classroom area in a school includes the space size, personnel flow, daily noise level, etc. These information helps to reasonably layout the audio acquisition sensors.
[0017] First, determine the set of areas in the target park that need to be monitored with priority according to the overall planning and security requirements analysis of the target park. For the convenience of subsequent description, use Q to represent the number of areas to be monitored with priority, where Q is an integer greater than or equal to 1. Then, based on the scene information of these areas to be monitored with priority (such as the size, shape, environmental noise level, etc. of the area), determine the optimal layout positions and quantities of audio acquisition sensors. For example, in a large conference room, sensors need to be arranged at the four corners to ensure full coverage. In this way, each area to be monitored with priority will have a dedicated audio acquisition sensor array to ensure that the sounds in that area can be captured.
[0018] Through the audio sensing layout module 10, audio acquisition sensors are specifically arranged in the key areas of the target park to ensure the acquisition of effective audio data in the areas to be monitored with priority, providing a basic data source for subsequent operations such as abnormal identification, and improving the effectiveness and pertinence of audio data acquisition.
[0019] The environmental fluctuation identification module 20 is used to collect Q sequences of regional environmental characteristic information of the Q areas to be monitored with priority within a preset monitoring window, and identify the degree of information fluctuation of the Q sequences of regional environmental characteristic information to obtain Q regional environmental fluctuation coefficients.
[0020] Specifically, the preset monitoring window is a preset time period during which the environmental characteristic information of the areas to be monitored with priority is collected and analyzed, and can be set to 1 day or 1 week, etc. The sequence of regional environmental characteristic information is a series of information about the environmental characteristics of the areas to be monitored with priority collected within the preset monitoring window, and these characteristics include noise intensity, sound frequency distribution, etc. The degree of information fluctuation refers to the fluctuation size of environmental characteristic information in the time series.
[0021] Within the preset monitoring window, use devices such as audio acquisition sensors to collect Q sequences of regional environmental characteristic information of the Q areas to be monitored with priority, and obtain Q sequences of regional environmental characteristic information. For example, in the playground area of a school campus, within the preset monitoring window of morning exercises, the sounds collected include environmental characteristic information such as the noise of students and the broadcast sound. Then, through video analysis software, the degree of fluctuation of the Q sequences of regional environmental characteristic information is identified to generate the corresponding Q regional environmental fluctuation coefficients. For example, compare the changes in noise intensity at different time points, calculate the fluctuation of the regional environmental characteristic information of each area to be monitored with priority within the preset monitoring window, and generate the corresponding regional environmental fluctuation coefficient. This regional environmental fluctuation coefficient is a quantitative representation of the degree of environmental fluctuation, reflecting the strength of sound fluctuation in the environment. A high fluctuation coefficient means that the environmental sound changes greatly.
[0022] The environmental fluctuation recognition module 20 can sense the environmental changes in the key monitoring areas in real time, accurately obtain the regional environmental fluctuation coefficients, provide a basis for the reasonable setting of alarm thresholds, and avoid false alarms or missed alarms caused by environmental fluctuations.
[0023] The alarm threshold configuration module 30 is used to determine Q regional alarm thresholds based on the magnitudes of the Q regional environmental fluctuation coefficients, where the regional environmental fluctuation coefficient is proportional to the regional alarm threshold.
[0024] Specifically, the regional alarm threshold is a numerical standard set for each key monitoring area. When the audio-related data exceeds this standard, an alarm will be triggered.
[0025] The Q regional alarm thresholds are determined according to the magnitudes of the Q regional environmental fluctuation coefficients obtained by the environmental fluctuation recognition module 20. The regional environmental fluctuation coefficient is proportional to the regional alarm threshold. For regions with a large fluctuation coefficient, their alarm thresholds will be correspondingly increased. For example, at the entrance of a campus with a large flow of people, the regional environmental fluctuation coefficient is large, and the alarm threshold may be set higher than that of a quiet library. Simple mathematical calculation tools or configuration software can be used to determine the alarm threshold according to the proportional relationship.
[0026] By dynamically setting the alarm threshold through the alarm threshold configuration module 30, the size of the alarm threshold can be dynamically adjusted according to the environmental conditions of different regions, adapting to the security monitoring requirements under different environments, reducing false alarms caused by environmental changes, and improving the accuracy of alarms.
[0027] The historical voice data acquisition module 40 is used to extract the voice data collected by the Q audio acquisition sensor deployment arrays within a preset historical window, and obtain Q historical voice data record sequences.
[0028] Specifically, the preset historical window is similar to the preset monitoring window and is a specific past time period for collecting historical voice data, such as the past 1 week or 1 month, etc. The voice data collected within the preset historical window is extracted from the Q audio acquisition sensor deployment arrays to obtain Q historical voice data record sequences. For example, in the library area of a school campus, within the past 1 month (preset historical window), voice data such as the conversations and page-turning sounds of readers are extracted from the voice data collected by the deployed sensors. The Q historical voice data record sequences provide a large amount of historical data for constructing the abnormal keyword library, enabling the device to learn and identify abnormal sound patterns based on historical data, and improving the accuracy of abnormal voice recognition.
[0029] The abnormal keyword library module 50 is used to traverse the Q historical voice data record sequences to construct an abnormal voice data keyword library, and obtain Q regional abnormal keyword libraries.
[0030] Specifically, the regional anomaly keyword library is a database containing keywords in historical voice data that are determined to be possibly related to abnormal situations. Traverse Q historical voice data record sequences, conduct data mining and voice analysis, and construct an abnormal voice data keyword library for each key monitoring area, obtaining Q regional anomaly keyword libraries. For example, in an industrial park, if words related to production anomalies such as "machine failure" appear multiple times in historical voice data, they are included in the regional anomaly keyword library. By constructing the regional anomaly keyword library through the anomaly keyword library module 50, it is possible to more accurately identify abnormal voices that may pose dangerous situations, providing a more accurate basis for abnormal voice recognition.
[0031] The voice anomaly analyzer module 60 is used to construct Q regional intelligent voice anomaly analyzers based on the Q regional anomaly keyword libraries and Q regional alarm thresholds.
[0032] Specifically, the regional intelligent voice anomaly analyzer is an analysis tool constructed by combining the regional anomaly keyword library and the regional alarm threshold, used to perform anomaly analysis on real-time voice data, determine whether it contains anomaly keywords, and then determine whether there are abnormal situations. Based on the Q regional anomaly keyword libraries and Q regional alarm thresholds, use programming algorithms, machine learning models, etc. to construct Q regional intelligent voice anomaly analyzers. For example, a program can be written in Python, inputting parameters such as the keyword library and the alarm threshold to construct the analyzer. It is also possible to train a machine learning model (such as a neural network) using the Q regional anomaly keyword libraries and Q regional alarm thresholds to construct the intelligent voice anomaly analyzer. The Q regional intelligent voice anomaly analyzers correspond one-to-one with the Q key monitoring areas, respectively used to identify abnormal voices in specific key monitoring areas, further improving the accuracy of the alarm.
[0033] By constructing the regional intelligent voice anomaly analyzer through the voice anomaly analyzer module 60, it is possible to perform targeted abnormal voice analysis according to the situations of different regions, improving the accuracy and effectiveness of abnormal voice analysis.
[0034] The anomaly recognition module 70 is used to collect real-time voice data from the Q audio acquisition sensor arrays, obtain Q real-time voice data, use the Q regional intelligent voice anomaly analyzers to perform anomaly recognition on the Q real-time voice data, and obtain alarm information based on the Q recognition results.
[0035] Specifically, real-time voice data is the voice data collected by audio acquisition sensors at the current moment, which can reflect the actual sound situation in the key monitoring area at present. The Q audio acquisition sensors are arranged in an array to collect real-time voice data, and Q real-time voice data are obtained. Then, the constructed Q regional intelligent voice anomaly analyzers are used to identify anomalies in these real-time voice data respectively. For example, if keywords such as "fire" or "explosion" are detected in a laboratory on a campus and exceed the alarm threshold of the area, the module will generate alarm information, such as specific light alarms, sound alarms, etc.
[0036] Through the anomaly recognition module 70, the real-time voice data can be effectively recognized, and it can accurately judge whether there is an abnormal situation in the key monitoring area, so as to send out alarm information in time and ensure the safety of the target park.
[0037] Furthermore, as Figure 2 shown, the abnormal keyword library module 50 is also used to perform the following steps:
[0038] Step P51: Extract the first historical voice data records located at the beginning from the Q historical voice data record sequences respectively, and obtain Q first historical voice data records.
[0039] Step P52: Extract the keywords related to abnormal situations and the sentences where the keywords are located from the Q first historical voice data records, and obtain Q first abnormal keyword sets and Q first abnormal key sentence sets.
[0040] Step P53: Add the Q first abnormal keyword sets and the Q first abnormal key sentence sets into the Q first memory banks respectively.
[0041] Step P54: Taking the Q first memory banks as indexes, perform keyword association extraction on the Q second historical voice data records located at the second place extracted from the Q historical voice data record sequences, and combine the sentences where the extracted keywords are located to obtain Q second abnormal keyword sets and Q second abnormal key sentence sets.
[0042] Step P55: Add the Q second abnormal keyword sets and the Q second abnormal key sentence sets into the first memory bank set respectively for memory bank update, and obtain Q second memory banks.
[0043] Step P56: Based on the Q second memory banks, perform keyword association recognition on the Q historical voice data record sequences until the end of the sequence, and obtain Q target memory banks.
[0044] Step P57: Obtain the Q regional abnormal keyword libraries based on the Q target memory banks.
[0045] Specifically, the memory bank is a library for storing the abnormal keyword set and the key sentence set, and is an intermediate storage structure in the process of constructing the abnormal keyword library.
[0046] In the Q historical voice data record sequences, the first voice data record of each sequence is located respectively and extracted, so as to obtain Q first historical voice data records. For example, there are 3 historical voice data record sequences, and the first voice data records ranked first in these 3 sequences are taken out respectively.
[0047] Perform a detailed analysis on the Q first historical voice data records, and extract the keywords related to abnormal situations and the sentences containing these keywords from them through natural language processing technologies (such as text analysis, semantic analysis, etc.) to form Q first abnormal keyword sets and Q first abnormal key sentence sets respectively. For example, using technologies such as word vector models, identify keywords such as "dangerous" and "accident" that may be related to abnormal situations, and find the sentences containing these keywords.
[0048] Store the previously obtained Q first abnormal keyword sets and Q first abnormal key sentence sets into Q first memory banks respectively. Each first memory bank contains the first abnormal keyword set and the first abnormal key sentence set extracted from the historical voice data record sequence corresponding to a key monitoring area.
[0049] Taking the Q first memory banks as indexes, perform keyword association extraction on the Q second historical voice data records. That is, according to the first abnormal keywords and first abnormal key sentences in the first memory bank, search for the related keywords in the second historical voice data records, and combine with the sentences where they are located to obtain Q second abnormal keyword sets and Q second abnormal key sentence sets. In the process of searching for abnormal keywords and abnormal key sentences, keyword matching algorithms, text similarity calculation algorithms (such as cosine similarity algorithms), etc. can be used to determine the association degree between keywords, so as to perform association extraction.
[0050] Add the Q second abnormal keyword sets and Q second abnormal key sentence sets into the Q first memory banks respectively, and update the first memory banks according to the update rules such as merging and deduplication to obtain Q second memory banks. For example, if there are new keywords in the second abnormal keyword set, add them to the corresponding structure of the first memory bank, and do not add them if there are duplicate keywords. Through this update operation, the memory bank can evolve and be perfected as new data is added.
[0051] Based on the Q second memory banks, continue to perform keyword association recognition on the Q historical voice data record sequences. According to the aforementioned association extraction method, gradually process the subsequent data in the historical voice data record sequences, continuously update the memory banks, and each time, based on the current memory bank, process the next historical voice data record. Loop this process until the entire sequence is processed, and finally obtain Q target memory banks.
[0052] According to the information in the Q target memory banks, extract the keywords in the target memory banks that can best represent abnormal situations, and sort out and screen the final Q regional abnormal keyword libraries.
[0053] By gradually associating and updating the memory banks to construct the regional abnormal keyword library, it is possible to more comprehensively and accurately mine the keywords related to abnormal situations from historical voice data. Compared with the simple traversal construction method, this method can consider the association relationships between keywords and continuously optimize the extraction of keywords by using the information in the previous memory banks, thereby improving the quality of the abnormal keyword library and further enhancing the accuracy of voice anomaly analysis.
[0054] Further, step P54 includes:
[0055] Step P541: Use the Q first abnormal keyword sets stored in the Q first memory banks as indexes to perform keyword similarity matching on the Q second historical voice data records.
[0056] Step P542: Extract the keywords whose matching results meet the preset similarity threshold to obtain Q matching second abnormal keyword sets.
[0057] Step P543: Use the Q first abnormal key sentence sets stored in the Q first memory banks as indexes to perform semantic similarity matching on the Q second historical voice data records.
[0058] Step P544: Extract the sentences whose matching results meet the preset semantic similarity threshold to obtain Q matching abnormal sentence sets.
[0059] Step P545: Respectively extract the keywords related to abnormal situations from the Q first abnormal key sentence sets to obtain Q semantically matching second abnormal keyword sets.
[0060] Step P546: Perform the union operation on the Q matching second abnormal keyword sets and the Q semantically matching second abnormal keyword sets to obtain the Q second abnormal keyword sets.
[0061] Specifically, the preset similarity threshold is a numerical standard preset when performing keyword similarity matching. Only when the similarity between keywords reaches this standard is it considered that the keywords are matched.
[0062] For each set of first abnormal keywords in the first memory bank, similarity is calculated between them and the corresponding keywords in the second historical voice data record. For example, a word vector model is used to convert the first abnormal keyword (such as "fire") and the keywords in the second historical voice data record (such as "fire" and "flame") into vector form, and then the similarity between the two vectors is calculated using the cosine similarity formula.
[0063] For the multiple keyword similarity results calculated, select keywords that meet a preset similarity threshold. For example, if the preset similarity threshold is 0.8, then extract all keywords with a similarity greater than or equal to 0.8. These keywords that meet the preset similarity threshold are aggregated to form a set of Q matching second abnormal keywords.
[0064] For each set of first abnormal key sentences in the first memory bank, semantic similarity is calculated between them and the corresponding sentences in the second historical voice data record. A deep learning-based semantic analysis model, such as a pre-trained BERT model, can be used to input the sentences into the model to obtain semantic vectors. The similarity between the vectors is then calculated to obtain multiple key sentence similarity results. From the multiple key sentence similarity results, sentences that meet a preset semantic similarity threshold are selected to obtain Q sets of matching abnormal sentences. These Q sets of matching abnormal sentences are used as the Q second sets of abnormal key sentences.
[0065] For each set of first abnormal key sentences, natural language processing techniques (such as part-of-speech tagging and named entity recognition) are used to extract keywords related to the abnormal situation. For example, in the sentence "I saw the library on fire," part-of-speech tagging identifies "on fire" as a keyword related to the abnormal situation, thereby obtaining Q semantically matching second abnormal keyword sets.
[0066] The keywords in the Q matching second abnormal keyword sets and the Q semantic matching second abnormal keyword sets are merged, and repeated keywords are removed to finally obtain Q second abnormal keyword sets.
[0067] By calculating keyword similarity and semantic similarity to construct a second set of abnormal keywords, we can more accurately extract keywords related to abnormal situations from the second historical voice data record. Keyword similarity matching directly identifies words similar to existing abnormal keywords, while semantic similarity matching mines potential related information at the sentence level. Combined with the keywords extracted from the first set of abnormal key sentences, the resulting second set of abnormal keywords more comprehensively and accurately reflects the vocabulary related to abnormal situations, laying the foundation for the subsequent construction of a more comprehensive regional abnormal keyword library, thereby improving the accuracy of speech anomaly recognition.
[0068] Further, step P57 includes:
[0069] Extract the abnormal keywords with the top m frequencies in the Q target memory banks respectively to obtain Q regional abnormal keyword libraries, where m is an integer greater than or equal to 3.
[0070] Specifically, for each of the Q target memory banks, count the occurrence frequency of each abnormal keyword. By traversing the data in the target memory bank, a counter can be used to record the number of times each keyword appears. Then, sort all the abnormal keywords in descending order of occurrence frequency. Finally, extract the top m abnormal keywords, which are considered to be the words that can best represent abnormal situations. Here, m is a preset parameter that can be adjusted according to actual needs to balance the coverage and accuracy of the keyword library. To ensure the number of regional abnormal keyword libraries, m is an integer greater than or equal to 3. Classify these high-frequency abnormal keywords into the corresponding Q regions respectively to form Q regional abnormal keyword libraries. For example, in a certain target memory bank in the school campus, it is counted that "fire" appears 10 times, "emergency" appears 7 times, and "rupture" appears 5 times. When m = 3, extract the abnormal keywords "fire", "emergency", and "rupture" to form a regional abnormal keyword library.
[0071] Extracting the regional abnormal keyword library according to the occurrence frequency can preferentially select the abnormal keywords with higher frequencies in the historical voice data. These high-frequency keywords can better represent the possible abnormal situations in the target park, making the regional abnormal keyword library more representative and targeted. In this way, in subsequent voice anomaly analysis, abnormal voices related to these high-frequency abnormal keywords can be more accurately identified, improving the efficiency and accuracy of voice anomaly recognition and alarm.
[0072] Further, as Figure 3 shown, the voice anomaly analyzer module 60 is also used to perform the following steps:
[0073] Step P61: Randomly generate interference data for the Q regional abnormal keyword libraries according to a preset perturbation scale to obtain Q abnormal regional voice data sets.
[0074] Step P62: Combine the Q regional alarm thresholds to perform anomaly recognition on the Q abnormal regional voice data sets to obtain Q anomaly recognition results.
[0075] Step P63: Use the Q abnormal regional voice data sets and the Q anomaly recognition results to perform supervised training on the network layer constructed based on the feedforward neural network until the training converges, and obtain the Q regional intelligent voice anomaly analyzers that are trained and completed.
[0076] Specifically, the preset perturbation scale is a predefined standard used to determine the degree or range of interference when generating interference data. For example, if the preset perturbation scale is 0.5, it means that the fluctuation amplitude when generating interference data is within 0.5 times of a certain benchmark. According to the preset perturbation scale, random generation of interference data is performed on the Q regional abnormal keyword libraries. This process can use data augmentation techniques, such as adding background noise, changing pitch or speed in voice data, etc., to simulate various situations that may occur in the real environment. Exemplary: For a campus area, if the abnormal keyword "emergency" appears in the voice data, through the generation of interference data, some background conversation sounds, traffic flow sounds or other common noises are added to simulate the voice situation in different environments. The interference data generated for each regional abnormal keyword library are respectively aggregated to obtain Q sets of voice data for abnormal regions. By generating data with interference, the training data can be made more diverse, simulating more complex actual situations, and improving the adaptability of the regional intelligent voice anomaly analyzer to various situations.
[0077] Compare each data in the Q sets of voice data for abnormal regions with the corresponding Q regional alarm thresholds to determine whether the alarm threshold is reached, thereby obtaining Q abnormal recognition results. For example, if the abnormal degree of a certain data in the voice data for the abnormal region exceeds the regional alarm threshold, then the abnormal recognition result of this data is abnormal, otherwise it is normal. Combining the regional alarm threshold for abnormal recognition provides accurate supervision data for the training of the regional intelligent voice anomaly analyzer, enabling the regional intelligent voice anomaly analyzer to accurately judge abnormal situations according to the set standard.
[0078] Using a deep learning framework, a feedforward neural network is constructed. For each network layer of the constructed feedforward neural network, the Q sets of voice data for abnormal regions are used as input data, and the Q abnormal recognition results are used as supervision labels (i.e., correct answers) to perform supervised training on the feedforward neural network. During the training process, the weights of the network layer are adjusted according to the difference between the prediction result and the correct label (such as using a loss function like mean square error), and this process is continuously iterated until the training converges (i.e., the loss function no longer significantly decreases), and finally Q regional intelligent voice anomaly analyzers after training are obtained. Through supervised training, the trained regional intelligent voice anomaly analyzers have higher accuracy and reliability, so that they can more accurately identify abnormal situations in actual voice anomaly analysis.
[0079] Furthermore, the audio sensing layout module 10 in the embodiment of the present application is further configured to perform the following steps:
[0080] Step P11: Extract features from the Q regional scene information according to the preset scene features to obtain Q regional scene features.
[0081] Step P12: Use the sensor layout plan recognizer to recognize the Q regional scene features, and obtain Q regional sensor layout plans.
[0082] Step P13: Perform regional sensor layout according to the layout positions and quantities of the audio acquisition sensors in the Q regional sensor layout plans, and obtain the Q audio acquisition sensor layout arrays.
[0083] Specifically, the preset scene features are some feature types or patterns related to the regional scene that are preset in advance, and are parameters used to describe and distinguish different monitoring regional features, such as spatial size, personnel flow direction, whether it is an enclosed space, etc. The regional scene features are the unique scene features of each region extracted from the Q regional scene information according to the preset scene features. For example, if a key monitoring region is a classroom in a school, its regional scene features may include a small space, concentrated personnel flow during breaks, and a semi-enclosed space, etc. The sensor layout plan recognizer is a tool or algorithm that can analyze and recognize regional scene features, and thus obtain a sensor layout plan suitable for this region. It can work based on predefined rules, models or machine learning algorithms. The regional sensor layout plan is a plan for how to layout the audio acquisition sensors for each region, including information such as the layout positions of the sensors (such as the four corners of the classroom, etc.) and the layout quantities (such as 3 sensors are needed, etc.).
[0084] Extract the scene features of the Q regions respectively according to the preset scene features. Data analysis tools and feature engineering algorithms can be used for scene feature extraction. For example, image recognition technology can be used to determine the number of people in the region, and sound analysis tools can be used to determine the background noise level of the region. Through feature extraction, the scene features of each region can be obtained, providing a basis for subsequent sensor layout.
[0085] Input the Q regional scene features into the sensor layout plan recognizer respectively. This sensor layout plan recognizer contains internal rule-based judgment logic or pre-trained machine learning models (such as decision tree models, etc.). Taking the school library region as an example, if the regional scene features are a large space, less personnel flow and quiet, the sensor layout plan recognizer judges a plan to layout a small number of sensors at several key positions in the library (such as the entrance, bookshelf aisles, etc.) according to the internal rules or models, and thus obtains the Q regional sensor layout plans.
[0086] According to the specified layout positions and quantities in the Q regional sensor layout plans, the actual layout work of audio acquisition sensors is carried out in each corresponding area. For example, according to the regional sensor layout plan of a certain meeting room, the specified number of audio acquisition sensors are installed at the corners and positions designated in the plan, and finally Q audio acquisition sensor layout arrays are obtained. Ensure that each area can obtain appropriate audio acquisition coverage, thereby improving the acquisition quality and coverage of sound data, enabling the alarm device to more accurately capture abnormal sound events, and improving the response ability and accuracy of the alarm.
[0087] Furthermore, the environmental fluctuation recognition module 20 is further configured to perform the following steps:
[0088] Step S21: Traverse the Q regional environmental feature information sequences according to a preset environmental feature extraction scale set to perform asynchronous information extraction, and obtain a set of Q asynchronous extraction information subsequence sets.
[0089] Step S22: Perform mean processing on the Q asynchronous extraction information subsequence sets respectively to obtain a set of Q asynchronous extraction information means.
[0090] Step S23: Perform cross-temporal information fluctuation recognition based on the set of Q asynchronous extraction information means to obtain the Q regional environmental fluctuation coefficients.
[0091] Specifically, the preset environmental feature extraction scale set is a set of pre-set standards used to determine the range or degree of asynchronous information extraction from the regional environmental feature information sequence. For example, this set can include different scale standards such as time scale (such as extracting information every 5 minutes) and space scale (such as environmental features of every 10 square meters area).
[0092] For each of the Q regional environmental feature information sequences, traverse according to the standards in the preset environmental feature extraction scale set. For example, if the preset environmental feature extraction scale set stipulates that asynchronous information extraction is performed according to a time interval of every 10 minutes, then data points are extracted from the regional environmental feature information sequence at intervals of 10 minutes to form multiple asynchronous information subsequences, and finally a set of Q asynchronous extraction information subsequence sets are obtained. Each subsequence contains information (such as data within 10 minutes) extracted from a specific regional environmental feature information sequence in a specific asynchronous manner.
[0093] For each of the Q asynchronous extraction information subsequence sets, calculate the average value of all elements therein to obtain a set of Q asynchronous extraction information means. Each asynchronous extraction information mean set contains multiple means, and each mean corresponds to an asynchronous extraction information subsequence and contains the mean value of all data in that subsequence.
[0094] Identify cross-temporal information fluctuations in the data of the Q asynchronous extracted information mean sets. For example, fluctuation calculation methods in time series analysis can be adopted, such as calculating the difference between adjacent means, or more complex statistical models (such as autoregressive moving average models) can be used to analyze the fluctuation of the data, and finally Q regional environmental fluctuation coefficients are obtained.
[0095] Through the above steps, the environmental fluctuation identification module 20 can accurately identify the fluctuations of the environmental characteristic information in each monitoring area. By quantifying the environmental fluctuations, the alarm threshold can be set more reasonably, reducing false alarms and missed alarms, and improving the reliability and effectiveness of the alarm device.
[0096] Further, step S23 includes:
[0097] Calculate the fluctuation variances for the Q asynchronous extracted information mean sets respectively to obtain Q fluctuation variances.
[0098] Traverse and calculate the ratio of each of the Q fluctuation variances to the sum of the Q fluctuation variances to obtain Q regional environmental fluctuation coefficients.
[0099] Specifically, when calculating the regional environmental fluctuation coefficients, for each of the Q asynchronous extracted information mean sets, first use the variance calculation formula to calculate to obtain Q fluctuation variances. Then, calculate the sum of the Q fluctuation variances. Divide each fluctuation variance by this sum of fluctuation variances to obtain Q regional environmental fluctuation coefficients.
[0100] By calculating the fluctuation variances, the fluctuation degree of each asynchronous extracted information mean set can be accurately quantified. Obtaining the regional environmental fluctuation coefficients by comparing the fluctuation variances with the sum can normalize the fluctuation degree of each region to a relative value, enabling direct comparison of the environmental fluctuation situations between different regions and facilitating the identification of regions with larger or smaller fluctuations in overall environmental monitoring and management.
[0101] In summary, the alarm device based on artificial intelligence voice analysis provided by the embodiments of the present application has the following technical effects:
[0102] Through the collaborative work of multiple modules, combining artificial intelligence voice analysis and environmental fluctuation recognition, accurate and intelligent abnormal voice detection and alarm functions are achieved. First, the audio sensor layout module 10 optimizes the sensor layout according to the regional scene information, ensuring that there is a suitable audio collection array in each key monitoring area, improving the coverage rate and accuracy of the collection. Then, the environmental fluctuation recognition module 20 conducts fluctuation analysis on the environmental feature information sequence to obtain the environmental fluctuation coefficients of each region, providing a basis for the adaptive adjustment of subsequent alarm thresholds and reducing the possibility of false alarms and missed alarms. The alarm threshold configuration module 30 then flexibly adjusts the alarm sensitivity of each region according to the change of the environmental fluctuation coefficient, ensuring that the alarm device can work properly in a dynamic environment. The abnormal keyword library module 50 establishes an accurate abnormal voice keyword library through keyword extraction and similarity matching of historical voice data, enabling the alarm device to identify complex voice abnormalities such as disputes and distress calls. The voice abnormality analyzer further improves the accuracy and response speed of abnormal detection by combining the trained intelligent model and identifying abnormalities in real-time voice data.
[0103] Overall, through the cooperation of multiple modules in the embodiments of this application, intelligent abnormal recognition based on environmental fluctuations and voice content is achieved, significantly improving the intelligence level and adaptability of the alarm device, enabling accurate triggering of alarms under different environmental conditions, and enhancing the alarm accuracy and response speed in public places.
[0104] The above description of the disclosed embodiments enables those skilled in the art to implement or use this application. Various modifications to these embodiments will be obvious to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of this application. Therefore, this application will not be limited to these embodiments shown herein, but rather will be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. The alarm device based on artificial intelligence voice analysis is characterized by: The device comprises: An audio sensor deployment module is configured to determine Q key monitoring areas of a target park, deploy audio acquisition sensors based on Q regional scene information of the Q key monitoring areas, and obtain Q audio acquisition sensor deployment arrays, where each audio acquisition sensor regional deployment array corresponds to one key monitoring area, and Q is an integer greater than or equal to 1; an environmental fluctuation identification module, configured to collect Q regional environmental characteristic information sequences of the Q key monitoring areas within a preset monitoring window, identify the degree of information fluctuation on the Q regional environmental characteristic information sequences, and obtain Q regional environmental fluctuation coefficients, wherein the regional environmental characteristic information sequences are information collected within the preset monitoring window on the environmental characteristics of the key monitoring areas, such characteristics including noise intensity and sound frequency distribution. The regional environmental fluctuation coefficient is a quantitative representation of the degree of environmental fluctuation, reflecting the strength of the sound fluctuation in the environment. A high fluctuation coefficient means that the environmental sound changes significantly. an alarm threshold configuration module, configured to determine Q regional alarm thresholds based on the magnitudes of the Q regional environmental fluctuation coefficients, wherein the regional environmental fluctuation coefficient is proportional to the regional alarm threshold; A historical voice data acquisition module is used to extract the voice data collected by the Q audio acquisition sensor arrays within a preset historical window to obtain Q historical voice data record sequences; An abnormal keyword library module is used to traverse the Q historical voice data record sequences to construct an abnormal voice data keyword library and obtain Q regional abnormal keyword libraries; A voice anomaly analyzer module is used to build Q-region intelligent voice anomaly analyzers based on Q-region abnormal keyword libraries and Q-region alarm thresholds; an anomaly recognition module, configured to collect real-time voice data from the array of the Q audio acquisition sensors to obtain Q real-time voice data, perform anomaly recognition on the Q real-time voice data using the Q regional intelligent voice anomaly analyzers, and obtain alarm information based on the Q recognition results; The execution steps of the environmental fluctuation identification module include: According to a preset environmental feature extraction scale set, traverse the Q regional environmental feature information sequences to perform asynchronous information extraction to obtain Q asynchronous extraction information subsequence sets; Performing mean processing on the Q asynchronously extracted information subsequence sets respectively to obtain Q asynchronously extracted information mean value sets; Identify cross-time series information fluctuations based on the Q asynchronously extracted information mean value sets to obtain the Q regional environmental fluctuation coefficients; Performing fluctuation variance calculations on the Q asynchronously extracted information mean value sets respectively to obtain Q fluctuation variances; The Q fluctuation variances are divided by the sum of the Q fluctuation variances to obtain Q regional environmental fluctuation coefficients.
2. The alarm device based on artificial intelligence voice analysis according to claim 1, characterized in that: The execution steps of the abnormal keyword library module include: Extracting the first historical voice data record at the top from the Q historical voice data record sequences respectively to obtain Q first historical voice data records; Extracting keywords related to abnormal situations and sentences containing the keywords from the Q first historical voice data records to obtain Q first abnormal keyword sets and Q first abnormal key sentence sets; Adding the Q first abnormal keyword sets and the Q first abnormal key sentence sets into Q first memory banks respectively; Using the Q first memory banks as indexes, performing keyword association extraction on the Q second historical voice data records located second in the Q historical voice data record sequences, and combining the sentences containing the extracted keywords to obtain Q second abnormal keyword sets and Q second abnormal key sentence sets; Add the Q second abnormal keyword sets and the Q second abnormal key sentence sets to the first memory bank set respectively to update the memory bank, and obtain Q second memory banks; Performing keyword association recognition on the Q historical voice data record sequences based on the Q second memory banks until reaching the end of the sequence, thereby obtaining Q target memory banks; The Q regional abnormal keyword libraries are obtained based on the Q target memory libraries.
3. The alarm device based on artificial intelligence voice analysis as claimed in claim 2, characterized in that: The execution steps of the abnormal keyword library module include: Using the Q first abnormal keyword sets stored in the Q first memory banks as indexes, performing keyword similarity matching on the Q second historical voice data records; Extract keywords whose matching results meet the preset similarity threshold and obtain a set of Q matching second abnormal keywords; Performing semantic similarity matching on the Q second historical voice data records using the Q first abnormal key sentence sets stored in the Q first memory banks as indexes; Extract sentences whose matching results meet the preset semantic similarity threshold and obtain a set of Q matching abnormal sentences; Extracting keywords related to the abnormal situation from the Q first abnormal key sentence sets respectively to obtain Q semantically matching second abnormal keyword sets; The Q matching second abnormal keyword sets and the Q semantic matching second abnormal keyword sets are unioned to obtain the Q second abnormal keyword sets.
4. The alarm device based on artificial intelligence voice analysis as claimed in claim 3, characterized in that: The execution steps of the abnormal keyword library module include: Extract the abnormal keywords with the top m occurrence frequencies in the Q target memory banks respectively to obtain Q regional abnormal keyword banks, where m is an integer greater than or equal to 3.
5. The alarm device based on artificial intelligence voice analysis according to claim 1, characterized in that: The execution steps of the voice anomaly analyzer module include: According to a preset disturbance scale, randomly generate interference data for the Q regional abnormal keyword libraries to obtain Q abnormal region speech data sets; Performing abnormality recognition on the Q abnormal area voice data sets in combination with the Q area alarm thresholds to obtain Q abnormality recognition results; The network layer constructed based on the feedforward neural network is supervised and trained using the Q abnormal area speech data sets and the Q abnormal recognition results respectively until the training converges, thereby obtaining the trained Q area intelligent speech anomaly analyzers.
6. The alarm device based on artificial intelligence voice analysis according to claim 1, characterized in that: The execution steps of the audio sensor layout module include: Extracting features from the Q regional scene information according to preset scene features to obtain Q regional scene features; Using a sensor layout plan identifier to identify the Q regional scene features, and obtain Q regional sensor layout plans; Perform regional sensor layout according to the layout positions and layout quantities of the audio collection sensors in the Q regional sensor layout schemes to obtain the Q audio collection sensor layout arrays.
Citation Information
Patent Citations
Equipment alarm method based on multiple sound feature analysis
CN115165079A
Smart home safety monitoring alarm system
CN118587840A