Sound-light alarm device integrated with voice broadcast

By designing signal detection, content matching, speech synthesis and sound-optical output modules in the sound-optical alarm device, combining the two-level dynamic classification architecture and emotional speech synthesis, the problems of insufficient voice broadcast synchronization, generation efficiency and multilingual support in the existing technology are solved, and efficient and flexible voice broadcasting effects are achieved.

CN120108094APending Publication Date: 2025-06-06CHINA ACADEMY OF RAILWAY SCI CORP LTD +2
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510273674.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-10
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

The existing audio-light alarm devices have shortcomings in the synchronization of voice broadcasting and alarm signals, voice content generation efficiency and multilingual support, and it is difficult to meet the diverse needs in complex scenarios.

Method used

A sound-optical alarm device integrating voice broadcast is designed, using signal detection module, content matching module, speech synthesis module and sound-optical output module. Through a two-level dynamic classification architecture and emotional voice synthesis, efficient voice content generation and multi-language support are achieved.

Benefits of technology

It significantly improves the quality and efficiency of voice broadcasts, enhances the warning effect and user perception and response speed, and adapts to diverse needs in complex scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120108094A_ABST
    Figure CN120108094A_ABST
Patent Text Reader

Abstract

The invention provides a sound-light alarm device integrated with voice broadcast. The sound-light alarm device comprises a signal detection module, a content matching module, a voice synthesis module and a sound-light output module, the signal detection module is used for receiving a multi-sensor original data stream; the content matching module is used for carrying out sensor coding on the multi-sensor original data stream, carrying out double-level classification on the multi-sensor original data stream based on a double-level dynamic classification architecture and the sensor coding, and matching voice contents of corresponding alarm keywords and emotion levels in a voice index database according to a classification result; and the voice synthesis module is used for selecting a pronunciation template with a corresponding emotion tone according to the emotion characteristics, and converting the voice content into alarm audio information. And the acousto-optic output module is used for playing the target audio information and lightening the alarm lamp. According to the invention, the synchronism of voice broadcast and alarm signals is improved; an efficient voice content generation mechanism is realized, various voice expression functions are supported, and the warning effect is enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention belongs to the technical field of sound and light alarms, and relates to a sound and light alarm device integrated with voice broadcast. Background Art

[0002] In modern security systems, traditional sound and light alarms mainly rely on sound and light to convey alarm information. However, traditional sound and light alarms can only provide simple visual and auditory warnings, lacking detailed voice prompts, which may not meet user needs in some special occasions. For example, in an emergency, relying solely on flashing lights and alarm sounds may not be enough to convey specific information to people around. Therefore, it is particularly important to develop a sound and light alarm device that can integrate voice broadcasting functions.

[0003] In the prior art, although some alarm devices have integrated voice broadcasting functions, their voice processing capabilities and broadcasting efficiency still need to be improved. The sound and light alarm devices currently on the market still have some shortcomings:

[0004] The synchronization between voice broadcast and alarm signal is not ideal;

[0005] The efficiency of voice content generation is low and it is difficult to meet the needs of real-time response;

[0006] The support for multiple languages ​​is limited and the adaptability is poor.

[0007] Therefore, how to provide an acousto-optic alarm device with integrated voice broadcast that can significantly improve the quality and efficiency of voice broadcast is a problem that those skilled in the art urgently need to solve. Summary of the invention

[0008] In view of this, the present invention proposes an acousto-optic alarm device with integrated voice broadcast, which improves the synchronization between the voice broadcast and the alarm signal; realizes an efficient voice content generation mechanism, supports multiple voice expression functions, and enhances the warning effect.

[0009] In order to achieve the above object, the present invention adopts the following technical solution:

[0010] The present invention discloses an acousto-optic alarm device with integrated voice broadcast, comprising: a signal detection module, a content matching module, a voice synthesis module and an acousto-optic output module; wherein:

[0011] The signal detection module is used to receive multi-sensor raw data streams;

[0012] The content matching module is used to perform sensor encoding on the multi-sensor raw data stream, perform dual-level classification on the multi-sensor raw data stream based on the sensor encoding based on the dual-level dynamic classification architecture, and match the voice content of the corresponding alarm keywords and emotion levels in the voice index database according to the classification results;

[0013] The speech synthesis module is used to select a pronunciation template with corresponding emotional timbre according to the emotional characteristics, and convert the speech content into alarm audio information.

[0014] The sound and light output module is used to play the target audio information and light up the alarm light.

[0015] Preferably, the signal detection module is also used to extract key features based on the multi-sensor raw data stream, and the key features include: sensor type features and sensor position features.

[0016] Preferably, the content matching module includes an encoding unit, configured to perform sensor encoding according to the key feature:

[0017] E 1 =f(S type ,S position ,time)

[0018] In the formula, E 1 is the sensor encoding vector, [type code, location code], S type is the sensor type characteristic, S position is the sensor location feature, time is the timestamp;

[0019] The encoding unit combines and encodes the same key features received within a set time interval according to the timestamp.

[0020] Preferably, the voice index database includes an index table and a storage area; the index table stores index addresses of multiple voice text packages and the sensor codes corresponding to the index addresses; the voice text package contains multiple voice text entries, and each of the voice text packages contains voice text entries corresponding to the same sensor type, and each of the voice entries is accompanied by an emotion level label.

[0021] Preferably, the content matching module includes a classifier 1 for classifying the content of the sensor encoding vector E according to the sensor encoding vector E. 1 Classify and filter voice packets in the voice index database:

[0022] Traversing and filtering the index address of the corresponding voice text package according to the type code;

[0023] The sensor encoding vector E 1Inputting a pre-trained classification model to generate a weight factor of a corresponding location code, wherein the weight factor is used to characterize the importance of the corresponding location sensor data;

[0024] Update the weight factor to the sensor encoding vector E 1 , get the sensor encoding vector E 2 .

[0025] Preferably, the content matching module includes a second classifier for classifying the content of the sensor according to the sensor encoding vector E 2 Categorize and filter speech text entries in speech text packages:

[0026] The sensor encoding vector E 2 Input the pre-trained classification model to generate the corresponding emotion level;

[0027] The voice text package is accessed according to the index address, and the corresponding voice text entries are traversed and filtered according to the emotion level.

[0028] Preferably, the content matching module includes a second classifier for classifying the content of the sensor according to the sensor encoding vector E 2 Categorize and filter speech text entries in speech text packages:

[0029] Obtaining the sensing characteristic parameters of the raw data stream of the sensor, comparing them with a preset characteristic parameter value range, and generating a risk factor;

[0030] The sensor encoding vector E 2 and the risk factors are input into a pre-trained classification model to generate a corresponding emotion level;

[0031] The voice text package is accessed according to the index address, and the corresponding voice text entries are traversed and filtered according to the emotion level.

[0032] Preferably, a key slot is provided in the voice text entry, and the voice synthesis module is used to perform the following steps:

[0033] Acquire sensor type features and sensor position features, and insert them into the key slot at the corresponding position as alarm keywords to obtain target voice text;

[0034] Calling a corresponding pronunciation template according to the emotion level, fusing the target speech text with the pronunciation template, and obtaining an alarm audio;

[0035] The alarm audio is input into a tone adjuster, and the tone adjuster predicts the tone parameters of the next speech frame using the tone parameters of the previous speech frame, and outputs the alarm audio information.

[0036] Preferably, the speech synthesis module has a configuration unit, which receives a configuration parameter modification instruction of the pronunciation template and updates the pronunciation template; the configuration parameters include: audio, speed, volume and sound effect.

[0037] Preferably, the sound and light output module includes a plurality of warning lights of different colors, and the warning lights of corresponding colors are lit according to the emotion level.

[0038] It can be seen from the above technical solution that, compared with the prior art, the beneficial effects of the present invention include:

[0039] The present invention extracts key features of multi-sensor data through a signal detection module, and combines it with a two-level dynamic classification architecture to achieve efficient processing and classification of multi-source data; it solves the limitations of traditional alarm devices in processing multi-sensor data and can adapt to the needs of multi-source data fusion in complex scenarios.

[0040] By introducing an online learning mechanism, the present invention enables the system to continuously optimize the classification model and speech synthesis strategy based on historical alarm data and user feedback, thereby improving alarm accuracy and speech naturalness.

[0041] The present invention combines emotional speech synthesis with dynamic tone adjustment to improve the expressiveness of alarm information. It solves the problem of single voice content and lack of emotional expression in traditional alarm devices, and can dynamically adjust the sense of urgency and emotional color of the voice according to the risk level. Through emotional speech synthesis, the communication effect of alarm information is significantly improved, and the user's perception and response speed are enhanced.

[0042] The multi-color lights of the present invention are linked with the emotion levels to enhance the visual warning effect; it solves the problem of monotonous voice output and lack of flexibility of traditional alarm devices, and through dynamic adjustment and personalized configuration, it significantly improves the naturalness and adaptability of voice broadcasting.

[0043] Compared with traditional alarm devices, the present invention has a higher level of intelligence and adaptability, and can meet diverse needs in complex scenarios. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art are briefly introduced below. Obviously, the drawings in the following description are only embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on the provided drawings without creative work.

[0045] Figure 1 A module connection diagram of an integrated voice broadcast sound and light alarm device provided in an embodiment of the present invention;

[0046] Figure 2 A flow chart of a two-level classification architecture provided in an embodiment of the present invention. DETAILED DESCRIPTION

[0047] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0048] The embodiment of the present invention discloses an acousto-optic alarm device with integrated voice broadcast, comprising: a signal detection module, a content matching module, a voice synthesis module and an acousto-optic output module; wherein:

[0049] The signal detection module is used to receive the raw data stream of multiple sensors;

[0050] The content matching module is used to perform sensor encoding on the multi-sensor raw data stream, and based on the two-level dynamic classification architecture, perform two-level classification on the multi-sensor raw data stream based on the sensor encoding, and match the voice content of the corresponding alarm keywords and emotion levels in the voice index database according to the classification results;

[0051] The speech synthesis module is used to select a pronunciation template with corresponding emotional timbre according to the emotional characteristics and convert the speech content into alarm audio information.

[0052] The sound and light output module is used to play the target audio information and light up the alarm light.

[0053] In one embodiment, the signal detection module is further used to extract key features according to the multi-sensor raw data stream, and the key features include: sensor type features and sensor position features.

[0054] In specific implementation, the input of the signal detection module is the raw data stream of multiple sensors (such as temperature, smoke, vibration, etc.). The processing steps are as follows:

[0055] Extract key features, including: sensor type feature T (such as temperature sensor coded as T01); sensor location feature L (such as location coded as L05); timestamp t (data collection time).

[0056] Merge sensor data of the same type and location within the same time period to reduce redundancy.

[0057] In one embodiment, the content matching module includes an encoding unit for performing sensor encoding according to the key features:

[0058] E 1=f(S type ,S position ,time)

[0059] In the formula, E 1 is the sensor encoding vector, [type code, location code], S type is the sensor type characteristic, S position is the sensor location feature, time is the timestamp;

[0060] The encoding unit combines and encodes the same key features received within a set time interval according to the timestamp.

[0061] In one embodiment, a speech index database includes an index table and a storage area; the index table stores index addresses of multiple speech text packages and sensor codes corresponding to the index addresses; the speech text package contains multiple speech text entries, and each speech text package contains speech text entries corresponding to the same sensor type, and each speech entry is attached with an emotion level label.

[0062] In this embodiment, the index table stores the index address of the voice text package and the corresponding sensor type code T. The storage area is used to store the voice text package, each voice text package contains multiple voice entries, and each entry is attached with an emotion level label (such as "The temperature is too high, please evacuate! - E3")

[0063] In one embodiment, the content matching module includes a classifier 1 for classifying the content of the sensor encoding vector E according to the sensor encoding vector E. 1 Classify and filter voice packets in the voice index database:

[0064] Traverse and filter the index address of the corresponding voice text package according to the type code;

[0065] The sensor encoding vector E 1 Input the pre-trained classification model to generate a weight factor of the corresponding location code, and the weight factor is used to characterize the importance of the corresponding location sensor data;

[0066] Update the weight factor to the sensor encoding vector E 1 , get the sensor encoding vector E 2 .

[0067] In the specific implementation, classifier 1 is the first level classification of the two-level classification architecture, and the execution steps include:

[0068] Voice packet screening: Match the index address of the corresponding voice text packet in the voice index database according to the type code T.

[0069] Weight factor calculation:

[0070] α=f model (V s )

[0071] Where α is the weight factor of the location code L, which is generated by the pre-trained classification model and represents the importance of the sensor location.

[0072] Encoding vector update: Update the weight factor α to V s , and get V' s =[T,L,t.α].

[0073] In this embodiment, the classifier is a neural network model based on the attention mechanism, and the specific architecture is as follows:

[0074] Input layer: Input sensor encoding vector [T, L, t]. Embedding is performed on T, L, t to convert them into a vector representation of fixed dimension.

[0075] Feature extraction layer: Use convolutional neural network (CNN) to extract features:

[0076] H = ReLU(W·V s +b)

[0077] Among them, W is the weight matrix, b is the bias term, and H is the extracted feature vector.

[0078] Attention mechanism layer: Introduce the self-attention mechanism and calculate the weight factor α of the position code L:

[0079]

[0080] In the formula, Q, K, and V are query, key, and value matrices respectively. k is the dimension of the vector.

[0081] Output layer: Output weight factor α and the filtered voice text package index address.

[0082] In one embodiment, the content matching module includes a second classifier for classifying the content of the sensor encoding vector E according to the sensor encoding vector E. 2 Categorize and filter speech text entries in speech text packages:

[0083] The sensor encoding vector E 2 Input the pre-trained classification model to generate the corresponding emotion level;

[0084] Access the speech text package according to the index address, and filter the corresponding speech text entries according to the emotion level.

[0085] In one embodiment, the content matching module includes a second classifier for classifying the content of the sensor encoding vector E according to the sensor encoding vector E. 2 Categorize and filter speech text entries in speech text packages:

[0086] Obtain the sensing characteristic parameters of the sensor raw data stream, compare them with the preset characteristic parameter value range, and generate risk factors;

[0087] The sensor encoding vector E 2 and risk factors input into the pre-trained classification model to generate the corresponding sentiment level;

[0088] Access the voice text package according to the index address, and traverse and filter the corresponding voice text entries according to the emotion level.

[0089] In this implementation, the second classifier is based on the classification model of the multi-layer perceptron (MLP), and the specific architecture is as follows:

[0090] Input layer: Input the updated sensor encoding vector Vs' = [T, L, t, α] and the risk factor β.

[0091] Feature fusion layer: Use a fully connected layer to fuse Vs' and β:

[0092] H'=ReLU(W 1 ·V' s +W 2 β+b)

[0093] Where W 1 , W 2 is the weight matrix and b is the bias term.

[0094] Classification layer: Use Softmax classifier to generate sentiment level E:

[0095] E=Softmax(W c ·H'+b c )

[0096] Among them, W c is the classification weight matrix, b c is the classification bias term.

[0097] Output layer: Outputs the emotion level E (such as 1-low risk, 2-medium risk, 3-high risk) and the corresponding speech text entries.

[0098] In the specific implementation, classifier 2 is the second level classification of the two-level classification architecture, and the execution steps include:

[0099] Emotional level determination process:

[0100] Method 1: Directly through V' s The input classification model generates a sentiment level E∈{1,2,3} (1-low risk, 2-medium risk, 3-high risk).

[0101] Method 2: Combine the characteristic parameter P (such as temperature value) of sensor data with the preset threshold to generate the risk factor β, and then calculate E=f(V' s ,β). According to the emotion level E, the corresponding words are selected in the speech text package.

[0102] The pre-trained models of classifiers 1 and 2 use an incremental learning algorithm to regularly update model parameters using newly collected sensor data and user feedback (such as false alarm correction). For example, if a location sensor frequently falsely alarms, the system automatically reduces its weight factor α.

[0103] Voice templates can adopt optimization strategies, record users' adjustments to voice emotion levels (such as changing high-risk voice from "urgent" to "calm"), dynamically update the pronunciation template library, and adapt to the needs of different scenarios.

[0104] In one embodiment, a key slot is provided in the speech text entry, and the speech synthesis module is used to perform the following steps:

[0105] Obtain sensor type features and sensor position features, and insert them into the key slots at corresponding positions as alarm keywords to obtain target voice text;

[0106] The corresponding pronunciation template is called according to the emotion level, and the target speech text is merged with the pronunciation template to obtain the alarm audio;

[0107] The alarm audio is input into the pitch adjuster, and the pitch adjuster predicts the pitch parameters of the next voice frame using the pitch parameters of the previous voice frame and outputs the alarm audio information.

[0108] In the specific implementation, the speech synthesis module inputs the target speech text and the emotion level E. The processing steps are as follows:

[0109] Keyword insertion step: Insert the sensor type T and location L into the key slot of the speech entry to generate the target text (such as "[T01] detected an anomaly at [L05]").

[0110] Emotional timbre selection steps: call the corresponding pronunciation template according to E (such as E3 uses a rapid timbre).

[0111] Steps for dynamic pitch adjustment:

[0112] F n+1 =F n +ΔF·γ

[0113] Among them, F n is the pitch of the current speech frame, and γ is the adjustment coefficient corresponding to the emotion level.

[0114] Configuration parameter update steps: Support user-defined audio parameters (sound speed, volume, etc.).

[0115] In one embodiment, the speech synthesis module has a configuration unit, which receives configuration parameter modification instructions of the pronunciation template and updates the pronunciation template; the configuration parameters include: audio, sound speed, volume and sound effect.

[0116] In one embodiment, the sound and light output module includes a plurality of warning lights of different colors, and the warning lights of corresponding colors are lit according to the emotion level.

[0117] Take the fire alarm scenario as an example:

[0118] The temperature sensor (T01) detects 50°C (threshold 40°C) at position L05, generating Vs = [T01, L05, t];

[0119] Classifier 1 calculates the weight factor α=0.9 and updates it to Vs′=[T01,L05,t,0.9];

[0120] Classifier 2 combines the risk factor β = 1.2 and determines E = 3;

[0121] The synthesized voice "The temperature in area L05 is too high, emergency evacuation!" is played and the red alarm light is turned on.

[0122] The above is a detailed introduction to the sound and light alarm device with integrated voice broadcast provided by the present invention. In this embodiment, specific examples are used to illustrate the principle and implementation method of the present invention. The description of the above embodiment is only used to help understand the method of the present invention and its core idea; at the same time, for those skilled in the art, according to the idea of ​​the present invention, there will be changes in the specific implementation method and application scope. In summary, the content of this specification should not be understood as limiting the present invention.

[0123] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present invention. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined in the present embodiments may be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to the embodiments shown in the present embodiment, but will conform to the widest range consistent with the principles and novel features disclosed in the present embodiment.

Claims

1. An acousto-optic alarm device with integrated voice broadcast, characterized in that: include: Signal detection module, content matching module, speech synthesis module and sound and light output module; wherein, The signal detection module is used to receive multi-sensor raw data streams; The content matching module is used to perform sensor encoding on the multi-sensor raw data stream, perform dual-level classification on the multi-sensor raw data stream based on the sensor encoding based on the dual-level dynamic classification architecture, and match the voice content of the corresponding alarm keywords and emotion levels in the voice index database according to the classification results; The speech synthesis module is used to select a pronunciation template with corresponding emotional timbre according to the emotional characteristics, and convert the speech content into alarm audio information. The sound and light output module is used to play the target audio information and light up the alarm light.

2. The sound and light alarm device with integrated voice broadcast according to claim 1, characterized in that: The signal detection module is also used to extract key features according to the multi-sensor raw data stream, and the key features include: sensor type features and sensor position features.

3. The sound and light alarm device with integrated voice broadcast according to claim 1, characterized in that: The content matching module includes an encoding unit, which is used to perform sensor encoding according to the key features: E1=f(S type ,S position ,time) Where E1 is the sensor encoding vector, [type code, location code], S type is the sensor type characteristic, S position is the sensor location feature, time is the timestamp; The encoding unit combines and encodes the same key features received within a set time interval according to the timestamp.

4. The sound and light alarm device with integrated voice broadcast according to claim 3, characterized in that: The speech index database includes an index table and a storage area; the index table stores index addresses of multiple speech text packages and the sensor codes corresponding to the index addresses; the speech text package contains multiple speech text entries, and each of the speech text packages contains speech text entries corresponding to the same sensor type, and each of the speech entries is attached with an emotion level label.

5. The sound and light alarm device with integrated voice broadcast according to claim 4, characterized in that: The content matching module includes a classifier 1, which is used to classify and filter voice packets in the voice index database according to the sensor encoding vector E1: Traversing and filtering the index address of the corresponding voice text package according to the type code; Input the sensor encoding vector E1 into a pre-trained classification model to generate a weight factor of the corresponding position code, wherein the weight factor is used to characterize the importance of the corresponding position sensor data; The weight factor is updated to the sensor encoding vector E1 to obtain the sensor encoding vector E2.

6. The sound and light alarm device with integrated voice broadcast according to claim 5, characterized in that: The content matching module includes a second classifier for classifying and filtering speech text entries in the speech text package according to the sensor encoding vector E2: Inputting the sensor encoding vector E2 into a pre-trained classification model to generate a corresponding emotion level; The voice text package is accessed according to the index address, and the corresponding voice text entries are traversed and filtered according to the emotion level.

7. The sound and light alarm device with integrated voice broadcast according to claim 5, characterized in that: The content matching module includes a second classifier for classifying and filtering speech text entries in the speech text package according to the sensor encoding vector E2: Obtaining the sensing characteristic parameters of the raw data stream of the sensor, comparing them with a preset characteristic parameter value range, and generating a risk factor; Input the sensor encoding vector E2 and the risk factor into a pre-trained classification model to generate a corresponding emotion level; The voice text package is accessed according to the index address, and the corresponding voice text entries are traversed and filtered according to the emotion level.

8. The sound and light alarm device with integrated voice broadcast according to claim 2, characterized in that: The speech text entry is provided with a key slot, and the speech synthesis module is used to perform the following steps: Acquire sensor type features and sensor position features, and insert them into the key slot at the corresponding position as alarm keywords to obtain target voice text; Calling a corresponding pronunciation template according to the emotion level, fusing the target speech text with the pronunciation template, and obtaining an alarm audio; The alarm audio is input into a tone adjuster, and the tone adjuster predicts the tone parameters of the next speech frame using the tone parameters of the previous speech frame, and outputs the alarm audio information.

9. The sound and light alarm device with integrated voice broadcast according to claim 8, characterized in that: The speech synthesis module has a configuration unit, which receives a configuration parameter modification instruction of the pronunciation template and updates the pronunciation template; the configuration parameters include: audio, sound speed, volume and sound effect.

10. The sound and light alarm device with integrated voice broadcast according to claim 1, characterized in that: The sound and light output module includes a plurality of warning lights of different colors, and the warning lights of corresponding colors are lit according to the emotion level.