Voice recognition system, device and method thereof

TW202636431AActive Publication Date: 2026-09-01ACER INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
TW114106974
Authority / Receiving Office
TW · TW
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-02-25
Publication Date
2026-09-01
Estimated Expiration
2045-02-24

AI Technical Summary

Technical Problem

Existing electronic devices fail to effectively assist individuals with impaired hearing by converting a wide range of sound cues into visual or tactile alerts, posing a risk of inconvenience or danger in situations where sound recognition is crucial.

Method used

A sound recognition system utilizing a pre-trained AI model to identify various sound events, connecting to a router to determine device usage, and sending predefined notifications to electronic devices for visual or tactile alerts when specific sounds are detected.

Benefits of technology

Enhances sound recognition accuracy and breadth, enabling real-time conversion of important sound events into visual or tactile notifications on connected devices, ensuring users with impaired hearing are alerted to relevant events or dangers.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure TWG2TA001073975_001
    Figure TWG2TA001073975_001
  • Figure TWG2TA001073975_002
    Figure TWG2TA001073975_002
  • Figure TWG2TA001073975_003
    Figure TWG2TA001073975_003
Patent Text Reader

Abstract

An aspect of the present disclosure features a voice recognition device. The voice recognition device includes: a voice detection module configured to receive an environmental voice; a voice recognition module including an AI voice recognition model, configured to recognize the environmental voice whether includes one of at least one sound event; a router management module coupled to a router, configured to read information of multiple electronic device coupled to the router, and configure notification settings, corresponding to the at least one sound event, of the multiple electronic device; and a notification management module coupled to the router. When the voice detection module determines that the environmental voice includes one of the at least one sound event, the notification management module sends a notification message to at least one of the multiple electronic devices according to the notification settings, and the notification message is displayed by a display unit of the at least one of the multiple electronic devices.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to a voice recognition system and apparatus, and to a voice recognition method using an AI voice recognition model. Prior Technology

[0002] With advancements in technology, various electronic devices in the environment provide sound effects to inform users of current situations or dangers. However, for users with limited or impaired hearing, such as the hearing-impaired or deaf, their inability to hear sounds as normally as others can cause inconvenience or danger when relevant sound cues occur. For example, they may not hear fire alarms, family members calling them, or vehicle horns. Currently, electronic products on the market can only use built-in functions to identify specific sounds and convert them into visual or other forms of alerts, displaying or vibrating them directly on the device. Therefore, there is a need for technologies that can expand the range and types of sound recognition to assist users in these situations. Summary of the Invention

[0003] The technology disclosed herein utilizes a pre-trained AI sound recognition model to identify various sound events in the environment, such as within a home. Users can train the model to recognize specific environmental sounds, thereby improving the accuracy and breadth of sound recognition. Upon detecting a sound event, the system connects to a router to determine the device the user may be using and sends a pre-defined notification message to that device. This visual notification is provided when the user is unable to hear the sound (e.g., for hearing-impaired individuals).

[0004] According to a first aspect of this disclosure, a sound recognition system is proposed. The sound recognition system includes a router. The sound recognition system also includes multiple electronic devices, each having a display unit, and each electronic device is coupled to the router. The sound recognition system also includes a main electronic device. The main electronic device includes a processor. The main electronic device also includes a sound detection module coupled to the processor for receiving ambient sound. The main electronic device also includes a sound recognition module coupled to the processor and includes an AI sound recognition model for identifying whether the ambient sound contains at least one of at least one sound event. The main electronic device also includes a router management module coupled to the router and the processor, the router management module being used to read information from the multiple electronic devices connected to the router and to set notification settings for the multiple electronic devices corresponding to at least one sound event. The main electronic device also includes a notification management module coupled to the router and the processor. When the sound detection module determines that the ambient sound contains at least one of the sound events, the notification management module sends a notification message to at least one of the multiple electronic devices according to the notification settings, and displays the notification message on the display unit of at least one of the multiple electronic devices.

[0005] According to a second aspect of this disclosure, a sound recognition device is proposed. The sound recognition device includes a processor. The sound recognition device also includes a sound detection module coupled to the processor for receiving ambient sound. The sound recognition device also includes a sound recognition module comprising an AI sound recognition model coupled to the processor for identifying whether the ambient sound contains at least one of at least one sound event. The sound recognition device also includes a router management module coupled to a router and the processor, the router management module for reading information from multiple electronic devices connected to the router and setting notification settings for the multiple electronic devices corresponding to at least one sound event. The sound recognition device also includes a notification management module coupled to the router and the processor. When the sound detection module determines that the ambient sound contains at least one of at least one sound event, the notification management module sends a notification message to at least one of the multiple electronic devices according to the notification settings, and displays the notification message on the display unit of at least one of the multiple electronic devices.

[0006] According to a third aspect of this disclosure, a sound recognition method is proposed. The sound recognition method includes receiving ambient sound via a sound detection module of a sound recognition device. The sound recognition method also includes identifying whether the ambient sound contains one of at least one sound event using an AI sound recognition model of the sound recognition module of the sound recognition device. The sound recognition method also includes reading information from multiple electronic devices connected to the router and setting notification settings for the multiple electronic devices corresponding to at least one sound event via a router management module of the sound recognition device coupled to the router. The sound recognition method also includes sending a notification message to at least one of the multiple electronic devices according to the notification settings via a notification management module of the sound recognition device coupled to the router when the sound detection module determines that the ambient sound contains one of at least one sound event, and displaying the notification message via a display unit of at least one of the multiple electronic devices.

[0007] The foregoing description is not intended to represent every embodiment or aspect of this disclosure. Rather, it provides only examples of some novel aspects and features set forth in this disclosure. The foregoing features and advantages, as well as other features and advantages, will become apparent from the following detailed description of representative embodiments and modes of implementation of the invention when taken in conjunction with the accompanying drawings and the claims. Other aspects of this disclosure will be apparent to those skilled in the art from the detailed description of various embodiments with reference to the drawings and the brief description provided below.

[0008] To provide a better understanding of the above and other aspects of the present invention, specific embodiments are described below in conjunction with the accompanying drawings: Simple Explanation of the Diagram

[0009] This disclosure, its advantages, and the drawings will be readily understood from the following description of representative embodiments in conjunction with the accompanying drawings. These drawings depict only representative embodiments and should not be considered as limiting the scope of the various embodiments or the claims. Figure 1 illustrates a functional block diagram of an example voice recognition system according to several embodiments of the present disclosure. Figure 2 illustrates a schematic diagram of the operation interface and audio file list of a sound recording unit according to various embodiments of this disclosure. Figure 3 illustrates a schematic diagram of the operation interface of a user interface setting module according to various embodiments of the present disclosure. Figure 4 is a schematic diagram illustrating the operation interface of a router management module according to various embodiments of the present disclosure. Figure 5 illustrates a flowchart of an example program for sound recognition according to various embodiments of this disclosure. Implementation

[0010] Figure 1 illustrates a functional block diagram of an example voice recognition system 1000 according to several embodiments of the present disclosure. The voice recognition system 1000 includes a router 300 and a plurality of electronic devices (electronic devices 400-1 to 400-n) coupled to the router 300. Each of the electronic devices 400-1 to 400-n has a display unit 401-1 to a display unit 401-n.

[0011] The voice recognition system 1000 also includes a main electronic device 200, which can also be called a voice recognition device. The main electronic device 200 includes a processor 210, a voice detection module 220, a voice recognition module 230, a router management module 240, a notification management module 250, a training module 260, and a user interface setting module 270.

[0012] The sound detection module 220 is coupled to the processor 210 and can be used to receive ambient sounds. Ambient sounds may include a variety of different sound events, such as sound events in a home, such as the music from a washing machine after washing clothes, a child crying, the sound of a kettle boiling water, a fire alarm, and a doorbell.

[0013] The sound recognition module 230 is coupled to the processor 210 and includes an AI sound recognition model 231. The AI ​​sound recognition model 231 is used to identify whether ambient sounds contain one of the aforementioned sound events, such as the music from a washing machine, a child crying, and a doorbell ringing. In some embodiments, the AI ​​sound recognition model 231 may use a depthwise separable convolutional architecture to classify sound events, such as the MobileNetV1 depthwise separable convolutional architecture.

[0014] In some implementations, the AI ​​voice recognition model 231 can be implemented using the YAMNet model. The YAMNet model is a voice recognition model developed by Google, and it can be trained using AudioSet-YouTube data, which contains over 2 million audio clips from YouTube videos. Furthermore, the YAMNet model can be customized using TensorFlow Hub with other voice data collected by the user to tailor the desired voice data for recognition.

[0015] Referring also to Figure 2, a schematic diagram illustrating the operation interface 261a and audio file list 261b of the sound recording unit according to various embodiments of the present disclosure is provided. In some embodiments, the training module 260 can be used to train the AI ​​sound recognition model 231. The training module 260 includes a sound recording unit 261, a noise reduction unit 262, a frequency conversion unit 263, a clipping unit 264, an AI training unit 265, and a checking unit 266. The recording unit 261 can be used to record ambient sounds through the sound detection module 220 to form multiple audio files and set file tags, as shown in the operation interface 261a and audio file list 261b of Figure 2. That is, the user can record and set tags for custom sound events, and the operation interface 261a and audio file list 261b can visually remind users who cannot or are inconvenient to hear sounds (e.g., hearing-impaired individuals) of the current operating status, such as recording, and the audio file list 261b will be displayed after pressing save on the operation interface 261a. In some implementations, to achieve a higher sound event recognition rate, the user can be required to record multiple audio files for the same sound event, as shown in audio file list 261b. In some implementations, the audio files can be processed by a noise reduction unit 262 to remove background noise, a frequency conversion unit 263 to convert the frequency of the audio files to 16kHz, and a segment of the corresponding sound event in the audio files can be extracted by a segmentation unit 264 to form an audio file format suitable for AI training. In audio file list 261b, after the user presses the AI ​​training button, the AI ​​training unit 265 can use these audio files to train the AI ​​sound recognition model 231, for example, by converting the audio files and data tags into a data format acceptable to the YAMNet model for training, so as to adjust it to a YAMNet model that meets the user's sound recognition needs.

[0016] In some embodiments, the checking unit 266 can be used to input at least a portion of similar sounds that are the same as those in the recording file used to train the AI ​​sound recognition model 231, and check whether the AI ​​sound recognition model 231 correctly identifies the similar sounds as corresponding to the same sound events, so as to confirm the recognition accuracy of the AI ​​sound recognition model 231. Based on the checking results, it can be decided whether to retrain the AI ​​sound recognition model 231 using the training module 260 according to the above process.

[0017] Referring also to Figures 1 and 3, Figure 3 illustrates a schematic diagram of the operation interface 271 of the user interface setting module 270 according to various embodiments of the present disclosure. The user interface setting module 270 is coupled to the processor 210 and allows setting the priority order, time range, notification message, and activation settings of the aforementioned sound events through its operation interface 271, which are then stored in a sound event database. The sound event database includes the sound number, event name, priority order, time range, notification message, and activation settings of each sound event, as shown in Table 1 below. Table 1: Sound Number Event Name Priority Time range Notification message activation 0x0001 The baby's cry High All The baby cried. yes 0x0002 fire alarm Emergency All Fire yes 0x0005 doorbell Medium 10:00 ~ 22:00 The doorbell rang yes 0x0006 Boil water High 10:00 ~ 20:00 The water has boiled yes 0x0007 washing machine Info 8:00 ~ 20:00 The clothes are washed no

[0018] The sound ID is a unique identifier automatically generated when a new sound event is added. The event name is a user-defined sound event, such as the music played when the washing machine finishes washing clothes. The priority order is the urgency of the sound event, from highest to lowest: Emergency, High, Medium, and Info. The time range sets the required notification time interval, such as All-day notification or start time to end time. The notification message is the text message to be displayed, such as "The doorbell rang." The enable setting determines whether the corresponding sound event is enabled, such as Yes (enabled) or No (disabled). When the priority order is set to Emergency, the time range is automatically set to All-day notification and cannot be changed; the enable setting is also automatically set to Yes (enabled) and cannot be canceled.

[0019] Referring also to Figures 1 and 4, Figure 4 illustrates a schematic diagram of the operation interface 241 of a router management module 240 according to various embodiments of the present disclosure. The router management module 240 is coupled to the router 300 and the processor 210. The router management module 240 can be used to read information from multiple electronic devices connected to the router 300 (e.g., electronic devices 400-1 to 400-n in Figure 1, or Pixel 7, Google Watch, Sony Smart TV, or Acer Aspire in Figure 4), and set notification settings for the multiple electronic devices corresponding to one of the aforementioned sound events, for example, through the operation interface 241 in Figure 4. The notification settings are used to set and store information for all electronic devices accessing the internet through the router 300 (e.g., Pixel 7, Google Watch, Sony Smart TV, or Acer Aspire in Figure 4), as shown in List 2 below. Table 2: Device number Device Name Device type Notification method activation flow 0x0001 Pixel 7 Smart Phone Vision, message, vibration yes 7 0x0002 Google Watch Smart Watch Vision, message, vibration yes 2 0x0005 Sony Smart TV Smart TV Visuals, messages no 115 0x0006 iPad Air Tablet Visuals, messages yes 5 0x0007 Acer Aspire Laptop Visuals, messages no 5

[0020] The device ID is the ID assigned to each electronic device when it is added to the router 300. The device name is the name automatically assigned or user-defined when the device is added, such as Acer Aspire. The device type is the type automatically assigned or selected by the user when the device is added, such as Smartphone, Smart Watch, Smart TV, Tablet, or Laptop. The notification method is the method of notification to the electronic device. Specifically, for users who cannot or are unable to receive audio messages (e.g., hearing-impaired or deaf individuals), the notification method can be visual, text message, or vibration. However, vibration is not supported by every type of electronic device (e.g., Smart TV, Tablet, or Laptop) and can be configured by the user. The enable setting determines whether notifications are enabled for the electronic device, such as Yes (enabled) or No (disabled). The traffic is the network traffic (data transfer volume) between the electronic device and the router, used to determine if the electronic device is being used by a user. Traffic can be, for example, the data transfer volume for 3 minutes, measured in MB.

[0021] In some implementations, similar functions of the sound detection module 220 can be achieved by the microphone unit of another electronic device external to and coupled to the main electronic device 200 (e.g., directly connected or connected via router 300), such as the microphone function of a mobile phone or other electronic device, thereby increasing the range of ambient sound monitoring. If other electronic devices receive ambient sounds, a notification can be displayed on their display unit, such as a flashing blue frame around the display unit, to inform the user that the electronic device has entered monitoring mode. Similarly, when ambient sounds are received, they can be transmitted to the AI ​​sound recognition model 231 of the sound recognition module 230 of the main electronic device 200 for identification and subsequent operations.

[0022] Referring to Figure 1, the notification management module 250 is coupled to the router 300 and the processor 210. When the sound detection module 220 determines that the ambient sound contains one of the aforementioned sound events, the notification management module 250 can send a notification message to the corresponding electronic device according to the aforementioned notification settings, and display the notification message on the display unit of the electronic device, such as displaying the notification message on the display of the Sony smart TV in Table 2, or displaying the notification message on the screen of the Google Watch.

[0023] Specifically, when the sound detection module 220 detects a sound event, the notification management module 250 decides which electronic device to send a notification message to. The determination process is as follows: First, the priority order of the sound events can be queried from the sound event database (as shown in Table 1). If it is the highest priority (Emergency), all possible notifications, such as visual, vibration, and text messages, will be sent to all electronic devices connected to the router 300. If it is another priority order, it can first determine whether it is within the set time interval, and then check which electronic devices are currently connected to the router 300. If multiple electronic devices are connected to the router 300 at the same time, the notification management module 250 can obtain the data traffic of the electronic devices connected to the router 300 through the router management module 240, such as the data transfer volume in 3 minutes, and then compare it with the data traffic set in the notification settings to determine whether any electronic device has a data traffic volume in 3 minutes that is greater than or equal to the set value. If so, it means that the electronic device is being used, and a notification message will be sent to this electronic device first. For example, assuming a Smart TV is used to watch a 1080p video, the download traffic for 3 minutes is 112.5 MB and the upload traffic for 3 minutes is 3 MB. Therefore, the total traffic is 112.5 MB + 3 MB = 115.5 MB, which is greater than the set traffic of 115 MB (as shown in Table 2). In this case, the notification management module 250 can determine that the Smart TV is being used and will prioritize sending a notification message to the Smart TV. In some embodiments, if the notification management module 250 cannot determine which electronic device is being used, it can send notifications to all electronic devices connected to the router 300 according to the notification settings.

[0024] Figure 5 illustrates a flowchart of an example procedure for sound recognition according to several embodiments of this disclosure. In step S510, ambient sound may be received, for example, by a sound detection module (e.g., sound detection module 220 of Figure 1) of the sound recognition device (or the main electronic device 200 of Figure 1). In some embodiments, the function of the sound detection module may be implemented by the microphone unit of another electronic device external to and coupled to the sound recognition device (e.g., directly connected or connected via a router), such as the microphone function of a mobile phone or other electronic device.

[0025] In step S520, the AI ​​sound recognition model of the sound recognition module of the sound recognition device (e.g., AI sound recognition model 231 of the sound recognition module 230 in Figure 1) can be used to identify whether the ambient sound contains at least one of the sound events, such as one of the sound event databases in Table 1. If not, proceed to step S510. If yes, continue to step S530.

[0026] In step S530, the router management module (e.g., the router management module 240 in Figure 1) coupled to the sound recognition device of the router can read information of multiple electronic devices connected to the router and set notification settings for multiple electronic devices corresponding to at least one sound event, such as the notification settings of different electronic devices in Table 2.

[0027] In step S540, a notification message (e.g., the notification management module 250 in Figure 1) can be sent to at least one of a plurality of electronic devices according to the notification settings via a notification management module of a voice recognition device coupled to the router.

[0028] In step S540, a notification message is displayed by the display unit of at least one of the multiple electronic devices.

[0029] In some specific settings, an example procedure for sound recognition may include: recording ambient sounds to form an audio file using the sound recording unit of the training module and setting file tags; removing background noise from the audio file using the noise reduction unit of the training module; converting the frequency of the audio file to 16kHz using the frequency conversion unit of the training module; extracting at least a portion of the audio file corresponding to at least one of the sound events using the extraction unit of the training module; training an AI sound recognition model using at least a portion of the audio file using the AI ​​training unit of the training module; and inputting a similar sound that is the same as at least a portion of the audio file using the checking unit of the training module, and checking whether the AI ​​sound recognition model correctly identifies the similar sound as one of the at least one sound events.

[0030] In certain specific settings, example procedures for voice recognition may include: setting the priority, time range, notification message, and enable settings for at least one voice event through a user interface setting module of the voice recognition device, and storing the settings in a voice event database. The voice event database includes the voice number, event name, priority, time range, notification message, and enable settings for at least one voice event.

[0031] In some specific settings, the notification settings also include, when the sound detection module determines that the ambient sound contains at least one of the sound events, causing at least one of the multiple electronic devices to vibrate and causing the display unit of at least one of the multiple electronic devices to provide a visual cue.

[0032] In certain settings, when the sound detection module determines that the ambient sound contains at least one of the sound events, it notifies the management module to determine which of the multiple electronic devices to send a notification message to based on the network usage traffic information of the multiple electronic devices in the router management module.

[0033] Accordingly, the sound recognition technology provided by the various embodiments disclosed herein can use an AI sound recognition model to identify ambient sounds and determine sound events, and can be trained automatically to handle suitable sound events. Furthermore, a router can be used to notify electronic devices connected to the router in real time via visual display or vibration. Therefore, important sound events in the environment can be converted from sound into visual or tactile (vibration) notifications, allowing users who cannot or are unable to receive sound messages (such as the hearing impaired or deaf) to be aware of relevant events or dangers requiring attention.

[0034] This disclosure and other examples can be implemented as one or more computer program products, for example, where one or more modules of computer program instructions encoded on a computer-readable medium are executed by or control the operation of a data processing device. The computer-readable medium can be a machine-readable storage device, a machine-readable storage substrate, a memory device, or a combination thereof. The term "data processing device" includes all means, apparatus, and machines for processing data, including, for example, a programmable processor, a computer, or multiple processors or computers. In addition to hardware, this means may include program code that establishes the execution environment of the computer program in question, such as program code constituting processor firmware, a protocol stack, a database management system, an operating system, or a combination thereof.

[0035] Computer programs (also known as programs, software, software applications, instruction code, or program code) can be written in any form of programming language, including compiled or interpreted languages, and can be deployed in any form, including as a standalone program, a module, component, subroutine, or other unit used in a computing environment. A computer program does not necessarily correspond to a file in a file system. A program can be stored as part of a file containing other programs or data (e.g., one or more instruction codes stored in a markup language document), in a single file dedicated to the program in question, or in multiple coordinated files (e.g., files containing one or more modules, subroutines, or portions of program code). Computer programs can be configured to execute on one or more computers. These computers may be located in one place or distributed across multiple locations and interconnected via a communication network.

[0036] The programs and logic flows described herein can be executed by one or more programmable processors that execute one or more computer programs to perform the functions described herein. The programs and logic flows can also be executed by special purpose logic circuitry, and the devices can also be implemented by special purpose logic circuitry, such as field programmable gate arrays (FPGAs) or application-specific integrated circuits (ASICs), or executed by a system-on-a-chip (SoC).

[0037] Processors suitable for executing computer programs include, for example, both general-purpose microprocessors and special-purpose microprocessors, and any one or more processors of any type of digital computer. Generally, a processor receives instructions and data from read-only memory or random access memory, or both. The basic components of a computer may include a processor for executing instructions and one or more memory devices for storing instructions and data. Typically, a computer may also include or be operatively coupled to one or more mass storage devices for storing data, to receive data from, or to transfer data to, or both of these mass storage devices. Examples of such mass storage devices are magnetic disks, magneto-optical disks, or optical disks. However, a computer does not necessarily need to have such devices. Computer-readable media suitable for storing computer program instructions and data may include all forms of non-volatile memory, media, and memory devices, including, for example, semiconductor memory devices, such as EPROM, EEPROM, and flash memory devices; magnetic disks. Processors and memory may be supplemented or integrated into special-purpose logic circuitry.

[0038] Various embodiments are described with reference to the accompanying drawings, wherein all drawings use the same element reference numerals to denote similar or equivalent elements. The drawings are not necessarily drawn to scale and are provided only to illustrate aspects and features of this disclosure. Numerous specific details, relationships, and methods are set forth to provide a comprehensive understanding of certain aspects and features of this disclosure, although those skilled in the art will recognize that these aspects and features can be implemented without one or more of the specific details, relationships, or methods. In some cases, well-known structures or operations are not shown in detail for illustrative purposes. The various embodiments disclosed herein are not necessarily limited to the order of the described actions or events, as some actions may occur in a different order and / or simultaneously with other actions or events. Furthermore, not all actions or events in the drawings are necessary to realize certain aspects and features of this disclosure.

[0039] Although the invention has been described and illustrated with respect to one or more embodiments, other skilled in the art will recognize or understand equivalent changes and modifications upon reading and understanding this specification and the accompanying drawings. Furthermore, while a particular feature of the invention may be disclosed only in one of several embodiments, this feature may be combined with one or more other features of other embodiments, as these features may be desired and advantageous for any given or particular application.

[0040] While various embodiments of the invention have been described above, it should be understood that they are presented by way of example only and not as limitation. Many changes may be made to the disclosed embodiments based on the disclosure without departing from the spirit or scope of the invention. Therefore, the breadth and scope of the invention should not be limited by any of the above embodiments. Rather, the scope of the invention should be defined according to the appended claims and their equivalents.

[0041] 1000: Voice Recognition System 200: Main electronic device 210: Processor 220: Sound Detection Module 230: Voice Recognition Module 231: AI Voice Recognition Model 240: Router Management Module 241, 261a, 271: User Interface 250: Notification Management Module 260: Training Module 261: Sound Recording Unit 261b: Audio File List 262: Noise Reduction Unit 263: Frequency conversion unit 264: Capture Unit 265: AI Training Unit 266: Inspection Unit 270: User Interface Settings Module 300: Router 400-1~400-n: Electronic devices 401-1~401-n: Display Unit

Claims

1. A voice recognition system, comprising: One router; A plurality of electronic devices, each having a display unit, and each of the electronic devices being coupled to the router; The system also includes a main electronic device comprising: a processor; a sound detection module coupled to the processor for receiving an ambient sound; a sound recognition module coupled to the processor and including an AI sound recognition model for identifying whether the ambient sound contains at least one of at least one sound event; a router management module coupled to the router and the processor, the router management module for reading information of the electronic devices connected to the router and setting a notification setting for the electronic devices corresponding to the at least one sound event; and a notification management module coupled to the router and the processor, wherein, when the sound detection module determines that the ambient sound contains one of the at least one sound event, the notification management module sends a notification message to at least one of the electronic devices according to the notification setting, and displays the notification message on the display unit of the at least one of the electronic devices. When the sound detection module determines that the ambient sound contains one of the at least one sound event, the notification management module determines which of the electronic devices to send the notification message to based on a network usage traffic information from the electronic devices in the router management module.

2. The voice recognition system as described in claim 1, wherein the main electronic device includes a training module for training the AI ​​voice recognition model, the training module comprising: A sound recording unit is used to record ambient sounds through the sound detection module to form an audio file and set a file tag; A noise reduction unit is used to remove background noise from the recording file; a frequency conversion unit is used to convert the frequency of the recording file to 16kHz; a truncating unit is used to truncate at least a portion of the recording file corresponding to at least one of the at least one sound events; an AI training unit is used to train the AI ​​sound recognition model using the at least a portion of the recording file; and a checking unit is used to input a similar sound that is the same as the at least a portion of the recording file and check whether the AI ​​sound recognition model correctly identifies the similar sound as one of the at least one sound events.

3. The voice recognition system as claimed in claim 1, wherein the main electronic device includes a user interface setting module coupled to the processor and for setting a priority order, a time interval, a notification message and an enable setting for each of the at least one voice event, and storing them in a voice event database, wherein the voice event database includes a voice number, an event name, the priority order, the time interval, the notification message and the enable setting for each of the at least one voice event.

4. The sound recognition system as claimed in claim 1, wherein the notification setting further includes, when the sound detection module determines that the ambient sound contains one of the at least one sound event, causing at least one of the electronic devices to vibrate and causing the display unit of at least one of the electronic devices to provide a visual prompt.

5. A voice recognition device, comprising: One processor; A sound detection module is coupled to the processor to receive an ambient sound; A sound recognition module, including an AI sound recognition model and coupled to the processor, for recognizing whether the ambient sound contains one of at least one sound event; a router management module, coupled to a router and the processor, the router management module for reading information from a plurality of electronic devices connected to the router and setting a notification setting for the electronic devices corresponding to the at least one sound event; The system also includes a notification management module coupled to the router and the processor. When the sound detection module determines that the ambient sound contains one of the at least one sound events, the notification management module sends a notification message to at least one of the electronic devices according to the notification settings, and displays the notification message on a display unit of the at least one of the electronic devices. When the sound detection module determines that the ambient sound contains one of the at least one sound events, the notification management module determines which of the electronic devices to send the notification message to based on a network usage traffic information in the router management module.

6. The voice recognition device as described in claim 5 further includes a training module for training the AI ​​voice recognition model, the training module comprising: A sound recording unit is used to record ambient sounds through the sound detection module to form an audio file and set a file tag; A noise reduction unit is used to remove background noise from the recording file; a frequency conversion unit is used to convert the frequency of the recording file to 16kHz; a truncating unit is used to truncate at least a portion of the recording file corresponding to at least one of the at least one sound events; an AI training unit is used to train the AI ​​sound recognition model using the at least a portion of the recording file; and a checking unit is used to input a similar sound that is the same as the at least a portion of the recording file and check whether the AI ​​sound recognition model correctly identifies the similar sound as one of the at least one sound events.

7. The voice recognition device as described in claim 5 further includes a user interface setting module coupled to the processor and for setting a priority order, a time interval, a notification message, and an enable setting for each of the at least one voice event, and storing them in a voice event database, wherein the voice event database includes a voice number, an event name, the priority order, the time interval, the notification message, and the enable setting for each of the at least one voice event.

8. The sound recognition device as claimed in claim 5, wherein the notification setting further includes, when the sound detection module determines that the ambient sound contains one of the at least one sound event, causing at least one of the electronic devices to vibrate and causing the display unit of at least one of the electronic devices to provide a visual prompt.

9. A sound recognition method, comprising: An ambient sound is received by a sound detection module of a sound recognition device; an AI sound recognition model of the sound recognition module of the sound recognition device identifies whether the ambient sound contains one of at least one sound event; a router management module of the sound recognition device coupled to a router reads information of a plurality of electronic devices connected to the router and sets a notification setting for the electronic devices corresponding to the at least one sound event; and when the sound detection module determines that the ambient sound contains one of the at least one sound event, a notification management module of the sound recognition device coupled to the router sends a notification message to at least one of the electronic devices according to the notification setting, and displays the notification message on a display unit of the at least one of the electronic devices, wherein when the sound detection module determines that the ambient sound contains one of the at least one sound event, the notification management module determines which of the electronic devices to send the notification message to based on a network usage traffic in the information of the electronic devices in the router management module.

10. The voice recognition method as described in claim 9 further includes: An ambient sound is recorded by a sound recording unit of a training module to form an audio file and a file tag is set; background noise is removed from the audio file by a noise reduction unit of the training module; the frequency of the audio file is converted to 16kHz by a frequency conversion unit of the training module; at least a portion of the audio file corresponding to at least one of the at least one sound event is extracted by a segmentation unit of the training module; an AI training unit of the training module uses the at least a portion of the audio file to train an AI sound recognition model; and a checking unit of the training module inputs a similar sound that is the same as the at least a portion of the audio file and checks whether the AI ​​sound recognition model correctly identifies the similar sound as one of the at least one sound event.

11. The voice recognition method as described in claim 9 further includes: The user interface setting module of the voice recognition device sets a priority order, a time interval, a notification message, and an enable setting for each of the at least one voice event, and stores them in a voice event database, wherein the voice event database includes a voice number, an event name, the priority order, the time interval, the notification message, and the enable setting for each of the at least one voice event.

12. The sound recognition method as claimed in claim 9, wherein the notification setting further includes, when the sound detection module determines that the ambient sound contains one of the at least one sound event, causing at least one of the electronic devices to vibrate and causing the display unit of at least one of the electronic devices to provide a visual prompt.