Home space scene regeneration and monitoring system and method in noise sensitive space environment
By integrating text-to-speech, surveillance camera footage, and environmental voice monitoring technologies, a home space scene reproduction and monitoring system for noise-sensitive environments was constructed. This system solves the real-time monitoring and communication problems in noise-sensitive environments and achieves comprehensive acoustic and visual scene reproduction.
Patent Information
- Application Number
- CN202511040654.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-28
- Publication Date
- 2025-12-09
AI Technical Summary
Existing home monitoring products cannot achieve effective real-time monitoring and communication in noise-sensitive environments, resulting in untimely or delayed information transmission and failing to meet the real-time communication needs of people.
By integrating text-to-speech, monitoring equipment, and monitoring technology systems, a comprehensive acoustic and visual scene reconstruction is constructed, including text-to-speech, real-time acquisition of surveillance camera footage, and voice monitoring of the monitoring environment, to meet the real-time monitoring and communication needs in noise-sensitive environments.
It enables real-time monitoring and communication of home spaces in noise-sensitive environments, meeting the special needs of working or active people and avoiding interference from voice calls in noise-sensitive spaces.
Smart Images

Figure CN121099008A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of environmental monitoring and communication, and in particular to a system and method for reproducing and monitoring home space scenes in noise-sensitive environments. Background Technology
[0002] Existing home monitoring typically involves users installing cameras in their homes and viewing the environment at any time via mobile phones or other devices. If users need to communicate or intervene with people in their homes, they can do so via voice communication on their mobile phones or other devices, and the voice information will be played back in the home environment.
[0003] However, due to their specific functions, places such as offices, schools, and hospitals have stricter requirements for sound. People in these places are limited by the noise-sensitive environment and cannot achieve interconnection and communication during home monitoring through voice calls.
[0004] Therefore, current home monitoring products on the market fail to consider the monitoring needs of people in noise-sensitive environments, resulting in untimely or delayed information transmission, making it impossible for people working or active in noise-sensitive spaces to meet their home communication needs. Summary of the Invention
[0005] This invention provides a home space scene reproduction and monitoring system and method in noise-sensitive environments. It integrates three technical systems: text-to-analog sound conversion, real-time acquisition of surveillance camera images, and voice monitoring of the monitoring environment. This system constructs a comprehensive acoustic and visual scene reproduction of the home environment, meeting the special needs of people working or active in noise-sensitive environments for real-time monitoring and communication of their home spaces.
[0006] To solve the above-mentioned technical problems, the technical solution adopted by the present invention is as follows:
[0007] Firstly, a method for reproducing and monitoring the scene of a home in a noise-sensitive environment is provided, including:
[0008] The server, audio equipment, and monitoring equipment are set up in the home space. The server is connected to the audio equipment and monitoring equipment. The server has a server-side application installed for home monitoring scene playback service.
[0009] A mobile terminal device and earphone in a noise-sensitive space environment, the mobile terminal device and earphone are connected, and the mobile terminal device has a client installed with a home monitoring scene playback service;
[0010] When a user in a noise-sensitive environment needs to monitor their home, the client receives the user's monitoring instructions and controls the server's monitoring equipment to acquire video information and the audio equipment to collect audio information.
[0011] The client displays video information via a mobile terminal device and sends audio information to headphones for output.
[0012] When a user needs to intervene in their home space, the client receives the user's non-voice information and sends it to the server.
[0013] The server converts non-voice information into voice information and sends it to the audio equipment for playback.
[0014] Furthermore, audio equipment includes loudspeakers and microphones;
[0015] The loudspeaker establishes a Bluetooth connection with the server;
[0016] The monitoring equipment is a camera, and the camera has monitoring software;
[0017] The camera establishes a video transmission link with the server.
[0018] Furthermore, monitoring commands include surveillance commands and listening commands;
[0019] The client receives monitoring commands from the user and controls the server's monitoring equipment to acquire video information and the audio equipment to collect audio information, including:
[0020] The client receives monitoring commands initiated by the user and sends the monitoring commands to the server.
[0021] The server controls the camera's monitoring software to acquire the monitoring screen according to the monitoring instructions; it uses the screenshot function to capture images of the monitoring screen at preset time intervals and sends the captured images to the client.
[0022] The client receives the listening command initiated by the user and sends the listening command to the server;
[0023] The server controls the microphone of the audio equipment to record audio files according to the monitoring instructions; and sends the audio files to the client according to a preset period.
[0024] Furthermore, the client receives non-voice information from the user and sends it to the server, including:
[0025] The client obtains the intervention text information input by the user through the text input module of the mobile terminal device, and generates non-voice information based on the intervention text information;
[0026] The client sends non-voice information to the server.
[0027] Furthermore, the server converts non-voice information into voice information and sends it to the audio equipment for playback, including:
[0028] The server parses the non-voice information to obtain the intervention text information, inputs the intervention text information into the text-to-speech program, and simulates the first voice information;
[0029] The server sends the first voice message to the audio equipment, which then plays it through the loudspeaker.
[0030] Furthermore, after the server sends the first voice message to the audio equipment and plays it through the loudspeaker, it also includes:
[0031] The server records the first playback sound corresponding to the first voice information through the microphone of the audio equipment;
[0032] The server converts the first played sound into the first text information and determines whether the first text information is consistent with the intervention text information;
[0033] If they match, the server determines that the text-to-speech program is working correctly.
[0034] If there is a discrepancy, the server will determine that the text-to-speech program is malfunctioning.
[0035] Furthermore, after the server determines that the text-to-speech program is malfunctioning, it also includes:
[0036] The server sends a program error message to the client;
[0037] The client disconnects from the text input module of the mobile terminal device based on the program error message and connects to the sensing control module of the mobile terminal device.
[0038] The client obtains the user's sensor control information through the sensor control module and sends it to the server;
[0039] The server obtains the pre-set second voice information corresponding to the sensing control information; and sends the second voice information to the audio equipment for playback through the loudspeaker.
[0040] Furthermore, the sensing control information can be either contact-based or non-contact-based.
[0041] Secondly, a method for reproducing and monitoring home scene in noise-sensitive environments is provided, applied to a home scene reproduction and monitoring system in noise-sensitive environments. The system includes a server, audio equipment, and monitoring equipment located in the home space. The server is connected to the audio equipment and monitoring equipment, and the server has a server-side application for home scene reproduction services installed. A mobile terminal device and headphones are located in the noise-sensitive environment. The mobile terminal device is connected to the headphones, and the mobile terminal device has a client-side application for home scene reproduction services installed. The method includes:
[0042] When a user in a noise-sensitive environment needs to monitor their home, the client receives the user's monitoring instructions and controls the server's monitoring equipment to acquire video information and the audio equipment to collect audio information.
[0043] The client displays video information via a mobile terminal device and sends audio information to headphones for output.
[0044] When a user needs to intervene in their home space, the client receives the user's non-voice information and sends it to the server.
[0045] The server converts non-voice information into voice information and sends it to the audio equipment for playback.
[0046] The beneficial effects achieved by this invention are as follows:
[0047] The system consists of a server, audio equipment, and monitoring equipment set up in the home space. The server is connected to the audio and monitoring equipment and has a server-side application for home monitoring scene reproduction service installed on it. A mobile terminal device and headphones are located in a noise-sensitive environment. The mobile terminal device is connected to the headphones and has a client-side application for the home monitoring scene reproduction service installed on it. When a user in a noise-sensitive environment needs to monitor their home space, the client receives the user's monitoring instructions and controls the server-side monitoring equipment to acquire video information and the audio equipment to collect audio information. The client displays the video information on the mobile terminal device and sends the audio information to the headphones for output. When a user needs to intervene in the home space, the client receives the user's non-voice information and sends it to the server. The server converts the non-voice information into voice information and sends it to the audio equipment for playback. By integrating three technical systems—text-to-analog conversion, real-time acquisition of surveillance camera footage, and voice monitoring of the monitoring environment—a comprehensive acoustic and visual scene reproduction system for the home environment is constructed, meeting the special needs of people working or active in noise-sensitive environments for real-time monitoring and communication. Attached Figure Description
[0048] Figure 1 This is a structural diagram of the home space scene reproduction and monitoring system in a noise-sensitive space environment according to the present invention;
[0049] Figure 2 This is a flowchart of the method for regenerating and monitoring home space scenarios in noise-sensitive environments according to the present invention. Detailed Implementation
[0050] The present invention will be further described below with reference to the accompanying drawings. The following embodiments are only used to more clearly illustrate the technical solution of the present invention, and should not be used to limit the scope of protection of the present invention.
[0051] like Figure 1 As shown, this embodiment of the invention provides a home space scene reproduction and monitoring system for noise-sensitive environments, including:
[0052] A server 11, an audio device 12, and a monitoring device 13 are installed in the home space. The server 11 is connected to the audio device 12 and the monitoring device 13. The server 11 is equipped with a server-side 111 for home monitoring scene playback service.
[0053] A mobile terminal device 21 and an earphone 22 are located in a noise-sensitive space environment. The mobile terminal device 21 is connected to the earphone 22. The mobile terminal device 21 has a client 211 installed with a home monitoring scene playback service.
[0054] When a user in a noise-sensitive environment needs to monitor their home space, the client 211 receives the user's monitoring instructions and controls the monitoring equipment 13 of the server 11 to obtain video information and the audio equipment 12 to collect audio information according to the monitoring instructions.
[0055] The client 211 displays video information through the mobile terminal device 21 and sends audio information to the headphones 22 for output;
[0056] When a user needs to intervene in their home space, the client 211 receives the user's non-voice information and sends the non-voice information to the server 111.
[0057] The server 111 converts the non-voice information into voice information and sends it to the audio device 12 for playback.
[0058] Noise-sensitive spaces refer to the spatial environment of buildings such as hospitals, schools, government agencies, and research institutions that require a quiet environment.
[0059] It should be noted that the audio equipment includes a loudspeaker and a microphone, and the loudspeaker establishes a Bluetooth connection with the server; the monitoring equipment is a camera, which has monitoring software; the camera establishes a video transmission link with the server.
[0060] Based on the above Figure 1 The system shown in this application includes the following software: a text-to-speech "home monitoring scene replay" program, NAT traversal software accessible from the external network, and monitoring software built into the camera. The NAT traversal software is a technology that enables the internal network server to receive external network data packets. NAT traversal, also known as intranet traversal, is used to ensure that data packets with a specific source IP address and port number are not blocked by NAT devices and are correctly routed to the internal network host. NAT traversal can be understood as a dedicated channel initiated by the internal network machine (Client) to the external network server (Server), with the purpose of allowing clients on the external network to access internal network services through the external network server.
[0061] The "Home Monitoring Scene Regeneration" software is primarily built on the Flask framework for its backend. It utilizes the Python open-source text-to-speech library pyttsx3 to write a conversion program, enabling the monitoring party to convert text input into speech and transmit it to the monitored environment. It also uses the open-source automated testing library pyautogui to write a program for real-time dynamic image capture, synchronizing the camera's built-in monitoring software feed to the client. Finally, it uses the open-source audio signal processing library pyaudio to write a monitoring environment program. The client application is built using HTML.
[0062] pyttsx3 is a cross-platform text-to-speech library that can run on Windows, Linux, and macOS without installing any additional dependencies. pyttsx3 uses the system's built-in TTS (text-to-speech) engine, thus ensuring high stability and usability across various operating systems.
[0063] The server can utilize a home laptop placed in a home environment, and install intranet penetration software (such as cpolar) to obtain an address and interface A that can be accessed from the external network; run the "Home Monitoring Scene Regeneration" program in the server's Python compiler, and set the interface to be consistent with interface A, that is, to remotely access the service at that address;
[0064] The server is connected to an audio system; in a home environment, the server is connected to an audio device, such as a Bluetooth speaker, via Bluetooth, and both the server's voice output and microphone are set as audio devices; the audio device is placed within the monitoring area to ensure that the sound can be heard in a timely manner by personnel monitoring security.
[0065] Install cameras; in a home environment, place the cameras in the monitored area; install the camera's software on the server side for real-time monitoring and enable the monitoring screen.
[0066] Combination Figure 1 In the embodiments shown, preferably, in some embodiments of the present invention, the monitoring instructions include surveillance instructions and listening instructions;
[0067] Client 211 receives monitoring commands from the user and controls server 111's monitoring equipment 13 to acquire video information and audio equipment 12 to collect audio information, including:
[0068] Client 211 receives the monitoring command initiated by the user and sends the monitoring command to server 111;
[0069] The server 111 controls the camera's monitoring software to acquire the monitoring screen according to the monitoring instructions; it uses the screenshot function to capture the monitoring screen at preset time intervals and sends the captured images to the client 211; specifically, it uses the screenshot function provided by pyautogui to capture the monitoring screen of the camera's software every 0.5 seconds and updates the client's display image in real time.
[0070] Client 211 receives the listening command initiated by the user and sends the listening command to server 111;
[0071] The server 111 controls the microphone of the audio device 12 to record audio files according to the monitoring instructions; the audio files are then sent to the client according to a preset period. Specifically, a recording program based on the pyaudio library is used to record and store the audio files on the server 111. The program updates the audio files every 5 seconds and automatically plays them on the client 211 after each update. Users can monitor the voice environment through headphones without interfering with noise-sensitive spaces.
[0072] Combination Figure 1 In the embodiments shown, preferably in some embodiments of the present invention, the client 211 receives non-voice information from the user and sends the non-voice information to the server 111, including:
[0073] Client 211 obtains the intervention text information input by the user through the text input module of the mobile terminal device, and generates non-voice information based on the intervention text information; for example, a user in a noise-sensitive environment uses a mobile phone to view the monitoring of the home space and finds that the gas stove has not been turned off after use. At this time, it is necessary to remind the people in the home space to turn off the gas stove, so a prompt needs to be issued. However, it is not appropriate for the user in the noise-sensitive environment to speak. At this time, the user can input the intervention text information in the text input module of the mobile phone, which could be "the gas stove has not been turned off";
[0074] Client 211 sends non-voice information to server 111;
[0075] The server 111 parses the non-voice information to obtain the intervention text information "Gas stove not turned off". The intervention text information "Gas stove not turned off" is input into the text-to-speech program to simulate the first voice information.
[0076] The server 111 sends the first voice information to the audio device 12, which plays it through the loudspeaker.
[0077] This allows users in noise-sensitive environments to intervene in their home space simply by entering text information on their mobile phones, which can then be converted into voice messages and played in their home.
[0078] It should be noted that since text-to-speech is performed by a program, conversion errors may occur. If the text-to-speech is incorrect, it cannot be used for intervention. For example, if "the gas stove is not turned off" is converted into "the TV is not turned off" as the speech message, it will lead to a misunderstanding by people in the home.
[0079] Therefore, after the server 111 sends the first voice information to the audio device 12 and plays it through the loudspeaker, it also includes:
[0080] The server 111 records the first playback sound corresponding to the first voice information through the microphone of the audio device 12;
[0081] The server 111 converts the first playback sound into the first text information and determines whether the first text information is consistent with the intervention text information;
[0082] If they match, then the server confirms that the text-to-speech program is working properly.
[0083] If there is a discrepancy, the server will determine that the text-to-speech program is malfunctioning.
[0084] In the event of an error in the text-to-speech program, server 111 sends an error message to client 211.
[0085] The client 211 disconnects from the text input module of the mobile terminal device according to the program error message and connects to the sensing control module of the mobile terminal device 21; the sensing control information is either contact sensing control information or non-contact sensing control information. The contact sensing control information is triggered by a pre-set trigger module on the touch screen of the mobile phone. When the user clicks the trigger module, the corresponding second voice information is triggered by default.
[0086] The non-contact sensing control information is gesture control information, that is, the user's gestures are monitored through the front camera of the mobile phone, and different gestures are pre-set with corresponding second voice information.
[0087] The client 211 obtains the user's sensing control information through the sensing control module and sends it to the server 111; the server 111 obtains the pre-set second voice information corresponding to the sensing control information; and sends the second voice information to the audio device 12 for playback through the loudspeaker.
[0088] The beneficial effects of the embodiments of the present invention are as follows:
[0089] By integrating three technical systems—text-to-analog audio conversion, real-time acquisition of surveillance camera footage, and voice monitoring of the surveillance environment—a comprehensive acoustic and visual scene reconstruction of the home environment is constructed, meeting the special needs of people working or active in noise-sensitive environments for real-time monitoring and communication of their home spaces.
[0090] Based on the above embodiments describing the method for regenerating and monitoring home space scenarios in noise-sensitive environments, the following embodiments illustrate the system for regenerating and monitoring home space scenarios in noise-sensitive environments.
[0091] like Figure 2 As shown, this embodiment of the invention provides a method for reproducing and monitoring home space scenes in a noise-sensitive environment, including:
[0092] 201. When a user in a noise-sensitive environment needs to monitor their home, the client receives the user's monitoring instructions and controls the server's monitoring equipment to acquire video information and the audio equipment to collect audio information according to the monitoring instructions.
[0093] 202. The client displays video information through the mobile terminal device and sends audio information to the headphones for output;
[0094] 203. When a user needs to intervene in their home space, the client receives the user's non-voice information and sends it to the server.
[0095] 204. The server converts the non-voice information into voice information and sends it to the audio equipment for playback.
[0096] The beneficial effects of the embodiments of the present invention are as follows:
[0097] By integrating three technical systems—text-to-analog audio conversion, real-time acquisition of surveillance camera footage, and voice monitoring of the surveillance environment—a comprehensive acoustic and visual scene reconstruction of the home environment is constructed, meeting the special needs of people working or active in noise-sensitive environments for real-time monitoring and communication of their home spaces.
[0098] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0099] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0100] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0101] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0102] The above are merely embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention are included within the scope of the claims of the present invention pending approval.
Claims
1. A home space scene reproduction and monitoring system for noise-sensitive environments, characterized in that, include: A server, audio equipment, and monitoring equipment are installed in a home space. The server is connected to the audio equipment and the monitoring equipment. The server is equipped with a server-side application for home monitoring scene playback service. A mobile terminal device and earphones in a noise-sensitive space environment, wherein the mobile terminal device is connected to the earphones and the mobile terminal device has the client of the home monitoring scene reproduction service installed; When a user in the noise-sensitive environment needs to monitor the home space, the client receives the user's monitoring instructions and controls the monitoring equipment on the server to acquire video information and the audio equipment to collect audio information according to the monitoring instructions; The client displays the video information through the mobile terminal device and sends the audio information to the headphones for output. When the user needs to intervene in the home space, the client receives the user's non-voice information and sends the non-voice information to the server; The server converts the non-voice information into voice information and sends it to the audio device for playback.
2. The home space scene reproduction and monitoring system in a noise-sensitive environment according to claim 1, characterized in that, The audio equipment includes a loudspeaker and a microphone; The loudspeaker establishes a Bluetooth connection with the server; The monitoring device is a camera, and the camera has monitoring software. The camera establishes a video transmission link with the server.
3. The home space scene reproduction and monitoring system in a noise-sensitive environment according to claim 2, characterized in that, The monitoring commands include surveillance commands and listening commands; The client receives monitoring instructions from the user, and controls the monitoring equipment on the server to acquire the video information and the audio equipment to collect the audio information according to the monitoring instructions, including: The client receives the monitoring command initiated by the user and sends the monitoring command to the server; The server controls the camera's monitoring software to acquire monitoring images according to the monitoring instructions; it uses the screenshot function to capture images of the monitoring images at preset time intervals and sends the captured images to the client. The client receives the listening instruction initiated by the user and sends the listening instruction to the server; The server controls the microphone of the audio device to record an audio file according to the monitoring command; and sends the audio file to the client according to a preset period.
4. The home space scene reproduction and monitoring system in a noise-sensitive environment according to claim 3, characterized in that, The client receives the user's non-voice information and sends the non-voice information to the server, including: The client obtains the intervention text information input by the user through the text input module of the mobile terminal device, and generates non-voice information based on the intervention text information; The client sends the non-voice information to the server.
5. The home space scene reproduction and monitoring system in a noise-sensitive space environment according to claim 4, characterized in that, The server converts the non-voice information into voice information and sends it to the audio device for playback, including: The server parses the non-voice information to obtain the intervention text information, inputs the intervention text information into a text-to-speech program, and simulates the first voice information; The server sends the first voice information to the audio device, which then plays it through the loudspeaker.
6. The home space scene reproduction and monitoring system in a noise-sensitive space environment according to claim 5, characterized in that, After the server sends the first voice information to the audio device and plays it through the loudspeaker, the system further includes: The server records a first playback sound corresponding to the first voice information through the microphone of the audio device. The server converts the first playback sound into first text information and determines whether the first text information is consistent with the intervention text information. If they match, the server determines that the text-to-speech program is working properly. If there is a discrepancy, the server determines that the text-to-speech program is malfunctioning.
7. The home space scene reproduction and monitoring system in a noise-sensitive environment according to claim 6, characterized in that, After the server determines that the text-to-speech program is malfunctioning, it also includes: The server sends a program error message to the client. The client disconnects from the text input module of the mobile terminal device according to the program error message and connects to the sensing control module of the mobile terminal device. The client obtains the user's sensing control information through the sensing control module and sends it to the server; The server obtains the second voice information that is pre-set and corresponds to the sensing control information; and sends the second voice information to the audio device for playback through the loudspeaker.
8. The home space scene reproduction and monitoring system in a noise-sensitive space environment according to claim 7, characterized in that, The sensing control information can be either contact sensing control information or non-contact sensing control information.
9. A method for reproducing and monitoring the scene of a home space in a noise-sensitive environment, characterized in that, A home space scene reproduction and monitoring system for noise-sensitive environments is provided. The home space scene reproduction and monitoring system for noise-sensitive environments includes a server, audio equipment and monitoring equipment set up in the home space. The server is connected to the audio equipment and the monitoring equipment. The server is equipped with a server-side application for home monitoring scene reproduction services. A mobile terminal device and an earphone in a noise-sensitive environment, wherein the mobile terminal device is connected to the earphone, and the mobile terminal device has a client installed with the home monitoring scene reproduction service, the method comprising: When a user in the noise-sensitive environment needs to monitor the home space, the client receives the user's monitoring instructions and controls the monitoring equipment on the server to acquire video information and the audio equipment to collect audio information according to the monitoring instructions; The client displays the video information through the mobile terminal device and sends the audio information to the headphones for output. When the user needs to intervene in the home space, the client receives the user's non-voice information and sends the non-voice information to the server; The server converts the non-voice information into voice information and sends it to the audio device for playback.