An intelligent internet of things security response method, device and system based on voice wake-up

By using multimodal data acquisition and adaptive listening technology, combined with natural language processing and environmental sensors, the system enables IoT devices to identify and respond efficiently to emergencies in closed or special environments. This solves the problems of low accuracy and response delay in existing technologies, and improves the initiative and emergency response efficiency of IoT security monitoring.

CN121056494BActive Publication Date: 2026-03-27XIAMEN TIANYU INTERNET OF THINGS TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-11-04
Publication Date
2026-03-27

AI Technical Summary

Technical Problem

Existing IoT devices lack proactive voice interaction capabilities in closed or special environments, making it impossible to identify emergencies in a timely manner and trigger effective responses. Traditional voice recognition has low accuracy in complex environments and lacks contextual understanding and intelligent decision-making capabilities, leading to delays or failures in safety responses during emergencies.

Method used

Employing a multi-response mechanism that integrates multimodal data acquisition, dynamic monitoring area construction, adaptive listening, multi-level wake-up and semantic analysis, scene adaptation and fusion, and edge-cloud collaboration, the system uses smart speakers, cameras, and sensors to collaboratively perceive and monitor voice input in real time, dynamically adjust the listening strategy, and combine natural language processing and environmental sensor data to accurately identify and efficiently handle emergencies.

Benefits of technology

It enables real-time and accurate identification and dynamic adaptation of emergencies in closed or special IoT scenarios, improving the initiative of IoT security monitoring and emergency response efficiency, and ensuring that users can communicate with the outside world in a timely manner and obtain targeted assistance in emergency scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121056494B_ABST
    Figure CN121056494B_ABST
Patent Text Reader

Abstract

The application provides a smart Internet of Things safety response method, device and system based on voice wake-up, relates to the technical field of smart Internet of Things safety monitoring, and the method comprises the following steps: performing real-time monitoring on voice input, and simultaneously cooperating with three fixed Internet of Things equipment nodes in the environment to perform sensing, so as to obtain multi-modal sensing data; the equipment nodes comprise a smart sound box, a camera and a sensor; based on the multi-modal sensing data, real-time position coordinates of each Internet of Things node are extracted, spatial operation is performed on the position coordinates through a polygon region construction algorithm, and a dynamic monitoring region is obtained. The application realizes accurate identification and efficient disposal of an emergency in a closed or special Internet of Things scene.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of intelligent Internet of Things security monitoring, in particular to an intelligent Internet of Things security response method, device and system based on voice wake-up. BACKGROUND

[0002] The security monitoring of existing Internet of Things devices mainly relies on video monitoring and sensor alarms, lacks active voice interaction and emergency response capabilities, and in closed or special environments such as elevators and kitchens, users often cannot trigger the alarm system in time or effectively communicate with the outside world when an emergency occurs. At the same time, the accuracy of traditional voice recognition systems is low in complex environments, and existing wake-up word detection technology is single, lacking context understanding and intelligent decision-making capabilities.

[0003] For example, when a resident is taking an elevator in a high-rise building, the elevator suddenly stops due to mechanical failure, there is no manual alarm device in the car, and the video monitoring outside the elevator fails to capture the abnormality in the car in real time due to the negligence of the monitoring personnel on duty. The device sensors provided by the elevator can only monitor the operating parameters and are not associated with the emergency scene of the trapped personnel. The resident shouted several times, but the traditional voice recognition system could not accurately recognize the emergency keywords due to the background noise in the elevator, such as the sound of the car ventilation equipment running and the mechanical friction sound in the shaft. Moreover, it could not determine that the current situation is an emergency scene in combination with the elevator outage. Neither the alarm was triggered nor the communication channel with the outside world was established, resulting in the resident being trapped for nearly 1 hour before being discovered by another resident. The case highlights the defects of current technology. In closed and special Internet of Things scenarios with environmental noise, there is a lack of fusion mechanism for anti-interference voice wake-up capability and scene state perception, which cannot capture the user's emergency needs through active voice interaction, and it is also difficult to achieve intelligent decision-making based on the device operating state, resulting in obvious delay or even failure of safety response in emergency situations. SUMMARY

[0004] The technical problem to be solved by the present application is to provide an intelligent Internet of Things security response method, device and system based on voice wake-up, which realizes accurate identification and efficient disposal of emergency situations in closed or special Internet of Things scenarios.

[0005] To solve the above technical problems, the technical solution of the present application is as follows:

[0006] In a first aspect, an intelligent Internet of Things security response method based on voice wake-up, the method comprising:

[0007] Real-time monitoring of voice input, while cooperating with three fixed Internet of Things device nodes in the environment for perception, obtaining multi-modal perception data, the device nodes including an intelligent sound box, a camera and a sensor;

[0008] Based on the multi-modal perception data, the real-time position coordinates of each Internet of Things node are extracted, and the position coordinates are subjected to spatial operation through a polygon area construction algorithm to obtain a dynamic monitoring area.

[0009] The dynamic monitoring area is subjected to grid division to obtain a plurality of analysis units; adaptive listening weight coefficients are calculated according to the dynamic change characteristics in each analysis unit; the voice listening strategy is dynamically adjusted based on the weight coefficients, and when a preset emergency keyword is detected in the voice stream, a first wake-up instruction is triggered and generated;

[0010] After receiving the first wake-up instruction, the voice data is subjected to semantic understanding through natural language processing to obtain semantic analysis results of the emergency degree and specific requirements.

[0011] Based on the semantic analysis results, scene adaptation and fusion analysis are performed in combination with real-time acquired environmental sensor data to determine and obtain a response level.

[0012] According to the response level, a corresponding multiple security response mechanism is started, alarm information is sent to the monitoring center, a two-way voice dialogue function is started, relevant rescue departments are automatically contacted, and the on-site situation is continuously monitored.

[0013] Further, the voice input is subjected to real-time listening, and at the same time, three fixed Internet of Things device nodes in the environment are cooperatively perceived to obtain multi-modal perception data, the device nodes including an intelligent sound box, a camera and a sensor, comprising:

[0014] The user voice input is subjected to real-time collection, and the collected original voice stream is subjected to buffering to obtain a voice data queue for analysis.

[0015] Based on the voice data queue, the three fixed-position Internet of Things device nodes including the intelligent sound box, the camera and the sensor are cooperatively scheduled, audio, video and environmental state data corresponding to the voice data are synchronously collected to form a multi-modal perception data set.

[0016] The multi-modal perception data set is subjected to timestamp alignment and format unification processing, voice features, image features and sensor reading features are extracted based on the processed data to obtain multi-modal perception data.

[0017] Further, based on the multi-modal perception data, the real-time position coordinates of each Internet of Things node are extracted, and the position coordinates are subjected to spatial operation through a polygon area construction algorithm to obtain a dynamic monitoring area, comprising:

[0018] Based on the multi-modal perception data, the real-time position coordinates of each Internet of Things node are extracted to obtain a node position data set.

[0019] Based on the node position data set, the spatial position points of the three Internet of Things device nodes are determined; the relative distance and azimuth angle between each position point are calculated;

[0020] Based on the calculated relative distance and azimuth angle, a polygonal region is formed by sequentially connecting each position point as a vertex according to the spatial position relationship; whether the polygonal region contains the monitoring range of all nodes is verified to form an initial monitoring region;

[0021] Based on the initial monitoring region, the region boundary is optimized in combination with the real-time working state of each node to obtain a dynamic monitoring region.

[0022] Further, the dynamic monitoring region is divided into a grid to obtain a plurality of analysis units; based on the dynamic change characteristics in each analysis unit, an adaptive listening weight coefficient is calculated; based on the weight coefficient, the voice listening strategy is dynamically adjusted, and when a preset emergency keyword is detected in the voice stream, a first wake-up instruction is triggered and generated, including:

[0023] Based on the dynamic monitoring region, the dynamic monitoring region is divided into a grid to obtain a plurality of analysis units;

[0024] According to the analysis unit, in combination with the dynamic change characteristics in each unit in the multi-modal perception data, an adaptive listening weight coefficient is calculated;

[0025] Based on the adaptive listening weight coefficient, the voice listening strategy is dynamically adjusted, and the voice data in the high-weight region is preferentially processed to obtain an adjusted voice listening strategy;

[0026] The voice stream is analyzed in real time through the adjusted voice listening strategy, and when a preset emergency keyword is detected in the voice stream, a first wake-up instruction is triggered and generated.

[0027] Further, after receiving the first wake-up instruction, the voice data is semantically understood through natural language processing to obtain a semantic analysis result of the emergency degree and specific demand, including:

[0028] After receiving the first wake-up instruction, the corresponding voice data is extracted from the cached voice data queue;

[0029] The extracted voice data is transmitted to the edge processing node, and the voice content is semantically analyzed through natural language processing technology to obtain a semantic analysis result;

[0030] Based on the semantic analysis result, the emergency degree feature and user demand feature in the voice content are identified;

[0031] According to the emergency degree feature and user demand feature, a structured semantic analysis result is obtained.

[0032] Further, based on the semantic analysis result, scene adaptation and fusion analysis are performed in combination with real-time environment sensor data to determine and obtain a response level, including:

[0033] Based on the structured semantic analysis result, scene state data collected in real time by the environment sensor is obtained;

[0034] The semantic analysis result and the scene state data are fused and analyzed to identify a current environment scene type and related event features;

[0035] According to the environment scene type and the related event features, in combination with a preset response rule library, a final response level is evaluated and determined.

[0036] Further, according to the response level, a corresponding multiple security response mechanism is started, alarm information is sent to a monitoring center, a two-way voice dialogue function is started, a related rescue department is automatically contacted, and a field status is continuously monitored, including:

[0037] Based on the response level, a corresponding multiple security response mechanism is selected from a preset response strategy library;

[0038] According to the selected response mechanism, alarm information containing the response level and a field position is first sent to the monitoring center;

[0039] At the same time of sending the alarm information, the two-way voice dialogue function is started to establish a voice communication connection with the user;

[0040] Based on a state of the voice communication connection, a rescue department corresponding to the response level is automatically contacted, and a field status is continuously monitored until an event is handled.

[0041] In a second aspect, an intelligent Internet of Things security response system based on voice wake-up includes:

[0042] An acquisition module is configured to listen to voice input in real time, and simultaneously sense three fixed Internet of Things device nodes in an environment to obtain multi-modal sensing data, the device nodes including an intelligent sound box, a camera, and a sensor;

[0043] A calculation module is configured to extract real-time position coordinates of each Internet of Things node based on the multi-modal sensing data, perform spatial operation on the position coordinates through a polygon area construction algorithm to obtain a dynamic monitoring area, divide the dynamic monitoring area into grids to obtain a plurality of analysis units, calculate adaptive listening weight coefficients according to dynamic change features in each analysis unit, and dynamically adjust a voice listening strategy based on the weight coefficients, and when it is detected that a preset emergency keyword is contained in a voice stream, trigger and generate a first wake-up instruction;

[0044] The analysis module is configured to, after receiving the first wake-up instruction, perform semantic understanding on the voice data through natural language processing to obtain semantic analysis results of the emergency degree and specific requirements.

[0045] The fusion module is configured to, based on the semantic analysis results, perform scene adaptation and fusion analysis by combining real-time acquired environment sensor data to determine and obtain a response level.

[0046] The processing module is configured to, according to the response level, start a corresponding multi-level security response mechanism, send alarm information to a monitoring center, start a two-way voice dialogue function, automatically contact a related rescue department, and continuously monitor a field situation.

[0047] In a third aspect, a computing device includes:

[0048] one or more processors;

[0049] a storage device storing one or more programs, when the one or more programs are executed by the one or more processors, the one or more processors implement the method.

[0050] In a fourth aspect, a computer-readable storage medium stores a program, when the program is executed by a processor, the method is implemented.

[0051] The above scheme of the present application at least has the following beneficial effects:

[0052] Because the multi-modal data acquisition, dynamic monitoring area construction, adaptive listening, multi-level wake-up and semantic analysis, scene adaptation and fusion, and end-to-end cloud collaborative multi-response mechanism are adopted, the technical problems of the prior art, such as the dependence on video monitoring and sensor alarm in the safety monitoring of the Internet of Things, the lack of active voice interaction and scene-based emergency response capability, the low accuracy of traditional voice recognition in complex environments, the single wake-up word detection technology and the lack of context understanding capability, and the inability to trigger effective response in emergency situations in time, are overcome, and the technical effects of realizing real-time and accurate identification of emergency situations, intelligent response level determination of dynamic adaptation to scenes, and rapid collaborative multi-emergency disposal in closed or special Internet of Things scenes such as elevators and kitchens are achieved, the initiative, identification accuracy and emergency response efficiency of safety monitoring of the Internet of Things are improved, and the user can timely establish communication with the outside world and obtain targeted rescue in an emergency situation. BRIEF DESCRIPTION OF DRAWINGS

[0053] Figure 1 is a flowchart of an intelligent Internet of Things safety response method based on voice wake-up provided by an embodiment of the present application.

[0054] Figure 2A schematic diagram of an intelligent Internet of Things security response system based on voice wake-up is provided by an embodiment of the present application. DETAILED DESCRIPTION

[0055] Exemplary embodiments of the present disclosure will be described in greater detail below with reference to the accompanying drawings. Although exemplary embodiments of the present disclosure are shown in the drawings, it is understood that the present disclosure can be implemented in various forms and should not be limited by the embodiments set forth herein. Rather, these embodiments are provided so that the present disclosure can be more thoroughly understood and the scope of the present disclosure can be accurately conveyed to those skilled in the art.

[0056] As Figure 1 shown, an embodiment of the present application proposes a voice wake-up based intelligent Internet of Things security response method, which comprises the following steps:

[0057] Step 1, real-time monitoring of voice input is performed, and at the same time, three fixed Internet of Things device nodes in the environment are sensed to obtain multi-modal sensing data, the device nodes including an intelligent sound box, a camera, and a sensor;

[0058] Step 2, based on the multi-modal sensing data, real-time position coordinates of each Internet of Things node are extracted, and spatial operation is performed on the position coordinates through a polygon region construction algorithm to obtain a dynamic monitoring region;

[0059] Step 3, the dynamic monitoring region is divided into a grid to obtain a plurality of analysis units; adaptive monitoring weight coefficients are calculated according to the dynamic change characteristics in each analysis unit; and a voice monitoring strategy is dynamically adjusted based on the weight coefficients, and when a preset emergency keyword is detected in the voice stream, a first wake-up instruction is triggered and generated;

[0060] Step 4, after receiving the first wake-up instruction, semantic understanding of the voice data is performed through natural language processing to obtain semantic analysis results of the emergency degree and specific requirements;

[0061] Step 5, based on the semantic analysis results, scene adaptation and fusion analysis are performed in combination with real-time acquired environmental sensor data to determine and obtain a response level;

[0062] Step 6, according to the response level, a corresponding multiple security response mechanism is started, alarm information is sent to a monitoring center, a two-way voice dialogue function is started, relevant rescue departments are automatically contacted, and the on-site situation is continuously monitored.

[0063] In the embodiment of the present application, because the three types of fixed Internet of Things device nodes of cooperative intelligent sound box, camera and sensor are adopted to synchronously collect voice, video and environmental state data to obtain multi-modal perception data, combined with the polygon area construction algorithm and the generation of dynamic monitoring area by the real-time working state of the node, the dynamic monitoring area is grid divided and the adaptive listening weight coefficient is calculated according to the dynamic change characteristics of the unit to adjust the listening strategy, after the first wake-up triggered by the preset emergency keyword detection, the natural language processing technology is used to analyze the emergency degree and specific demand of the voice data, combined with the semantic analysis result, environmental sensor data and preset response rule library to determine the response level, and then the multiple safety response mechanisms such as alarm, two-way voice dialogue, contact with rescue departments and continuous monitoring are started according to the response level, so that the technical problems of the existing Internet of Things safety monitoring relying on video monitoring and sensor alarm, lacking active voice interaction and scene-based emergency response capability, low accuracy of traditional voice recognition in complex environment, single wake-up word detection technology and no context understanding ability, which leads to the inability to trigger effective response in emergency situation in time, are overcome, and the technical effects of realizing real-time accurate identification of emergency situation, scene-based intelligent response level determination and fast and cooperative multiple emergency disposal in closed or special Internet of Things scenes such as elevators and kitchens are achieved, the initiative of Internet of Things safety monitoring and the accuracy of voice recognition are improved, the user can timely establish communication with the outside world and obtain targeted rescue in emergency scene, and the efficiency of Internet of Things safety emergency response is improved.

[0064] In a preferred embodiment of the present application, step 1 can include:

[0065] Step 1.1, real-time collection of user voice input, buffering of the collected original voice stream to obtain a voice data queue for analysis, specifically including: through the voice collection device deployed in the target monitoring scene such as elevator and kitchen, such as the multi-channel microphone array carried by the intelligent sound box connected with the Internet of Things system, all voice inputs issued by the user in the scene are continuously and real-time collected, the original voice stream collected each time without filtering and processing is arranged in order according to the actual collection time of the voice signal to form an ordered and callable voice data queue for analysis.

[0066] Step 1.2, based on the voice data queue, cooperatively scheduling three fixed-position Internet of Things device nodes, including a smart speaker, a camera, and a sensor, to synchronously collect audio, video, and environmental state data corresponding to the voice data, forming a multi-modal perception data set, specifically including: after forming the voice data queue to be analyzed, according to the collection time stamp corresponding to each piece of voice data in the voice data queue, sending a cooperative scheduling instruction to the three Internet of Things device nodes pre-installed at different positions in the monitoring scene; wherein the smart speaker node receives the scheduling instruction and immediately collects audio data completely synchronized with the voice data collection time, which is used to supplement the feature dimension of voice information; the camera node receives the scheduling instruction and synchronously collects scene video data within the same time period, which is used to capture visual information such as personnel state and device position in the scene; the sensor node receives the scheduling instruction and synchronously collects contemporaneous environmental state data, such as elevator speed in the elevator scene, temperature in the car, gas concentration in the kitchen scene, and device operating parameters; then the audio data, video data, and environmental state data collected by the three types of device nodes, namely the smart speaker, the camera, and the sensor, are integrated together to form a multi-modal perception data set covering multiple information dimensions of sound, vision, and environmental parameters.

[0067] Step 1.3, time stamp alignment and format unification processing of the multi-modal perception data set, extracting voice features, image features, and sensor reading features based on the processed data to obtain multi-modal perception data, specifically including: for the multi-modal perception data set that has been formed, first extract the collection time stamp of each data in the set, and calibrate and adjust the time stamps of audio data, video data, and environmental state data based on the collection time stamp of voice data, to ensure that the three types of data are completely corresponding and matched in the time dimension, eliminating the time sequence deviation caused by the difference in collection response speed of different device nodes; then, according to the pre-set unified data format standard, convert the different format heterogeneous data output by the smart speaker, camera, and sensor into standardized data format that can be directly processed, avoiding the influence of subsequent data fusion analysis due to non-uniform data format; finally, based on the standardized time series data after time stamp alignment and format unification processing, extract voice features reflecting voice key information such as voiceprint frequency and intonation change, image features reflecting scene visual changes such as dynamic regions and personnel actions in video pictures, and sensor reading features reflecting environmental state such as parameter value size and change trend monitored by the sensor, to finally form structured and complete multi-modal perception data.

[0068] In the embodiment of the present application, because the real-time collection of user voice input and the caching as a to-be-analyzed voice data queue are adopted, the three types of fixed Internet of Things nodes of intelligent sound box, camera and sensor are cooperatively scheduled based on the queue to synchronously collect corresponding audio, video and environmental state data to form a multi-modal perception data set, and after the data set is timestamp aligned and format unified, the voice features, image features and sensor reading features are extracted, the technical means overcome the technical problems in the prior art that the data collection in the Internet of Things safety perception is single, the multi-source data is not synchronized in time sequence, the formats are not unified, the subsequent analysis cannot be effectively fused, and the environment and user state cannot be comprehensively captured through multi-dimensional data, and thus the multi-dimensional high-quality perception data including voice, image and environmental parameters that are consistent in time sequence, standardized in format and provided for subsequent steps such as dynamic monitoring area construction and adaptive listening are achieved.

[0069] In a preferred embodiment of the present application, step 2 can include:

[0070] Step 2.1, based on the multi-modal perception data, real-time position coordinates of each Internet of Things node are extracted to obtain a node position data set, specifically including: the obtained multi-modal perception data is called to filter out information related to the positions of the three Internet of Things device nodes, such as real-time coordinate data fed back by a positioning device of the sensor, position parameters recorded after installation of the intelligent sound box and the camera combined with real-time feedback of environmental position-related information, the information is arranged, and real-time position coordinates of each Internet of Things node in the current monitoring scene, such as an elevator car or a kitchen space, are extracted, and the real-time position coordinates of the three nodes are recorded according to the node type to form a node position data set containing node identifiers and corresponding real-time coordinates.

[0071] Step 2.2, based on the node position data set, the spatial position points of the three Internet of Things device nodes are determined; the relative distance and azimuth angle between each position point are calculated, specifically including: based on the obtained node position data set, the spatial position points corresponding to the three Internet of Things device nodes are first determined, such as the position point of the intelligent sound box on the car side wall, the position point of the camera on the car top and the position point of the sensor beside the car control panel in the elevator scene; then the straight-line distance between any two position points is calculated according to the coordinates of each position point, such as the distance between the intelligent sound box position point and the camera position point, the distance between the camera position point and the sensor position point, and the distance between the sensor position point and the intelligent sound box position point; at the same time, the azimuth angle between any two position points is determined according to the directional relationship of the coordinates of each position point, such as the specific direction of the intelligent sound box position point relative to the camera position point and the specific direction of the sensor position point relative to the intelligent sound box position point, so as to clearly determine the relative position relationship of the three nodes in space.

[0072] Step 2.3, based on the calculated relative distance and azimuth angle, the position points are connected in turn to form a polygonal region according to the spatial position relationship; verify whether the polygonal region contains the monitoring range of all nodes to form an initial monitoring region, which specifically includes: according to the calculated relative distance and azimuth angle between the position points, the spatial position points corresponding to the three Internet of Things device nodes are taken as the three vertices of the polygon, and the three vertices are connected in turn with straight lines according to the actual spatial position relationship of the three points in the current monitoring scene to form a closed polygonal region; then check the audio monitoring coverage of the smart speaker, the visual monitoring coverage of the camera, and the environmental parameter sensing coverage of the sensor, and confirm whether the monitoring ranges of the three nodes are all contained in the polygonal region. If the monitoring range of a certain node exceeds the polygonal region, adjust the connection mode of the vertices or fine-tune the position of the vertices until the polygonal region completely covers the monitoring range of all nodes, and finally form the initial monitoring region.

[0073] Step 2.4, based on the initial monitoring region, optimize the region boundary in combination with the real-time working state of each node to obtain a dynamic monitoring region, which specifically includes: first, obtain the real-time working state of the three Internet of Things device nodes, such as checking whether the smart speaker can normally collect audio, whether the camera can clearly capture pictures, and whether the sensor can accurately read environmental data, to determine whether each node is in normal working state; if a node fails, such as the camera lens being blocked resulting in reduced monitoring range, or the sensor failing resulting in invalid sensing range, then based on the initial monitoring region, exclude the region that the faulty node cannot effectively monitor from the initial monitoring region, and adjust the region boundary according to the actual monitoring ability of the normal nodes; if all nodes are working normally, then combine the slight changes in the real-time monitoring range of each node to make small corrections to the boundary of the initial monitoring region, and finally obtain a dynamic monitoring region that completely adapts to the current node working state.

[0074] In the embodiment of the present application, because the real-time position coordinates of each Internet of Things node are extracted from the multi-modal perception data to obtain a node position data set, the spatial position points of the three Internet of Things device nodes are determined based on the data set, the relative distance and azimuth between each position point are calculated, the polygon region is formed by sequentially connecting each position point as the vertex according to the spatial position relationship, and it is verified whether it contains all node monitoring ranges to form an initial monitoring region, and then the initial monitoring region boundary is optimized in combination with the real-time working state of each node, so the technical problems that the monitoring region in the existing Internet of Things security monitoring is mostly fixedly set and cannot be dynamically adjusted according to the actual position and working state of the device node, resulting in that the monitoring range does not match the actual demand in the elevator, kitchen and other closed or special scenes, and the working state of the node is not considered, resulting in insufficient monitoring reliability, are overcome, and a dynamic monitoring region that is accurately adapted to the actual position of the Internet of Things node, covers all node monitoring ranges and can be dynamically adjusted according to the working state of the node is constructed.

[0075] In a preferred embodiment of the present application, step 3 can include:

[0076] Step 3.1, based on the dynamic monitoring region, the dynamic monitoring region is grid divided to obtain a plurality of analysis units, specifically including: first obtaining the dynamic monitoring region, determining the grid division precision according to the actual space size and scene characteristics of the dynamic monitoring region, for example, in the elevator car and other narrow and closed scenes, the dynamic monitoring region is uniformly divided into square grid with edge length of tens of centimeters in horizontal and vertical directions, in the kitchen and other slightly larger scenes, the dynamic monitoring region is divided into square grid with edge length of about one meter, ensuring that each grid can accurately cover the local region and will not waste computing resources due to too fine division; strictly follow the boundary range of the dynamic monitoring region during the division process, do not exceed the region and do not miss any part in the region, finally divide the dynamic monitoring region into a plurality of independent and continuous space regions, each space region is an analysis unit.

[0077] Step 3.2, according to the analysis unit, combined with the dynamic change characteristics in each unit in the multi-modal perception data, the adaptive monitoring weight coefficient is calculated, specifically including: first, the multi-modal perception data is retrieved, which contains the extracted speech features, image features and sensor reading features; for each analysis unit obtained, the feature data corresponding to each analysis unit is selected from the multi-modal perception data, for example, the speech signal intensity change in the analysis unit, the moving frequency and amplitude of the personnel in the unit in the image, and the fluctuation of the environmental parameters in the unit monitored by the sensor, which together constitute the dynamic change characteristics in each unit; then, according to the significant degree of the dynamic change characteristics, the calculation rule is set, if the speech signal intensity in a certain analysis unit frequently increases, the personnel moving amplitude is large or the environmental parameter fluctuation is obvious, it means that the unit dynamic change characteristics, give higher weight score in calculation, otherwise give lower score, through the rule, the dynamic change characteristics of each analysis unit are quantitatively calculated, and finally the adaptive monitoring weight coefficient corresponding to each analysis unit is obtained.

[0078] Step 3.3, based on the adaptive monitoring weight coefficient, dynamically adjust the speech monitoring strategy, preferentially process the speech data of high weight area, get the adjusted speech monitoring strategy, specifically including: statistics of the adaptive monitoring weight coefficient of all analysis units, the coefficient value is sorted in order from high to low, and the high weight area, the medium weight area and the low weight area are divided; based on the sorting result, adjust the speech monitoring strategy, allocate more system computing resources to the high weight area, for example, shorten the collection interval of the area speech data, preferentially carry out noise reduction processing and feature extraction on the speech data of the area, ensure that the speech information of the high weight area can be quickly processed; for the medium weight area, monitor according to the normal resource configuration, maintain the normal speech data collection and processing rhythm; for the low weight area, appropriately reduce the allocation of computing resources, prolong the speech data collection interval and reduce the processing priority, avoid meaningless resource consumption; through such resource allocation adjustment, the adjusted speech monitoring strategy with focus and efficient resource utilization is formed.

[0079] Step 3.4, real-time analysis of the voice stream by the adjusted voice monitoring strategy, when detecting that the voice stream contains preset emergency keywords, triggering and generating a first wake-up instruction, specifically including: starting voice stream real-time analysis according to the adjusted voice monitoring strategy, preferentially processing voice streams in high-weight areas, loading a preset emergency keyword library in the processing process, the keyword library containing life-saving, being trapped, failure, gas leakage and other emergency keywords for elevator, kitchen and other scenes; through voice feature matching technology, the real-time collected voice stream is compared with the keywords in the keyword library, and at the same time, noise reduction processing technology is used to filter the interference noise in the environment, such as elevator ventilation equipment operation sound and kitchen appliance working sound, reducing the influence of noise on feature comparison; if the content in the voice stream is detected to have a feature matching degree reaching a preset threshold with any keyword in the keyword library, and the content comes from valid voice, not environmental noise or invalid dialogue, then the instruction generation mechanism is triggered immediately to generate a first wake-up instruction.

[0080] In the embodiment of the application, the dynamic monitoring area is grid-divided to obtain a plurality of analysis units, the adaptive monitoring weight coefficient is calculated based on the dynamic change characteristics in each analysis unit in the structured multi-modal perception data, the voice monitoring strategy is dynamically adjusted according to the weight coefficient to preferentially process voice data in high-weight areas, and the real-time analysis of the voice stream is performed through the adjusted strategy, and the first wake-up instruction is triggered when the preset emergency keywords are detected. Therefore, the technical problems of the prior art Internet of Things voice monitoring, such as no differentiated monitoring of the monitoring area, unreasonable resource allocation, inability to pay attention to high-dynamic-change or high-risk areas, low detection efficiency of preset emergency keywords in complex environments, easy to miss detection, and thus affecting the timeliness of emergency wake-up, are overcome, and the technical problems are further solved. Therefore, the voice monitoring resources are accurately allocated, the focus is on the high-weight areas with dynamic changes, the detection speed and accuracy of the preset emergency keywords in complex scenes are improved, and the first wake-up instruction can be quickly and accurately triggered.

[0081] In a preferred embodiment of the application, the above-mentioned step 4 can include:

[0082] Step 4.1, after receiving the first wake-up instruction, extracting the corresponding voice data from the cached voice data queue, specifically including: after receiving the first wake-up instruction, first acquiring the time information when the first wake-up instruction is triggered, and searching in the cache formed voice data queue to be analyzed based on this time; since the voice data queue is stored in the order of the collection time of the original voice stream, by matching the first wake-up instruction trigger time with the collection time of each piece of voice data in the voice data queue, one or more complete voice data corresponding to the wake-up instruction is located, and then the located voice data is extracted from the queue, ensuring that the extracted voice data completely reflects the user's voice expression content when the wake-up instruction is triggered, avoiding missing key voice information due to time matching deviation.

[0083] Step 4.2, transmitting the extracted voice data to the edge processing node, and performing semantic analysis on the voice content through natural language processing technology to obtain the semantic analysis result, specifically including: sending the extracted voice data to the edge processing node deployed in the monitoring scene through a special data transmission channel. The edge processing node has low delay and high real-time data processing capability, and can quickly respond to voice data analysis requirements; in the edge processing node, the received voice data is first subjected to targeted noise reduction processing to filter out environmental interference noise, such as the running sound of the ventilation equipment in the elevator car, the mechanical friction sound in the shaft, or the noise of the electrical appliances in the kitchen, etc.; then the noise-reduced voice data is converted into text information through voice-to-text technology; and then the natural language processing technology is used to analyze the syntax and semantic association of the text information, such as identifying the subject, predicate, and object structure in the text, understanding the core intent of the user's expression, for example, the user says I am trapped in the elevator, the door cannot be opened, through analysis it can be determined that the user is in the elevator and faces the situation that the door cannot be opened, and finally the semantic analysis result containing the text content and core intent association information is formed.

[0084] Step 4.3, based on the semantic analysis result, identifying the emergency level feature and user demand feature in the voice content, specifically including: based on the obtained semantic analysis result, first identifying the emergency level feature in the voice content, by analyzing the emotional related expressions in the text of the analysis result, such as whether it contains words with urgency such as fast, immediately, save me, etc., whether there are repeated expressions, and at the same time combining the information of voice tone change, speech speed, etc. recorded in the process of voice to text, to judge the emergency level of the user's current situation, for example, the user repeatedly says fast save me, the elevator is not moving, which corresponds to a higher emergency level. Then identify the user demand feature, extract the user's explicit or implicit expression from the semantic analysis result, such as the user says the elevator is broken, I can't get out, which can extract the event type of elevator failure and the core demand of needing help to get out of the elevator, or the user says there is a smell of coal in the kitchen, which can extract the event type of gas leakage and the core demand of needing to check the leakage and handle it, so as to determine the emergency level feature and the user demand feature respectively.

[0085] Step 4.4, according to the emergency level feature and the user demand feature, obtaining the structured semantic analysis result, specifically including: structuring the identified emergency level feature and user demand feature; first set the fixed format of the structured result, the format includes emergency level field, demand type field and event description field, in the emergency level field, according to the identified emergency level feature, mark three levels of high, medium and low, for example, the user's urgent call for help corresponds to high level, and ordinary fault feedback corresponds to medium level; in the demand type field, according to the user demand feature, mark the specific demand category, such as being trapped rescue, equipment maintenance, safety hazard handling, etc.; in the event description field, integrate the key information in the semantic analysis result, record the scene, specific situation, core content of user expression, etc. of the event occurrence, for example, the scene is high-rise residential elevator, the situation is that the elevator stops suddenly and the door cannot be opened, and the user expresses that he is trapped and cannot contact the outside world, through such arrangement, the scattered feature information is converted into structured semantic analysis result with clear arrangement and unified format.

[0086] In the embodiment of the present application, because the corresponding voice data is extracted from the cached voice data queue after receiving the first level wake-up instruction, the extracted voice data is transmitted to the edge side processing node, and the voice content is semantically analyzed by natural language processing technology, the emergency degree feature and the user demand feature in the voice content are identified based on the semantic analysis result, and the structured semantic analysis result is obtained according to the two types of features. Therefore, the technical problem that the traditional voice recognition system can only detect the emergency keyword but cannot associate the corresponding complete voice data, lacks the deep semantic understanding ability, and cannot accurately judge the severity of the emergency and the specific demand of the user, and then it is difficult to carry out targeted emergency response is overcome. Therefore, the technical effect that the complete voice information corresponding to the first level wake-up is accurately obtained, the severity of the emergency and the actual demand of the user are determined through the efficient semantic analysis of the edge side, the clear and accurate data support is provided for the subsequent scene adaptation and response level judgment by using the structured semantic analysis result, and the deviation of the emergency response direction caused by the fuzzy semantic understanding is avoided.

[0087] In a preferred embodiment of the present application, step 5 can include:

[0088] Step 5.1, based on the structured semantic analysis result, acquiring the scene state data collected by the environment sensor in real time, specifically including: first reading the obtained structured semantic analysis result, extracting the key information from the result, including the scene related expression and the event type mentioned by the user; according to the key information, sending data acquisition instruction to the environment sensor in the corresponding scene, for example, if the semantic analysis result shows that the user is in the elevator scene and mentions being trapped, sending instruction to the running state sensor in the elevator and the environment sensor in the car; if the semantic result points to the kitchen scene and mentions the odor, send instructions to the gas concentration sensor, temperature sensor and smoke sensor in the kitchen; after receiving the instruction, the environment sensor immediately collects the real-time data at the current time.

[0089] Step 5.2, fuse the semantic analysis results with the scene state data for analysis, identify the current environment scene type and related event characteristics, including: the structured semantic analysis results are associated with the real-time scene state data obtained for matching and fusion analysis; first, for scene type identification, if the user explicitly mentions the elevator in the semantic results, and the sensor feedback is that the elevator is out of service, the door cannot be normally opened, and there is a personnel activity signal in the car, then the current environment scene type is determined to be an elevator trapped scene; if the user mentions the kitchen in the semantic results, and the gas concentration sensor feedback is that the concentration exceeds the safety threshold, the temperature sensor data is normal, and the smoke sensor has no alarm, then the scene type is determined to be a kitchen gas leakage scene; then identify the related event characteristics, extract the core attributes of the event combining semantic and sensor data, for example, in the elevator trapped scene, the event characteristics include elevator mechanical failure leading to stop, people trapped in the car, and the door cannot be manually opened; in the kitchen gas leakage scene, the event characteristics include excessive gas concentration, no open fire or smoke, and the user perceives the odor and feedback. Through the above fusion analysis, the current scene type and key characteristics of the event are clearly defined.

[0090] Step 5.3, according to the environment scene type and related event characteristics, combining the preset response rule library, evaluate and determine the final response level, including: first, call the preset response rule library, the rule library stores the corresponding association rules of different scene types, event characteristics and response levels, for example, the rule library clearly states that the elevator trapped scene plus the user's urgent emotion plus the elevator completely stops running and no emergency door opening signal corresponds to the highest response level, the kitchen gas slightly exceeds the standard plus the user only feedbacks the odor without mentioning discomfort corresponds to the medium response level, and the device slightly fails to feedback plus the sensor monitoring device can still operate normally corresponds to the lower response level; then compare the identified current environment scene type and related event characteristics with the rules in the response rule library one by one, find the completely matched or closest rule item; according to the matched rule item, determine the corresponding response level, for example, in the elevator trapped scene and when the highest level rule condition is met, the final response level is determined to be the first level response; when the kitchen gas slightly exceeds the standard scene matches the medium level rule condition, the response level is determined to be the second level response, ensuring that the response level accurately adapts to the emergency level and risk size of the event.

[0091] In the embodiment of the present application, because the structured-based semantic analysis result is adopted to obtain the scene state data collected in real time by the environmental sensor, the semantic analysis result and the scene state data are fused and analyzed to identify the current environmental scene type and the related event features, and the preset response rule library is combined to evaluate and determine the final response level, the technical problems of the prior art, such as lack of fusion mechanism of voice semantics and scene state perception, inability to accurately identify the scene type and event nature in combination with the real-time state of the environmental equipment, and lack of ability to formulate an adaptive response level according to the scene and event features, thereby causing deviation or failure of emergency response, are overcome, and the technical effects of accurately associating the user semantic demand with the real-time environmental state in the elevator, kitchen and other closed or special Internet of Things scenes, clearly identifying the current scene type and event core features, and matching a scientific and reasonable response level according to the rule library to avoid insufficient or excessive response are achieved.

[0092] In a preferred embodiment of the present application, step 6 can include:

[0093] Step 6.1, based on the response level, selecting the corresponding multiple safety response mechanism from the preset response strategy library, specifically including: first calling the preset response strategy library, and the strategy library pre-stores the association rules of different response levels and corresponding multiple safety response mechanisms, for example, the response level is divided into first emergency, second important and third routine, wherein the first emergency corresponds to scenes such as elevator personnel being trapped and kitchen gas being seriously leaked, and the response mechanism includes real-time high-priority alarm to the monitoring center, starting a two-way voice dialogue, linking a fire or professional rescue team, and continuously high-frequency on-site monitoring; the second important corresponds to scenes such as elevator minor failure without personnel being trapped and kitchen gas being slightly over-standard, and the response mechanism includes regular alarm to the monitoring center, starting a two-way voice consultation, linking equipment maintenance personnel, and continuously regular-frequency on-site monitoring; the third routine corresponds to scenes such as device minor abnormal feedback and small amplitude fluctuation of environmental parameters, and the response mechanism includes low-priority prompt to the monitoring center, starting one-way voice notification, linking property inspection personnel, and on-demand on-site monitoring. After the system obtains the determined final response level, the multiple safety response mechanisms completely corresponding to the level are screened out from the response strategy library through rule matching, so as to ensure that the selected mechanism can accurately adapt to the emergency degree and risk size of the current event, and avoid insufficient or excessive response.

[0094] Step 6.2, according to the selected response mechanism, first send alarm information containing response level and field location to the monitoring center, including: according to the selected response mechanism, first integrate event key information, in addition to response level and field location, also including identified environmental scene type, such as elevator trapped scene, kitchen gas leakage scene, related event characteristics, such as elevator downtime, gas concentration value, form a complete alarm information package; then through the special data transmission channel between the monitoring center, the alarm information package is sent to the monitoring center in real-time push mode, data encryption is used in the transmission process to ensure information security; after receiving the alarm information, the monitoring center will trigger the sound and light prompt of the center terminal, and automatically pop up the alarm information window on the monitoring interface, clearly show the response level, field location, such as XX community 3# building 2# elevator, XX unit 101 room kitchen and scene type and event characteristics, so that the monitoring center staff can accurately master the on-site emergency situation at the first time, avoid the problem that the abnormal situation cannot be found in time due to information loss or delay in traditional monitoring.

[0095] Step 6.3, at the same time of sending alarm information, start the two-way voice dialogue function and establish voice communication connection with the user, including: at the same time of sending alarm information to the monitoring center, automatically activate the intelligent sound box voice interaction module deployed on the scene, establish voice transmission link through real-time data channel between Internet of Things devices; one end of the link connects the microphone and speaker of the on-site intelligent sound box, the other end connects the voice processing terminal of the monitoring center or the system preset intelligent voice response unit, forming a two-way voice communication channel; after the channel is established, automatically send voice prompt to the user, such as your help information has been received, rescue is being contacted, please explain the specific situation on the scene, the user can speak directly through the intelligent sound box, the voice signal is transmitted to the monitoring center or the intelligent voice response unit in real time through the channel, the staff or the system can directly respond to the user, such as telling the rescue progress, soothing the user's emotion, asking for supplementary information, such as whether there are old people or children in the elevator, whether the kitchen has closed the gas valve, so as to solve the problem that the user cannot establish effective communication channel with the outside world in traditional technology.

[0096] Step 6.4, based on the state of the voice communication connection, automatically contact the rescue department corresponding to the response level, and continuously monitor the on-site situation until the event is handled, specifically including: first monitoring the state of the two-way voice communication connection, if the user confirms the on-site emergency in the voice communication, such as explicitly indicating that the elevator door cannot be opened, being trapped for more than 20 minutes, or judging through voice content analysis that the event needs external rescue intervention, then according to the selected response level, the contact number of the corresponding rescue organization is retrieved from the preset rescue department directory, for example, the contact number of the fire prevention rescue department and the elevator professional maintenance company for the first emergency response, and the contact number of the gas company maintenance department and the equipment after-sales team for the second important response; through automatic dialing or sending a dispatch instruction containing event information, contact the rescue department to inform the emergency situation and the specific location that needs to be handled; at the same time, continuously call the on-site camera and sensor to collect real-time on-site pictures, such as the state of the people in the elevator, whether there is smoke in the kitchen, and environmental data such as changes in temperature and gas concentration in the car, and transmit the data to the monitoring center in real time, so that the staff can master the on-site dynamics in real time; until the rescue department arrives at the scene and handles it, for example, the elevator door is opened, the trapped person is rescued, or the kitchen gas leakage point is repaired and the concentration is restored to a safe value, the system stops continuous monitoring, and records the event handling result for archiving.

[0097] In the embodiment of the application, because the corresponding multiple security response mechanisms are selected from the preset response strategy library based on the response level, the alarm information containing the response level and the on-site location is sent to the monitoring center according to the selected mechanism, the two-way voice dialogue function is started at the same time as the alarm information is sent to establish a voice communication connection with the user, and the rescue department corresponding to the response level is automatically contacted based on the state of the voice communication connection and the on-site situation is continuously monitored until the event is handled, the technical problems of the prior art Internet of Things security monitoring lacking targeted emergency response strategies, being unable to accurately transmit event information to the monitoring center in an emergency, lacking effective voice communication channels between the user and the outside world, rescue department linkage being not timely, and the event handling process lacking continuous monitoring, resulting in low emergency response efficiency, the user's needs being unable to be met in time, and even response failure, are overcome, and the technical effects of achieving precise and multi-level security response matching the emergency level in elevator, kitchen and other closed or special Internet of Things scenes, allowing the monitoring center to quickly grasp the key information of the event, the user to communicate with the outside world in time to obtain comfort and feedback details, the rescue department to accurately intervene, and the event to be fully controllable through continuous monitoring until it is properly handled, and improving the timeliness, targeting and effectiveness of Internet of Things security emergency response are achieved.

[0098] As shown in Figure 2 The embodiment of the application also provides an intelligent Internet of Things security response system based on voice wake-up, which comprises:

[0099] An acquisition module is configured to listen to a voice input in real time, and cooperates with three fixed Internet of Things device nodes in an environment to obtain multi-modal sensing data, the device nodes including an intelligent sound box, a camera, and a sensor;

[0100] A calculation module is configured to extract real-time position coordinates of each Internet of Things node based on the multi-modal sensing data, perform spatial operation on the position coordinates through a polygon region construction algorithm, and obtain a dynamic monitoring region; the dynamic monitoring region is divided into a plurality of analysis units through gridding; adaptive listening weight coefficients are calculated according to dynamic change characteristics in each analysis unit; a voice listening strategy is dynamically adjusted based on the weight coefficients; when a preset emergency keyword is detected in a voice stream, a first wake-up instruction is triggered and generated;

[0101] An analysis module is configured to, after receiving the first wake-up instruction, perform semantic understanding on voice data through natural language processing, and obtain semantic analysis results of an emergency degree and specific requirements;

[0102] A fusion module is configured to, based on the semantic analysis results, perform scene adaptation and fusion analysis in combination with real-time acquired environmental sensor data, and determine and obtain a response level;

[0103] A processing module is configured to, according to the response level, start a corresponding multiple security response mechanism, send alarm information to a monitoring center, start a two-way voice dialogue function, automatically contact a related rescue department, and continuously monitor a site condition.

[0104] The above describes preferred embodiments of the present application. It should be noted that, for those skilled in the art, without departing from the principles of the present application, a number of improvements and refinements can be made, and these improvements and refinements should also be considered as falling within the scope of protection of the present application.

Claims

1. A voice wake-up based smart Internet of Things security response method, characterized in that, The method comprises: Real-time monitoring of voice input, while cooperating with three fixed Internet of Things device nodes in the environment for sensing, obtaining multi-modal sensing data, the device nodes including a smart speaker, a camera and a sensor; Based on the multi-modal sensing data, the real-time position coordinates of each Internet of Things node are extracted, and the position coordinates are subjected to spatial operation through a polygon region construction algorithm to obtain a dynamic monitoring region; The dynamic monitoring region is divided into grids to obtain a plurality of analysis units; adaptive listening weight coefficients are calculated according to the dynamic change characteristics in each analysis unit; and the voice listening strategy is dynamically adjusted based on the weight coefficients, and when it is detected that the voice stream contains a preset emergency keyword, a first wake-up instruction is triggered and generated, including: Based on the dynamic monitoring region, the dynamic monitoring region is divided into grids to obtain a plurality of analysis units; According to the analysis unit, the dynamic change characteristics in each unit in the multi-modal sensing data are combined to calculate adaptive listening weight coefficients, wherein if the voice signal strength in the analysis unit frequently increases, the personnel movement amplitude is large or the environmental parameter fluctuation is obvious, it indicates that the unit dynamic change characteristics, and a higher weight score is given during calculation, otherwise a lower score is given, the dynamic change characteristics of each analysis unit are quantitatively calculated through rules, and finally the adaptive listening weight coefficients corresponding to each analysis unit are obtained; Based on the adaptive listening weight coefficients, the voice listening strategy is dynamically adjusted, and the voice data in the high-weight region is preferentially processed to obtain an adjusted voice listening strategy; According to the adjusted voice listening strategy, the voice stream real-time analysis is started, and the voice stream in the high-weight region is preferentially processed, a preset emergency keyword library is loaded in the processing process, the voice stream collected in real time is compared with the keywords in the keyword library through voice feature matching technology, and the interference noise in the environment is filtered through noise reduction processing technology, if the content in the voice stream is detected to have a feature matching degree reaching a preset threshold with any one of the keywords in the keyword library, and the content comes from valid voice, not environmental noise and invalid dialogue, then the instruction generation mechanism is triggered immediately to generate a first wake-up instruction; After receiving the first wake-up instruction, the voice data is subjected to semantic understanding through natural language processing to obtain semantic analysis results of the emergency degree and specific demand; Based on the semantic analysis results, scene adaptation and fusion analysis are performed in combination with the real-time acquired environmental sensor data to determine and obtain a response level; According to the response level, the corresponding multiple security response mechanisms are started, alarm information is sent to the monitoring center, a two-way voice dialogue function is started, relevant rescue departments are automatically contacted, and the on-site situation is continuously monitored.

2. The method of claim 1, wherein the method further comprises: Real-time monitoring of voice input, while cooperating with three fixed Internet of Things device nodes in the environment for sensing, obtaining multi-modal sensing data, the device nodes including a smart speaker, a camera and a sensor, including: Real-time collection of user voice input, and buffering of the collected original voice stream to obtain a voice data queue for analysis; Based on the voice data queue, three fixed position Internet of Things device nodes including a smart speaker, a camera and a sensor are cooperatively scheduled, audio, video and environmental state data corresponding to the voice data are synchronously collected to form a multi-modal perception data set; The multi-modal perception data set is timestamp aligned and format unified, voice features, image features and sensor reading features are extracted based on the processed data to obtain multi-modal perception data. 3.The method of claim 2, wherein, Based on the multi-modal perception data, real-time position coordinates of each Internet of Things node are extracted, and spatial operation is performed on the position coordinates through a polygon area construction algorithm to obtain a dynamic monitoring area, including: Based on the multi-modal perception data, real-time position coordinates of each Internet of Things node are extracted to obtain a node position data set; Based on the node position data set, the spatial position points of the three Internet of Things device nodes are determined; the relative distance and azimuth angle between each pair of position points are calculated; Based on the calculated relative distance and azimuth angle, a polygon area is formed by connecting the position points in sequence according to the spatial position relationship; it is verified whether the polygon area contains the monitoring range of all nodes to form an initial monitoring area; Based on the initial monitoring area, the area boundary is optimized in combination with the real-time working state of each node to obtain a dynamic monitoring area.

4. The intelligent Internet of Things security response method based on voice wake-up according to claim 3, characterized in that, After receiving the first wake-up instruction, the semantic understanding of the voice data is performed through natural language processing to obtain semantic analysis results of the urgency and specific needs, including: After receiving the first wake-up instruction, the corresponding voice data is extracted from the cached voice data queue; The extracted voice data is transmitted to the edge processing node, and the semantic analysis result is obtained by performing semantic analysis on the voice content through natural language processing technology; Based on the semantic analysis result, the urgency feature and user demand feature in the voice content are identified; According to the urgency feature and user demand feature, a structured semantic analysis result is obtained.

5. The intelligent IoT security response method based on voice wake-up according to claim 4, characterized in that, Based on the semantic analysis result, scene adaptation and fusion analysis are performed in combination with the real-time environmental sensor data to determine and obtain the response level, including: Based on the structured semantic analysis result, scene state data collected in real time by the environmental sensor is obtained; The semantic analysis result and the scene state data are analyzed to identify the current environmental scene type and related event features; According to the environmental scene type and related event features, the final response level is evaluated and determined in combination with the preset response rule library.

6. The intelligent Internet of Things security response method based on voice wake-up according to claim 5, characterized in that, According to the response level, the corresponding multiple security response mechanism is started, the alarm information is sent to the monitoring center, the two-way voice dialogue function is started, the related rescue departments are automatically contacted, and the on-site situation is continuously monitored, including: Based on the response level, the corresponding multiple security response mechanism is selected from the preset response strategy library; According to the selected response mechanism, the alarm information containing the response level and the on-site location is first sent to the monitoring center; At the same time of sending the alarm information, the two-way voice dialogue function is started to establish a voice communication connection with the user; Based on the state of the voice communication connection, the rescue department corresponding to the response level is automatically contacted, and the on-site situation is continuously monitored until the event is handled.

7. An intelligent voice wake-up based IoT security response system, the system implements the method of any one of claims 1 to 6, characterized in that, including: An acquisition module is configured to listen to voice input in real time, and to obtain multi-modal sensing data by cooperating with three fixed Internet of Things device nodes in an environment, the device nodes including an intelligent sound box, a camera, and a sensor; A calculation module is configured to extract real-time position coordinates of each Internet of Things node based on the multi-modal sensing data, to perform spatial operation on the position coordinates by a polygon region construction algorithm, and to obtain a dynamic monitoring region; the dynamic monitoring region is divided into a plurality of analysis units by grid division; and adaptive listening weight coefficients are calculated based on dynamic change characteristics in each analysis unit; A voice listening strategy is dynamically adjusted based on the weight coefficients, and a first wake-up instruction is triggered and generated when a preset emergency keyword is detected in a voice stream; An analysis module is configured to perform semantic understanding on voice data by natural language processing after receiving the first wake-up instruction, and to obtain semantic analysis results of an emergency degree and specific requirements; A fusion module is configured to perform scene adaptation and fusion analysis based on the semantic analysis results and real-time environment sensor data, and to determine and obtain a response level; A processing module is configured to start a corresponding multi-level security response mechanism according to the response level, to send alarm information to a monitoring center, to start a two-way voice dialogue function, to automatically contact a related rescue department, and to continuously monitor a site condition.

8. A computing device, comprising: One or more processors; A storage device is configured to store one or more programs, and when the one or more programs are executed by the one or more processors, the one or more processors implement the method in any one of claims 1 to 6. The computer readable storage medium stores a program, and the program is executed by the processor to implement the method in any one of claims 1 to 6.

9. A computer-readable storage medium, characterized in that, ​

Citation Information

Patent Citations

  • Intelligent monitoring method and system based on Internet of Things, medium and program product

    CN119135741A

  • Home safety monitoring system and method

    CN120783454A