An integrated operation and maintenance management platform based on the Internet of Things

Through the Internet of Things acquisition module and analysis module, the speech signal is characterized and combined with the line of sight direction and mouth data, the problems of speech analysis accuracy and control command accuracy in campus scenarios are solved, and efficient speech control is achieved in noisy environments.

CN119544751BActive Publication Date: 2025-07-22GUANGZHOU SHENG YE INFORMATION TECH CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202411724829.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-28
Publication Date
2025-07-22
Estimated Expiration
2044-11-28

AI Technical Summary

Technical Problem

When the existing comprehensive operation and maintenance management platform performs speech analysis in campus scenarios, especially when the number of speakers suddenly increases during the get out of class period, the analysis speed decreases, resulting in a decrease in recognition accuracy and the accuracy of control instructions.

Method used

The Internet of Things acquisition module is used to obtain voice signals and monitoring data, and the instruction characteristics of the voice signals are analyzed through the data monitoring module, including the fluctuation values of the semantic correlation degree and time domain dimensions. The analysis module calculates the instruction characterization parameters. The instruction control module determines the clear category of the voice signals, and determines the control instructions based on the matching of the line of sight and mouth data. Finally, the instruction response module controls the equipment in the classroom.

Benefits of technology

It improves the accuracy and management efficiency of speech analysis, ensures the recognition accuracy of the monitoring target voice signal in noisy environments, and improves the accuracy of control instructions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119544751B_ABST
    Figure CN119544751B_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of Internet of Things management, and in particular to an integrated operation and maintenance management platform based on the Internet of Things. The present invention is provided with an Internet of Things acquisition module, a data monitoring module, an analysis module, an instruction control module, and an instruction response module. The data monitoring module is used to analyze the voice signals in each time domain segment to determine the instruction characteristics of the voice signals. The analysis module is used to calculate the instruction representation parameters for the voice signals in the time domain segment based on the instruction characteristics to determine the instruction clarity category of the voice signals. The instruction control module is used to analyze the voice signals based on the instruction clarity category of the voice signals in the time domain segment to determine the control instructions. The instruction response module is used to control the predetermined devices in the classroom based on the determined control instructions. By determining the control instructions, the present invention improves the accuracy of device control, the accuracy of voice analysis, and the management efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of Internet of Things management, and particularly to an integrated operation and maintenance management platform based on the Internet of Things. Background Art

[0002] The integrated operation and maintenance management platform is a platform that integrates a variety of operation and maintenance tools and technologies. Its application scenarios are very extensive, including but not limited to manufacturing, e-commerce, finance, telecommunications, Internet, etc. They all rely on a stable IT system to support their business continuity and customer experience, improve the efficiency and quality of operation and maintenance work, reduce the occurrence of faults and accidents, and provide comprehensive security management functions. Especially with the continuous development of technologies such as cloud computing, big data, and artificial intelligence, the integrated operation and maintenance management platform is also constantly innovating and improving, and has great application prospects in the campus environment. It can uniformly manage and maintain various information resources on campus, including infrastructure such as hardware, software, and network, and improve management efficiency and management quality.

[0003] Chinese Patent Publication No.: CN118574148A, discloses an integrated operation and maintenance management system and method based on the Internet of Things. The system includes: a number of monitoring systems, each monitoring system is deployed in a communication base station for real-time monitoring of the power parameters and environmental parameters of the communication base station; a monitoring station, the monitoring station includes: at least one main station, each main station is communicatively connected to a number of slave stations, and each slave station is deployed within a pre-divided area range; each slave station is used to collect the power parameters and environmental parameters monitored by all monitoring systems within the area range where the slave station is located, determine whether the power parameters and / or environmental parameters are abnormal, analyze the abnormal power parameters and / or environmental parameters, generate an abnormal report and a warning message, and upload the abnormal report and the warning message to the main station. The present invention enhances the monitoring function of the power environment of the communication base station and can effectively improve the automation of the operation and maintenance management of the communication base station.

[0004] Chinese Patent Publication No.: CN118863868A, discloses a system operation and maintenance integrated management platform, including: an energy-consuming equipment monitoring system for obtaining the operation data and operation status of each energy-consuming equipment of the power dispatching data network; a full-process operation and maintenance control module for intelligent scheduling of operation work; an energy management system for real-time detection, automatic statistics and analysis of energy utilization by region, category, time, and item based on the Internet of Things and big data. In the present invention, the energy management system has obvious effects in load optimization, fault prediction accuracy, reduction of power outage frequency, and improvement of energy efficiency. Based on the integrated application of artificial intelligence, Internet of Things, and big data analysis technologies, scientific extraction and intelligent governance of data are realized at the data processing layer, and intelligent inspection, proactive operation and maintenance, and intelligent diagnosis are realized at the application layer, and business aggregation, process integration, and data fusion of the energy operation and maintenance governance function in industrial parks.

[0005] However, there are still the following problems in the prior art:

[0006] The existing integrated operation and maintenance management platform is mostly used in industrial parks, but it has a wide range of uses and is also applicable in school scenarios. When the integrated operation and maintenance management platform is used to manage campus scenarios, when voice analysis is required, more than one person's voice signals may be included, especially at break time, when the number of people speaking increases suddenly. In this case, the platform still analyzes all voice signals, which will reduce the analysis speed, resulting in lower recognition accuracy and lower accuracy of voice analysis, affecting the accuracy of the analyzed control instructions. Summary of the invention

[0007] To this end, the present invention provides a comprehensive operation and maintenance management platform based on the Internet of Things, which is used to solve the problem that when voice analysis is required, more than one person's voice signals may be included, especially at get out of class time, when the number of people speaking increases suddenly. In this case, the platform still analyzes all voice signals, which will reduce the analysis speed, resulting in lower recognition accuracy, lower accuracy of voice analysis, and affecting the accuracy of the analyzed control instructions.

[0008] To achieve the above objectives, the present invention provides a comprehensive operation and maintenance management platform based on the Internet of Things, which includes:

[0009] The Internet of Things acquisition module includes a voice acquisition unit arranged in a target area of the classroom for acquiring voice signals and a monitoring unit for acquiring monitoring data;

[0010] A data monitoring module, which is connected to the Internet of Things acquisition module, is used to analyze the voice signal in each time domain segment, and determine the instruction characteristics of the voice signal based on the analysis result, including the semantic relevance corresponding to the voice signal and the fluctuation value of the voice signal in the time domain dimension;

[0011] An analysis module connected to the data monitoring module, for calculating a command characterization parameter for the voice signal in a time domain segment based on the command feature, so as to determine a command clarity category of the voice signal;

[0012] An instruction control module is connected to the analysis module and is used to analyze the voice signal based on the instruction clarity category of the voice signal in the time domain segment to determine the control instruction, including:

[0013] Determine the sight direction of each monitoring target based on the monitoring data, determine the key monitoring target based on the angle between the sight direction and the predetermined receiving direction, synchronously determine the lip shape data of the key monitoring target, match the lip shape data with the voice signal to obtain a text signal to determine the control instruction;

[0014] Or, directly convert the voice signal into a text signal to determine the control instruction;

[0015] An instruction response module, which is connected to the Internet of Things acquisition module and the instruction control module, is used to control the predetermined devices in the classroom based on the determined control instructions.

[0016] Further, the data monitoring module determines the semantic correlation degree corresponding to the voice signal, including,

[0017] Used to convert the voice signal in each time domain segment into text information;

[0018] Used to extract the keywords in the text information;

[0019] Used to determine the semantic correlation degree of the text information according to the keywords.

[0020] Further, the data monitoring module determines the fluctuation value of the voice signal in the time domain dimension, including,

[0021] Used to determine the peak value of the voice signal in each time domain segment;

[0022] Used to determine that the variance of each peak value in the time domain segment is the fluctuation value.

[0023] Further, the analysis module calculates the instruction characterization parameters for the voice signal in the time domain segment, including,

[0024] Used to determine the ratio of the semantic correlation degree to the reference semantic correlation degree as the semantic correlation degree influence parameter;

[0025] Used to determine the ratio of the reference fluctuation value to the fluctuation value as the fluctuation value influence parameter;

[0026] Used to determine the weighted sum value of the semantic correlation degree influence parameter and the fluctuation value influence parameter as the instruction characterization parameter.

[0027] Further, the analysis module determines the instruction clarity category of the voice signal, where,

[0028] If the instruction characterization parameter is less than the instruction characterization parameter threshold, determine that the instruction clarity category is a weak clarity instruction;

[0029] If the instruction characterization parameter is greater than or equal to the instruction characterization parameter threshold, determine that the instruction clarity category is a strong clarity instruction.

[0030] Further, the instruction control module analyzes the voice signal based on the instruction clarity category of the voice signal in the time domain segment to determine the control instruction, where,

[0031] If the command clarity category is a weakly clear command, the sight direction of each monitoring target is determined based on the monitoring data, the key monitoring target is determined based on the angle between the sight direction and the predetermined receiving direction, the lip shape data of the key monitoring target is determined synchronously, and the lip shape data is matched with the voice signal to obtain a text signal to determine the control command;

[0032] If the command clarity category is a strong clear command, the voice signal is directly converted into a text signal to determine the control command.

[0033] Furthermore, the command control module determines the sight direction of each monitoring target, including:

[0034] Determine a facial reference point group of a monitoring target and construct a facial plane according to the facial reference point group;

[0035] Used to determine the direction of the vector perpendicular to the facial plane as the sight direction of the monitoring target;

[0036] The facial reference point group is three points within the facial contour area.

[0037] Furthermore, the command control module determines the key monitoring target based on the angle between the line of sight direction and the predetermined receiving direction, including:

[0038] Used to determine the angle between the line of sight direction corresponding to the monitoring target and the predetermined receiving direction;

[0039] If the degree of the angle is less than the degree of the reference angle, it is determined that the monitoring target corresponding to the line of sight direction is a key monitoring target.

[0040] Furthermore, the instruction control module matches the lip shape data with the voice signal, including:

[0041] Determine a number of associated lip shape keywords based on the lip shape data;

[0042] To match each keyword in the voice signal with the lip shape keyword;

[0043] Converting a speech signal into a text signal and removing unmatched keywords from the text signal;

[0044] If the keyword in the voice signal does not match any of the lip shape keywords, the keyword is determined to be an unmatched keyword.

[0045] Furthermore, the instruction response module controls the predetermined device in the classroom based on the determined control instruction, including:

[0046] Used to determine control instructions according to text signals;

[0047] to determine the predetermined device in the classroom corresponding to the control instruction;

[0048] for sending the control instruction to the predetermined device.

[0049] Compared with the prior art, the present invention is provided with an Internet of Things acquisition module, a data monitoring module, an analysis module, an instruction control module and an instruction response module. The data monitoring module is used to analyze the voice signals in each time domain segment to determine the instruction features of the voice signals, including the semantic correlation degree corresponding to the voice signals and the fluctuation value of the voice signals in the time domain dimension. The analysis module is used to calculate the instruction representation parameters for the voice signals in the time domain segment based on the instruction features to determine the instruction clarity category of the voice signals. The instruction control module is used to analyze the voice signals based on the instruction clarity category of the voice signals in the time domain segment to determine the control instruction. The instruction response module is used to control the predetermined device in the classroom based on the determined control instruction. By determining the control instruction, the present invention improves the accuracy of device control, and improves the accuracy of voice analysis and management efficiency.

[0050] In particular, the present invention analyzes the voice signals in each time domain segment, converts the voice signals into text information and audio information, and then determines the instruction features of the voice signals. In actual use, when receiving the voice signals of the monitoring target, the target device is controlled to complete operations such as turning on and off according to the voice signals. For a single monitoring target, the fluctuation value of the voice signals it emits in the time domain dimension is small. However, in the campus environment, especially during the class break time, the background is noisy, and the voice signals of the monitoring target may be interfered, which is reflected in the increased difficulty of distinguishing the voice signals of the monitoring target subsequently, and then leads to errors in the recognized control instructions and the inability to effectively control the target device. Based on this, the present invention considers overall analysis of the collected voice signals to determine the semantic correlation degree corresponding to the voice signals and the fluctuation value of the voice signals in the time domain dimension, providing a data basis for calculating the instruction representation parameters, making the recognition of the voice signals of the monitoring target more accurate, and improving the accuracy of determining the control instruction.

[0051] In particular, the present invention determines the instruction clarity category of the voice signals by calculating the instruction representation parameters for the voice signals in the time domain segment. When analyzing the voice signals, there will be relatively clear voice signals. At this time, if complex analysis is still performed on all the received voice signals, this non-adaptive adjustment analysis method results in low analysis efficiency for the voice signals. Based on this, the present invention calculates the instruction representation parameters, classifies the voice signals, directly determines the control instruction for the voice signals with strong clear instructions, and performs signal matching on the voice signals with weak clear instructions, making the recognition of the voice signals of the monitoring target more accurate, and improving the accuracy of determining the control instruction.

[0052] In particular, the present invention determines control instructions based on key monitoring targets and lip movement data. For the situation of receiving multiple voice signals in a noisy environment, it is necessary to analyze the voice signals to obtain accurate control instructions. Based on this, the present invention considers real-time monitoring of the target area in the classroom, determines key monitoring targets according to the monitoring data, matches the lip movement data of the key monitoring targets with the voice signals, and then determines control instructions, making the recognition of the voice signals of the monitoring targets more accurate and improving the accuracy of determining control instructions. BRIEF DESCRIPTION OF THE DRAWINGS

[0053] Figure 1 is a schematic structural diagram of an Internet of Things-based integrated operation and maintenance management platform according to an embodiment of the present invention;

[0054] Figure 2 is a logic block diagram for determining the instruction clarity category of the voice signal according to an embodiment of the present invention;

[0055] Figure 3 is a logic block diagram for determining control instructions according to an embodiment of the present invention;

[0056] Figure 4 is a logic block diagram for determining key monitoring targets according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0057] In order to make the objectives and advantages of the present invention clearer, the present invention will be further described below in conjunction with embodiments; it should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0058] The preferred embodiments of the present invention will be described below with reference to the accompanying drawings. Those skilled in the art should understand that these embodiments are only used to explain the technical principles of the present invention and do not limit the protection scope of the present invention.

[0059] It should be noted that in the description of the present invention, unless otherwise clearly specified and limited, the terms "installed", "connected", and "connected" should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be directly connected, or indirectly connected through an intermediate medium, and it can be the communication inside two components. For those skilled in the art, the specific meanings of the above terms in the present invention can be understood according to specific situations.

[0060] Please refer to Figures 1 - 4 shown Figure 1 is a schematic structural diagram of an Internet of Things-based integrated operation and maintenance management platform according to an embodiment of the present invention, Figure 2 is a logic block diagram for determining the instruction clarity category of the voice signal according to an embodiment of the present invention, Figure 3It is a logic block diagram for determining control instructions according to an embodiment of the present invention. Figure 4 It is a logic block diagram for determining key monitoring targets according to an embodiment of the present invention. An integrated operation and maintenance management platform based on the Internet of Things according to the present invention includes:

[0061] An Internet of Things acquisition module, which includes a voice acquisition unit arranged in the target area of the classroom for acquiring voice signals and a monitoring unit for acquiring monitoring data;

[0062] Specifically, the target area of the classroom is the visible area of the integrated operation and maintenance management platform. It can be understood that only by performing voice control on the integrated operation and maintenance management platform in the target area can the integrated operation and maintenance management platform be mobilized to run, which will not be elaborated here.

[0063] Specifically, the acquisition method of the voice acquisition unit is not limited. For example, it can be a voice collector. There are mainly the following types of voice collectors:

[0064] Digital microphone: The digital microphone adds a clock circuit, a memory, and a digital processor, which can minimize noise and restore human voices, and is suitable for environments with high performance requirements.

[0065] Microphone array: The microphone array utilizes the spatial domain filtering characteristics, forms a directional pickup beam by positioning the angle of the waking person, and suppresses the noise outside the beam to ensure high recording quality.

[0066] Audio acquisition and playback system: This type of system includes the conditioning, acquisition, storage, and playback of voice signals. Based on the principle of digital storage and recovery of voice signals, it uses A / D, D / A conversion technologies and voice signal interpolation compression algorithms to implement the digital storage and playback functions of voice signals.

[0067] The above-mentioned voice collectors can all acquire voice signals, which will not be elaborated here.

[0068] Specifically, the acquisition method of the monitoring unit is not limited. For example, it can be a camera, as long as it can ensure that the lip movements of any target in the target area can be monitored. This is the prior art and will not be elaborated here.

[0069] Specifically, the specific content of the voice signal is not limited. In this embodiment, the voice signal can be a sentence or an instruction phrase. For example, "Please help me turn on the air conditioner and adjust the temperature to 26 degrees", "Turn on the air conditioner", "Set the temperature to 26 degrees", etc., which will not be elaborated here.

[0070] A data monitoring module, which is connected to the IoT acquisition module, is used to analyze the voice signals in each time domain segment, and determine the instruction features of the voice signals based on the analysis results, including the semantic correlation degree corresponding to the voice signals and the fluctuation value of the voice signals in the time domain dimension;

[0071] An analysis module, which is connected to the data monitoring module, is used to calculate the instruction representation parameters for the voice signals in the time domain segment based on the instruction features to determine the instruction clarity category of the voice signals;

[0072] An instruction control module, which is connected to the analysis module, is used to analyze the voice signals based on the instruction clarity category of the voice signals in the time domain segment to determine control instructions, including,

[0073] Determine the line-of-sight directions of each monitoring target based on the monitoring data, determine the key monitoring targets based on the angle between the line-of-sight direction and the predetermined receiving direction, synchronously determine the lip movement data of the key monitoring targets, match the lip movement data with the voice signals to obtain text signals, so as to determine control instructions;

[0074] Or, directly convert the voice signals into text signals to determine control instructions;

[0075] It can be understood that determining the control instructions based on the voice signals can convert the voice signals into text signals, and determine the control instructions corresponding to the voice signals through the pre-established association relationship between the text signals and the control instructions. For example, taking the text signal corresponding to the voice signal as "turn on the air conditioner" as an example, there is a pre-established association relationship between "turn on the air conditioner" and the control instructions, and then determine the corresponding control instructions to control the air conditioner. Of course, other methods can also be used, and those skilled in the art can choose according to needs, which will not be elaborated here.

[0076] An instruction response module, which is connected to the IoT acquisition module and the instruction control module, is used to control the predetermined devices in the classroom based on the determined control instructions.

[0077] Specifically, the data monitoring module determines the semantic correlation degree corresponding to the voice signals, including,

[0078] Used to convert the voice signals in each time domain segment into text information;

[0079] Used to extract the keywords in the text information;

[0080] Used to determine the semantic correlation degree of the text information according to the keywords.

[0081] Specifically, the conversion method of converting the voice signals into text information is not limited, as long as the voice signals can be converted. For example, it can be;

[0082] Automatic Speech Recognition (ASR) system: Using computer software to recognize and understand speech commands and convert them into text, which is the prior art and can perform real-time conversion, improving the efficiency of determining semantic relevance.

[0083] Speech recognition APIs and SDKs: In the prior art, speech recognition APIs and software development kits (SDKs) can easily integrate speech recognition functions into applications. For example, Google Cloud Speech-to-Text, IBM Watson Speech to Text, Amazon Transcribe, etc.

[0084] Specifically, before extracting keywords, the keywords need to be determined. In the prior art, the methods for determining keywords can be based on word segmentation tools, TF-IDF, and keyword extraction tools, etc. In this implementation, TF-IDF is used to extract keywords, and the steps are as follows.

[0085] Step S1, clean the text information, including removing punctuation marks, stop words (such as common words like "de", "shi", etc.), and may also include stemming or lemmatization.

[0086] Step S2, create a vocabulary to store historical text information, and the vocabulary contains all non-repeating words.

[0087] Step S3, calculate the term frequency of each word in the vocabulary, that is, the number of times the word appears in the historical text information.

[0088] Step S4, for each word in the vocabulary, calculate its inverse document frequency in all documents, and the inverse document frequency is calculated by formula (1).

[0089]

[0090] Where N represents the total number of historical text information, and |{d∈D:t∈d}| represents the number of historical text information containing the word t.

[0091] Step S5, calculate the TF-IDF value of each word in the historical text information, and the TF-IDF value is calculated by formula (2).

[0092] TF-IDF(t,d) = TF(t,d) × IDF(t) (2)

[0093] Where TF(t,d) represents the term frequency of the word t in the historical text information d, and IDF(t) represents the inverse document frequency of the word t.

[0094] Step S6: Create a matrix where the rows represent historical text information, the columns represent the words in the vocabulary, and the values in the matrix are the TF-IDF values of the corresponding words in the corresponding historical text information.

[0095] Step S7: For the text information, select the word with the highest TF-IDF value as the keyword. It can be understood that a keyword determination threshold can be set to select the top N highest words as the keywords of the text information.

[0096] Regarding the method for extracting keywords, those skilled in the art can select according to the actual situation, which will not be elaborated here.

[0097] Specifically, there is no limitation on the calculation method of the semantic relevance degree of the text information, as long as it can achieve the purpose of calculating the semantic relevance degree. For example, the following methods can be used:

[0098] Vector space model method: Convert the keywords into vectors, calculate the similarity between the vectors by methods such as cosine similarity, and determine the average similarity of the vectors corresponding to each word in the text information as the semantic relevance degree.

[0099] Semantic network method: Construct a relationship network between the keywords, measure the relevance between the words through indicators such as path length, and determine the average relevance between the words in the text information as the semantic relevance degree.

[0100] Of course, other forms can also be adopted, which will not be elaborated here. In this embodiment, the cosine similarity method is used to calculate the semantic relevance degree.

[0101] Specifically, the data monitoring module determines the fluctuation value of the speech signal in the time domain dimension, including

[0102] To determine the peak value of the speech signal in each time domain segment;

[0103] To determine that the variance of each peak value in the time domain segment is the fluctuation value.

[0104] Specifically, there is no limitation on the detection method of the speech signal peak value. For example, it can be:

[0105] Sliding window method: That is, divide the continuous speech signal into several segments, and find the maximum amplitude of the signal intensity, that is, the peak value, in each segment of the speech signal. It can be understood that the length of the speech signal after division depends on the length of the sliding window, and several speech signal segments can be divided through the sliding window.

[0106] Peak detection basic formula principle method: The speech signal is divided into several frames of fixed length. For each frame, calculate the absolute value of the signal intensity of all sample points within the frame to obtain a vector. Perform peak detection on this vector to find the maximum value in the vector, which is the peak of the frame. Repeat the above steps to obtain several peaks.

[0107] It can be understood that the above methods can be used alone or in combination to improve the accuracy and robustness of peak detection.

[0108] Specifically, calculating the variance of each peak within the time domain segment is a prior art, and those skilled in the art can choose according to the actual situation, as long as the variance of each peak within the time domain segment can be obtained, which will not be elaborated here.

[0109] Specifically, by analyzing the speech signals within each time domain segment, converting the speech signals into text information and audio information, and then determining the command characteristics of the speech signals. In actual use, when receiving the speech signals of the monitoring target, control the target device to complete operations such as turning on and off according to the speech signals. For a single monitoring target, the fluctuation value of the speech signals it emits in the time domain dimension is small, but in the campus environment, especially during the class dismissal period, the background is noisy, and the speech signals of the monitoring target may be interfered, which is reflected in the increased difficulty of distinguishing the speech signals of the monitoring target subsequently, and then leads to errors in the recognized control commands and inability to effectively control the target device. Based on this, the present invention considers conducting an overall analysis of the collected speech signals to determine the semantic correlation degree corresponding to the speech signals and the fluctuation value of the speech signals in the time domain dimension, providing a data basis for calculating the command representation parameters, making the recognition of the speech signals of the monitoring target more accurate, and improving the accuracy of determining the control commands.

[0110] Specifically, the analysis module calculates the command representation parameters for the speech signals within the time domain segment, including,

[0111] Determining the ratio of the semantic correlation degree to the reference semantic correlation degree as the semantic correlation degree influence parameter;

[0112] Determining the ratio of the reference fluctuation value to the fluctuation value as the fluctuation value influence parameter;

[0113] Determining the weighted sum value of the semantic correlation degree influence parameter and the fluctuation value influence parameter as the command representation parameter.

[0114] Specifically, the reference semantic correlation degree is pre-calculated. First, obtain several semantic correlation degrees of the speech signals emitted by a single monitoring target, calculate the average value of the several semantic correlation degrees, and determine 0.85 times of this average value as the reference semantic correlation degree.

[0115] Specifically, the reference fluctuation value is pre-calculated. A number of voice signals in the classroom environment are acquired in advance, and the fluctuation values are analyzed. The average value of the fluctuation values is calculated as the reference fluctuation value.

[0116] Specifically, the weight coefficient of the semantic association degree influence parameter is 0.64, and the weight coefficient of the fluctuation value influence parameter is 0.36.

[0117] Specifically, the analysis module determines the instruction clarity category of the voice signal, where,

[0118] If the instruction representation parameter is less than the instruction representation parameter threshold, it is determined that the instruction clarity category is a weak clarity instruction;

[0119] If the instruction representation parameter is greater than or equal to the instruction representation parameter threshold, it is determined that the instruction clarity category is a strong clarity instruction.

[0120] Specifically, the instruction representation parameter threshold is selected within the interval [1.05, 1.15].

[0121] Specifically, the instruction control module analyzes the voice signal based on the instruction clarity category of the voice signal in the time domain segment to determine the control instruction, where,

[0122] If the instruction clarity category is a weak clarity instruction, the line-of-sight directions of each monitoring target are determined based on the monitoring data. The key monitoring target is determined based on the angle between the line-of-sight direction and the predetermined reception direction. The lip data of the key monitoring target is synchronously determined. The lip data is matched with the voice signal to obtain the text signal to determine the control instruction;

[0123] If the instruction clarity category is a strong clarity instruction, the voice signal is directly converted into a text signal to determine the control instruction.

[0124] Specifically, the present invention determines the instruction clarity category of the voice signal by calculating the instruction representation parameter for the voice signal in the time domain segment. When analyzing the voice signal, there will be relatively clear voice signals. At this time, if complex analysis is still performed on all received voice signals, this non-adaptive adjustment analysis method results in low analysis efficiency for the voice signal. Based on this, the present invention calculates the instruction representation parameter, classifies the voice signal, directly determines the control instruction for the voice signal with a strong clarity instruction, and performs signal matching on the voice signal with a weak clarity instruction, making the recognition of the voice signal of the monitoring target more accurate and improving the accuracy of determining the control instruction.

[0125] Specifically, the instruction control module determines the line-of-sight directions of each monitoring target, including,

[0126] To determine the facial reference point group of the monitoring target and construct a facial plane according to the facial reference point group;

[0127] The direction of the vector perpendicular to the facial plane is used as the line-of-sight direction of the monitoring target;

[0128] Wherein, the facial reference point group is three points within the facial contour area.

[0129] It can be understood that three points can form a plane. In some possible implementations, the eyes, the tip of the nose, and the corners of the mouth can be determined as the reference points.

[0130] Specifically, the method of identifying the facial reference points is not limited. For example, the face feature extraction of Dlib can be used, including the following steps.

[0131] Step S01, use Dlib to determine the face in the monitoring data. Usually, this method will return the position information of the face, including the left boundary, the upper boundary, the right boundary, and the lower boundary of the face;

[0132] Step S02, use the shape_predictor class of Dlib to predict the key feature points of the face. This usually requires a pre-trained model file (such as shape_predictor68_face_landmarks.dat), and this model can identify 68 key points on the face, including the eyes and the tip of the nose;

[0133] Step S03, label the identified feature points, and the coordinate data of each feature point can be extracted and saved for subsequent analysis and processing, such as operations like constructing a plane.

[0134] Specifically, the vector perpendicular to the facial plane is also the normal vector of the facial plane. It can be understood that the normal vector can be obtained by calculating the cross product of two vectors of the facial plane. In practical applications, the calculation of the normal vector is usually automatically completed by a computer vision library or 3D graphics software. Those skilled in the art only need to call the corresponding functions or methods to obtain the normal vector of the facial plane, which will not be elaborated here.

[0135] Specifically, the instruction control module determines the key monitoring target based on the angle between the line-of-sight direction and the predetermined receiving direction, including

[0136] To determine the angle between the monitoring target corresponding to the line-of-sight direction and the predetermined receiving direction;

[0137] If the degree of the angle is less than the reference angle degree, determine the monitoring target corresponding to the line-of-sight direction as the key monitoring target.

[0138] Specifically, the predetermined receiving direction is determined according to the voice collection unit and the monitoring unit. It can be understood that the voice collection unit can receive voice signals within an area, and the monitoring unit can obtain monitoring data within the area. Usually, the monitoring unit and the voice collection unit are set within 30 cm.

[0139] Due to the positional relationship between the monitoring unit and the voice collection unit, the predetermined receiving direction is determined based on the shooting direction of the monitoring unit, and the shooting direction of the monitoring unit is determined as the predetermined receiving direction.

[0140] Specifically, the purpose of setting the angle is to identify whether the line of sight direction of the monitored target is significantly offset from the predetermined receiving direction. Therefore, the angle is selected within the interval [30°, 90°].

[0141] Specifically, the instruction control module matches the lip data with the voice signal, including:

[0142] Determine a number of associated lip shape keywords based on the lip shape data;

[0143] To match each keyword in the voice signal with the lip shape keyword;

[0144] Converting the speech signal into a text signal, and removing unmatched keywords from the text signal;

[0145] If the keyword in the voice signal does not match any of the lip shape keywords, the keyword is determined to be an unmatched keyword.

[0146] It is understandable that there may be more than one key monitoring target, so in this implementation, lip-keyword matching is performed on the key monitoring target to make the recognition of the monitoring target voice signal more accurate.

[0147] Specifically, there is no limitation on the method of determining the lip-sync keywords. For example, it can be:

[0148] Deep learning technology: Using a deep learning framework, the speaker’s face is identified from the image, and the features of the mouth shape changes during continuous speaking are extracted. These features are input into the lip reading recognition model to identify the corresponding pronunciation, and finally determine a number of mouth shape keywords associated with the pronunciation. It is understandable that for mouth shape, the same mouth shape may be associated with multiple mouth shape keywords, which will not be elaborated here.

[0149] Dataset application: You can use special lip reading recognition datasets, such as LRW, LRW-1000 and Ou l uVS2. These datasets contain a large amount of video and audio data. Using the dataset, you can identify several lip shape keywords.

[0150] Wav2Lip Technology: A deep learning-based audio-visual synchronization technology that achieves high-precision lip synchronization effects by analyzing audio signals and video frames. The input audio is converted into a spectrogram, the best lip position is matched in the video frame, and lip transformation is performed at this position according to the audio signal. Several lip movement keywords are identified based on the lip transformation.

[0151] The above method determines the lip movements from different perspectives and technical means to identify the speaking content. It combines visual information and speech information, improving the accuracy and robustness of recognition. It can be understood that accurate lip movement keywords may not be determined through lip movement data, and those skilled in the art can select from the above methods according to the application environment.

[0152] Specifically, the present invention determines control instructions based on key monitoring targets and lip movement data. For the situation of receiving multiple speech signals in a noisy environment, it is necessary to analyze the speech signals to obtain accurate control instructions. Based on this, the present invention considers real-time monitoring of the target area in the classroom, determines the key monitoring targets according to the monitoring data, matches the lip movement data of the key monitoring targets with the speech signals, and then determines the control instructions, making the recognition of the speech signals of the monitoring targets more accurate and improving the accuracy of determining the control instructions.

[0153] Specifically, the instruction response module controls the predetermined devices in the classroom based on the determined control instructions, including,

[0154] being used to determine control instructions according to text signals;

[0155] being used to determine the predetermined devices in the classroom corresponding to the control instructions;

[0156] being used to send the control instructions to the predetermined devices.

[0157] Specifically, the control instruction is an electrical signal, which changes according to the change of the control target, and only needs to be able to control the device to perform the corresponding action, which will not be elaborated here.

[0158] So far, the technical solution of the present invention has been described in combination with the preferred embodiments shown in the drawings. However, those skilled in the art can easily understand that the protection scope of the present invention is obviously not limited to these specific embodiments. Without departing from the principle of the present invention, those skilled in the art can make equivalent changes or substitutions to the relevant technical features, and the technical solutions after these changes or substitutions will fall within the protection scope of the present invention.

[0159] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention; for those skilled in the art, various modifications and variations can be made to the present invention. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. An integrated operation and maintenance management platform based on the Internet of Things, characterized in that Including: An Internet of Things acquisition module, which includes a voice acquisition unit arranged in the target area of the classroom for acquiring voice signals and a monitoring unit for acquiring monitoring data; A data monitoring module, which is connected to the Internet of Things acquisition module, for analyzing voice signals in each time domain segment, and determining the instruction characteristics of the voice signals based on the analysis results, including the semantic correlation degree corresponding to the voice signals and the fluctuation value of the voice signals in the time domain dimension; An analysis module, which is connected to the data monitoring module, for calculating instruction representation parameters for the voice signals in the time domain segment based on the instruction characteristics to determine the instruction clarity category of the voice signals; An instruction control module, which is connected to the analysis module, for analyzing the voice signals based on the instruction clarity category of the voice signals in the time domain segment to determine control instructions, including, If the instruction clarity category is a weakly clear instruction, determining the line-of-sight directions of each monitoring target based on the monitoring data, determining the key monitoring target based on the included angle between the line-of-sight direction and the predetermined receiving direction, synchronously determining the lip movement data of the key monitoring target, and matching the lip movement data with the voice signal to obtain a text signal to determine the control instruction; If the instruction clarity category is a strongly clear instruction, directly converting the voice signal into a text signal to determine the control instruction; An instruction response module, which is connected to the Internet of Things acquisition module and the instruction control module, for controlling the predetermined devices in the classroom based on the determined control instructions; The data monitoring module determines the semantic correlation degree corresponding to the voice signal, including, For converting voice signals in each time domain segment into text information; For extracting keywords in the text information; For determining the semantic correlation degree of the text information according to the keywords; The data monitoring module determines the fluctuation value of the voice signal in the time domain dimension, including, For determining the peak value of the voice signal in each time domain segment; For determining the variance of each peak value in the time domain segment as the fluctuation value; The analysis module calculates the instruction representation parameters for the voice signals in the time domain segment, including, For determining the ratio of the semantic correlation degree to the reference semantic correlation degree as the semantic correlation degree influence parameter; For determining the ratio of the reference fluctuation value to the fluctuation value as the fluctuation value influence parameter; For determining the weighted sum value of the semantic correlation degree influence parameter and the fluctuation value influence parameter as the instruction representation parameter; Wherein, the reference semantic correlation degree and the reference fluctuation value are pre-calculated.

2. The integrated operation and maintenance management platform based on the Internet of Things according to claim 1, characterized in that The analysis module determines the instruction clarity category of the voice signal, wherein, If the instruction representation parameter is less than the instruction representation parameter threshold, determining that the instruction clarity category is a weakly clear instruction; If the instruction representation parameter is greater than or equal to the instruction representation parameter threshold, determining that the instruction clarity category is a strongly clear instruction.

3. The integrated operation and maintenance management platform based on the Internet of Things according to claim 1, characterized in that, The instruction control module determines the line-of-sight directions of each monitoring target, including, For determining the facial reference point group of the monitoring target and constructing a facial plane according to the facial reference point group; For determining the direction of the vector perpendicular to the facial plane as the line-of-sight direction of the monitoring target; Wherein, the facial reference point group is three points in the facial contour area.

4. The integrated operation and maintenance management platform based on the Internet of Things according to claim 1, wherein The command control module determines the key monitoring target based on the angle between the line of sight and the predetermined receiving direction. include, Used to determine the angle between the line of sight direction corresponding to the monitoring target and the predetermined receiving direction; If the degree of the angle is less than the degree of the reference angle, it is determined that the monitoring target corresponding to the line of sight direction is a key monitoring target.

5. The integrated operation and maintenance management platform based on the Internet of Things according to claim 1, characterized in that The command control module matches the lip shape data with the voice signal, including: Determine a number of associated lip shape keywords based on the lip shape data; To match each keyword in the voice signal with the lip shape keyword; Converting a speech signal into a text signal and removing unmatched keywords from the text signal; If the keyword in the voice signal does not match any of the lip shape keywords, the keyword is determined to be an unmatched keyword.

6. The integrated operation and maintenance management platform based on the Internet of Things according to claim 1, characterized in that The instruction response module controls the predetermined device in the classroom based on the determined control instruction, including: Used to determine control instructions according to text signals; to determine the predetermined device in the classroom corresponding to the control instruction; Used to send the control instruction to the predetermined device.

Citation Information

Patent Citations

  • Comprehensive operation and maintenance management system and method based on Internet of Things

    CN118574148A

  • System operation and maintenance integrated management platform

    CN118863868A

  • Voice recognition method, mobile terminal and computer readable storage medium

    CN107799125A

  • Teaching equipment control method, control equipment, teaching system and storage medium

    CN114779922A

  • First-aid kit capable of intelligently outputting diagnosis strategy

    CN118356301A