A device intelligent management method, system, storage medium and device
By monitoring indoor voice signals in real time and determining whether the equipment has entered sleep mode based on the voice signal ratio, the problem of resource consumption caused by prolonged equipment operation is solved, thus extending the equipment's lifespan.
Patent Information
- Application Number
- CN202210584193.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-05-27
- Publication Date
- 2025-11-21
- Estimated Expiration
- 2042-05-27
AI Technical Summary
Keeping equipment running for extended periods leads to higher resource consumption and a shorter lifespan.
By acquiring the status data of indoor devices, it is determined whether the devices are in working condition. Voice signals in the environment are acquired at a preset frequency, and the voice data ratio in the voice signals is extracted. If the ratio is lower than the threshold and continues for a certain period of time, the device is controlled to enter a sleep state.
It improves the accuracy of human voice data detection, reduces equipment resource consumption, and extends equipment lifespan.
Smart Images

Figure CN115019835B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of device management, in particular to a device intelligent management method and system, a storage medium and a device. BACKGROUND
[0002] With the development of electronic device technology, according to the work requirements, indoor synchronous recording and video recording host, lighting, switches, large screens, display and other electronic devices need to be configured, the use of devices is related to the work activities of personnel, when personnel enter the scene, the devices need to be turned on in time for on-site recording.
[0003] In the prior art, technical personnel need to turn on the devices before personnel enter the scene, and technical personnel need to manually pause or turn off the devices during the intermission or when personnel leave the scene, or many places choose not to turn off or forget to turn off the devices, which makes it difficult to turn off the devices in time. When the devices are in an open state for a long time, the resource cost is large, which leads to the shortening of the service life of the devices. SUMMARY
[0004] Therefore, the purpose of the present application is to provide a device intelligent management method and system, a storage medium and a device to solve the problem of long-time opening of devices in the background art, large resource cost, and shortening of the service life of the devices.
[0005] In one aspect, the present application provides a device intelligent management method, which comprises:
[0006] Obtaining state data of indoor devices, and determining whether the devices are in a working state according to the state data;
[0007] If the devices are in a working state, obtaining voice signals in the environment at a first preset frequency, extracting voice data in the voice signals, calculating the ratio of the voice data to the voice signals, and determining whether the ratio is lower than a preset threshold;
[0008] If the ratio is lower than the preset threshold, it is determined that there is no human voice data in the indoor environment, and the timing is started, the first detection time is recorded, and it is determined whether the first detection time is greater than a first preset time;
[0009] If the first detection time is greater than the first preset time, the devices are controlled to enter a sleep state.
[0010] The device intelligent management method in the application, by acquiring the voice signal in the environment at a first preset frequency, extracting the voice data of the voice signal, when the ratio of the voice data occupying the voice signal is lower than a preset threshold, judging that there is no person data in the room, improving the accuracy of the voice data detection, if there is no voice data, starting timing, judging whether there is voice data within the first preset time, if not, controlling the device to enter the sleep state, avoiding the device consumption, by monitoring the voice signal in the room in real time, timely adjusting the running state of the device according to the judgment result, reducing the resource consumption of the device, solving the problem that the device in the background technology is in the open state for a long time, the resource cost consumption is large, and the service life of the device is shortened.
[0011] Further, the step of extracting the voice data in the voice signal and calculating the ratio of the voice data occupying the voice signal comprises:
[0012] Acquiring the voice signal in the environment, after frame processing of the voice signal, obtaining multiple frames of voice signals, and performing feature extraction on the voice signal;
[0013] Classifying the voice signal after feature extraction to obtain voice data and mute data corresponding to the voice signal, and calculating the ratio of the voice data occupying the voice signal.
[0014] Further, the step of controlling the device to enter the sleep state further comprises:
[0015] Acquiring the voice signal in the environment at a second preset frequency, judging whether there is voice data in the room, if there is no voice data, starting timing, recording the second detection time, and judging whether the second detection time is greater than the second preset time;
[0016] If the second detection time is greater than the second preset time, the device is controlled to enter the closed state.
[0017] Further, the step of controlling the device to enter the sleep state further comprises:
[0018] When the entrance face data is acquired, the detection range of the voice signal in the environment is expanded, and the indoor environment data is detected to judge whether there is voice data in the room.
[0019] Further, the step of acquiring the voice signal in the environment at a first preset frequency further comprises:
[0020] The voice signal is input into a deep neural network model, the voice feature in the voice signal is extracted, the feature description vector corresponding to the voice feature is obtained, and the feature description vector is pre-stored as a standard feature vector.
[0021] Further, the step of controlling the device to enter the sleep state further comprises:
[0022] Acquire a current voice signal in an environment, input the current voice signal into a deep neural network model, extract a current voice feature in the current voice signal, obtain a current feature description vector corresponding to the current voice feature,
[0023] Compare the current feature description vector with a standard feature vector, and determine whether the comparison is successful,
[0024] If yes, control the device to enter a start state.
[0025] Further, if the device is in a working state, the step of acquiring the voice signal in the environment at the first preset frequency further comprises:
[0026] When the off-site face data is acquired, the first preset frequency is increased to a third preset frequency, the voice signal in the environment is acquired at the third preset frequency, and it is determined whether there is voice data indoors.
[0027] Another aspect of the present application provides a device intelligent management system, the system comprising:
[0028] A first detection module is configured to control the device to enter a start state, detect whether there is voice data in the environment at a first preset frequency, start timing if no voice data is detected, and determine whether voice data is detected within a first preset time;
[0029] A sleep control module is configured to control the device to enter a sleep state if no voice data is detected;
[0030] A second detection module is configured to detect whether there is voice data in the environment at a second preset frequency, start timing if no voice data is detected, and determine whether voice data is detected within a second preset time;
[0031] A shutdown control module is configured to control the device to enter a shutdown state if no voice data is detected.
[0032] Another aspect of the present application provides a computer readable storage medium having a computer program stored thereon, the program being executed by a processor to implement the device intelligent management method of any one of the above.
[0033] Another aspect of the present application also provides a data processing device comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, the processor executing the program to implement the device intelligent management method of any one of the above. BRIEF DESCRIPTION OF DRAWINGS
[0034] Figure 1 A flowchart of the device intelligent management method in the first embodiment of the present application;
[0035] Figure 2The flow chart of the device intelligent management method in the second embodiment of the present application is shown in the figure.
[0036] Figure 3 The block diagram of the device intelligent management system in the third embodiment of the present application is shown in the figure.
[0037] The following detailed description will further explain the present application in combination with the above figures. DETAILED DESCRIPTION
[0038] In order to facilitate the understanding of the present application, the present application will be described more fully below with reference to the accompanying drawings. The figures show several embodiments of the present application. However, the present application can be realized in many different forms and is not limited to the embodiments described herein. On the contrary, the purpose of providing these embodiments is to make the disclosure of the present application more thorough and comprehensive.
[0039] It should be noted that when an element is referred to as being "fixedly attached" to another element, it can be directly on the other element or there can be intervening elements. When an element is referred to as being "connected" or "coupled" to another element, it can be directly connected or coupled to the other element or intervening elements can be present. As used herein the terms "vertical", "horizontal", "left", "right", and the like are merely used for the purpose of illustration.
[0040] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used in the description of the application herein is for the purpose of describing particular embodiments only and is not intended to be limiting of the application. As used herein, the term "and / or" includes any and all combinations of one or more of the associated listed items.
[0041] The device intelligent management method in the present embodiment can be applied indoors to intelligently manage a plurality of electronic devices in the indoor environment, wherein the electronic devices include recording, video recording and video devices, and the method can also be applied in other places where electronic device management is required. By judging whether the number of people entering and leaving the place is consistent, it is determined that all the people in the place have left, and thus the devices are automatically controlled to enter a sleep state, reducing the wear and tear of the devices and achieving the effect of energy saving and economy.
[0042] Embodiment One
[0043] Referring to Figure 1 , a device intelligent management method in the first embodiment of the present application is shown, which includes steps S11-S14.
[0044] S11, acquire the state data of the indoor device, and determine whether the device is in a working state according to the state data.
[0045] If the device is in a working state, step S12 is performed.
[0046] First, the personnel need to register personnel information when entering, and pass through the access control verification. The face information of the entering personnel is identified by the face recognition device, and the face information of the staff and the parties is recorded by the system. When the face recognition is consistent with the pre-stored face information in the system, the control device enters the starting state, including video recording, audio recording equipment, etc. The device starts to work after starting, records the information in the field, and feeds back the device state to the background in real time. Optionally, the staff can start the device by themselves.
[0047] The state of the device is acquired in real time, including the starting state, the sleep state, or the shutdown state. It is judged whether the device is in the starting state, and if so, step S12 is executed.
[0048] S12, acquire the speech signal in the environment at a first preset frequency, extract the speech data in the speech signal, calculate the ratio of the speech data occupying the speech signal, and judge whether the ratio is lower than a preset threshold.
[0049] If the ratio is lower than the preset threshold, step S13 is executed.
[0050] The speech signal in the environment is acquired at a first preset frequency, which can be 2 min / time. When detecting human voice in the environment, there is usually environmental noise interference, so the environmental noise needs to be removed to judge whether there is human voice data in the speech signal.
[0051] The speech signal extracted in the preset frequency includes silence data, speech data, and other audio data. By extracting the speech data in the speech signal, the ratio of the speech data in the entire speech signal is calculated. When the ratio is lower than the preset frequency, it can be judged that the human voice data occupies a lower space in the preset frequency, that is, it is judged that there is no human voice data in the frequency.
[0052] S13, it is judged that there is no human voice data in the room, and the timing starts to record the first detection time, and it is judged whether the first detection time is greater than the first preset time.
[0053] If the first detection time is greater than the first preset time, step S14 is executed.
[0054] When it is judged that there is no human voice data in the room, the timing starts to continuously detect whether there is human voice in the field, and the duration of the non-existing human voice data is recorded as the first detection time. It is judged whether the first detection time is greater than the first preset time. When the first detection time is greater than the first preset time, step S14 is executed. The first preset time can be 10 min, that is, whether the time of continuously detecting no human voice in the room reaches 10 min.
[0055] S14, control the device to enter the sleep state.
[0056] If it is judged that there is no human voice data in the first preset time, it indicates that the in-field may enter the halftime break state, and the device is automatically controlled to enter the hibernation state to reduce the power consumption of the device. Specifically, it includes entering the hibernation state of the video recording, audio recording, and camera device.
[0057] In summary, the device intelligent management method in the above embodiments of the present application acquires the voice signal in the environment at a first preset frequency, extracts the voice data of the voice signal, and judges that there is no human data in the room when the ratio of the voice data occupying the voice signal is lower than the preset threshold, thereby improving the accuracy of human voice data detection. If there is no human voice data, the timing is started, and it is judged whether there is human voice data in the first preset time. If not, the device is controlled to enter the hibernation state, thereby avoiding device consumption. By monitoring the voice signal in the room in real time, the running state of the device is adjusted in time according to the judgment result, the resource consumption of the device is reduced, and the problem of long-time opening of the device in the background technology, large resource cost consumption, and short service life of the device is solved.
[0058] Embodiment two
[0059] Please refer to Figure 2 , which is a device intelligent management method in the second embodiment of the present application, including steps S21-S27.
[0060] S21, acquire the state data of the indoor device, and judge whether the device is in the working state according to the state data.
[0061] If the device is in the working state, step S22 is executed.
[0062] Firstly, the personnel entering the field need to register personnel information when entering the field, and pass through the access control verification. The face recognition device recognizes the face information of the entering personnel, and the system records the face information of the staff and the parties. When the face recognition is consistent with the pre-stored face information in the system, the device is controlled to enter the starting state, including video recording, audio recording device, etc. After the device is started, it enters the working state, records the information in the field, and feeds back the device state to the background in real time. Optionally, the staff can start the device by themselves.
[0063] The state of the device is acquired in real time, including the starting state, the hibernation state or the shutdown state. It is judged whether the device is in the starting state, and if so, step S22 is executed.
[0064] S22, acquire the voice signal in the environment at a first preset frequency, extract the voice data in the voice signal, calculate the ratio of the voice data occupying the voice signal, and judge whether the ratio is lower than the preset threshold.
[0065] If the ratio is lower than the preset threshold, step S23 is executed.
[0066] The voice signal in the environment is acquired at a first preset frequency, which can be 2 min / second. When detecting human voice in the environment, there is usually environmental noise interference, so the environmental noise needs to be removed to determine whether there is human voice data in the voice signal.
[0067] Specifically, in the embodiment, the processing of the voice signal includes, first, performing frame processing on the voice signal, and the first step is to pass the voice signal through a high-pass filter with a cutoff frequency of about 200 Hz. The purpose of this step is to remove the direct current bias component and some low-frequency noise in the signal. Although there is still part of the voice signal below 200 Hz, it will not have a great impact on the voice signal.
[0068] Before feature extraction, we first frame the voice signal with a length of 20-40 ms, and the overlap between frames is generally 10 ms. For example, if the sampling rate of our voice signal is 16 kHz, and the window size is 25 ms, in this case, the data points contained in each frame of data are: 0.025*16000 = 400 sampling points. Let the overlap between frames be 10 ms, the starting point of the data of the first frame is sample0, and the starting point of the data of the second frame is sample160.
[0069] Extract features from each frame of data; after frame processing, features can be extracted from each frame of data. Feature extraction is to reduce the influence of irrelevant information in the voice signal, reduce the amount of data to be processed in the subsequent recognition stage, and generate feature parameters representing the speaker information carried in the voice signal. According to the different uses of voice features, different feature parameters need to be extracted, so as to ensure the accuracy of voice recognition. In the following discussion, x(n) is a frame of audio data, where n ranges from 1 to L (L is the length of each frame of data). The following five features are extracted for each frame of data:
[0070] (1) Log frame energy:
[0071]
[0072] (2) Zero-crossing rate: the number of times each frame of data crosses zero.
[0073] (3) Standardized autocorrelation coefficient at a delay position:
[0074]
[0075] (4) P th The first coefficient of the P-order linear prediction.
[0076] (5) Pth pairs of linear prediction errors.
[0077] A classification model is trained on a set of data frames in which a known speech and silence signal region, wherein the speech signal includes voiced and unvoiced sounds. A Bayesian classification model is used, which includes computing the mean and variance of the features corresponding to silence and the mean and variance of the features corresponding to speech. To determine the label of a frame that is not visible, we compute its likelihood of coming from each label, assuming that the data is distributed according to a multivariate Gaussian distribution. The model that gives the higher likelihood is chosen as the frame label.
[0078] The frames are classified as speech or silence using the trained classification model, the classified speech data is extracted, and it is determined whether the ratio of the speech data in the entire speech signal is greater than a preset threshold within a preset frequency. If not, step S23 is performed to determine that there is no human voice data, and the accuracy of human voice detection can be improved by extracting features from the speech signal in the environment.
[0079] S23, it is determined that there is no human voice data in the room, and timing is started to record the first detection time. It is determined whether the first detection time is greater than a first preset time.
[0080] If the first detection time is greater than the first preset time, step S24 is performed.
[0081] When it is determined that there is no human voice data in the room, timing is started, and it is continuously detected whether there is human voice in the field. The duration of the absence of human voice data is recorded as the first detection time. When the first detection time is greater than a first preset time, step S24 is performed, wherein the first preset time can be 10 minutes.
[0082] Optionally, the access control is provided with a face recognition device at the entrance / exit. When the face recognition device at the exit of the access control obtains the exit face data, the human voice data in the environment is detected at a third preset frequency. The third preset frequency is greater than the first preset frequency. When it is detected that someone exits, the detection frequency is accelerated to improve the efficiency of detecting human voice data and timely adjust the state of the device.
[0083] S24, the device is controlled to enter a sleep state, and the indoor environment data is continuously detected at a second preset frequency to determine whether there is human voice data in the room.
[0084] If there is human voice data, step S25 is performed.
[0085] If there is no human voice data, step S26 is performed.
[0086] If it is determined that there is no human voice data in the room within the first preset time, it indicates that the field may enter the halftime break state, and the device is automatically controlled to enter the hibernation state to reduce the power consumption of the device. Specifically, it includes entering the hibernation state of the video recording, audio recording, and camera device.
[0087] After the device enters the hibernation state, the human voice data in the environment is detected at a second preset frequency to detect whether the human voice reappears in the field. The second preset frequency can be greater than the first preset frequency, for example, 1 min / time.
[0088] Optionally, in some other optional embodiments, when the face recognition device at the entrance of the access control reacquires the face data of the person entering the field, it indicates that someone has reentered the field, and the voice signal in the environment is detected at a fourth preset frequency, wherein the fourth preset frequency is greater than the second preset frequency. By speeding up the voice detection frequency, the human voice data in the environment is quickly detected, and if human voice data is detected at this time, the device can be quickly started to avoid affecting the indoor working state.
[0089] Optionally, in some other embodiments, when the face recognition device at the entrance of the access control reacquires the face data of the person entering the field, the detection range of the voice signal in the environment is expanded, for example, the detection range of the voice signal in the environment is expanded from the center of the room to the entrance of the access control, so as to timely detect the voice signal emitted at the current entrance of the access control and timely judge whether there is human voice data. If human voice data is detected in the room at this time, step S25 is executed to quickly start the device.
[0090] Optionally, in some other embodiments, before the device enters the hibernation state, the voice signal in the environment is acquired, the extracted voice signal is input into the deep neural network model, the human voice features in the voice signal are extracted, a group of feature description vectors corresponding to the human voice features are obtained, different human voices have different feature description vectors, and the extracted different feature description vectors are prestored as standard feature vectors.
[0091] After the device hibernates, the current voice signal in the environment is continuously extracted, and the current voice signal is input into the deep neural network model for processing to extract the current human voice features in the current voice to obtain a group of current feature description vectors corresponding to the current human voice features, which are the human voice feature description vectors in the hibernation state.
[0092] The current feature description vectors in the hibernation state are compared with the prestored standard feature vectors, for example, the Euclidean distance is calculated, when the distance is greater than a preset value, it is considered that the sound is from different people, and when the distance is less than the preset value, it is considered that the sound is from the same person. After recognizing that the sound is from the same person, the device is immediately started, and there is no need to detect whether there is human voice in the room at a preset frequency, thereby improving the efficiency of starting the device.
[0093] S25, control the device to enter the start state.
[0094] Immediately control the device to enter the start state.
[0095] S26, start timing, record the second detection time, and judge whether the second detection time is greater than the second preset time.
[0096] If the second detection time is greater than the second preset time, step S27 is executed.
[0097] S27, control the device to enter the off state.
[0098] If no voice data is detected within the second preset frequency, timing is restarted, the time when no voice data is detected in the environment is recorded as the second detection time, and it is judged whether the second detection time is greater than the second preset time. Optionally, the second preset time is 15 min, that is, after 15 min of continuous detection of no voice data, the device is controlled to enter the off state.
[0099] If voice data is detected within the second preset time, the device is automatically controlled to enter the start state and re-enter the working state.
[0100] In summary, the device intelligent management method in the above embodiments of the present application acquires voice signals in the environment at a first preset frequency, extracts voice data of the voice signals, and judges that no voice data exists indoors when the ratio of the voice data in the voice signals is lower than a preset threshold. The accuracy of voice data detection is improved. If no voice data exists, timing is started, and it is judged whether voice data exists within the first preset time. If not, the device is controlled to enter the sleep state, avoiding device consumption. By monitoring the voice signals in the room in real time, the running state of the device is adjusted in a timely manner according to the judgment result, reducing resource consumption of the device, and solving the problem of long-term opening of the device in the background technology, large resource cost consumption, and shortening of the service life of the device.
[0101] Embodiment three
[0102] The present application also provides a device intelligent management system, please refer to Figure 3 , the device intelligent management method system block diagram in the embodiment is shown, the system includes:
[0103] A device state monitoring module is used to acquire state data of indoor devices and judge whether the device is in a working state according to the state data.
[0104] The first judging module is configured to, if the device is in a working state, acquire a voice signal in an environment at a first preset frequency, extract voice data in the voice signal, calculate a ratio of the voice data in the voice signal, and judge whether the ratio is lower than a preset threshold.
[0105] The second judging module is configured to, if the ratio is lower than the preset threshold, determine that there is no human voice data in the indoor environment, start timing, record a first detection time, and judge whether the first detection time is greater than the first preset time.
[0106] The hibernation control module is configured to, if the first detection time is greater than the first preset time, control the device to enter a hibernation state.
[0107] Further, in some other optional embodiments, the first judging module comprises:
[0108] The voice detection unit is configured to acquire a voice signal in an environment, perform frame processing on the voice signal to obtain a plurality of frames of voice signals, and perform feature extraction on the voice signal.
[0109] The voice signal after feature extraction is classified to obtain voice data and mute data corresponding to the voice signal, and a ratio of the voice data in the voice signal is calculated.
[0110] Further, in some other optional embodiments, the system further comprises:
[0111] The third judging module is configured to acquire a voice signal in an environment at a second preset frequency, judge whether there is human voice data in the indoor environment, if there is no human voice data, start timing, record a second detection time, and judge whether the second detection time is greater than the second preset time.
[0112] If the second detection time is greater than the second preset time, the device is controlled to enter an off state.
[0113] Further, in some other optional embodiments, the system further comprises:
[0114] The voice detection expansion module is configured to, when the entry face data is acquired, expand the detection range of the voice signal in the environment, and detect indoor environment data to judge whether there is human voice data in the indoor environment.
[0115] Further, in some other optional embodiments, the system further comprises:
[0116] The human voice feature pre-storage module is configured to
[0117] The voice signal is input into a deep neural network model, human voice features in the voice signal are extracted, a feature description vector corresponding to the human voice features is obtained, and the feature description vector is pre-stored as a standard feature vector.
[0118] Further, in some other optional embodiments, the system further comprises:
[0119] A human voice feature comparison module is configured to
[0120] A current voice signal in the environment is obtained, the current voice signal is input into a deep neural network model, current human voice features in the current voice signal are extracted, a current feature description vector corresponding to the current human voice features is obtained,
[0121] The current feature description vector is compared with the standard feature vector, and it is determined whether the comparison is successful,
[0122] If yes, the control device enters an activation state.
[0123] Further, in some other optional embodiments, the system further comprises:
[0124] A fourth determination module is configured to, when the off-site face data is obtained, increase the first preset frequency to a third preset frequency, obtain a voice signal in the environment at the third preset frequency, and determine whether there is human voice data in the room.
[0125] The functions or operation steps realized when the above modules and units are executed are generally the same as those of the above method embodiments, and will not be described here.
[0126] In summary, the device intelligent management system in the above embodiments increases the accuracy of human voice data detection by obtaining a voice signal in the environment at a first preset frequency, extracting voice data of the voice signal, and determining that there is no human data in the room when the ratio of the voice data in the voice signal is lower than a preset threshold. If there is no human voice data, timing is started, and it is determined whether there is human voice data within a first preset time. If not, the device enters a sleep state, thereby avoiding device consumption. By monitoring the voice signal in the room in real time and adjusting the operation state of the device according to the determination result in a timely manner, the resource consumption of the device is reduced, and the problem of a shortened service life of the device due to a large resource cost consumption and a long time of the device in an open state in the background technology is solved.
[0127] The embodiment of the application further provides a computer readable storage medium having a computer program stored thereon, and the program is executed by a processor to implement the steps of the device intelligent management method in the above embodiments.
[0128] Embodiment four
[0129] In another aspect, the present application provides a device, which comprises a memory, a processor, and a computer program stored in the memory and capable of running on the processor, and the processor implements the intelligent management method of the device in the above embodiments when executing the program. In some embodiments, the processor can be an Electronic Control Unit (ECU), a Central Processing Unit (CPU), a controller, a microcontroller, a microprocessor, or other data processing chip, which is used to run the program code or process data stored in the memory, such as executing the access restriction program.
[0130] The memory comprises at least one type of readable storage medium, such as a flash memory, a hard disk, a multimedia card, a card-type memory (e.g., an SD or DX memory, etc.), a magnetic memory, a magnetic disk, an optical disk, etc. In some embodiments, the memory can be an internal storage unit of the device, such as a hard disk of the device. In other embodiments, the memory can also be an external storage device of the device, such as a plug-in hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc. equipped on the device. Further, the memory can comprise both the internal storage unit and the external storage device of the device. The memory can be used not only to store application software and various data installed on the device, but also to temporarily store data that has been output or will be output.
[0131] The logic and / or steps represented in the flowcharts or otherwise described herein, for example, can be considered as a list of executable instructions for implementing logic functions, which can be specifically embodied in any computer-readable medium for use by or in connection with an instruction execution system, apparatus or device, such as a computer-based system, a system including a processor or other system that can fetch the instructions from the instruction execution system, apparatus or device and execute the instructions. For the purpose of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate or transport the program for use by or in connection with the instruction execution system, apparatus or device or in conjunction with these instruction execution systems, apparatus or devices.
[0132] More specific examples (a non-exhaustive list) of the computer-readable medium include the following: an electrical connection (electronic) having one or more wires, a portable computer diskette (magnetic), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, and a portable compact disc read-only memory (CDROM). Additionally, the computer-readable medium can also be paper or another suitable medium upon which the program is printed, as the program can be electronically captured, for example, via an optical scanner, then compiled, interpreted, or otherwise processed, and stored in a computer memory in a form that is then suitable for use by the computer. The computer-readable medium can also be a memory of a mobile device into which the program is downloaded. The memory of the mobile device can load the program into the mobile device such that the program becomes executable by the mobile device.
[0133] It should be understood that aspects of the application can be implemented in hardware, software, firmware or combinations thereof. In the embodiments described above, various steps or methods can be implemented, for example, by software or firmware stored in memory and executed by a suitable instruction execution system. For example, if implemented in hardware, and in another embodiment, any of the following techniques, which are well known in the art, can be used to implement the application: a hybrid of the techniques mentioned above; a combination of one or more of the techniques mentioned above; or one or more other techniques suitable for use in the art.
[0134] In the description of the specification, the description of the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples" and the like means that the specific features, structures, materials or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the application. In the specification, the illustrative description of the above terms does not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in any appropriate manner in one or more embodiments or examples.
[0135] The above-described embodiments only express several implementation manners of the present application, which are described in a more specific and detailed manner, but cannot be understood as a limitation on the patent scope of the present application. It should be noted that for those skilled in the art, several modifications and improvements can be made without departing from the concept of the present application, which are all within the protection scope of the present application. Therefore, the patent protection scope of the present application should be subject to the appended claims.
Claims
1. A method for intelligent equipment management, characterized in that, The method includes: The system acquires the status data of indoor equipment, determines whether the equipment is in working condition based on the status data, acquires the information of the personnel entering the venue, identifies the facial information of the personnel entering the venue through the facial recognition device, and records the facial information of the staff and the person involved. When the facial information is recognized by the facial recognition device and matches the facial information stored in the system, the system controls the equipment to enter the start state. After the equipment is started, it begins to work, records the information in the venue, and feeds back the equipment status to the background in real time. If the device is in operation, it acquires the speech signal in the environment at a first preset frequency, performs frame processing on the speech signal, passes the speech signal through a high-pass filter to remove the DC bias component and some low-frequency noise in the signal, extracts features from each frame of data, extracts different feature parameters according to the different uses of the speech features, uses the trained classification model to classify the frame as speech or silence, extracts the classified speech data, calculates the ratio of the speech data to the speech signal within a preset frequency, and determines whether the ratio is lower than a preset threshold. If the ratio is lower than the preset threshold, it is determined that there is no human voice data in the room, and timing begins to record the first detection time. It is then determined whether the first detection time is greater than the first preset time. If the first detection time is greater than the first preset time, the control device enters a sleep state; The step of acquiring the voice signal in the environment at a first preset frequency further includes: The speech signal is input into a deep neural network model to extract human voice features from the speech signal, and a feature description vector corresponding to the human voice features is obtained. The feature description vector is then pre-stored as a standard feature vector. After the step of the control device entering sleep mode, the following is also included: The current speech signal in the environment is acquired, and the current speech signal is input into a deep neural network model to extract the current human voice features from the current speech signal, thereby obtaining a current feature description vector corresponding to the current human voice features. The current feature description vector is compared with the standard feature vector to determine whether the comparison was successful. If so, the control device will enter the start state.
2. The intelligent equipment management method according to claim 1, characterized in that, The step of extracting speech data from the speech signal and calculating the ratio of the speech data to the speech signal includes: Acquire speech signals from the environment, perform frame segmentation on the speech signals to obtain multi-frame speech signals, and extract features from the speech signals. The extracted speech signals are classified to obtain speech data and silence data corresponding to the speech signals, and the ratio of the speech data to the speech signals is calculated.
3. The intelligent equipment management method according to claim 1, characterized in that, After the step of the control device entering sleep mode, the following is also included: Acquire voice signals in the environment at a second preset frequency, determine whether there is human voice data indoors, if there is no human voice data, start timing, record the second detection time, and determine whether the second detection time is greater than the second preset time; If the second detection time is greater than the second preset time, the control device will enter the off state.
4. The intelligent equipment management method according to claim 3, characterized in that, After the step of the control device entering sleep mode, the following is also included: Once the facial data of the person entering the venue is acquired, the detection range of voice signals in the environment is expanded, and indoor environmental data is detected to determine whether there is human voice data indoors.
5. The intelligent equipment management method according to claim 4, characterized in that, If the device is in a working state, the step of acquiring the voice signal in the environment at a first preset frequency further includes: When facial data of a departing person is acquired, the first preset frequency is increased to the third preset frequency, and the voice signal in the environment is acquired at the third preset frequency to determine whether there is human voice data in the room.
6. An intelligent equipment management system, characterized in that, The system includes: The equipment status monitoring module is used to acquire the status data of indoor equipment, determine whether the equipment is in working condition based on the status data, acquire the information of personnel entering the venue, identify the facial information of personnel entering the venue through facial recognition equipment, and record the facial information of staff and the person involved. When the facial recognition information matches the facial information stored in the system, the system controls the equipment to enter the start state. After the equipment is started, it begins to work, records the information in the venue, and feeds back the equipment status to the background in real time. The first judgment module is used to acquire the voice signal in the environment at a first preset frequency if the device is in working state, perform frame processing on the voice signal, pass the voice signal through a high-pass filter to remove the DC bias component and some low-frequency noise in the signal, extract features from each frame of data, extract different feature parameters according to the different uses of the voice features, classify the frame into voice or silence using a trained classification model, extract the classified voice data, calculate the ratio of the voice data to the voice signal within a preset frequency, and determine whether the ratio is lower than a preset threshold. The second judgment module is used to determine that there is no human voice data in the room if the ratio is lower than the preset threshold, and to start timing, record the first detection time, and determine whether the first detection time is greater than the first preset time. A sleep control module is used to control the device to enter a sleep state if the first detection time is greater than the first preset time. The step of acquiring the voice signal in the environment at a first preset frequency further includes: The speech signal is input into a deep neural network model to extract human voice features from the speech signal, and a feature description vector corresponding to the human voice features is obtained. The feature description vector is then pre-stored as a standard feature vector. After the step of the control device entering sleep mode, the following is also included: The current speech signal in the environment is acquired, and the current speech signal is input into a deep neural network model to extract the current human voice features from the current speech signal, thereby obtaining a current feature description vector corresponding to the current human voice features. The current feature description vector is compared with the standard feature vector to determine whether the comparison was successful. If so, the control device will enter the start state.
7. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the intelligent device management method as described in any one of claims 1-5.
8. A data processing apparatus, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the intelligent device management method as described in any one of claims 1-5.
Citation Information
Patent Citations
Classroom intelligent control energy-saving method, device and equipment and readable storage medium
CN113777949A