Low-power voice recognition method and device for photovoltaic camera
By adjusting the working state of the recognition unit in the photovoltaic camera based on real-time power supply and historical data, the problem of inaccurate energy consumption control of photovoltaic cameras in different environments is solved, and the battery life and recognition accuracy are improved.
Patent Information
- Application Number
- CN202411267897.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-11
- Publication Date
- 2025-11-21
- Estimated Expiration
- 2044-09-11
AI Technical Summary
现有技术无法在不同使用环境下对光伏摄像头的能耗进行精确控制,影响其续航时间。
The first recognition unit determines the first recognition information of the target voice. The recognition units with different power consumptions are woken up according to the real-time photovoltaic power supply of the photovoltaic camera to perform voice recognition. The recognition units include a second recognition unit and a third recognition unit. The power consumption of the second recognition unit is greater than that of the third recognition unit. The working state of the recognition unit is adjusted according to the real-time power supply and historical power supply data.
It achieves precise energy consumption control under different photovoltaic power supply conditions, improves the battery life of photovoltaic cameras and the accuracy of voice recognition, and avoids unstable working results caused by frequent switching.
Smart Images

Figure CN119495295B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of voice recognition control technology, and more specifically, to a low-power voice recognition method and device for a photovoltaic camera. Background Technology
[0002] A photovoltaic (PV) camera is a camera device that integrates photovoltaic (PV) power generation technology. This technology provides a continuous power supply to the camera, constantly converting sunlight into electrical energy. PV cameras do not require an external power source and can be deployed in environments lacking a stable power supply, such as outdoor and remote monitoring. PV cameras have broad application prospects and are increasingly used in modern surveillance systems due to their excellent environmental friendliness and low maintenance costs. For solar-powered cameras, it is necessary to reduce the power consumption of video and audio processing to improve the camera's battery life and overall performance.
[0003] Patent CN111755002B (application number: CN202010567523.X) provides a speech recognition device. After acquiring and storing the original speech signal collected by the microphone and the reference speech signal played by the speaker in the audio memory, the first core, in a low-power state, determines whether to trigger a first-level wake-up state based on the original speech signal and the reference speech signal. In the first-level wake-up state, a wake-up model is run. When the wake-up model recognizes wake-up based on the original speech signal, a second-level wake-up state is triggered, activating the second core. In the second-level wake-up state, the second core runs a speech recognition model to perform speech recognition on the original speech signal. The speech recognition device in patent CN111755002B, after acquiring the original speech signal and the reference speech signal, achieves the goal of reducing energy consumption while maintaining high-performance speech recognition by activating different components step by step to perform speech recognition on the original speech signal. However, the method in patent CN111755002B cannot precisely control energy consumption under different usage environments. Summary of the Invention
[0004] The purpose of this application is to provide a low-power voice recognition method and device for photovoltaic cameras, which solves the technical problem of not being able to accurately control energy consumption in different usage environments, and achieves the technical effect of accurately controlling energy consumption based on the usage environment of photovoltaic cameras.
[0005] This application provides a low-power speech recognition method for a photovoltaic camera. The method includes: determining first recognition information of target speech through a first recognition unit; when the first recognition information meets a preset first recognition condition, acquiring the real-time photovoltaic power supply of the photovoltaic camera; when the real-time photovoltaic power supply is greater than a preset real-time photovoltaic power supply, waking up a second recognition unit to recognize subsequent speech of the target speech; when the real-time photovoltaic power supply is less than or equal to the preset real-time photovoltaic power supply, waking up a third recognition unit to recognize subsequent speech of the target speech; wherein the power consumption of the second recognition unit is greater than the power consumption of the third recognition unit.
[0006] In one possible implementation, the method further includes: when the real-time photovoltaic power supply is less than or equal to a preset real-time photovoltaic power supply, acquiring the historical photovoltaic power supply of the photovoltaic camera at multiple times within a preset first time period from the current time; when the historical photovoltaic power supply at multiple times within the preset first time period meets the preset photovoltaic power supply condition, recognizing the subsequent speech of the target speech through a second recognition unit, and when the second recognition unit recognizes the subsequent speech of the target speech, a third recognition unit is turned off; wherein, the preset photovoltaic power supply condition includes a preset number of historical photovoltaic power supplies within the preset first time period that are greater than a preset first photovoltaic power supply threshold; or the average power of the historical photovoltaic power supply at multiple times within the preset first time period is greater than a preset second photovoltaic power supply threshold; or the historical photovoltaic power supply at multiple times within the preset first time period gradually increases in time, and the maximum historical photovoltaic power supply among the historical photovoltaic power supplies at multiple times within the preset first time period is greater than a preset third photovoltaic power supply threshold.
[0007] In another possible implementation, when the historical photovoltaic power supply power at multiple moments within a preset first time period meets a preset photovoltaic power supply power condition, the subsequent speech of the target speech is recognized by the second recognition unit, including: acquiring the historical photovoltaic power supply power of the photovoltaic camera at multiple moments within a preset second time period from the current time; wherein the preset second time period is less than the preset first time period; when the historical photovoltaic power supply power at multiple moments within the preset second time period is greater than or equal to a preset fourth photovoltaic power supply threshold, the subsequent speech of the target speech is recognized by the second recognition unit; when the historical photovoltaic power supply power at multiple moments within the preset second time period is less than the preset fourth photovoltaic power supply threshold, the subsequent speech of the target speech is recognized by the third recognition unit, and the second recognition unit is turned off when the third recognition unit recognizes the subsequent speech of the target speech.
[0008] In another possible implementation, the method further includes: when the real-time photovoltaic power supply is less than or equal to the preset real-time photovoltaic power supply, obtaining the remaining power value of the photovoltaic camera; when the remaining power value is greater than the preset power value, waking up the second recognition unit to recognize the subsequent speech of the target speech; and when the second recognition unit recognizes the subsequent speech of the target speech, the third recognition unit is turned off.
[0009] In another possible implementation, when the remaining battery power is greater than a preset battery power value, the second recognition unit is activated to recognize the subsequent speech of the target speech, including: obtaining the maximum volume value of the subsequent speech of the target speech at multiple moments within a preset third time period starting from the current moment; when the maximum volume value at multiple moments within the preset third time period is greater than a preset volume threshold, the third recognition unit is activated to recognize the subsequent speech of the target speech, and the second recognition unit is turned off when the third recognition unit recognizes the subsequent speech of the target speech; when the maximum volume value at multiple moments within the preset third time period is less than or equal to the preset volume threshold, or when the maximum volume value at multiple moments within the preset third time period gradually decreases, the second recognition unit is activated to recognize the subsequent speech of the target speech, and the third recognition unit is turned off when the second recognition unit recognizes the subsequent speech of the target speech.
[0010] In another possible implementation, the method further includes: when the maximum volume value at multiple moments within a preset third time period is greater than a preset volume threshold, obtaining the noise value in the subsequent speech when the third recognition unit recognizes the subsequent speech of the target speech; when the noise value is greater than a preset noise value, storing the subsequent speech of the target speech, and when the real-time photovoltaic power supply is greater than a preset real-time photovoltaic power supply, waking up the second recognition unit to recognize the stored subsequent speech of the target speech.
[0011] This application also provides a low-power voice recognition device for a photovoltaic camera, including a unit for performing the method as described in any of the preceding claims.
[0012] This application also provides a low-power voice recognition device for a photovoltaic camera, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the method described in any of the preceding claims.
[0013] This application also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the method described in any of the preceding claims.
[0014] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the method described in any of the preceding claims.
[0015] This application also provides a photovoltaic camera, which, when in operation, implements the steps described in any of the preceding methods.
[0016] The beneficial effects of the embodiments in this application compared with the prior art are:
[0017] This application provides a low-power speech recognition method for a photovoltaic camera. The method includes: determining first recognition information of target speech through a first recognition unit; when the first recognition information meets preset conditions, acquiring the real-time photovoltaic power supply of the photovoltaic camera; when the real-time photovoltaic power supply is greater than a preset real-time photovoltaic power supply, waking up a second recognition unit to recognize subsequent speech of the target speech; when the real-time photovoltaic power supply is less than or equal to the preset real-time photovoltaic power supply, waking up a third recognition unit to recognize subsequent speech of the target speech; wherein the power consumption of the second recognition unit is greater than the power consumption of the third recognition unit. The low-power speech recognition method for a photovoltaic camera in this application embodiment can perform speech recognition through recognition units with different energy consumption under different photovoltaic power supply conditions, achieving precise control of energy consumption and improving the battery life of the photovoltaic camera. Attached Figure Description
[0018] To more clearly illustrate the technical solutions in the embodiments of this application, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0019] Figure 1 A flowchart illustrating a low-power voice recognition method for a photovoltaic camera provided in an embodiment of this application;
[0020] Figure 2 This is a schematic diagram of the structure of a photovoltaic camera used in a low-power voice recognition method for a photovoltaic camera according to an embodiment of this application.
[0021] Figure 3 This is a schematic diagram illustrating a time range of a preset first duration and a preset second duration in an embodiment of this application.
[0022] Figure 4 A schematic diagram of the logic structure of a low-power voice recognition device for a photovoltaic camera provided in an embodiment of this application;
[0023] Figure 5 This is a schematic diagram of the physical structure of a low-power voice recognition device for a photovoltaic camera, provided in an embodiment of this application. Detailed Implementation
[0024] It should be understood that, when used in this application specification and the appended claims, the term "comprising" indicates the presence of the described features, integrals, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components and / or a collection thereof.
[0025] It should also be understood that the term “and / or” as used in this application specification and the appended claims means any combination of one or more of the associated listed items and all possible combinations, and includes such combinations.
[0026] As used in this application specification and the appended claims, the term "if" may be interpreted, depending on the context, as "when," "once," "in response to determination," or "in response to detection." Similarly, the phrase "if determined" or "if detected [the described condition or event]" may be interpreted, depending on the context, as meaning "once determined," "in response to determination," "once detected [the described condition or event]," or "in response to detection [the described condition or event]."
[0027] Furthermore, in the description of this application and the appended claims, the terms "first," "second," "third," etc., are used only to distinguish descriptions and should not be construed as indicating or implying relative importance.
[0028] References to "one embodiment" or "some embodiments" as described in this specification mean that one or more embodiments of this application include a specific feature, structure, or characteristic described in connection with that embodiment. Therefore, the phrases "in one embodiment," "in some embodiments," "in other embodiments," "in still other embodiments," etc., appearing in different parts of this specification do not necessarily refer to the same embodiment, but rather mean "one or more, but not all, embodiments," unless otherwise specifically emphasized. The terms "comprising," "including," "having," and variations thereof mean "including but not limited to," unless otherwise specifically emphasized.
[0029] Existing speech recognition devices, after acquiring the original speech signal and the reference speech signal, activate different components step by step to perform speech recognition on the original speech signal, achieving the goal of reducing energy consumption while maintaining high-performance speech recognition. However, they cannot precisely control energy consumption in different usage environments.
[0030] Based on the above reasons, this application provides a low-power voice recognition method for a photovoltaic camera. The method includes: determining first recognition information of target speech through a first recognition unit; when the first recognition information meets preset conditions, acquiring the real-time photovoltaic power supply of the photovoltaic camera; when the real-time photovoltaic power supply is greater than a preset real-time photovoltaic power supply, waking up a second recognition unit to recognize subsequent speech of the target speech; when the real-time photovoltaic power supply is less than or equal to the preset real-time photovoltaic power supply, waking up a third recognition unit to recognize subsequent speech of the target speech; wherein the power consumption of the second recognition unit is greater than the power consumption of the third recognition unit. The low-power voice recognition method for a photovoltaic camera in this application can perform voice recognition through recognition units with different energy consumption under different photovoltaic power supply conditions, achieving precise control of energy consumption and improving the battery life of the photovoltaic camera.
[0031] In some scenarios, the low-power voice recognition method for photovoltaic cameras according to the embodiments of this application can be applied to the voice recognition control of photovoltaic cameras, which can improve the battery life of photovoltaic cameras.
[0032] The following describes in detail, with specific examples, a low-power voice recognition method for a photovoltaic camera provided in this application embodiment.
[0033] Figure 1 A flowchart illustrating a low-power voice recognition method for a photovoltaic camera provided in this application embodiment is shown below. Figure 1 As shown, this method includes S110 to S120, and S110 to S120 will be described in detail below.
[0034] S110. First recognition information of the target voice is determined by the first recognition unit. When the first recognition information meets the preset first recognition conditions, the real-time photovoltaic power supply of the photovoltaic camera is obtained.
[0035] Figure 2 This is a schematic diagram of the photovoltaic camera used in a low-power voice recognition method for a photovoltaic camera according to an embodiment of this application, as shown below. Figure 2 As shown, the photovoltaic camera in this embodiment includes a photovoltaic solar panel 1 and a camera 2. The camera 2 is equipped with a microphone, which can collect sound information and perform voice recognition to realize intelligent security monitoring functions.
[0036] When the low-power voice recognition method of the photovoltaic camera in this application embodiment is working, it can first collect the target voice through the microphone, and then perform preliminary recognition of the target voice through the first recognition unit to determine the first recognition information of the target voice. Based on the first recognition information, it can be determined that further voice recognition is needed.
[0037] When the first recognition information meets the preset first recognition conditions, it means that the target voice needs to be further recognized. At this time, the real-time photovoltaic power supply of the photovoltaic camera can be obtained, and the voice recognition method can be determined according to the real-time photovoltaic power supply.
[0038] For example, the preset first recognition condition can be to determine that the target speech is a human voice, contains keywords, or other judgment conditions. The preset first recognition condition is a judgment condition that determines that the target speech needs to be further recognized.
[0039] S120. When the real-time photovoltaic power supply is greater than the preset real-time photovoltaic power supply, the second recognition unit is activated to recognize the subsequent speech of the target speech. When the real-time photovoltaic power supply is less than or equal to the preset real-time photovoltaic power supply, the third recognition unit is activated to recognize the subsequent speech of the target speech. The power consumption of the second recognition unit is greater than that of the third recognition unit.
[0040] After obtaining the real-time photovoltaic power supply, if the real-time photovoltaic power supply is greater than the preset real-time photovoltaic power supply, it means that the photovoltaic power generation of the photovoltaic camera is large and there is sufficient power to supply the photovoltaic camera. At this time, the second recognition unit can be activated to recognize the subsequent speech of the target speech. The second recognition unit has a larger model scale and higher recognition accuracy, making it suitable for use in high-precision recognition scenarios.
[0041] After obtaining the real-time photovoltaic power supply, if the real-time photovoltaic power supply is less than or equal to the preset real-time photovoltaic power supply, it indicates that the photovoltaic power generation of the photovoltaic camera is small and there is not enough real-time photovoltaic power to supply the photovoltaic camera. The third recognition unit can be activated to recognize the subsequent speech of the target speech. The model of the second recognition unit is smaller and the recognition accuracy is generally lower, making it suitable for use in low-energy consumption scenarios.
[0042] It should be noted that the power consumption of the second recognition unit is greater than that of the third recognition unit, so that the target voice can be recognized by the third recognition unit when the power of the photovoltaic camera is low, thus significantly reducing the energy consumption of the voice recognition function of the photovoltaic camera.
[0043] It should be noted that the second identification unit and the third identification unit can be located in different chips to fully isolate the second identification unit and the third identification unit and control power consumption.
[0044] For example, the difference between the second recognition unit and the third recognition unit is that the second recognition unit can extract more keyword information, emotional information, etc., and can play a more powerful role in security monitoring compared to the third recognition unit.
[0045] The beneficial effect of the above implementation method is that it enables voice recognition through different energy consumption recognition units under different photovoltaic power supply conditions, thereby achieving precise control of energy consumption and improving the battery life of photovoltaic cameras.
[0046] In some implementations, the above method also includes S210 to S220, which will be described in detail below.
[0047] S210. When the real-time photovoltaic power supply is less than or equal to the preset real-time photovoltaic power supply, acquire the historical photovoltaic power supply of the photovoltaic camera at multiple times within a preset first time period from the current time.
[0048] During the use of a photovoltaic camera, when the real-time photovoltaic power supply is less than or equal to the preset real-time photovoltaic power supply, it indicates that the real-time photovoltaic power generation of the photovoltaic camera is low. However, the photovoltaic camera may temporarily reduce its power generation due to cloud cover. In order to avoid the frequent switching of the voice recognition function of the photovoltaic camera and affect the working effect of the camera, the historical photovoltaic power supply of the photovoltaic camera at multiple moments within a preset first time period can be obtained, and the working status of the photovoltaic camera can be controlled according to the historical photovoltaic power supply of multiple moments within the preset first time period.
[0049] For example, the preset first duration can be 1 hour, 1.5 hours, or 2.5 hours.
[0050] For example, when obtaining the historical photovoltaic power supply at multiple moments within a preset first time period, the historical photovoltaic power supply at multiple moments can be obtained at equal time intervals within the preset first time period.
[0051] S220. When the historical photovoltaic power supply power at multiple moments within a preset first time period meets the preset photovoltaic power supply power condition, the second recognition unit recognizes the subsequent speech of the target speech. The third recognition unit is turned off when the second recognition unit recognizes the subsequent speech of the target speech. The preset photovoltaic power supply power condition includes: a preset number of historical photovoltaic power supplies within the preset first time period exceeding a preset first photovoltaic power supply threshold; or the average power of the historical photovoltaic power supply power at multiple moments within the preset first time period exceeding a preset second photovoltaic power supply threshold; or the historical photovoltaic power supply power at multiple moments within the preset first time period gradually increases in time, and the maximum historical photovoltaic power supply power among the historical photovoltaic power supplies at multiple moments within the preset first time period exceeds a preset third photovoltaic power supply threshold.
[0052] When the photovoltaic camera is working, if the historical photovoltaic power supply at multiple moments within a preset first time period meets the preset photovoltaic power supply conditions, it means that the historical photovoltaic power supply at multiple moments within the preset first time period meets the predetermined conditions and can achieve the expected power. The subsequent speech of the target speech can be recognized by the second recognition unit to improve the speech recognition accuracy. At the same time, it can ensure the continuous working state of the second recognition unit and avoid the discontinuity of speech recognition quality.
[0053] When the second recognition unit is working, the third recognition unit is turned off while the second recognition unit recognizes the subsequent speech of the target speech, which can further reduce the energy consumption of the photovoltaic camera.
[0054] In some implementations, the preset photovoltaic power supply condition may include a preset number of historical photovoltaic power supplies exceeding a preset first photovoltaic power supply threshold within a preset first time period. When a preset number of historical photovoltaic power supplies exceed the preset first photovoltaic power supply threshold within a preset first time period, it indicates that the lighting environment of the photovoltaic camera is good, and the photovoltaic power generation is only temporarily reduced. At this time, the subsequent speech of the target speech can still be recognized by the second recognition unit to ensure the continuous working state of the second recognition unit.
[0055] For example, the preset quantity can be 50% to 80% of the total historical photovoltaic power generation within the first time period.
[0056] In some implementations, the preset photovoltaic power supply condition may include the average power of historical photovoltaic power supply at multiple times within a preset first time period being greater than a preset second photovoltaic power supply threshold. When the average power of historical photovoltaic power supply at multiple times within a preset first time period is greater than the preset second photovoltaic power supply threshold, it indicates that the photovoltaic power generation of the photovoltaic camera is relatively large within the preset first time period. At this time, voice recognition can still be performed through the second recognition unit, ensuring the continuous working state of the second recognition unit.
[0057] In some implementations, the preset second photovoltaic power supply threshold can be 50% to 60% of the maximum photovoltaic power supply power of the day.
[0058] When designing a photovoltaic camera, the overall operating power of the photovoltaic camera is generally adapted to the maximum photovoltaic power supply of the day under different seasons and weather conditions. By setting the preset second photovoltaic power supply threshold to the maximum photovoltaic power supply of the day as a reference, the working state of the second recognition unit can be controlled based on the photovoltaic power supply of the day. This allows the control of the working state of the second recognition unit to adapt to the different light conditions of the day and season, ensuring that the photovoltaic camera can perform high-precision voice recognition through the second recognition unit. This avoids the second recognition unit being in a closed state for a long time, thus improving the use effect of the photovoltaic camera.
[0059] In some implementations, the preset photovoltaic power supply conditions may include the historical photovoltaic power supply at multiple moments within a preset first time period gradually increasing in time sequence, and the maximum historical photovoltaic power supply among the historical photovoltaic power supply at multiple moments within the preset first time period being greater than a preset third photovoltaic power supply threshold.
[0060] During operation, the preset photovoltaic power supply conditions can include a gradual increase in historical photovoltaic power supply at multiple moments within a preset first time period. When the historical photovoltaic power supply at multiple moments within the preset first time period gradually increases, it indicates that the historical photovoltaic power supply at multiple moments within the preset first time period is gradually increasing during a period of enhanced solar irradiance, and the decrease in photovoltaic power supply is only temporary due to shading. Simultaneously, if the maximum historical photovoltaic power supply at multiple moments within the preset first time period exceeds a preset third photovoltaic power supply threshold, it indicates that the historical photovoltaic power supply at multiple moments within the preset first time period has reached a relatively high power supply level. At this point, the second identification unit can be activated to ensure a sufficient power supply.
[0061] For example, the preset third photovoltaic power supply threshold can be 60% to 70% of the maximum photovoltaic power supply power of the day, so that the preset third photovoltaic power supply threshold can adapt to the daytime lighting conditions of the photovoltaic camera and improve the use effect of the photovoltaic camera.
[0062] The beneficial effect of the above implementation method is that when the movement of clouds obstructs the camera and causes changes in the camera's power, it can avoid the frequent switching of the voice recognition function of the photovoltaic camera, which would affect the working effect of the camera.
[0063] The beneficial effect of the above implementation method is that it can ensure the continuous operation of the second recognition unit when the photovoltaic power generation power fluctuates temporarily, avoid the interruption of speech recognition quality, and ensure the speech recognition effect.
[0064] In some implementations, when the second recognition unit recognizes the subsequent speech of the target speech, the third recognition unit is turned off after a preset waiting time; when the third recognition unit recognizes the subsequent speech of the target speech, the second recognition unit is turned off after a preset waiting time.
[0065] When the second recognition unit recognizes the subsequent speech of the target speech, the third recognition unit shuts down after a preset waiting time. This allows the third recognition unit to wait for the preset waiting time before shutting down, thus avoiding the impact on the lifespan of the photovoltaic camera caused by frequent power-on and power-off of the third recognition unit.
[0066] When the third recognition unit recognizes the subsequent speech of the target speech, the second recognition unit shuts down after a preset waiting time. This allows the second recognition unit to wait for the preset waiting time before shutting down, thus avoiding the impact on the lifespan of the photovoltaic camera caused by frequent power switching on and off of the second recognition unit.
[0067] The beneficial effect of the above implementation method is that when the working state of the second identification unit and the third identification unit is switched, they can be turned off after waiting for a preset time. This can delay the switching when the photovoltaic power generation conditions change, improve the service life of the photovoltaic camera, and avoid the increase in energy consumption caused by the frequent switching of the second identification unit and the third identification unit.
[0068] In some implementations, in the above-mentioned S220, when the historical photovoltaic power supply power at multiple times within a preset first time period meets the preset photovoltaic power supply power condition, the subsequent speech of the target speech is recognized by the second recognition unit, including S221 to S222. S221 to S222 will be explained in detail below.
[0069] S221. Obtain the historical photovoltaic power output of the photovoltaic camera at multiple times within a preset second time period from the current time. The preset second time period is shorter than a preset first time period.
[0070] Figure 3 This is a schematic diagram illustrating a time range of a preset first duration and a preset second duration in an embodiment of this application, as shown below. Figure 3 As shown, during operation, when the historical photovoltaic power supply power at multiple moments within a preset first time period meets the preset photovoltaic power supply power condition, it indicates that the photovoltaic power supply power within the preset first time period from the current moment meets the requirements. In order to more accurately judge the photovoltaic power supply effect, the historical photovoltaic power supply power of the photovoltaic camera at multiple moments within a preset second time period from the current moment can be obtained. The preset second time period is shorter than the preset first time period, so the photovoltaic power supply status can be more accurately judged by the historical photovoltaic power supply power at multiple moments within the preset second time period from the current moment.
[0071] For example, the preset second duration can be 20% to 30% of the preset first duration.
[0072] For example, when acquiring the historical photovoltaic power supply at multiple moments within a preset second time period, the historical photovoltaic power supply at multiple moments within the preset second time period can be acquired at equal time intervals.
[0073] S222. When the historical photovoltaic power supply at multiple moments within a preset second time period is greater than or equal to a preset fourth photovoltaic power supply threshold, the second recognition unit recognizes the subsequent speech of the target speech. When the historical photovoltaic power supply at multiple moments within a preset second time period is less than the preset fourth photovoltaic power supply threshold, the third recognition unit recognizes the subsequent speech of the target speech. The second recognition unit is turned off when the third recognition unit recognizes the subsequent speech of the target speech.
[0074] When the historical photovoltaic power supply power at multiple moments within the preset second time period is greater than or equal to the preset fourth photovoltaic power supply threshold, it indicates that the historical photovoltaic power supply power at multiple moments within the preset second time period is relatively large. At this time, the subsequent speech of the target speech can be recognized by the second recognition unit to ensure the accuracy of speech recognition.
[0075] For example, the preset fourth photovoltaic power supply threshold can be 40% to 50% of the maximum photovoltaic power supply power of the day.
[0076] For example, the preset fourth photovoltaic power supply threshold can be lower than the preset third photovoltaic power supply threshold, so that the photovoltaic power supply threshold within the preset second time period closer to the current time is lower. This allows voice recognition to be performed by the third recognition unit when the photovoltaic power supply is low in a local time period, avoiding the use of the third recognition unit with lower power consumption when the photovoltaic power supply is low within the preset second time period closer to the current time. This prevents the photovoltaic camera from consuming too much energy and the photovoltaic power supply from affecting the battery life.
[0077] When the historical photovoltaic power supply power at multiple times within the preset second time period is less than the preset fourth photovoltaic power supply threshold, it indicates that the photovoltaic power supply power within the preset second time period is low. At this time, the subsequent speech of the target speech can be recognized by the third recognition unit to reduce the energy consumption during speech recognition.
[0078] While the third recognition unit is recognizing subsequent speech of the target speech, the second recognition unit is turned off to reduce the energy consumption of the photovoltaic camera.
[0079] The beneficial effect of the above implementation method is that when the working state of the photovoltaic camera is controlled by the photovoltaic power supply within a preset first time period, the voice recognition function is further controlled by the photovoltaic power supply within a preset second time period that is closer to the current time, which improves the accuracy of energy consumption control of voice recognition and increases the battery life of the photovoltaic camera.
[0080] The beneficial effect of the above implementation method is that the preset fourth photovoltaic power supply threshold is lower than the preset third photovoltaic power supply threshold. When the photovoltaic power supply is low in a local time period, voice recognition can be performed through the third recognition unit. This avoids the situation where the photovoltaic power supply is low in the preset second time period closer to the current time, and the third recognition unit with lower power consumption is used for voice recognition. This avoids the photovoltaic camera's excessive energy consumption and low photovoltaic power supply affecting the battery life.
[0081] In some implementations, the above method also includes S310 to S320, which will be described in detail below.
[0082] S310. When the real-time photovoltaic power supply is less than or equal to the preset real-time photovoltaic power supply, obtain the remaining power value of the photovoltaic camera.
[0083] During operation, the voice recognition function can also be controlled by the remaining power value of the photovoltaic camera. When the real-time photovoltaic power supply is less than or equal to the preset real-time photovoltaic power supply, it indicates that the photovoltaic power supply is low. At this time, the remaining power value of the photovoltaic camera can be obtained, and the voice recognition function can be controlled according to the remaining power value of the photovoltaic camera.
[0084] S320. When the remaining power value is greater than the preset power value, the second recognition unit is activated to recognize the subsequent speech of the target speech. When the second recognition unit recognizes the subsequent speech of the target speech, the third recognition unit is turned off.
[0085] After obtaining the remaining power value of the photovoltaic camera, if the remaining power value is greater than the preset power value, it indicates that the photovoltaic camera has a large remaining power. At this time, the second recognition unit can be activated to recognize the subsequent speech of the target voice. The second recognition unit achieves high-precision speech recognition. The third recognition unit is turned off while the second recognition unit is recognizing the subsequent speech of the target voice.
[0086] For example, the preset power value can be 50% to 60% of the rated power value of the photovoltaic camera.
[0087] The beneficial effect of the above implementation method is that it can control the voice recognition function according to the remaining power value of the photovoltaic camera, so as to achieve high-precision voice recognition when the remaining power of the photovoltaic camera is high, thereby improving the use effect of the photovoltaic camera.
[0088] In some implementations, in the above-mentioned S320, when the remaining power value is greater than the preset power value, the second recognition unit is activated to recognize the subsequent speech of the target speech, including S321 to S322. S321 to S322 will be explained in detail below.
[0089] S321. Obtain the maximum volume value of the subsequent speech of the target speech within a preset third duration starting from the current moment.
[0090] When the second recognition unit is activated to recognize the subsequent speech of the target speech, since the second recognition unit consumes a lot of energy, the expected working time of the second recognition unit can be predicted, and the speech recognition function of the photovoltaic camera can be controlled according to the expected working time of the second recognition unit to ensure the continuity of the recognition unit used in speech recognition and avoid malfunctions in the speech recognition process.
[0091] In order to predict the expected working time of the second recognition unit during operation, the maximum volume value of the subsequent speech of the target speech at multiple moments within a preset third time period can be obtained, and the expected working time of the second recognition unit can be predicted based on the maximum volume value.
[0092] For example, the preset third duration can be 1 minute, 3 minutes or 5 minutes.
[0093] S322. When the maximum volume value at multiple moments within a preset third time period is greater than a preset volume threshold, the third recognition unit is activated to recognize the subsequent speech of the target speech. While the third recognition unit is recognizing the subsequent speech of the target speech, the second recognition unit is deactivated. When the maximum volume value at multiple moments within a preset third time period is less than or equal to the preset volume threshold, or when the maximum volume value at multiple moments within a preset third time period gradually decreases, the second recognition unit is activated to recognize the subsequent speech of the target speech. While the second recognition unit is recognizing the subsequent speech of the target speech, the third recognition unit is deactivated.
[0094] After obtaining the maximum volume values at multiple moments within a preset third time period, if the maximum volume values at multiple moments within the preset third time period are all greater than the preset volume threshold, it indicates that the subsequent speech of the target speech may last for a long time. In order to reduce the energy consumption of speech recognition, the third recognition unit can be woken up to recognize the subsequent speech of the target speech to ensure that speech recognition can continue for a long time. The second recognition unit can be turned off when the third recognition unit recognizes the subsequent speech of the target speech.
[0095] After obtaining the maximum volume values at multiple moments within a preset third time period, if the maximum volume values at multiple moments within the preset third time period are all less than or equal to a preset volume threshold, or if the maximum volume values at multiple moments within the preset third time period gradually decrease, it indicates that the target speech is gradually decreasing or the preset volume threshold is low and the recognition requirement is low. This suggests that the recognition time for the target speech may be short, and the second recognition unit can be activated to recognize the subsequent speech of the target speech. This avoids switching the recognition unit used for speech recognition during the speech recognition process, ensuring the continuity of speech recognition and avoiding speech recognition errors. Simultaneously, the third recognition unit is turned off while the second recognition unit is recognizing the subsequent speech of the target speech.
[0096] The beneficial effect of the above implementation method is that it can control the voice recognition function of the photovoltaic camera according to the expected working time of the second recognition unit, thereby ensuring the continuity of the recognition unit used in the voice recognition and avoiding failure in the voice recognition process.
[0097] The beneficial effect of the above implementation method is that when the remaining power value is greater than the preset power value, the second recognition unit is woken up to recognize the subsequent speech of the target speech. When the subsequent speech of the target speech may last for a long time, the third recognition unit is woken up to recognize the subsequent speech of the target speech, which can reduce the energy consumption of speech recognition.
[0098] The beneficial effect of the above implementation method is that when the remaining power value is greater than the preset power value, the second recognition unit is woken up to recognize the subsequent speech of the target speech. When the recognition time of the target speech may be short, the second recognition unit is woken up to recognize the subsequent speech of the target speech, thus avoiding switching the recognition unit used for speech recognition during the speech recognition process.
[0099] In some implementations, the above method also includes S410 to S420, which are described in detail below.
[0100] S410. When the maximum volume value at multiple moments within a preset third time period is greater than the preset volume threshold, obtain the noise value in the subsequent speech when the third recognition unit recognizes the subsequent speech of the target speech.
[0101] When the maximum volume value at multiple moments within the preset third time period is greater than the preset volume threshold, the third recognition unit can be activated to recognize the subsequent speech of the target speech. Since the recognition accuracy of the third recognition unit is low, measures can be taken to supplement the recognition effect of the third recognition unit.
[0102] When supplementing the recognition effect of the third recognition unit, the noise value in the subsequent speech when the third recognition unit recognizes the target speech can be obtained. The noise value in the subsequent speech represents the recognition effect of the third recognition unit.
[0103] For example, noise values in subsequent speech can be determined using voiceprint features. Since speech recognition involves more computational complexity and data processing requirements, the hardware power consumption of speech recognition is usually higher than that of voiceprint feature extraction, making the power consumption for voiceprint feature judgment lower and facilitating the judgment of noise values.
[0104] S420. When the noise value is greater than the preset noise value, the subsequent speech of the target speech is stored. When the real-time photovoltaic power supply is greater than the preset real-time photovoltaic power supply, the second recognition unit is woken up to recognize the subsequent speech of the stored target speech.
[0105] When the noise value is greater than the preset noise value, it indicates that there is a lot of noise in the subsequent speech of the target speech, and the recognition effect of the third recognition unit may be poor. At this time, the subsequent speech of the target speech can be stored. When the real-time photovoltaic power supply is greater than the preset real-time photovoltaic power supply, the second recognition unit is woken up to recognize the stored subsequent speech of the target speech. The second recognition unit ensures the recognition effect of speech with a large noise value.
[0106] The beneficial effect of the above implementation method is that when there is a lot of noise in the subsequent speech of the target speech, the recognition effect of the third recognition unit may be poor. The second recognition unit can ensure the recognition effect of speech with large noise values.
[0107] This application also provides a low-power voice recognition device for a photovoltaic camera, including a unit for performing the method as described in any of the preceding claims.
[0108] Figure 4 This application provides a schematic diagram of the logic structure of a low-power voice recognition device for a photovoltaic camera, as shown in one embodiment. Figure 4 As shown, the apparatus 3 in this embodiment includes a processing unit 31, a storage unit 32, and a transceiver unit 33. The processing unit 31 is used to process data, the storage unit 32 is used to store data, and the transceiver unit 33 is used to send and receive data. The processing unit 31, the storage unit 32, and the transceiver unit 33 cooperate with each other to implement the above-described method. The beneficial effects brought about by the embodiments of this application have been described in the above-described method and will not be repeated here.
[0109] This application also provides a low-power voice recognition device for a photovoltaic camera, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the method described in any of the preceding claims.
[0110] Figure 5 This is a schematic diagram of the physical structure of a low-power voice recognition device for a photovoltaic camera, provided in one embodiment of this application. Figure 5 As shown, the device 4 of this embodiment includes: at least one processor 40 ( Figure 5 Only one processor 40, a memory 41, and a computer program 42 stored in the memory 41 and executable on the at least one processor 40 are shown. When the processor 40 executes the computer program 42, it implements the steps in any of the above-described method embodiments. The beneficial effects of the embodiments of this application have been described in the above-described methods and will not be repeated here.
[0111] It should be noted that the information interaction and execution process between the above-mentioned devices / units are based on the same concept as the method embodiments of this application. For details on their specific functions and technical effects, please refer to the method embodiments section, and they will not be repeated here.
[0112] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is merely an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiments can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit. Furthermore, the specific names of the functional units and modules are only for easy differentiation and are not intended to limit the scope of protection of this application. The specific working process of the units and modules in the above system can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.
[0113] This application also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps described in the various method embodiments above.
[0114] This application provides a computer program product that, when run on a mobile terminal, enables the mobile terminal to implement the steps described in the above-described method embodiments.
[0115] This application also provides a photovoltaic camera, which, when in operation, can perform the steps described in any of the preceding methods.
[0116] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, all or part of the processes in the methods of the above embodiments of this application can be implemented by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. The computer-readable medium can include at least: any entity or device capable of carrying computer program code to a photographing device / terminal device, a recording medium, a computer memory, a read-only memory (ROM), a random access memory (RAM), an electrical carrier signal, a telecommunication signal, and a software distribution medium. Examples include USB flash drives, portable hard drives, magnetic disks, or optical disks. In some jurisdictions, according to legislation and patent practice, computer-readable media cannot be electrical carrier signals or telecommunication signals.
[0117] In the above embodiments, the descriptions of each embodiment have different focuses. For parts that are not described in detail or recorded in a certain embodiment, please refer to the relevant descriptions of other embodiments.
[0118] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0119] In the embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative. For instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.
[0120] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0121] The above-described embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application, and should all be included within the protection scope of this application.
Claims
1. A low-power voice recognition method for a photovoltaic camera, characterized in that, The method includes: The first recognition unit determines the first recognition information of the target speech. When the first recognition information meets the preset first recognition conditions, the real-time photovoltaic power supply of the photovoltaic camera is obtained. The preset first recognition conditions are human voice or contain keywords. When the real-time photovoltaic power supply is greater than the preset real-time photovoltaic power supply, the second recognition unit is activated to recognize the subsequent speech of the target speech. When the real-time photovoltaic power supply is less than or equal to the preset real-time photovoltaic power supply, the third recognition unit is activated to recognize the subsequent speech of the target speech; and the historical photovoltaic power supply of the photovoltaic camera at multiple times within a preset first time period from the current time, and the remaining power value of the photovoltaic camera are obtained respectively. When the historical photovoltaic power supply at multiple moments within a preset first time period meets the preset photovoltaic power supply conditions, the following steps are executed: The system acquires the historical photovoltaic power output of the photovoltaic camera at multiple times within a preset second time period from the current time; wherein the preset second time period is shorter than a preset first time period. When the historical photovoltaic power supply power at multiple moments within a preset second time period is greater than or equal to a preset fourth photovoltaic power supply threshold, the second recognition unit recognizes the subsequent speech of the target speech, and the third recognition unit is turned off; when the historical photovoltaic power supply power at multiple moments within a preset second time period is less than the preset fourth photovoltaic power supply threshold, the third recognition unit recognizes the subsequent speech of the target speech, and the second recognition unit is turned off; wherein, the power consumption of the second recognition unit is greater than the power consumption of the third recognition unit; The preset photovoltaic power supply conditions include: a preset number of historical photovoltaic power supplies within a preset first time period are greater than a preset first photovoltaic power supply threshold; or the average power of historical photovoltaic power supplies at multiple moments within a preset first time period is greater than a preset second photovoltaic power supply threshold; or the historical photovoltaic power supplies at multiple moments within a preset first time period gradually increase in time sequence, and the maximum historical photovoltaic power supply among the historical photovoltaic power supplies at multiple moments within a preset first time period is greater than a preset third photovoltaic power supply threshold. When the remaining battery level is greater than the preset battery level, perform the following steps: Obtain the maximum volume values of subsequent speech of the target speech at multiple times within a preset third duration starting from the current time; When the maximum volume value at multiple moments within the preset third time period is greater than the preset volume threshold, the third recognition unit is activated to recognize the subsequent speech of the target speech, and the second recognition unit is turned off; when the maximum volume value at multiple moments within the preset third time period is less than or equal to the preset volume threshold, or when the maximum volume value at multiple moments within the preset third time period gradually decreases, the second recognition unit is activated to recognize the subsequent speech of the target speech, and the third recognition unit is turned off. When the maximum volume value at multiple moments within the preset third time period is greater than the preset volume threshold, the noise value in the subsequent speech when the third recognition unit recognizes the subsequent speech of the target speech is obtained. When the noise value is greater than the preset noise value, the subsequent speech of the target speech is stored. When the real-time photovoltaic power supply is greater than the preset real-time photovoltaic power supply, the second recognition unit is activated to recognize the subsequent speech of the stored target speech.
2. A low-power voice recognition device for a photovoltaic camera, characterized in that, Includes units for performing the method of claim 1.
3. A low-power voice recognition device for a photovoltaic camera, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the method as described in claim 1.
4. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the method as described in claim 1.
5. A photovoltaic camera, characterized in that, When the photovoltaic camera is working, it can perform the steps of the method described in claim 1.
Citation Information
Patent Citations
Speech recognition devices, electronic devices and speech recognition methods
CN111755002B
Method and device in support of automatic switching of wakeup modes
CN107277672A
Control method and device of new energy camera equipment and storage medium
CN114039398A