Intelligent camera voice interaction control system and method based on open source gap system

By configuring the open source Hongmeng driver interface and a voice interaction control system with pre-configured user permissions in the smart camera, the problem of cameras being unable to be inspected and repaired in a remote state without a network is solved, and convenient and efficient voice control effects are achieved.

CN120264125APending Publication Date: 2025-07-04UNIONMANTECH +1
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510468723.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-15
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

Existing smart cameras cannot be repaired in a remote air condition without network when the network fails, resulting in difficult and inefficient troubleshooting.

Method used

The intelligent camera voice interaction control system based on the open source Hongmeng system is adopted. By configuring the open source Hongmeng driver interface in the camera, user voice data is collected, and combined with pre-configured user operation permissions and keyword configuration tables, voice command control without network dependencies is realized.

Benefits of technology

It realizes convenient, simple and efficient voice interaction control of smart cameras in a networkless state, and improves the convenience and efficiency of troubleshooting and operation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120264125A_ABST
    Figure CN120264125A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of voice interaction control, and discloses an intelligent camera voice interaction control system and method based on an open source gap system. The system comprises an application layer, a system framework layer and an open-source gap driver layer, wherein the open-source gap driver layer is used for calling an open-source gap driver interface configured on an intelligent camera to obtain voice data of a target user; the system framework layer is used for analyzing the voice data to obtain text information and voiceprint information; according to the text information, the voiceprint information, a pre-configured operation authority of each user and a preset keyword configuration table, executing a business process within an operation authority range of the target user; and the application layer is used for configuring the operation authority of each user, and executing the interactive operation corresponding to the business process in the operation authority range of the target user by controlling the system framework layer. According to the invention, the voice instruction remote control can be carried out on the intelligent camera without depending on the network and professional knowledge of the user, and the method is convenient, simple and efficient.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of voice interaction control, and particularly to an intelligent camera voice interaction control system based on the open-source HarmonyOS system and an intelligent camera voice interaction control method based on the open-source HarmonyOS system. Background Art

[0002] Existing intelligent cameras need to be controlled through an application under network connection conditions. In particular, when performing configuration operations, a wired network is also required to connect the intelligent camera to a router, or the application is used for network configuration with the assistance of Bluetooth, or the platform management system provided by the supplier is deployed in the internal network for management. When a network failure occurs in the intelligent camera, the application or the platform management system cannot be connected to the intelligent camera, and only professional engineers can perform on-site maintenance. However, since the camera is often installed at a high position, it is extremely inconvenient to connect the computer with a network cable, and the troubleshooting work is difficult to execute, and the maintenance efficiency is low.

[0003] It can be seen that when a network failure occurs in the intelligent camera in the prior art, there is a problem that the intelligent camera cannot be remotely maintained in a network-free state. Summary of the Invention

[0004] The purpose of the present invention is to overcome the problem that when a network failure occurs in an intelligent camera in the prior art, the intelligent camera cannot be remotely maintained in a network-free state, and to provide an intelligent camera voice interaction control system and method based on the open-source HarmonyOS system.

[0005] To achieve the above purpose, on the one hand, the present invention provides an intelligent camera voice interaction control system based on the open-source HarmonyOS system. The system mainly includes an application layer, a system framework layer, and an open-source HarmonyOS driver layer. Among them, the open-source HarmonyOS driver layer is used to call the open-source HarmonyOS driver interface configured in the intelligent camera to obtain the voice data of the target user. The system framework layer is used to obtain the voice data, parse the voice data to obtain text information and voiceprint information, and execute the business process within the operation authority of the target user according to the text information, the voiceprint information, the operation authorities of each user pre-configured in the application layer, and a preset keyword configuration table. The application layer is used to configure the operation authorities of each user and execute the pending instructions corresponding to the business process within the operation authority of the target user by controlling the system framework layer.

[0006] Optionally, the OpenHarmony driver layer includes an OpenHarmony audio driver module and an OpenHarmony network driver module; wherein, the OpenHarmony audio driver module is used to call the audio input / output module configured in the smart camera to obtain the voice data of the target user; the OpenHarmony network driver module is used to perform network configuration operations when initializing and configuring the smart camera, and to feedback the network status information of the smart camera when the smart camera is running normally.

[0007] Optionally, the system framework layer includes a control execution module, and the application layer includes a permission management module and a voice control module; wherein, the permission management module is used to configure the operation permissions corresponding to the voiceprint information of each user; the voice control module is used to, when initializing and configuring the smart camera, in response to detecting the keyword corresponding to the initialization configuration and receiving a valid initialization password, record the voiceprint information of the target user and configure the operation permission of the target user as the initialization configuration permission; or, in response to receiving the control parameters corresponding to the instruction to be processed, control the control execution module to execute the service process corresponding to the instruction to be processed.

[0008] Optionally, the system framework layer includes a voice collection module, an AI voice module, and an instruction parsing module; wherein, the voice collection module is used to obtain the voice data of the OpenHarmony audio driver module and perform preprocessing operations to update the voice data; the AI voice module is used to parse the voice data into text information and voiceprint information respectively; the instruction parsing module is used to, when the smart camera is running normally, extract the keywords in the text information, and use a preset keyword configuration table to find the instruction to be processed corresponding to the keyword and the operation permissions required to execute the instruction to be processed.

[0009] Optionally, the AI voice module is further used to determine the identity of the target user corresponding to the voiceprint information based on the voiceprint information of each user parsed by itself and a preset similarity threshold;

[0010] The permission management module is further used to determine the operation permissions of the target user based on the operation permissions of each user.

[0011] Optionally, the control execution module includes a composite instruction execution unit and a simple instruction execution unit; wherein, the simple instruction execution unit is used to directly execute the service process corresponding to the instruction to be processed when the instruction to be processed is a simple instruction; the composite instruction execution unit is used to determine all the control parameters in the instruction to be processed and execute the service process corresponding to the instruction to be processed based on all the control parameters.

[0012] Optionally, the composite instruction execution unit includes a parameter verification subunit and a parameter set verification subunit; wherein, the parameter verification subunit is configured to verify the legality of the values of the control parameters in response to receiving the values of the control parameters, and obtain the legal values of the control parameters; the parameter set verification subunit is configured to construct a parameter set by using all the control parameters in response to receiving the legal values of the control parameters, and execute the service process corresponding to the to-be-processed instruction when the parameter set is legal.

[0013] A second aspect of the present invention provides an intelligent camera voice interaction control method based on the open-source HarmonyOS system, and the method includes:

[0014] Obtain the voice data of the target user;

[0015] Parse the voice data to obtain text information and voiceprint information;

[0016] Execute the to-be-processed instruction corresponding to the service process within the operation authority of the target user according to the text information, the voiceprint information, the pre-configured operation authorities of each user, and the preset keyword configuration table.

[0017] A third aspect of the present invention provides an electronic device, and the electronic device includes a memory for storing executable instructions; a processor for calling and running the executable instructions in the memory to implement the steps of the above-mentioned intelligent camera voice interaction control method based on the open-source HarmonyOS system.

[0018] A fourth aspect of the present invention provides a computer-readable storage medium, and program instructions are stored in the computer-readable storage medium. When the program instructions are run by a processor, the steps of the above-mentioned intelligent camera voice interaction control method based on the open-source HarmonyOS system are implemented.

[0019] Compared with the prior art, the beneficial effects of the present solution are as follows:

[0020] The present invention can collect the voice data of users by configuring an open-source HarmonyOS driver interface in an intelligent camera, and can realize the air control of voice instructions for the intelligent camera without relying on the network and the professional knowledge of users through pre-configuring the operation authorities of users and the keyword configuration table in the intelligent camera, which is convenient, simple and efficient.

[0021] Other features and advantages of the embodiments of the present invention will be described in detail in the subsequent specific implementation part. Description of the Drawings

[0022] The accompanying drawings are used to provide a further understanding of the embodiments of the present invention and constitute a part of the specification. Together with the following specific embodiments, they are used to explain the embodiments of the present invention, but do not constitute a limitation to the embodiments of the present invention. In the accompanying drawings:

[0023] Figure 1 It is a schematic diagram of the architecture of the intelligent camera voice interaction control system based on OpenHarmony of the present invention;

[0024] Figure 2 It is an operation flowchart of initializing the intelligent camera of the present invention;

[0025] Figure 3 It is a schematic diagram of the operation of the business process of the present invention;

[0026] Figure 4 It is an operation flowchart of the composite instruction of the present invention;

[0027] Figure 5 It is a method flowchart of the intelligent camera voice interaction control system based on OpenHarmony of the present invention;

[0028] Figure 6 It is a schematic diagram of the structure of the electronic device of the present invention. Specific Embodiments

[0029] Next, with reference to the accompanying drawings of the embodiments of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative work belong to the scope of protection of the present invention.

[0030] In the following description, many specific details are set forth in order to fully understand the present invention. However, the present invention can also be implemented in other ways different from those described herein. Those skilled in the art can make similar extensions without departing from the connotation of the present invention. Therefore, the present invention is not limited by the specific embodiments disclosed below.

[0031] Please refer to Figure 1 , an intelligent camera voice interaction control system based on OpenHarmony is proposed in the embodiments of the present invention. The system mainly includes an application layer, a system framework layer, and an OpenHarmony driver layer.

[0032] Among them, the OpenHarmony driver layer is used to call the OpenHarmony driver interface configured in the smart camera to obtain the voice data of the target user. The voice data refers to the voice data of the target user collected at a preset sampling rate and greater than a preset decibel value within a certain distance range. The target user refers to a human living body who wants to control the smart camera. For example, spectral features and / or phase information, or even endpoint detection of neural networks are used to identify whether the target user corresponding to the collected voice data is a human living body.

[0033] The system framework layer is used to obtain the voice data, parse the voice data to obtain text information, so as to determine the specific text content included in the voice data through the text information, including but not limited to keywords for controlling the smart camera, etc.; and use the voiceprint recognition method to parse the voice data into voiceprint information, so as to obtain the voiceprint information of the target user, so as to facilitate the subsequent identification of the operation authority of the target user through the voiceprint information. Among them, the voiceprint information is a digital fingerprint used to reflect the speaking characteristics of the human living body corresponding to the voice data. The system framework layer is also used to obtain the operation authorities of each user pre-configured in the application layer, and execute the business process within the operation authority of the target user according to the text information, voiceprint information, the pre-configured operation authorities of each user, and the preset keyword configuration table. Among them, when the text information contains a keyword in the keyword configuration table, the operation instruction with a mapping relationship in the keyword configuration table is found according to the keyword, and when the operation authority of the target user meets the operation authority requirements of the operation instruction, the business process for controlling the smart camera corresponding to the operation instruction is executed. Among them, the keyword configuration table is used to store keywords and the operation instructions corresponding to each keyword, that is, the mapping relationship table between keywords and operation instructions. It is easy to understand that the mapping relationship between keywords and operation instructions can be one-to-one, one-to-many or many-to-one, and can be flexibly set according to control requirements. By setting the keyword configuration table, the control instruction can be embedded in the smart camera, and the user only needs to say a sentence containing the keyword to execute the corresponding control instruction, which can greatly improve the control efficiency and operation simplicity of the smart camera. The business process refers to the operation process that can perform voice interaction control on the smart camera, including but not limited to business processes such as power on, power off, adjusting the direction and light brightness of the camera.

[0034] The application layer is used to configure the operation permissions of each user, that is, the administrator assigns operation permissions to each authorized user, stores the voice sample information of each authorized user, and identifies each authorized user according to the voiceprint information contained in the voice sample information. It is easy to understand that the operation permissions of authorized users can be changed under the operation of the administrator, and the voice sample information of each authorized user can be updated under the operation of the administrator. The application layer is also used to parse and identify the text information, voiceprint information, and operation permissions of the target user through the control system framework layer, and control the control system framework layer to execute the pending instructions corresponding to the business processes within the operation permissions of the target user.

[0035] Compared with the traditional method that requires professional engineers to repair cameras through wired networks, in this embodiment, by configuring an open-source HarmonyOS driver interface in the smart camera to collect the user's voice data, and by pre-configuring the operation permissions of the user and the keyword configuration table in the smart camera, it is possible to achieve wireless control of voice commands for the smart camera without relying on the network and the user's professional knowledge, which is convenient, simple, and efficient.

[0036] In a preferred embodiment, the open-source HarmonyOS driver layer includes an open-source HarmonyOS audio driver module and an open-source HarmonyOS network driver module. Among them, the open-source HarmonyOS audio driver module is used to obtain the voice data of the target user by calling the audio input and output module configured in the smart camera when initializing and configuring the smart camera or during the operation of the smart camera. The open-source HarmonyOS network driver module is used to perform network configuration operations on the smart camera when initializing and configuring the smart camera. For example, wired network static IP configuration, wired network DHCP automatic acquisition, wireless network (such as Wi-Fi) connection configuration, or configuration of advanced network parameters such as subnet mask, gateway, and DNS. The open-source HarmonyOS network driver module is also used to feedback the network status information of the smart camera when the smart camera is running normally, including real-time operation status monitoring information and fault monitoring information. For example, the real-time operation status monitoring information includes continuously monitoring the network connection status (connected or disconnected), measuring network quality indicators (delay, jitter, or packet loss rate, etc.), and / or tracking bandwidth usage, etc.; the fault monitoring information includes automatically detecting network anomaly information and / or identifying relevant information about common network problems such as IP conflicts and unreachable gateways. It can be seen that in this embodiment, by setting the open-source HarmonyOS network driver module, the system can achieve the initialization configuration and network status monitoring of the smart camera under the condition of no network, which helps to improve the convenience and timeliness of voice control of the smart camera.

[0037] In a preferred embodiment, the system framework layer includes a voice collection module, an AI voice module, an instruction parsing module, and a control execution module. Among them, the voice collection module is used to obtain the voice data of the target user by calling the interface of the open-source HarmonyOS driver module configured in the smart camera. The AI voice module is used to parse the voice data into text information and voiceprint information respectively by using the AI large model. The instruction parsing module is used to obtain the text information from the AI voice module, extract the keywords in the text information, and use the preset keyword configuration table to find the to-be-processed instructions corresponding to the keywords, the parameters included in the to-be-processed instructions, and the operation permissions required to execute the to-be-processed instructions. The control execution module is used to execute various instructions sent by the voice control module to implement the functions corresponding to various service processes.

[0038] In this embodiment, through the interaction among the voice collection module, the AI voice module, and the instruction parsing module in the system framework layer, it is possible to achieve voice interaction control of the smart camera in a network-free environment.

[0039] In a preferred embodiment, the application layer includes a permission management module and a voice control module. Among them, the permission management module is used to configure the operation permissions corresponding to the voiceprint information of each user. For the setting process of the operation permissions of each user, specifically, it includes: according to the need, input several voice data of this user, and after being collected by the open-source HarmonyOS driver layer and processed by the system framework layer, generate a voiceprint information, and the administrator sets the operation permissions corresponding to these voiceprint information, that is, sets the operation permissions of this user.

[0040] The voice control module is used to, when initializing and configuring the smart camera, in response to detecting the keywords of the initialization configuration permission (such as administrator permission), call the audio input and output module configured in the smart camera to ask for the initialization password of the smart camera. In response to receiving a valid initialization password, record the voiceprint information of the target user and configure the operation permission of the target user as the initialization configuration permission.

[0041] For example, such as Figure 2As shown, when entering the operation process of adding the administrator's voiceprint information for the first time, the keyword configuration table preset in the voice interaction control system of the intelligent camera becomes effective only after the administrator is added to the system. Before the administrator is added to the system, it only waits to receive the keyword for indicating the addition of the administrator in the voice data of the user. After receiving this keyword, it does not check the voiceprint and operation permissions, but reminds the user to say the default password (the default password is usually printed on the device body or inside the packaging box). If the password said by the user does not match the default password, it means that the password said by the user is incorrect, then it prompts an incorrect password and receives the password said by the user again; if the password said by the user matches the default password, it means that the password said by the user is correct, then it stores the voiceprint information of this user in the system, configures the operation permissions of this user as an administrator, and invalidates the default password at the same time, that is, the default password is only valid when adding the administrator for the first time and automatically becomes invalid after the administrator is added. It can be seen that in this embodiment, by embedding the default password of the intelligent camera in the voice interaction control system and pre-configuring the operation instructions corresponding to the keyword for indicating the addition of the administrator, it is possible to realize the initialization configuration of the intelligent camera under the condition of no network.

[0042] The voice control module is also used to, during the normal operation process after the initialization configuration is completed, after receiving the control parameters corresponding to the instruction to be processed, transmit the control parameters to the control execution module in the system framework layer, and control each module in the system framework layer to execute the interaction functions corresponding to the service process, specifically including:

[0043] After the intelligent camera is started, it detects the audio input / output module configured on the intelligent camera. If it detects that the audio input / output module is abnormal or fails, the intelligent camera cannot operate normally; if it detects that the audio input / output module is normal, it selects to use the input / output module to receive the voice data of the target user, and checks whether there is an administrator's voiceprint. If not, it means that the operation of adding the administrator has not been completed, then it executes the above operation process of adding the administrator's voiceprint information for the first time. If there is, it means that the operation of adding the administrator has been completed, then it loads the keyword configuration table in the instruction parsing module pre-configured in the system application layer, loads the voiceprint information pre-stored in the system and the operation permissions of each user corresponding to the pre-configured voiceprint information, to wait for the user to input voice data through the input / output module. When receiving the voice data input by the user, it transmits the voice data to the voice acquisition module in the system framework layer. Then, under the control of the voice control module in the application layer, it uses the AI voice module in the system framework layer to parse the voice data into text information and voiceprint information respectively, then uses the instruction parsing module to extract the keywords in the text information, and uses the preset keyword configuration table to find the instruction to be processed corresponding to the keyword, the parameters included in the instruction to be processed, and the operation permissions required to execute the instruction to be processed, specifically as Figure 3 shown.

[0044] In this embodiment, during the initialization configuration and normal operation of the smart camera, the voice acquisition module, the AI voice module, and the instruction parsing module in the system framework layer are controlled by the voice control module to interact, enabling voice interaction control in the normal working state of the smart camera.

[0045] In a preferred implementation manner, by comparing the operation permissions corresponding to the voiceprint information of the target user with the operation permissions required to execute the to-be-processed instruction corresponding to the voice data of the target user, it is determined whether the control operation that the target user wants to execute is within the scope of their operation permissions. Only when the control operation that the target user wants to execute is within the scope of their operation permissions can the corresponding business process be executed. The specific implementation process is as follows:

[0046] The AI voice module determines the identity of the target user based on the voiceprint information of each user parsed by itself and a preset similarity threshold. Specifically, it includes: comparing the voiceprint information of the target user with each pre-configured voiceprint information one by one, screening out one pre-configured voiceprint information with the highest similarity as the to-be-matched voiceprint information; calculating the similarity between the to-be-matched voiceprint information and the voiceprint information of the target user. Only when the preset similarity threshold is reached is it determined that the identity of the target user is the user identity corresponding to the to-be-matched voiceprint information. Then, the permission management module determines the operation permissions of the target user based on the operation permissions of each user. Specifically, it includes: determining the operation permissions of the user identity corresponding to the target user based on the stored operation permissions of each user as the operation permissions of the target user. It is easy to understand that the identity of the user in this embodiment is only used to establish a one-to-one mapping relationship between the voiceprint information and the user to reflect that different voiceprint information corresponds to different user identities.

[0047] When the operation permissions of the target user match the operation permissions required to execute the to-be-processed instruction, the voice control module sends the to-be-processed instruction to the control execution module to control the control execution module to execute the business process corresponding to the voice data, that is, the target user can only control the smart camera to execute the business processes within the scope of their operation permissions through voice.

[0048] In this embodiment, by comparing the operation permissions corresponding to the voiceprint information of the target user with the operation permissions required to execute the to-be-processed instruction corresponding to the voice data of the target user, it is determined whether the control operation that the target user wants to execute is within the scope of their operation permissions, which can effectively avoid malicious control of the smart camera by users and help improve the security of controlling the smart camera.

[0049] In a preferred embodiment, the control execution module includes a composite instruction execution unit, which is configured to determine all control parameters in the to-be-processed instruction and execute the service process corresponding to the to-be-processed instruction based on all the control parameters when the operation permission of the target user matches the operation permission required by the to-be-processed instruction and the to-be-processed instruction is a composite instruction.

[0050] In this embodiment, the to-be-processed instructions are divided into simple instructions and composite instructions. When the operation permission of the target user matches the operation permission required by the to-be-processed instruction and the to-be-processed instruction is a simple instruction, the simple instruction execution unit in the control execution module is directly controlled to execute the service process corresponding to the simple instruction, and the feedback information is output through the OpenHarmony network driver module. When the operation permission of the target user matches the operation permission required by the to-be-processed instruction and the to-be-processed instruction is a composite instruction, the composite instruction execution unit in the control execution module is used to determine all control parameters in the to-be-processed instruction and execute the service process corresponding to the to-be-processed instruction based on all the control parameters. Among them, a simple instruction refers to a control instruction that does not contain control parameters and can be executed without interacting with the user. A composite instruction refers to a control instruction that contains at least one control parameter and needs to ask the user for the values of each control parameter and can be executed after receiving the values of each control parameter. In this embodiment, the to-be-processed instructions are divided into two types: simple instructions and composite instructions according to whether they contain control parameters, and different execution processes are adopted for different types of instructions, which can effectively improve the execution efficiency of the instructions.

[0051] In a preferred embodiment, the composite instruction execution unit includes a parameter verification subunit and a parameter set verification subunit. Among them, the parameter verification subunit is configured to, in response to receiving the values of each control parameter received through the OpenHarmony audio driver module, verify whether the values of each control parameter are legal. If they are not legal, the values of the corresponding control parameters are repeatedly asked until the values of all control parameters are verified to be legal, and the legal values of each control parameter are obtained. The parameter set verification subunit is configured to, in response to receiving the legal values of all control parameters, construct a parameter set using all the control parameters in a preset format, and when the parameter set conforms to the preset format, it indicates that the parameter set is legal, and then execute the service process corresponding to the to-be-processed instruction. The preset format refers to the standard combination specification that conforms to each control parameter in the composite instruction. This embodiment sets a rigorous execution process for composite instructions, asks the user for the values of each control parameter in turn, and verifies the legality of each control parameter and the legality of the combination format of all control parameters, which can ensure the effective execution of composite instructions.

[0052] In an exemplary example, the detailed process of voice control of an intelligent camera using the voice interaction control system designed by the present invention is as Figure 4 shown as follows:

[0053] Assume that the intelligent camera has successfully added the administrator's voiceprint and is in a normal operating state. After receiving the sentence "Configure the wired network with a static IP" spoken by the target user through the OpenHarmony audio driver module, it is transmitted to the voice acquisition module; after the voice acquisition module extracts the effective voice segment from the voice data, it outputs to the AI voice module; the AI voice module parses the received voice data into text information and voiceprint information representing the user identity respectively, and then transmits them to the instruction parsing module and the control execution module; based on the operation permissions corresponding to the voiceprint information of each user pre-stored in the permission management module, the AI voice module checks whether there is the voiceprint information of the target user. If it is checked that the voiceprint information of the target user is not stored, the system feedbacks that the voice data is illegal. If it is checked that the voiceprint information of the target user exists, it judges the operation permissions corresponding to the voiceprint information of the target user, extracts the keywords in the text information through the instruction parsing module, and uses the preset keyword configuration table to find the to-be-processed instruction corresponding to the keyword and the operation permissions required to execute the to-be-processed instruction. If the operation permissions of the target user do not match the operation permissions required to execute the to-be-processed instruction, the system feedbacks that the voice data is illegal; if the operation permissions of the target user match the operation permissions required to execute the to-be-processed instruction, indicating that the voice interaction operation is legal, then it judges whether the to-be-processed instruction is a composite instruction. If it is not a composite instruction, it can be directly executed; if it is a composite instruction, it clarifies all the control parameters included in the composite instruction. For example, the instruction code of the to-be-processed instruction corresponding to the keyword in the text information is 213, and this instruction includes four control parameters: IP, subnet mask, gateway, and DNS. At this time, the voice control module switches the system mode to the instruction parameter input mode and asks the target user: "What do you want to set the IP to?" The user answers: "192.168.1.102", and the voice control module transmits this IP content to the AI voice module for parsing into text and checks whether the IP address represented by this text is legal. If it is legal, then it asks the user: "What do you want to set the subnet mask to?" The user answers: "255.255.255.0", and so on, asking for the values of the gateway and DNS in the same way. Among them, the voice control module can only receive voice data related to control parameters in the instruction parameter input mode, and the user cannot input data other than the values of the control parameters until exiting the instruction parameter input mode.

[0054] After receiving the values of each control parameter, it is determined one by one whether the control parameter is legal. If there is an illegal situation in the received value of the control parameter, for example, the IP address is an IPv4 address and the gateway is an IPv6 address, if it does not meet the requirements, it means it is illegal and needs to be repeatedly queried until the preset upper limit of the query times is reached or a legal value is received. Then, all control parameters are combined in the format of a wired static IP network configuration instruction to construct a parameter set, and the legality of the parameter set is judged. If the parameter set is legal, the composite instruction and the values of the above four control parameters are passed to the control execution module for processing, and the processing result is reported, and then the instruction parameter input mode is exited; otherwise, feedback information indicating that the parameter set is illegal is output.

[0055] Please refer to Figure 5 , the present invention provides a method for voice interaction control of an intelligent camera based on open source HarmonyOS, and the method includes:

[0056] Step S100: Obtain the voice data of the target user;

[0057] Step S200: Parse the voice data to obtain text information and voiceprint information;

[0058] Step S300: According to the text information, the voiceprint information, the operation permissions of each user pre-configured, and the preset keyword configuration table, execute the to-be-processed instruction corresponding to the service process within the operation permission range of the target user.

[0059] Specifically, in this embodiment, the specific functions of the above method can also refer to the corresponding descriptions in the above voice interaction control system of the intelligent camera based on open source HarmonyOS, and will not be elaborated here.

[0060] Based on the above embodiments, the present invention also provides an electronic device, and its principle block diagram can be as Figure 6 shown. This electronic device can be used to execute the steps of the method for voice interaction control of an intelligent camera based on open source HarmonyOS provided in the above embodiments. For the sake of brevity, it will not be elaborated here. This electronic device includes: a processor, the processor is coupled with a memory, the memory is used to store computer programs or instructions, and the processor is used to execute the computer programs or instructions stored in the memory, so that the method in the above method embodiments is executed.

[0061] The present invention also provides a computer-readable storage medium, on which computer instructions for implementing the method in the above method embodiments are stored.

[0062] For example, when the computer program is executed by a computer, the computer can implement the method in the above method embodiments.

[0063] The embodiments of the present application also provide a computer program product including instructions, which when executed by a computer cause the computer to implement the methods in the above method embodiments.

[0064] Those of ordinary skill in the art can realize that the units and algorithm steps of the examples described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can adopt different methods for each specific application to implement the described functions, but such implementation should not be considered to exceed the scope of this application.

[0065] Those skilled in the art can clearly understand that for the convenience and conciseness of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments, and will not be elaborated herein.

[0066] In several embodiments provided by the present application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division, and there can be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of devices or units can be in electrical, mechanical, or other forms.

[0067] The units described as separate components may or may not be physically separated. The components shown as units may or may not be physical units, that is, they can be located in one place, or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0068] In addition, in each embodiment of the present application, the functional units can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit.

[0069] If the described functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of this application. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs that can store program codes.

Claims

1. An intelligent camera voice interaction control system based on OpenHarmony, characterized in that, The system mainly includes an application layer, a system framework layer, and an OpenHarmony open-source driver layer; among them, the OpenHarmony open-source driver layer is used to call the OpenHarmony open-source driver interface configured in the smart camera to obtain the voice data of the target user; the system framework layer is used to obtain the voice data, parse the voice data to obtain text information and voiceprint information; according to the text information, the voiceprint information, the operation permissions of each user pre-configured in the application layer, and a preset keyword configuration table, execute the business process within the operation permissions of the target user; the application layer is used to configure the operation permissions of each user, and execute the to-be-processed instructions corresponding to the business process within the operation permissions of the target user by controlling the system framework layer.

2. The intelligent camera voice interaction control system based on the open-source HarmonyOS according to claim 1, characterized in that, The OpenHarmony open-source driver layer includes an OpenHarmony audio driver module and an OpenHarmony network driver module; among them, the OpenHarmony audio driver module is used to call the audio input and output module configured in the smart camera to obtain the voice data of the target user; the OpenHarmony network driver module is used to perform network configuration operations when initializing and configuring the smart camera, and feedback the network status information of the smart camera when the smart camera is running normally.

3. The intelligent camera voice interaction control system based on the open-source HarmonyOS according to claim 2, characterized in that, The system framework layer includes a control execution module, and the application layer includes a permission management module and a voice control module; among them, the permission management module is used to configure the operation permissions corresponding to the voiceprint information of each user; the voice control module is used to, when initializing and configuring the smart camera, in response to detecting the keyword corresponding to the initialization configuration and receiving a valid initialization password, record the voiceprint information of the target user and configure the operation permission of the target user as the initialization configuration permission; or, in response to receiving the control parameters corresponding to the to-be-processed instructions, control the control execution module to execute the business process corresponding to the to-be-processed instructions.

4. The intelligent camera voice interaction control system based on the open-source HarmonyOS according to claim 3, wherein The system framework layer includes a voice collection module, an AI voice module, and an instruction parsing module; among them, the voice collection module is used to obtain the voice data of the OpenHarmony audio driver module and perform preprocessing operations to update the voice data; the AI voice module is used to parse the voice data into text information and voiceprint information respectively; the instruction parsing module is used to, when the smart camera is running normally, extract the keywords in the text information, and use the preset keyword configuration table to find the to-be-processed instructions corresponding to the keywords and the operation permissions required to execute the to-be-processed instructions.

5. The intelligent camera voice interaction control system based on the open-source HarmonyOS according to claim 4, wherein, The AI voice module is also used to determine the identity of the target user corresponding to the voiceprint information based on the voiceprint information of each user parsed by itself and a preset similarity threshold; the permission management module is also used to determine the operation permissions of the target user based on the operation permissions of each user.

6. The intelligent camera voice interaction control system based on the open-source HarmonyOS according to claim 5, wherein The control execution module includes a composite instruction execution unit, and the composite instruction execution unit is used to, when the to-be-processed instruction is a composite instruction, determine all the control parameters in the to-be-processed instruction, and execute the business process corresponding to the to-be-processed instruction based on all the control parameters.

7. The intelligent camera voice interaction control system based on the open-source HarmonyOS according to claim 6, characterized in that, The composite instruction execution unit includes a parameter verification subunit and a parameter set verification subunit; among them, The parameter verification subunit is configured to verify the legality of the values of the control parameters in response to receiving the values of the control parameters, and obtain the legal values of the control parameters; The parameter set verification subunit is configured to construct a parameter set by using all the control parameters in response to receiving the legal values of the control parameters, and execute the service process corresponding to the to-be-processed instruction when the parameter set is legal.

8. A voice interaction control method for an intelligent camera based on OpenHarmony, characterized in that, The method includes: Obtaining voice data of a target user; Parsing the voice data to obtain text information and voiceprint information; Executing the to-be-processed instruction corresponding to the service process within the operation authority of the target user according to the text information, the voiceprint information, the pre-configured operation authorities of each user, and the preset keyword configuration table.

9. An electronic device, characterized in that, It includes: A memory for storing executable instructions; A processor for calling and running the executable instructions in the memory to implement the steps of the method for voice interaction control of an intelligent camera based on OpenHarmony as claimed in claim 8.

10. A computer-readable storage medium, characterized in that, Program instructions are stored in the computer-readable storage medium, and when the program instructions are run by the processor, the steps of the method for voice interaction control of an intelligent camera based on OpenHarmony as claimed in claim 8 are implemented.

Citation Information

Cited By

  • OpenHarmony-oriented end-side speech recognition optimization system

    CN122135722A

  • An edge-side speech recognition optimization system for OpenHarmony

    CN122135722B