Off-line voice interaction control method and device
By detecting the user's personnel feature information and analyzing the sound feature information, determining the target control instructions and parameters, the accuracy and reliability of the voice interaction control of the existing offline AI voice scheme in new scenarios is solved, and a higher user experience is achieved.
Patent Information
- Application Number
- CN202311679277.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-07
- Publication Date
- 2025-06-10
AI Technical Summary
Existing offline AI voice solutions are difficult to achieve high accuracy and reliability voice interaction control in practical applications, especially in new scenarios, where false wake-up or low recognition rate is prone to problems.
By detecting the user's personnel feature information, and after obtaining the sound feature information, the corresponding target control instructions and parameters are determined to perform matching functions. The method includes steps such as detecting personnel feature information, sound reception detection, sound feature information analysis, and target control instruction determination.
It improves the accuracy of offline voice detection and the reliability and accuracy of offline voice interaction control operations, thereby improving the user's experience in voice interaction control.
Smart Images

Figure CN120126466A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of voice interaction, and in particular, to an offline voice interaction control method and device. Background Art
[0002] Currently, the application of AI voice solutions has been very extensive. In the field of whole-house intelligence, considering the deployment scenarios of products, many products choose to deploy offline AI voice solutions. The offline AI voice solution enables the affiliated devices to have certain voice interaction control capabilities without relying on network data interaction. However, the offline AI solution also faces many problems in actual applications. Since the offline voice solution first conducts the learning and training of the pre-research model and then deploys the optimized model to the device side. The AI voice model trained based on the existing limited data sets cannot demonstrate good application capabilities in actual scenarios or some new scenarios. For example, in certain scenarios, the device is easily misawakened or cannot be awakened, or the recognition rate of obtaining user instructions is relatively low, making it difficult to achieve effective human-computer interaction. Therefore, it is particularly important to provide a method that can improve the accuracy of voice interaction control in offline scenarios. Summary of the Invention
[0003] The technical problem to be solved by the present invention is to provide an offline voice interaction control method and device, which can not only improve the accuracy of offline voice detection, but also enhance the reliability and accuracy of offline voice interaction control operations, thereby being beneficial to enhancing the user experience in voice interaction control.
[0004] To solve the above technical problem, in the first aspect of the present invention, an offline voice interaction control method is disclosed, and the method includes:
[0005] Detect whether the personnel feature information of a certain user is obtained;
[0006] When it is detected that the personnel feature information is obtained, perform a voice reception detection operation on the user to obtain a reception detection result;
[0007] When the reception detection result is used to indicate that the voice feature information of the user is received, determine a target control instruction corresponding to the voice feature information according to the voice feature information;
[0008] Generate a target control parameter corresponding to the target control instruction according to the target control instruction, and execute a function matching the target control parameter.
[0009] As an optional implementation manner, in the first aspect of the present invention, the determining a target control instruction corresponding to the voice feature information according to the voice feature information includes:
[0010] Analyze the voice feature information to obtain the analyzed voice feature information;
[0011] Based on the analyzed voice feature information, determine whether the analyzed voice feature information contains target voice information; the target voice information is voice information that matches at least one control instruction in a pre-stored control instruction set;
[0012] When it is determined that the analyzed voice feature information contains the target voice information, determine the target control instruction corresponding to the voice feature information according to all the control instructions corresponding to the target voice information in the analyzed voice feature information.
[0013] As an optional implementation manner, in the first aspect of the present invention, before analyzing the voice feature information to obtain the analyzed voice feature information, the method further includes:
[0014] Detect whether the voice feature information contains a confirmation voice instruction; the confirmation voice instruction is a voice instruction that matches at least one confirmation instruction in a pre-stored confirmation instruction set;
[0015] When it is detected that the voice feature information contains the confirmation voice instruction, trigger the operation of analyzing the voice feature information to obtain the analyzed voice feature information.
[0016] As an optional implementation manner, in the first aspect of the present invention, the personnel feature information includes the personnel position information of the user;
[0017] Before performing the voice reception detection operation on the user to obtain a reception detection result, the method further includes:
[0018] According to the personnel position information, determine the area parameters of the sound reception area that matches the personnel position information;
[0019] According to the area parameters, determine a target sound reception device that matches the area parameters from a preset set of sound reception devices, and determine the sound reception demand state of the target sound reception device that matches the area parameters;
[0020] Wherein, performing the voice reception detection operation on the user to obtain a reception detection result includes:
[0021] Based on the sound reception demand state of the target sound reception device, perform a voice reception detection operation on the user to obtain a reception detection result.
[0022] As an optional implementation manner, in the first aspect of the present invention, before detecting whether the personnel feature information of a certain user is obtained, the method further includes:
[0023] Detect whether there is a target set of sensing devices in a set of preset sensing devices for receiving personal characteristic information of a certain user that can be in a normal working state, and obtain a device status detection result; the normal working state means the working state in which the sensing device can normally receive the personal characteristic information.
[0024] Determine a target sensing demand state according to the device status detection result.
[0025] Among them, the detection of whether the personal characteristic information of a certain user is obtained includes:
[0026] Detect whether the personal characteristic information of a certain user is obtained according to the target sensing demand state.
[0027] As an optional implementation manner, in the first aspect of the present invention, the determining the target sensing demand state according to the device status detection result includes:
[0028] When the device status detection result indicates that there is a target set of sensing devices in the set of sensing devices that can be in the normal working state, obtain multi-dimensional scene parameters of the target scene where the target set of sensing devices is located; the multi-dimensional scene parameters of the target scene include one or a combination of environmental parameters, spatial layout parameters, network quality parameters, associated device parameters, and overall energy consumption parameters of the target scene.
[0029] Based on the multi-dimensional scene parameters of the target scene where the target set of sensing devices is located, determine target sensing devices that meet a preset first sensing condition from the target set of sensing devices, and determine the sensing demand state of the target sensing devices as the target sensing demand state.
[0030] As an optional implementation manner, in the first aspect of the present invention, the determining the target sensing demand state according to the device status detection result includes:
[0031] When the device status detection result indicates that there is no target set of sensing devices in the set of sensing devices that can be in the normal working state, determine a target relay sensing device set; the target relay sensing device set includes at least one target relay sensing device that can be used for personal characteristic sensing.
[0032] Determine the sensing demand state corresponding to each target relay sensing device as the target sensing demand state.
[0033] As an optional implementation manner, in the first aspect of the present invention, the determining the target relay sensing device set includes:
[0034] Obtain the multi-dimensional device information of all preset relay sensing devices, where the multi-dimensional device information includes at least one of the working status information, working range information, sensing type information, and transmission type information corresponding to each relay sensing device;
[0035] According to the multi-dimensional device information corresponding to each relay sensing device, determine all relay sensing devices that meet the preset second sensing conditions from all the relay sensing devices, as the target relay sensing device set, where the second sensing conditions include at least one of the personnel feature sensing condition, working idle state condition, working priority condition, and information transmission condition.
[0036] The second aspect of the present invention discloses an offline voice interaction control device, and the device includes:
[0037] A detection module, configured to detect whether the personnel feature information of a certain user is obtained; when it is detected that the personnel feature information is obtained, perform a voice reception detection operation on the user to obtain a reception detection result;
[0038] A first determination module, configured to, when the reception detection result is used to indicate that the voice feature information of the user is received, determine a target control instruction corresponding to the voice feature information according to the voice feature information;
[0039] An execution module, configured to generate a target control parameter corresponding to the target control instruction according to the target control instruction, and execute a function matching the target control parameter.
[0040] As an optional implementation manner, in the second aspect of the present invention, the first determination module includes:
[0041] An analysis sub-module, configured to analyze the voice feature information to obtain the analyzed voice feature information;
[0042] A judgment sub-module, configured to judge whether the analyzed voice feature information includes target voice information according to the analyzed voice feature information; the target voice information is voice information matching at least one control instruction in the pre-stored control instruction set;
[0043] A determination sub-module, configured to, when the judgment sub-module determines that the analyzed voice feature information includes the target voice information, determine a target control instruction corresponding to the voice feature information according to all the control instructions corresponding to the target voice information in the analyzed voice feature information.
[0044] As an optional implementation manner, in the second aspect of the present invention, the first determination module further includes:
[0045] A detection sub-module, configured to detect whether the voice feature information contains a confirmation voice command before the parsing sub-module parses the voice feature information to obtain the parsed voice feature information; when it is detected that the voice feature information contains the confirmation voice command, trigger the parsing sub-module to perform the operation of parsing the voice feature information to obtain the parsed voice feature information; the confirmation voice command is a voice command that matches at least one confirmation command in a pre-stored confirmation command set.
[0046] As an optional implementation manner, in the second aspect of the present invention, the personnel feature information includes the personnel location information of the user;
[0047] The device further includes:
[0048] A second determination module, configured to determine area parameters of a sound reception area that match the personnel location information according to the personnel location information before the detection module performs a sound reception detection operation on the user to obtain a reception detection result; according to the area parameters, determine a target sound reception device that matches the area parameters from a preset set of sound reception devices, and determine the sound reception demand state of the target sound reception device that matches the area parameters;
[0049] Wherein, the specific manner in which the detection module performs a sound reception detection operation on the user to obtain a reception detection result includes:
[0050] Based on the sound reception demand state of the target sound reception device, perform a sound reception detection operation on the user to obtain a reception detection result.
[0051] As an optional implementation manner, in the second aspect of the present invention, the detection module is further configured to:
[0052] Before detecting whether the personnel feature information of a certain user is obtained, detect whether there is a target perception device set that can be in a normal working state in a preset perception device set for receiving the personnel feature information of a certain user, and obtain a device state detection result; the normal working state means the working state in which the perception device can normally receive the personnel feature information;
[0053] The device further includes:
[0054] A third determination module, configured to determine a target perception demand state according to the device state detection result;
[0055] Wherein, the specific manner in which the detection module detects whether the personnel feature information of a certain user is obtained includes:
[0056] According to the target perception demand state, detect whether the personnel feature information of a certain user is obtained.
[0057] As an optional implementation manner, in the second aspect of the present invention, the specific manner in which the third determination module determines the target perception demand state according to the device state detection result includes:
[0058] When the device state detection result indicates that there is a target perception device set that can be in the normal working state in the perception device set, obtain the multi-dimensional scene parameters of the target scene where the target perception device set is located; the multi-dimensional scene parameters of the target scene include one or a combination of the environmental parameters, spatial layout parameters, network quality parameters, associated device parameters, and overall energy consumption parameters of the target scene;
[0059] Based on the multi-dimensional scene parameters of the target scene where the target perception device set is located, determine the target perception devices that meet the preset first perception condition from the target perception device set, and determine the perception demand state of the target perception devices as the target perception demand state.
[0060] As an optional implementation manner, in the second aspect of the present invention, the specific manner in which the third determination module determines the target perception demand state according to the device state detection result includes:
[0061] When the device state detection result indicates that there is no target perception device set that can be in the normal working state in the perception device set, determine a target relay perception device set; the target relay perception device set includes at least one target relay perception device that can be used for personnel feature perception;
[0062] Determine the perception demand state corresponding to each target relay perception device as the target perception demand state.
[0063] As an optional implementation manner, in the second aspect of the present invention, the specific manner in which the third determination module determines the target relay perception device set includes:
[0064] Obtain the multi-dimensional device information of all preset relay perception devices, and the multi-dimensional device information includes at least one of the working state information, working range information, perception type information, and transmission type information corresponding to each relay perception device;
[0065] Determine all relay sensing devices that meet the preset second sensing conditions from all the relay sensing devices according to the multi-dimensional device information corresponding to each relay sensing device, as the target relay sensing device set, where the second sensing conditions include at least one of a personnel feature sensing condition, a work idle state condition, a work priority condition, and an information transmission condition.
[0066] The third aspect of the present invention discloses another offline voice interaction control device, and the device includes:
[0067] A memory storing executable program code;
[0068] A processor coupled to the memory;
[0069] The processor calls the executable program code stored in the memory to execute the offline voice interaction control method disclosed in the first aspect of the present invention.
[0070] The fourth aspect of the present invention discloses a computer storage medium, and the computer storage medium stores computer instructions, which are used to execute the offline voice interaction control method disclosed in the first aspect of the present invention when called.
[0071] Compared with the prior art, the embodiments of the present invention have the following beneficial effects:
[0072] In the embodiments of the present invention, it is detected whether the personnel feature information of a certain user is obtained; when it is detected that the personnel feature information is obtained, a voice reception detection operation is performed to obtain a reception detection result; when the reception detection result is used to indicate that the voice feature information is received, according to the voice feature information, a target control instruction corresponding to the voice feature information is determined; according to the target control instruction, a target control parameter corresponding to the target control instruction is generated, and a function matching the control parameter is executed. It can be seen that implementing the present invention can perform a voice reception detection operation when it is detected that the personnel feature information is obtained, and determine the corresponding target control instruction according to the received voice feature information. In this way, not only can the accuracy of offline voice detection be improved, but also the reliability and accuracy of offline voice interaction control operations can be enhanced, which is beneficial to enhancing the user experience in voice interaction control. BRIEF DESCRIPTION OF THE DRAWINGS
[0073] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0074] Figure 1It is a schematic diagram of the scenario applicable to the offline voice interaction control method disclosed in the embodiments of the present invention;
[0075] Figure 2 It is a schematic flowchart of an offline voice interaction control method disclosed in the embodiments of the present invention;
[0076] Figure 3 It is a schematic flowchart of another offline voice interaction control method disclosed in the embodiments of the present invention;
[0077] Figure 4 It is a schematic structural diagram of an offline voice interaction control device disclosed in the embodiments of the present invention;
[0078] Figure 5 It is a schematic structural diagram of another offline voice interaction control device disclosed in the embodiments of the present invention;
[0079] Figure 6 It is a schematic structural diagram of yet another offline voice interaction control device disclosed in the embodiments of the present invention. Detailed implementation manners
[0080] In order to enable those skilled in the art to better understand the solution of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0081] The terms "first", "second", etc. in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish different objects, rather than to describe a specific order. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or terminal that includes a series of steps or units is not limited to the listed steps or units, but optionally further includes steps or units not listed, or optionally further includes other steps or units inherent to these processes, methods, products or terminals.
[0082] Referring to "embodiments" herein means that the specific features, structures or characteristics described in connection with the embodiments can be included in at least one embodiment of the present invention. The phrase appears in various places in the specification and does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those skilled in the art will explicitly and implicitly understand that the embodiments described herein can be combined with other embodiments.
[0083] The present invention discloses an offline voice interaction control method and device, which can not only improve the accuracy of offline voice detection, but also enhance the reliability and accuracy of offline voice interaction control operations, thus being beneficial to improving the user experience in voice interaction control. The following will be described in detail respectively.
[0084] To better understand an offline voice interaction control method and device described in the present invention, first, the applicable scenarios of an offline voice interaction control method are described. Specifically, the scenario can be as Figure 1 shown. Figure 1 is a schematic diagram of a home scenario applicable to an offline voice interaction control method disclosed in an embodiment of the present invention. As Figure 1 shown, the schematic diagram of the home scenario includes a user, a smart TV, a sofa, and a smart air conditioner with offline voice interaction control function. Among them, after the perception module integrated inside the smart air conditioner obtains the personal characteristic information of the user, the sound collection module integrated inside the smart air conditioner starts to receive the sound characteristic information, determines the control instruction according to the received sound characteristic information, and executes the corresponding function. Optionally, there can be multiple perception modules integrated inside the smart air conditioner, and the appropriate perception module can be determined through judgment conditions such as spatial position parameters (such as the position of the sofa). Optionally, when the perception module of the smart air conditioner is in a non-operating state, a relay perception device that meets the conditions can be screened out from the scenario. Exemplarily, if the screened relay perception device is integrated on the smart TV in this schematic diagram of the scenario, after the relay perception device corresponding to the smart TV obtains the personal characteristic information of the user, it can notify the smart air conditioner through Bluetooth or network, and then the sound collection module of the smart air conditioner starts to receive the sound characteristic information and determines the control instruction according to the received sound characteristic information to execute the corresponding function.
[0085] It should be noted that Figure 1 the shown schematic diagram of the scenario is only used to represent one of the scenarios applicable to the offline voice interaction control method and device, and it does not limit other scenarios applicable to the offline voice interaction control method and device. And Figure 1 the types of smart devices, the types of furniture, shapes, sizes, installation positions, etc. involved are not limited.
[0086] The above gives an example description of one of the scenarios applicable to the offline voice interaction control method and device. Next, the offline voice interaction control method and device will be described in detail.
[0087] Embodiment 1
[0088] Please refer to Figure 2 , Figure 2 is a schematic flowchart of an offline voice interaction control method disclosed in an embodiment of the present invention. Among them, Figure 2The described method can be applied to various scenario types, such as office scenarios, entertainment scenarios, home scenarios, etc., and the embodiments of the present invention do not make limitations. Further, this method can be applied to control various devices through voice recognition and interaction control in the foregoing scenario types, such as air conditioners, washing machines, refrigerators, televisions, floor cleaning robots, table lamps, surveillance cameras, speakers, water heaters, air purifiers, projectors, routers, rice cookers, etc. And optionally, this method can be implemented by an offline voice interaction control device, where the offline voice interaction control device can be a local server for processing the offline voice interaction control process, or a control module for implementing the offline voice interaction control process, and the embodiments of the present invention do not make limitations. As Figure 2 shown, the offline voice interaction control method includes the following operations:
[0089] 101. Detect whether the personnel feature information of a certain user is obtained.
[0090] In the embodiments of the present invention, the personnel feature information may include, but is not limited to, one or more of user presence judgment information, user identity feature information, user motion state information, and user motion state information. The user can be one or more users corresponding to the pre-set identity feature information, or a user without restricting the identity feature information; the detection action in step 101 can be continuously performed or performed according to a preset time, and the embodiments of the present invention do not make limitations. Optionally, when it is not detected that the personnel feature information of a certain user is obtained within the preset detection time range for whether the personnel feature information of a certain user is obtained, the relevant (sub) modules for functions such as sound collection and voice parsing can be controlled to enter the sleep state.
[0091] 102. When it is detected that the personnel feature information of the user is obtained, perform a sound reception detection operation on the user to obtain a reception detection result.
[0092] In the embodiments of the present invention, the sound reception detection operation on the user is performed by a set of sound collection devices. The set of sound collection devices includes at least one sound collection device. The sound collection device can be deployed on the outer surface or inside of the device with offline voice interaction control function in a centralized or distributed manner, or within the preset space range of the device / system / server with offline voice interaction control function, or at any position where the device / system / server with offline voice interaction control function needs to perform the sound collection function, and the embodiments of the present invention do not make limitations. Optionally, the sound collection devices in the set of sound collection devices can be a capacitive microphone, a crystal microphone, a carbon microphone, and a dynamic microphone, etc., or other acoustic sensors, and the embodiments of the present invention do not make limitations. Optionally, when it is not detected within the preset time range that the perception module obtains the personnel feature information of the user, higher-precision parameters can be used for perception.
[0093] 103. When the received detection result is used to indicate that the voice feature information of the user is received, determine the target control instruction corresponding to the voice feature information according to the voice feature information.
[0094] In an embodiment of the present invention, by way of example, the target control instruction may be an instruction for controlling an air conditioner to implement functions such as temperature change and wind direction adjustment, or an instruction for controlling a smart TV to implement functions such as volume adjustment and channel switching, or an instruction for controlling a smart home system to implement functions such as curtain opening and closing, lighting adjustment, and music switching. The embodiments of the present invention do not make limitations.
[0095] 104. Generate the target control parameter corresponding to the target control instruction according to the target control instruction, and execute the function matching the target control parameter.
[0096] In an embodiment of the present invention, by way of example, the target control parameter corresponding to the target control instruction may be the target control parameter corresponding to the instruction for controlling an air conditioner to implement functions such as temperature change and wind direction adjustment, or the target control parameter corresponding to the instruction for controlling a smart TV to implement functions such as volume adjustment and channel switching, or the target control parameter corresponding to the instruction for controlling a smart home system to implement functions such as curtain opening and closing, lighting adjustment, and music switching. The embodiments of the present invention do not make limitations.
[0097] In an embodiment of the present invention, optionally, the new offline voice interaction control method may be implemented based on a sensing module, a voice module, an MCU, and an electronic control unit. The sensing module senses the surrounding environment, the voice module performs sound collection and parsing, and the electronic control unit controls to implement the specified function. The data interaction between each functional unit is completed by the MCU. The MCU may be a physical entity or a logical unit. Further optionally, the new offline voice interaction control method may be implemented based on a sensing module, a voice module, and an MCU. The sensing module senses the environment where the offline voice device is located, the voice module performs sound collection and parsing, and at least one or a combination of the sensing module, the voice module, and the MCU controls to implement the specified function. The data interaction between each functional unit is completed by the MCU. Further optionally, the new offline voice interaction control method may be implemented based on a sensing module and a voice module. The sensing module senses the surrounding environment, the voice module performs sound collection and parsing, and at least one or a combination of the sensing module and the voice module controls to implement the specified function.
[0098] It can be seen that the method described in the embodiment of the present invention can perform a sound reception detection operation when personal characteristic information is detected and obtained, and determine the corresponding target control instruction based on the received sound characteristic information. This can not only improve the accuracy of offline voice detection, but also improve the reliability and accuracy of offline voice interaction control operations, thereby helping to improve the user experience in voice interaction control.
[0099] In an optional embodiment, determining the target control instruction corresponding to the sound feature information according to the sound feature information in the above step 103 may include:
[0100] Parsing the sound feature information to obtain parsed sound feature information;
[0101] According to the analyzed sound feature information, determining whether the analyzed sound feature information contains target voice information;
[0102] When it is determined that the analyzed sound feature information contains the target voice information, the target control instruction corresponding to the sound feature information is determined based on all control instructions corresponding to the target voice information in the analyzed sound feature information.
[0103] In this optional embodiment, the above-mentioned step of parsing the sound feature information is implemented by an AI offline voice model pre-stored locally. Optionally, the AI offline voice model can be formed by pre-training a deep neural network model based on a Tandem structure, a Tandem structure or a Grapheme structure, or can be formed based on a generative pre-trained Transformer model, or can be formed based on other methods, which is not limited in the embodiments of the present invention.
[0104] In this optional embodiment, the target voice information is voice information that matches at least one control instruction in the pre-stored control instruction set; illustratively, the pre-stored control instruction set may include an instruction set on the air-conditioning equipment for controlling the air-conditioning to realize functions such as temperature change and wind direction adjustment; on the smart TV equipment, it may include an instruction set for controlling the smart TV to realize functions such as volume adjustment and channel switching; in the smart home system, it may include an instruction set for controlling functions such as curtain opening and closing, lighting adjustment, and music switching, which is not limited to the embodiments of the present invention.
[0105] It can be seen that this optional embodiment can parse the voice feature information and determine whether the parsed voice feature information contains the target voice information. When the determination result is yes, all control instructions corresponding to the target voice information in the parsed voice feature information are used to determine the target control instruction corresponding to the voice feature information. In this way, the accuracy of determining the target control instruction based on the voice feature information can be improved, thereby improving the accuracy of offline voice interaction control and facilitating the improvement of the user experience.
[0106] In another optional embodiment, before parsing the voice feature information in the above step to obtain the parsed voice feature information, the method further includes:
[0107] Detect whether the voice feature information contains a confirmation voice command;
[0108] When it is detected that the voice feature information contains a confirmation voice command, trigger the operation of parsing the voice feature information to obtain the parsed voice feature information.
[0109] In this optional embodiment, the confirmation voice command is a voice command that matches at least one confirmation command in the pre-stored confirmation command set; optionally, the confirmation voice command can be a voice command preset by the user, or a voice command preset by the manufacturer or operator of the device / system / server integrated with the offline voice interaction control function. The embodiments of the present invention do not make limitations. Exemplarily, the user can customize a common nickname as the confirmation voice command, and the manufacturer or operator can pre-set "Hello" as the confirmation voice command. Optionally, when no confirmation voice command is detected within the preset confirmation voice command detection time range, the relevant (sub) modules for functions such as voice interaction can be controlled to enter the sleep state.
[0110] It can be seen that this optional embodiment can detect whether the voice feature information contains a confirmation voice command. Under the condition that it is detected that the voice feature information contains a confirmation voice command, parsing the voice feature information can confirm whether the voice feature information comes from the user and whether the voice feature information is the voice for the user to implement the voice interaction function. In this way, the occurrence of false wake-up situations in complex environments can be reduced, the accuracy of offline voice interaction control can be further improved, and the user experience can be facilitated to be improved.
[0111] In yet another optional embodiment, before detecting whether the voice feature information contains a confirmation voice command in the above step, the method further includes:
[0112] Detect whether the voiceprint parameters corresponding to the voice feature information match at least one voiceprint parameter in the pre-set voiceprint parameter set;
[0113] When it is detected that the voiceprint parameter corresponding to the voice feature information matches at least one voiceprint parameter in the preset voiceprint parameter set, an operation of detecting whether the voice feature information contains a confirmation voice command is triggered.
[0114] It can be seen that in this optional embodiment, under the condition that it is detected that the voiceprint parameter corresponding to the voice feature information matches at least one voiceprint parameter in the preset voiceprint parameter set, it is possible to detect whether the voice feature information contains a confirmation voice command, which can further confirm whether the voice feature information comes from a preset user, restricts the interaction between non-designated users and the system / device / server integrated with the offline voice interaction control function, can further improve the accuracy of offline voice interaction control, and thus is beneficial to improving the user experience.
[0115] In yet another optional embodiment, the personnel feature information includes the personnel location information of the user; before the operation of detecting the voice reception of the user in step 102 above to obtain a reception detection result, the method further includes:
[0116] According to the personnel location information, determine the area parameters of the sound reception area that matches the personnel location information;
[0117] According to the area parameters, determine the target sound reception device that matches the area parameters from the preset set of sound reception devices, and determine the sound reception demand state of the target sound reception device that matches the area parameters;
[0118] Among them, the operation of detecting the voice reception of the user in step 102 above to obtain a reception detection result includes:
[0119] Based on the sound reception demand state of the target sound reception device, perform an operation of detecting the voice reception of the user to obtain a reception detection result.
[0120] In this optional embodiment, by way of example, the target sound reception device that matches the area parameters may be the sound reception device in the set of sound reception devices that matches the corresponding direction of the area parameters; correspondingly, the determination of the sound reception demand state of the target sound reception device that matches the area parameters in the above step may be to increase the sound reception sensitivity of the sound reception sensor corresponding to the target sound reception device, and based on the increased sound reception sensitivity of the sound reception sensor corresponding to the target sound reception device, determine the sound reception demand state; correspondingly, the operation of detecting the voice reception of the user based on the sound reception demand state of the target sound reception device in the above step may be to perform an operation of detecting the voice reception of the user based on the sound reception demand state determined based on the adjusted sound reception sensitivity.
[0121] It can be seen that the optional embodiment can determine the regional parameters of the sound collection area matching the personnel position information according to the personnel position information, determine the target sound collection device matching the regional parameters from the preset sound collection device set, and perform a voice reception detection operation on the user based on the sound collection requirement state of the target sound collection device, which can improve the reception success rate and accuracy of the sound collection device during the reception process of sound feature information, reduce the noise interference in complex environments, and further improve the accuracy of offline voice interaction control, thus being beneficial to improving the user experience.
[0122] Embodiment 2
[0123] Please refer to Figure 3 , Figure 3 which is a schematic flowchart of another offline voice interaction control method disclosed in the embodiments of the present invention. Among them, Figure 3 the described method can be applied to various scenario types, such as office scenarios, entertainment scenarios, home scenarios, etc., and the embodiments of the present invention do not make limitations. Further, the method can be applied to control various devices through voice recognition and interaction control in the foregoing scenario types, such as air conditioners, washing machines, refrigerators, TVs, floor sweeping robots, table lamps, surveillance cameras, speakers, water heaters, air purifiers, projectors, routers, rice cookers, etc. And optionally, the method can be implemented by an offline voice interaction control device, where the offline voice interaction control device can be a local server for processing the offline voice interaction control process, or a control module for implementing the offline voice interaction control process, and the embodiments of the present invention do not make limitations. As Figure 3 shown, the offline voice interaction control method includes the following operations:
[0124] 201. Detect whether there is a target sensing device set that can be in a normal working state in the preset sensing device set for receiving the personnel feature information of a certain user, and obtain a device state detection result.
[0125] In the embodiments of the present invention, the set of sensing devices is a set of all sensing devices preset for receiving the personnel feature information of a certain user, and the set of target sensing devices is a set of all sensing devices that can be in a normal working state. The normal working state means the working state in which the sensing device can normally receive the personnel feature information. Optionally, all the sensing devices in the set of sensing devices can be deployed on the outer surface or inside of a device with voice interaction control function in a centralized or distributed manner, or within a preset space range of a device / system / server with voice interaction control function, or at any position where the device / system / server with voice interaction control function needs to perform the sensing function. The embodiments of the present invention do not make any limitations. Further optionally, the sensing devices in the set of sensing devices can be low-power passive sensors or modules, such as PIR (Passive Infrared Sensor) etc.; or low-cost active sensing sensors, such as modules based on the CSI (Channel State Information) sensing ability of WIFI etc.; or multifunctional active sensing sensors, such as indoor microwave radar sensors etc. The embodiments of the present invention do not make any limitations. Optionally, when the device state detection result indicates that there is no set of target sensing devices that can be in a normal working state in the set of sensing devices, directly start receiving the voice feature information. Further optionally, the working range of the sensing device is greater than the collection range of the sound collection of the sound receiving device, or through the parameter design of the sensing device and / or the sound receiving device, the sensing device detects the user before the sound receiving device works.
[0126] 202. Determine the target sensing requirement state according to the device state detection result.
[0127] In the embodiments of the present invention, the target sensing requirement state represents the state that needs to perform personnel sensing determined based on the device state detection result.
[0128] 203. According to the target sensing requirement state, detect whether the personnel feature information of a certain user is obtained.
[0129] 204. When it is detected that the personnel feature information is obtained, perform a voice reception detection operation on the user to obtain a reception detection result.
[0130] 205. When the reception detection result is used to indicate that the voice feature information of the user is received, determine the target control instruction corresponding to the voice feature information according to the voice feature information.
[0131] 206. According to the target control instruction, generate the target control parameter corresponding to the target control instruction, and execute the function matching the target control parameter.
[0132] In the embodiments of the present invention, for other descriptions of steps 204, 205 and 206, please refer to the detailed descriptions of steps 102, 103 and 104 in Embodiment 1, and the embodiments of the present invention will not be elaborated herein.
[0133] It can be seen that implementing the embodiments of the present invention can detect whether there is a target set of sensing devices that can be in a normal working state in the set of sensing devices before detecting whether the personnel feature information is obtained, determine the target sensing demand state, and detect whether the personnel feature information of a certain user is obtained according to the target sensing demand state. In this way, the acquisition operation of the personnel feature information can cope with diverse sensing scenarios, improving the stability and applicability of the acquisition operation of the personnel feature information, thereby improving the reliability and accuracy of the offline voice interaction control, and thus being beneficial to improving the user experience in the voice interaction control.
[0134] In an optional embodiment, the determining the target sensing demand state according to the device state detection result in step 202 above may include:
[0135] When the device state detection result indicates that there is a target set of sensing devices that can be in a normal working state in the set of sensing devices, obtain the multi-dimensional scene parameters of the target scene where the target set of sensing devices is located;
[0136] Based on the multi-dimensional scene parameters of the target scene where the target set of sensing devices is located, determine the target sensing devices that meet the preset first sensing condition from the target set of sensing devices, and determine the sensing demand state of the target sensing devices as the target sensing demand state.
[0137] In this optional embodiment, optionally, the multi-dimensional scene parameters of the target scene may include one or more combinations of the environmental parameters, spatial layout parameters, network quality parameters, associated device parameters, and overall energy consumption parameters of the target scene, and the embodiments of the present invention do not make any limitations.
[0138] Further optionally, the first sensing condition may include one or more combinations of the environmental parameter condition, spatial layout parameter condition, network quality parameter condition, associated device parameter condition, and overall energy consumption parameter condition, and the embodiments of the present invention do not make any limitations. Exemplarily, if the first sensing condition is the spatial layout parameter condition, the determining the target sensing devices that meet the preset first sensing condition from the target set of sensing devices in the above steps can be understood as determining the target sensing devices that meet the spatial layout parameter condition from the target set of sensing devices.
[0139] It can be seen that the optional embodiment can determine the target perception demand state according to the device state detection result. When the device state detection result indicates that there is a target perception device set that can be in a normal working state in the perception device set, multi-dimensional scene parameters of the target scene where the target perception device set is located are obtained. Based on the multi-dimensional scene parameters of the target scene where the target perception device set is located, target perception devices that meet the preset first perception condition are determined from the target perception device set, so that the acquisition of personnel feature information can be obtained based on the perception device that best matches the target scene, improving the applicability and accuracy of personnel feature information acquisition in complex scenes, and further improving the reliability, accuracy, and efficiency of offline voice interaction control.
[0140] In another optional embodiment, the determining the target perception demand state according to the device state detection result in step 202 above may include:
[0141] When the device state detection result indicates that there is no target perception device set that can be in a normal working state in the perception device set, a target relay perception device set is determined.
[0142] The perception demand state corresponding to each target relay perception device is determined as the target perception demand state.
[0143] In this optional embodiment, optionally, the target relay perception device set includes at least one target relay perception device that can be used for personnel feature perception. The target relay perception device can be understood as a device within the application scenario of this offline voice interaction control method that can implement the personnel perception function or a device integrated with a device that can implement the personnel perception function. Exemplarily, if there is no target perception device set that can be in a normal working state in an air conditioner integrated with the offline voice interaction control function, devices such as a smart TV and a floor sweeper integrated with the device that can implement the personnel perception function can be determined from the application scenario as the target relay perception device set.
[0144] In this optional embodiment, further optionally, the interaction between the target relay perception device and the device / system / server integrated with the offline voice interaction control function that the user wants to interact with can be carried out by means of Bluetooth or network, etc., which is not limited in the embodiments of the present invention.
[0145] It can be seen that when the device status detection result indicates that there is no target sensing device set that can be in a normal working state in the set of sensing devices, the target relay sensing device set can be determined, and the sensing requirement status corresponding to each target relay sensing device is determined as the target sensing requirement status, so that the acquisition of personnel feature information can continue in special cases such as when the set of sensing devices cannot work properly, improving the stability and reliability of the acquisition of personnel feature information, and further improving the applicability, accuracy and reliability of offline voice interaction control.
[0146] In another alternative embodiment, determining the target relay sensing device set in the above steps may include:
[0147] Obtain the multi-dimensional device information of all preset relay sensing devices;
[0148] According to the multi-dimensional device information corresponding to each relay sensing device, determine all relay sensing devices that meet the preset second sensing conditions from all relay sensing devices as the target relay sensing device set.
[0149] In this alternative embodiment, optionally, the multi-dimensional device information may include one or more combinations of the working status information, working range information, sensing type information, and transmission type information corresponding to each relay sensing device, which is not limited in this embodiment. Further optionally, the second sensing conditions include one or more combinations of personnel feature sensing conditions, working idle state conditions, working priority conditions, and information transmission conditions, which is not limited in this embodiment.
[0150] In this alternative embodiment, by way of example, when there is no target sensing device set that can be in a normal working state in the set of sensing devices corresponding to an air conditioner integrated with an offline voice interaction control function that the user wants to interact with, the sensing devices corresponding to a smart TV that meet the second sensing conditions can be determined as relay sensing devices; the second sensing conditions may be a combination of personnel sensing conditions, working idle state conditions, and information transmission conditions. The personnel sensing device can be understood as whether the relay sensing device or the system / device / server integrating the relay sensing device can perform personnel information sensing. The working idle state condition can be understood as whether the relay sensing device or the system / device / server integrating the relay sensing device is idle and can be used as a relay sensing device for personnel information sensing. The information transmission condition can be understood as whether the relay sensing device or the system / device / server integrating the relay sensing device can perform data transmission of information such as personnel sensing information with the system / device / server integrating the offline voice interaction control function that the user wants to interact with via Bluetooth, network, etc.
[0151] It can be seen that this optional embodiment can further obtain multi-dimensional device information of all preset relay sensing devices; according to the multi-dimensional device information corresponding to each relay sensing device, all relay sensing devices that meet the preset second sensing condition are determined from all relay sensing devices as the target relay sensing device set, so that when there is no target sensing device set that can be in a normal working state in the sensing device set, the acquisition of personnel feature information can be based on the most matching relay sensing device, improving the reliability of the acquisition of personnel feature information, and further improving the applicability, accuracy, and efficiency of offline voice interaction control, thereby enhancing the user experience in voice interaction control.
[0152] Embodiment III
[0153] Please refer to Figure 4 , Figure 4 which is a schematic structural diagram of an offline voice interaction control device disclosed in an embodiment of the present invention. Among them, Figure 4 the described device can be applied to various scenario types, such as office scenarios, entertainment scenarios, home scenarios, etc., and the embodiments of the present invention do not make limitations. Further, the device can be applied to control various devices through voice recognition and interaction control in the foregoing scenario types, such as air conditioners, washing machines, refrigerators, televisions, floor sweeping robots, table lamps, monitoring cameras, speakers, water heaters, air purifiers, projectors, routers, rice cookers, etc. As Figure 4 shown, the offline voice interaction control device may include:
[0154] A detection module 301, configured to detect whether personnel feature information of a certain user is obtained; when it is detected that the personnel feature information is obtained, perform a voice reception detection operation on the user to obtain a reception detection result;
[0155] A first determination module 302, configured to, when the reception detection result is used to indicate that voice feature information of the user is received, determine a target control instruction corresponding to the voice feature information according to the voice feature information;
[0156] An execution module 303, configured to generate target control parameters corresponding to the target control instruction according to the target control instruction, and execute a function matching the target control parameters.
[0157] It can be seen that implementing the device described in the embodiments of the present invention can perform a voice reception detection operation when it is detected that the personnel feature information is obtained, and determine a corresponding target control instruction according to the received voice feature information. This can not only improve the accuracy of offline voice detection, but also enhance the reliability and accuracy of offline voice interaction control operations, thereby being beneficial to enhancing the user experience in voice interaction control.
[0158] In an optional embodiment, as Figure 5 shown, the first determination module 302 includes:
[0159] A parsing sub-module 3021, configured to parse the voice feature information to obtain the parsed voice feature information;
[0160] A judgment sub-module 3022, configured to judge whether the parsed voice feature information contains target voice information according to the parsed voice feature information;
[0161] A determination sub-module 3023, configured to, when the judgment sub-module 3022 determines that the parsed voice feature information contains target voice information, determine a target control instruction corresponding to the voice feature information according to all control instructions corresponding to the target voice information in the parsed voice feature information.
[0162] In this optional embodiment, the target voice information is voice information that matches at least one control instruction in the pre-stored control instruction set;
[0163] It can be seen that the device described in implementing this optional embodiment can parse the voice feature information, and judge whether the parsed voice feature information contains target voice information. When the judgment result is yes, according to all control instructions corresponding to the target voice information in the parsed voice feature information, determine the target control instruction corresponding to the voice feature information. In this way, the accuracy of determining the target control instruction based on the voice feature information can be improved, and then the accuracy of offline voice interaction control is improved, which is beneficial to enhancing the user experience.
[0164] In another optional embodiment, as Figure 5 shown, the first determination module 302 further includes:
[0165] A detection sub-module 3024, configured to detect whether the voice feature information contains a confirmation voice instruction before the parsing sub-module 3021 parses the voice feature information to obtain the parsed voice feature information; when it is detected that the voice feature information contains a confirmation voice instruction, trigger the parsing sub-module 3021 to perform the operation of parsing the voice feature information to obtain the parsed voice feature information.
[0166] In this optional embodiment, the confirmation voice instruction is a voice instruction that matches at least one confirmation instruction in the pre-stored confirmation instruction set.
[0167] It can be seen that the device described in this alternative embodiment can detect whether the voice feature information contains a confirmation voice command. Under the condition that it is detected that the voice feature information contains a confirmation voice command, the voice feature information can be parsed to confirm whether the voice feature information comes from the user and whether the voice feature information is the voice for the user to implement the voice interaction function. In this way, the occurrence of false wake-up situations in complex environments can be reduced, the accuracy of offline voice interaction control can be further improved, and thus the user experience can be enhanced.
[0168] In yet another alternative embodiment, the personnel feature information includes the personnel location information of the user;
[0169] The device further includes:
[0170] A second determination module 304, configured to, before the detection module 301 performs a voice reception detection operation on the user to obtain a reception detection result, determine area parameters of a sound reception area that match the personnel location information according to the personnel location information; according to the area parameters, determine a target sound reception device that matches the area parameters from a preset set of sound reception devices, and determine the sound reception requirement status of the target sound reception device that matches the area parameters;
[0171] Among them, the specific manner in which the detection module 301 performs a voice reception detection operation on the user to obtain a reception detection result includes:
[0172] Performing a voice reception detection operation on the user based on the sound reception requirement status of the target sound reception device to obtain a reception detection result.
[0173] It can be seen that the device described in this alternative embodiment can determine area parameters of a sound reception area that match the personnel location information according to the personnel location information, determine a target sound reception device that matches the area parameters from a preset set of sound reception devices, and perform a voice reception detection operation on the user based on the sound reception requirement status of the target sound reception device, which can improve the reception success rate and accuracy of the sound reception device during the process of receiving voice feature information, reduce noise interference in complex environments, and further improve the accuracy of offline voice interaction control, thereby being beneficial to enhancing the user experience.
[0174] In yet another alternative embodiment, as Figure 5 shown, the detection module 301 is further configured to:
[0175] Before detecting whether the personnel feature information of a certain user is obtained, detect whether there is a target set of sensing devices that can be in a normal working state in a preset set of sensing devices for receiving the personnel feature information of a certain user to obtain a device state detection result;
[0176] The device further includes:
[0177] A third determination module 305, configured to determine a target perception requirement state according to the device state detection result;
[0178] Among them, the specific manner in which the detection module 301 detects whether the personnel feature information of a certain user is obtained includes:
[0179] According to the target perception requirement state, detect whether the personnel feature information of a certain user is obtained.
[0180] In this optional embodiment, the normal working state represents the working state in which the perception device can normally receive the personnel feature information;
[0181] It can be seen that the device described in this optional embodiment can detect whether there is a target perception device set that can be in a normal working state in the perception device set before detecting whether the personnel feature information is obtained, determine the target perception requirement state, and according to the target perception requirement state, detect whether the personnel feature information of a certain user is obtained. In this way, the acquisition operation of the personnel feature information can cope with diverse perception scenarios, improve the stability and applicability of the acquisition operation of the personnel feature information, and further improve the reliability and accuracy of the offline voice interaction control, thereby being beneficial to improving the user experience in the voice interaction control.
[0182] In another optional embodiment, as Figure 5 shown, the specific manner in which the third determination module 305 determines the target perception requirement state according to the device state detection result includes:
[0183] When the device state detection result indicates that there is a target perception device set that can be in a normal working state in the perception device set, obtain the multi-dimensional scene parameters of the target scene where the target perception device set is located;
[0184] Based on the multi-dimensional scene parameters of the target scene where the target perception device set is located, determine the target perception devices that meet the preset first perception condition from the target perception device set, and determine the perception requirement state of the target perception devices as the target perception requirement state.
[0185] In this optional embodiment, the multi-dimensional scene parameters of the target scene include one or a combination of the environmental parameters, spatial layout parameters, network quality parameters, associated device parameters, and overall energy consumption parameters of the target scene.
[0186] It can be seen that the device described in implementing this optional embodiment can determine the target perception demand state according to the device state detection result. When the device state detection result indicates that there is a target perception device set in the perception device set that can be in a normal working state, multi-dimensional scene parameters of the target scene where the target perception device set is located are obtained. Based on the multi-dimensional scene parameters of the target scene where the target perception device set is located, target perception devices that meet the preset first perception condition are determined from the target perception device set, so that the acquisition of personnel feature information can be obtained based on the perception device that best matches the target scene, improving the applicability and accuracy of personnel feature information acquisition in complex scenes, and further improving the reliability, accuracy, and efficiency of offline voice interaction control.
[0187] In another optional embodiment, as Figure 5 shown, the specific manner in which the third determination module 305 determines the target perception demand state according to the device state detection result includes:
[0188] When the device state detection result indicates that there is no target perception device set in the perception device set that can be in a normal working state, a target relay perception device set is determined;
[0189] The perception demand state corresponding to each target relay perception device is determined as the target perception demand state.
[0190] In this optional embodiment, the target relay perception device set includes at least one target relay perception device that can be used for personnel feature perception.
[0191] It can be seen that the device described in implementing this optional embodiment can determine a target relay perception device set when the device state detection result indicates that there is no target perception device set in the perception device set that can be in a normal working state, and determine the perception demand state corresponding to each target relay perception device as the target perception demand state, so that the acquisition of personnel feature information can continue in special cases such as when the perception device set cannot work properly, improving the stability and reliability of personnel feature information acquisition, and further improving the applicability, accuracy, and reliability of offline voice interaction control.
[0192] In another optional embodiment, as Figure 5 shown, the specific manner in which the third determination module 305 determines the target relay perception device set includes:
[0193] Obtain the multi-dimensional device information of all preset relay perception devices; according to the multi-dimensional device information corresponding to each relay perception device, determine all relay perception devices that meet the preset second perception condition from all relay perception devices as the target relay perception device set.
[0194] In this optional embodiment, the multi-dimensional device information includes at least one of the working status information, working range information, sensing type information, and transmission type information corresponding to each relay sensing device; the second sensing condition includes at least one of the personnel feature sensing condition, working idle status condition, working priority condition, and information transmission condition.
[0195] It can be seen that the device described in implementing this optional embodiment can further obtain the multi-dimensional device information of all preset relay sensing devices; according to the multi-dimensional device information corresponding to each relay sensing device, all relay sensing devices that meet the preset second sensing condition are determined from all relay sensing devices as the target relay sensing device set, so that when there is no target sensing device set that can be in a normal working state in the sensing device set, the acquisition of personnel feature information can be based on the most matching relay sensing device, improving the reliability of the acquisition of personnel feature information, and further improving the applicability, accuracy, and efficiency of offline voice interaction control, thereby enhancing the user experience in voice interaction control.
[0196] Embodiment Four
[0197] Please refer to Figure 6 , Figure 6 which is a schematic structural diagram of another offline voice interaction control device disclosed in the embodiments of the present invention. As Figure 6 shown, the offline voice interaction control device may include:
[0198] A memory 401 storing executable program code;
[0199] A processor 402 coupled to the memory 401;
[0200] The processor 402 calls the executable program code stored in the memory 401 and executes the steps in the offline voice interaction control method described in Embodiment One or Embodiment Two of the present invention.
[0201] Embodiment Five
[0202] The embodiments of the present invention disclose a computer storage medium that stores computer instructions, which are used to execute the steps in the offline voice interaction control method described in Embodiment One or Embodiment Two of the present invention when the computer instructions are called.
[0203] Embodiment Six
[0204] The embodiments of the present invention disclose a computer program product, which includes a non-transitory computer-readable storage medium storing a computer program, and the computer program is operable to cause a computer to execute the steps in the offline voice interaction control method described in Embodiment One or Embodiment Two.
[0205] The system embodiments described above are merely illustrative. The modules described as separate components may or may not be physically separated. The components shown as modules may or may not be physical modules, that is, they may be located in one place or distributed to multiple network modules. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. A person of ordinary skill in the art can understand and implement it without creative work.
[0206] Through the specific descriptions of the above embodiments, those skilled in the art can clearly understand that each implementation manner can be realized by means of software plus a necessary general hardware platform, and of course, it can also be realized by hardware. Based on such an understanding, the essence of the above technical solution, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium. The storage medium includes read-only memory (ROM), random access memory (RAM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), one-time programmable read-only memory (OTPROM), electrically-erasable programmable read-only memory (EEPROM), compact disc read-only memory (CD-ROM) or other optical disc memories, magnetic disk memories, tape memories, or any other computer-readable medium capable of carrying or storing data.
[0207] Finally, it should be noted that: The offline voice interaction control method and device disclosed in the embodiments of the present invention only disclose the preferred embodiments of the present invention, which are only used to illustrate the technical solutions of the present invention, rather than to limit them; Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: They can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; And these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. An offline voice interaction control method, characterized in that, the method includes: detecting whether personnel feature information of a certain user is obtained; when it is detected that the personnel feature information is obtained, performing a voice reception detection operation on the user to obtain a reception detection result; when the reception detection result is used to indicate that voice feature information of the user is received, determining a target control instruction corresponding to the voice feature information according to the voice feature information; generating a target control parameter corresponding to the target control instruction according to the target control instruction, and executing a function matching the target control parameter.
2. The offline voice interaction control method according to claim 1, characterized in that, the determining a target control instruction corresponding to the voice feature information according to the voice feature information includes: analyzing the voice feature information to obtain analyzed voice feature information; judging whether the analyzed voice feature information contains target voice information according to the analyzed voice feature information; the target voice information is voice information matching at least one control instruction in a pre-stored control instruction set; when it is judged that the analyzed voice feature information contains the target voice information, determining a target control instruction corresponding to the voice feature information according to all the control instructions corresponding to the target voice information in the analyzed voice feature information.
3. The offline voice interaction control method according to claim 2, characterized in that, before the analyzing the voice feature information to obtain analyzed voice feature information, the method further includes: detecting whether the voice feature information contains a confirmation voice instruction; the confirmation voice instruction is a voice instruction matching at least one confirmation instruction in a pre-stored confirmation instruction set; when it is detected that the voice feature information contains the confirmation voice instruction, triggering the operation of analyzing the voice feature information to obtain analyzed voice feature information.
4. The offline voice interaction control method according to any one of claims 1-3, characterized in that, the personnel feature information includes the personnel position information of the user; before the performing a voice reception detection operation on the user to obtain a reception detection result, the method further includes: determining area parameters of a sound collection area matching the personnel position information according to the personnel position information; determining a target sound collection device matching the area parameters from a preset sound collection device set, and determining a sound collection requirement state of the target sound collection device matching the area parameters; wherein, the performing a voice reception detection operation on the user to obtain a reception detection result includes: performing a voice reception detection operation on the user based on the sound collection requirement state of the target sound collection device to obtain a reception detection result.
5. The offline voice interaction control method according to claim 1, characterized in that, before the detecting whether personnel feature information of a certain user is obtained, the method further includes: Detect whether there is a target set of sensing devices in the set of sensing devices preset to receive the personal characteristic information of a certain user that can be in a normal working state, and obtain a device status detection result; the normal working state means the working state in which the sensing device can normally receive the personal characteristic information; Determine the target sensing demand state according to the device status detection result; Among them, the detection of whether the personal characteristic information of a certain user is obtained includes: According to the target sensing demand state, detect whether the personal characteristic information of a certain user is obtained.
6. The offline voice interaction control method according to claim 5, characterized in that, The determination of the target sensing demand state according to the device status detection result includes: When the device status detection result indicates that there is a target set of sensing devices in the set of sensing devices that can be in the normal working state, obtain the multi-dimensional scene parameters of the target scene where the target set of sensing devices is located; the multi-dimensional scene parameters of the target scene include one or more combinations of environmental parameters, spatial layout parameters, network quality parameters, associated device parameters, and overall energy consumption parameters of the target scene; Based on the multi-dimensional scene parameters of the target scene where the target set of sensing devices is located, determine the target sensing devices that meet the preset first sensing conditions from the target set of sensing devices, and determine the sensing demand state of the target sensing devices as the target sensing demand state.
7. The offline voice interaction control method according to claim 5, characterized in that, The determination of the target sensing demand state according to the device status detection result includes: When the device status detection result indicates that there is no target set of sensing devices in the set of sensing devices that can be in the normal working state, determine a target set of relay sensing devices; the target set of relay sensing devices includes at least one target relay sensing device that can be used for personal characteristic sensing; Determine the sensing demand state corresponding to each target relay sensing device as the target sensing demand state.
8. The offline voice interaction control method according to claim 7, characterized in that, The determination of the target set of relay sensing devices includes: Obtain the multi-dimensional device information of all preset relay sensing devices, and the multi-dimensional device information includes at least one of the working state information, working range information, sensing type information, and transmission type information corresponding to each relay sensing device; According to the multi-dimensional device information corresponding to each relay sensing device, determine all relay sensing devices that meet the preset second sensing conditions from all the relay sensing devices as the target set of relay sensing devices, and the second sensing conditions include at least one of personal characteristic sensing conditions, working idle state conditions, working priority conditions, and information transmission conditions.
9. An offline voice interaction control device, characterized in that, The device includes: A detection module, configured to detect whether the personnel feature information of a certain user is obtained; when it is detected that the personnel feature information is obtained, a voice reception detection operation is performed on the user to obtain a reception detection result; A first determination module, configured to, when the reception detection result is used to indicate that the voice feature information of the user is received, determine a target control instruction corresponding to the voice feature information according to the voice feature information; An execution module, configured to generate a target control parameter corresponding to the target control instruction according to the target control instruction, and execute a function matching the target control parameter.
10. An offline voice interaction control device, characterized in that, the device includes: a memory storing executable program code; a processor coupled to the memory; the processor calls the executable program code stored in the memory and executes the offline voice interaction control method according to any one of claims 1-8.