Intelligent device control method and apparatus, computer device, medium, and product
By acquiring the reception time and formant of voice wake-up information, determining the formant of voice control information, counting the number of sub-information items, and performing semantic analysis, the problem of misrecognition in multi-person dialogues is solved, and the recognition accuracy is improved.
Patent Information
- Application Number
- CN202511279553.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-09
- Publication Date
- 2025-11-28
- Estimated Expiration
- 2045-09-09
AI Technical Summary
In multi-person conversations, speech recognition systems struggle to distinguish the voices of individual speakers, increasing the risk of misidentification.
By acquiring the reception time and formant of voice wake-up information, the formant of voice control information is determined, and voice control information is acquired within a preset time. The number of sub-information items is counted, semantic analysis is performed to determine the control semantics, and then converted into matching control commands.
It reduces the probability of false recognition, improves recognition accuracy, and ensures that control commands are highly matched with user intentions.
Smart Images

Figure CN120766680B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of intelligent control, in particular to an intelligent device control method and device, computer equipment, medium and product. BACKGROUND
[0002] In the existing intelligent device control technology, voice control has become an extremely convenient and widely used interaction method. Users can issue instructions without physical contact with the device, freeing their hands and enhancing the safety and convenience of use.
[0003] However, in a multi-person conversation, different speakers may speak at the same time, causing the voice signals to overlap. In this case, the voice recognition system has difficulty distinguishing the voices of each speaker, thereby increasing the risk of misrecognition. SUMMARY
[0004] Therefore, it is necessary to provide an intelligent device control method, device, computer equipment, medium and product to solve the technical problem of misrecognition in traditional multi-person conversations.
[0005] In a first aspect, the present application provides an intelligent device control method. The method comprises:
[0006] obtaining voice wake-up information received by an intelligent device in an environment;
[0007] determining a receiving time of the voice wake-up information and a first formant corresponding to the voice wake-up information;
[0008] obtaining voice control information received by the intelligent device in the environment within a preset time after the receiving time;
[0009] determining a second formant corresponding to the voice control information; the voice control information comprises a plurality of sub-information; the second formant comprises a sub-formant of each of the sub-information;
[0010] in a case where each of the sub-formants does not match the first formant, respectively counting the number of information contents of each of the sub-information based on the information contents, and determining a sub-information quantity corresponding to each of the information contents;
[0011] if a maximum quantity in the sub-information quantities satisfies a preset quantity condition, performing semantic analysis on an information content corresponding to the maximum quantity to determine a control semantic;
[0012] converting the control semantic into a control instruction matching the voice control information; the control instruction is used to drive the intelligent device to perform a corresponding control operation.
[0013] In one of the embodiments, the method further comprises:
[0014] In the case that there is a target resonance peak matching the first resonance peak in each of the sub-resonance peaks, performing semantic analysis on target sub-information corresponding to the target resonance peak to determine control semantics of the target sub-information;
[0015] Converting the control semantics into a control instruction matching the voice control information.
[0016] In one of the embodiments, the performing semantic analysis on the target sub-information corresponding to the target resonance peak to determine control semantics of the target sub-information comprises:
[0017] Extracting a semantic feature vector of the target sub-information corresponding to the target resonance peak;
[0018] Performing similarity matching between the semantic feature vector and a candidate instruction template in a preset instruction template library, and taking a target instruction template satisfying a similarity condition as the control semantics of the target sub-information.
[0019] In one of the embodiments, the method further comprises:
[0020] If the maximum number of the sub-information quantity does not satisfy the preset quantity condition, performing semantic analysis on each of the sub-information to determine information semantics of the sub-information;
[0021] Respectively performing quantity statistics on each of the information semantics to determine sub-information quantities corresponding to each of the information semantics;
[0022] In the case that there is a first information semantics in each of the information semantics, and the corresponding sub-information quantity of the first information semantics satisfies the preset quantity condition, converting the first information semantics into a control instruction matching the voice control information.
[0023] In one of the embodiments, the method further comprises:
[0024] In the case that each of the sub-resonance peaks does not match the first resonance peak, performing semantic analysis on each of the sub-information to determine information semantics of the sub-information;
[0025] Determining a sub-information weight of the sub-information based on an information volume of the sub-information;
[0026] For each of the information semantics, determining a semantic proportion of the information semantics according to the sub-information weight of the sub-information in the information semantics;
[0027] In a case where a semantic proportion of the second information semantics is greater than a preset proportion, the second information semantics is converted into a control instruction matching the voice control information.
[0028] In a second aspect, the present application further provides a smart device control apparatus. The apparatus comprises:
[0029] a voice wake-up information acquisition module configured to acquire voice wake-up information received by a smart device in an environment;
[0030] a first formant determination module configured to determine a receiving time of the voice wake-up information and a first formant corresponding to the voice wake-up information;
[0031] a voice control information acquisition module configured to acquire voice control information received by the smart device in the environment within a preset time after the receiving time;
[0032] a second formant determination module configured to determine a second formant corresponding to the voice control information; the voice control information comprises a plurality of sub-information; the second formant comprises a plurality of sub-formants respectively corresponding to the sub-information;
[0033] a quantity statistics module configured to, in a case where each of the sub-formants does not match the first formant, respectively perform quantity statistics on information content of each of the sub-information based on the information content, and determine a quantity of each of the information content corresponding to the sub-information;
[0034] a control semantics determination module configured to, if a maximum quantity in the quantities of the sub-information satisfies a preset quantity condition, perform semantic analysis on information content corresponding to the maximum quantity, and determine control semantics;
[0035] a control instruction determination module configured to convert the control semantics into a control instruction matching the voice control information; the control instruction is used to drive the smart device to perform a corresponding control operation.
[0036] In one of the embodiments, the apparatus further comprises a second control instruction determination module, comprising:
[0037] a control semantics determination unit configured to, in a case where a target formant matching the first formant exists in each of the sub-formants, perform semantic analysis on target sub-information corresponding to the target formant, and determine control semantics of the target sub-information;
[0038] a conversion unit configured to convert the control semantics into a control instruction matching the voice control information.
[0039] In a third aspect, the present application provides a computer device. The computer device comprises a memory and a processor, the memory stores a computer program, and the processor implements the steps of the method described above when executing the computer program.
[0040] In a fourth aspect, the present application provides a computer readable storage medium. The computer readable storage medium stores a computer program, and the computer program implements the steps of the method described above when executed by a processor.
[0041] In a fifth aspect, the present application provides a computer program product. The computer program product comprises a computer program, and the computer program implements the steps of the method described above when executed by a processor.
[0042] The intelligent device control method, device, computer device, medium and product described above set an explicit starting point for the subsequent voice control process by first acquiring voice wake-up information, reducing the false triggering caused by long-time indiscriminate reception of sound. Then, the receiving time of the voice wake-up information and the first formant corresponding to the voice wake-up information are determined. Recording the receiving time can provide a time reference for subsequent acquisition of voice control information within a preset time, and determining the first formant can extract the key features of the voice wake-up information. Within the preset time after the receiving time, the voice control information received by the intelligent device in the environment can enable the intelligent device to focus on the time period in which the user may issue voice instructions after waking up the device, avoiding the influence of irrelevant time period voice information on the recognition result. Further, the sub-formants of each sub-information in the voice control information are determined, wherein the sub-formants collectively constitute a second formant. When the sub-formant and the first formant do not match, the number of sub-information corresponding to each sub-information content is counted. If the maximum number meets the preset condition, the semantic analysis of the information content is performed to determine the control semantics, and finally converted into a matching control instruction to drive the intelligent device to perform the corresponding control operation. This can ensure that the determined control instruction is highly matched with the user's real intention, greatly reducing the probability of misrecognition and improving the recognition accuracy. BRIEF DESCRIPTION OF DRAWINGS
[0043] Figure 1 An application environment diagram of the intelligent device control method in an embodiment;
[0044] Figure 2 A flowchart of the intelligent device control method in an embodiment;
[0045] Figure 3 A flowchart of the control instruction determination step in an embodiment;
[0046] Figure 4 A flowchart of the semantic analysis step in an embodiment;
[0047] Figure 5 Flowchart of the control instruction determining step in another embodiment;
[0048] Figure 6 Flowchart of the control instruction determining step in another embodiment;
[0049] Figure 7 Flowchart of the intelligent device control method in another embodiment;
[0050] Figure 8 Structural block diagram of the intelligent device control apparatus in an embodiment;
[0051] Figure 9 Internal structure diagram of the computer device in an embodiment. DETAILED DESCRIPTION
[0052] In order to make the purposes, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and should not be used to limit the present application.
[0053] The intelligent device control method provided by the embodiments of the present application can be applied to an application environment as shown in Figure 1 . In the application environment, the server 102 communicates with the intelligent device 104 through a network. A data storage system can store data required to be processed by the server 102. The data storage system can be integrated on the server 102, or placed on a cloud or other network server. The server 102 can be implemented by an independent server or a server cluster composed of multiple servers. The intelligent device 104 refers to an electronic device that integrates advanced sensor technology, data processing technology, communication technology and artificial intelligence algorithm, can perceive environmental information, understand user intent, and make corresponding decisions and actions according to preset rules or autonomous learning results, and realize intelligent operation and management. For example, the intelligent device can be a smart switch, a smart toilet, a smart curtain, etc. For example, the server can be a server arranged on the intelligent device, or an independently running server.
[0054] In an embodiment, as shown in Figure 2 , an intelligent device control method is provided. Taking the server 102 in Figure 1 as an example for illustration, it should be understood that the method can also be executed by the intelligent device 104, or executed by the intelligent device 104 and the server 102 interactively. In this embodiment, the method includes the following steps:
[0055] Step S202, obtaining voice wake-up information received by the smart device in the environment.
[0056] The voice wake-up information is a specific voice signal issued by a user, used to wake up a smart device in a standby or low-power state, so that the smart device enters an active state in which it can receive subsequent voice instructions, for example, saying "Xiao X classmate" to a smart speaker.
[0057] Specifically, during normal operation, the smart device continuously turns on its built-in voice receiving module, which can always listen to the sound in the surrounding environment. When a user needs to use the smart device, the user will issue specific voice wake-up information to wake up the smart device. Once the voice receiving module of the smart device captures a voice signal that meets the preset wake-up condition, it will record it as voice wake-up information for subsequent processing. For example, when the smart speaker is in standby state, the user says "Xiao X classmate", the microphone of the speaker will receive this voice signal and recognize it as voice wake-up information.
[0058] Step S204, determining the receiving time of the voice wake-up information and the first formant corresponding to the voice wake-up information.
[0059] The receiving time refers to the specific time when the smart device receives the voice wake-up information, accurate to a certain time unit, used for subsequent correlation analysis in the time dimension with other voice information. The time unit can be seconds or milliseconds, for example. The formant is an important feature parameter of the voice signal in frequency spectrum analysis, which reflects the basic resonance frequency of the vocal tract when pronouncing. It can be understood that the formant distribution and characteristics of different users' voices are different. In this embodiment, by extracting the first formant corresponding to the voice wake-up information, the key acoustic characteristics of the voice wake-up information can be obtained.
[0060] Specifically, after obtaining the voice wake-up information, the specific time when the information is received can be recorded, providing a time reference for subsequent search for related voice control information within a certain time range. At the same time, the server can also process the voice wake-up information, decompose the voice signal in the frequency domain, and find its main formant frequency characteristics, and extract the first formant. By recording the receiving time and extracting the first formant, the preparation for subsequent accurate recognition and processing of voice control information is completed.
[0061] Step S206, obtaining voice control information received by the smart device in the environment within a preset time after the receiving time.
[0062] The preset time is a time length set in advance, and the time is counted from the time when the voice wake-up information is received. In this time period, the smart device will focus on and receive possible voice control information, avoiding interference and resource waste caused by long-time indiscriminate reception of sound. The voice control information is a voice instruction issued by the user after the smart device is woken up by voice, which is used to control the device to perform a specific operation, such as "set the air conditioner temperature to 25 degrees".
[0063] Specifically, the server can start a preset time counting window according to the recorded receiving time of the voice wake-up information. In this preset time, the voice receiving module of the smart device will remain highly sensitive and continuously monitor the sound in the surrounding environment, focusing on obtaining possible voice control information. The setting of the preset time is based on the actual use scene and user habits, which needs to ensure that the voice control instruction issued by the user can be captured, and also needs to avoid receiving too much irrelevant sound for a long time, which increases the processing burden and the probability of misrecognition. For example, if the preset time is set to 5 seconds, then within 5 seconds after receiving the voice wake-up information, the smart device will continuously capture and analyze the voice control information in the audio stream through beamforming, noise reduction, and voice activity detection technologies.
[0064] Step S208, determining a second formant corresponding to the voice control information.
[0065] The second formant is also a feature parameter in the frequency spectrum analysis of the voice signal, and like the first formant, it is also a manifestation of the resonance frequency when the vocal tract is pronouncing, but it focuses on reflecting other acoustic characteristics of the voice. By extracting the second formant, the key acoustic characteristics of the voice control information can be obtained. The voice control information includes multiple sub-information. The second formant includes the sub-formant of each sub-information. It can be understood that the sub-formant of each sub-information can be the same or different. Each sub-information is information issued by the user. For example, the multiple sub-information included in the voice control information can be: turn off the light, turn on the light, it is a little dark, turn on the light, etc. The sub-formant is the respective formant feature corresponding to each sub-information in the frequency spectrum analysis.
[0066] Specifically, when the smart device obtains the voice control information within the preset time, it will also use the frequency spectrum analysis technology to process this voice signal and extract the second formant. Like the first formant, the second formant is an important acoustic characteristic of the voice signal, and it can reflect the pronunciation characteristics and acoustic properties of the voice. By extracting the second formant, the smart device can obtain the key features of the voice control information, which provides data support for subsequent correlation analysis with the first formant of the voice wake-up information.
[0067] Step S210, in the case that each sub-resonance peak does not match the first resonance peak, the number of each sub-information corresponding to each information content is determined based on the information content of each sub-information.
[0068] The number of sub-information is the number of sub-information containing each information content in all sub-information.
[0069] Specifically, in the case that each sub-resonance peak does not match the first resonance peak, the information content of all sub-information is sorted and analyzed, and the number of each different information content appearing in the sub-information is counted, that is, the number of sub-information corresponding to each information content is determined. For example, the voice control information has three sub-information "turn on the air conditioner", "turn up the air conditioner temperature", and "turn on the TV", and the information content is "turn on the air conditioner", "turn up the air conditioner temperature", and "turn on the TV". After counting, the number of sub-information corresponding to "turn on the air conditioner" is 1, the number of sub-information corresponding to "turn up the air conditioner temperature" is 1, and the number of sub-information corresponding to "turn on the TV" is 1.
[0070] Step S212, if the maximum number in the number of each sub-information meets the preset number condition, the information content corresponding to the maximum number is subjected to semantic analysis to determine the control semantics.
[0071] The preset number condition is a number threshold set in advance, which is used to determine whether the number of sub-information corresponding to a certain information content is sufficient to determine whether it is the main operation intention.
[0072] The control semantics is obtained after semantic analysis, which can accurately express the semantic information of the user's operation intention on the intelligent device.
[0073] Specifically, after obtaining the number of sub-information corresponding to each information content, the maximum value in these numbers is found. Then, the maximum number is compared with the preset number condition. If the maximum number meets the preset number condition, that is, reaches or exceeds the threshold set in advance, it means that there are many sub-information containing this information content, which is likely to be the user's main operation intention. At this time, the intelligent device will perform semantic analysis on the information content corresponding to the maximum number to deeply understand the operation intention that the user wants to achieve, and determine the control semantics. For example, if the preset number condition is 2, the voice control information has three sub-information "turn on the air conditioner", "turn on the air conditioner heating mode", and "turn up the TV volume", and the number of sub-information related to "turn on the air conditioner" is 2, which meets the condition, so the semantic analysis is performed on the "turn on the air conditioner" related content.
[0074] Step S214, converting the control semantics into a control instruction matching the voice control information.
[0075] The control instruction is used to drive the intelligent device to perform the corresponding control operation.
[0076] Specifically, after determining the control semantics, the smart device converts the control semantics into specific control instructions according to the preset instruction conversion rules in the system.
[0077] The smart device control method sets an explicit starting point for the subsequent voice control process by first obtaining the voice wake-up information, reducing the risk of false triggering caused by long-term indiscriminate reception of sound. Then, the receiving time of the voice wake-up information and the first formant corresponding to the voice wake-up information are determined. Recording the receiving time provides a time reference for subsequent acquisition of voice control information within a preset time, and determining the first formant extracts the key features of the voice wake-up information. Within the preset time after the receiving time, the voice control information received by the smart device in the environment is acquired, which enables the smart device to focus on the time period when the user may issue voice instructions after waking up the device, avoiding the influence of irrelevant voice information on the recognition result in unrelated time periods. Further, the sub-formants of each sub-information in the voice control information are determined, wherein the sub-formants collectively constitute the second formant. When the sub-formant and the first formant do not match, the number of sub-information corresponding to each sub-information content is counted. If the maximum number meets the preset condition, the semantic analysis of the information content is performed to determine the control semantics, which is finally converted into a matching control instruction to drive the smart device to perform the corresponding control operation. This ensures that the determined control instruction is highly matched with the user's real intention, greatly reduces the probability of misrecognition, and improves the recognition accuracy.
[0078] In one embodiment, in the case where there is a target formant matching the first formant among the sub-formants, the control instruction matching the voice control information is determined based on the sub-information corresponding to the target formant. Specifically, as shown in Figure 3 The smart device control method further includes:
[0079] Step S302, in the case where there is a target formant matching the first formant among the sub-formants, the target sub-information corresponding to the target formant is subjected to semantic analysis to determine the control semantics of the target sub-information.
[0080] The control semantics is obtained after semantic analysis and can accurately express the semantic information of the user's operation intention for the smart device.
[0081] Specifically, when it is judged that there is a target resonance peak matching the first resonance peak in each sub-resonance peak, it indicates that the target sub-information corresponding to the target resonance peak and the voice wake-up information are most likely from the same user's instruction. At this time, the server can deeply analyze the target sub-information corresponding to the target resonance peak, and parse the vocabulary, syntax structure, etc. in the target sub-information, and extract the semantic content related to the operation intention that the user wants to realize, that is, determine the control semantics of the target sub-information. For example, the user says "it's cold, lower the temperature of the air conditioner a little", and the semantic analysis will identify the key semantic information such as "temperature adjustment". It can be understood that the number of target resonance peaks is at least one.
[0082] Step S304, converting the control semantics into a control instruction matching the voice control information.
[0083] The control instruction is converted from the control semantics and can directly drive the smart device to execute a specific command of the corresponding control operation. The form of the control instruction may be, for example, a code or a digital instruction that can be recognized by the smart device.
[0084] Specifically, after determining the control semantics of the target sub-information, the server will convert the control semantics into a specific control instruction according to the preset instruction conversion rule. These preset rules are set in advance in the device system, which clearly shows the device operation instructions corresponding to different control semantics. For example, when the control semantics is "air conditioner temperature adjustment", it will be converted into the corresponding control instruction according to the rule, which can be directly sent to the air conditioner device to drive it to execute the corresponding operation.
[0085] In this embodiment, by first judging the resonance peak matching to ensure the relevance of the source of the voice instruction, and then performing semantic analysis and instruction conversion, the user's intention can be more accurately recognized, the understanding and execution accuracy of the smart device to the voice control information can be improved, the misoperation can be reduced, and the user experience can be improved.
[0086] In one embodiment, as shown in Figure 4 the semantic analysis of the target sub-information corresponding to the target resonance peak to determine the control semantics of the target sub-information, comprising:
[0087] Step S402, extracting a semantic feature vector of the target sub-information corresponding to the target resonance peak.
[0088] The semantic feature vector is a group of numerical vectors that can represent the semantic core features extracted by analyzing and processing the target sub-information, and is used to quantify the semantic content of the target sub-information.
[0089] Specifically, after receiving the target sub-information, the server can analyze the target sub-information to mine key features that can represent the core semantics thereof. Then, the key features are quantitatively processed according to certain rules and orders, converted into a set of numerical values, and formed into a semantic feature vector.
[0090] In step S404, the semantic feature vector is subjected to similarity matching with candidate instruction templates in a preset instruction template library, and a target instruction template that satisfies a similarity condition is taken as the control semantics of the target sub-information.
[0091] The preset instruction template library is a database pre-stored in the intelligent device system and containing templates corresponding to various common voice control instructions, which provides a reference standard for subsequent semantic matching. The candidate instruction templates are various instruction templates in the preset instruction template library and used for similarity comparison with the extracted semantic feature vector. The similarity condition is a threshold standard pre-set for measuring the similarity degree between the semantic feature vector and the candidate instruction template. Only when the similarity reaches or exceeds the standard, is the matching considered successful. The target instruction template is a candidate instruction template that satisfies the similarity condition in the similarity matching process.
[0092] Specifically, the server can perform similarity calculation on the extracted semantic feature vector and each candidate instruction template in the preset instruction template library, obtain a similarity value between them by comparing the numerical differences between the semantic feature vector and the candidate instruction template in each dimension, and then compare the calculated similarity value with the pre-set similarity condition. If the similarity value of a certain candidate instruction template satisfies the similarity condition, i.e., reaches or exceeds the set threshold, this candidate instruction template is determined as the target instruction template and taken as the control semantics of the target sub-information. For example, there is a template of "turn on the living room lighting device" in the preset instruction template library, and when it is calculated that the similarity between the template and the semantic feature vector of "turn on the light in the living room" satisfies the condition, the template is taken as the control semantics.
[0093] In this embodiment, by extracting the semantic feature vector and performing similarity matching with the preset instruction template library, the user's ambiguous or diverse target sub-information can be quickly and accurately converted into standard and normative control semantics, the understanding ability of the intelligent device for voice instructions is improved, the operation errors caused by semantic understanding deviation are reduced, and the efficiency and accuracy of the user's interaction with the intelligent device are improved.
[0094] In one embodiment, as shown in FIG. 4, Figure 5 the method further includes:
[0095] Step S502, if the maximum number of the number of sub-information does not satisfy the preset number condition, then for each sub-information, the sub-information is subjected to semantic analysis to determine the information semantics of the sub-information.
[0096] The information semantics is obtained by semantic analysis of the sub-information and can accurately express the semantic content of the operation intention contained in the sub-information.
[0097] Specifically, when the number of sub-information corresponding to each information content in the voice control information is counted, it is found that the maximum number does not satisfy the preset number condition, which means that the main operation intention of the user cannot be determined only according to the information content. At this time, the server will change the processing mode, and for each sub-information, the vocabulary, syntax structure and semantic relationship of the sub-information are analyzed in depth, so as to determine the information semantics of each sub-information, that is, to determine what operation each sub-information specifically wants to implement. For example, the voice control information has three sub-information "turn on the air conditioner" "the sky is a little dark" "turn on the bedroom light", and the maximum number of sub-information corresponding to the information content does not reach the preset condition, so each sub-information is analyzed to obtain the information semantics "turn on the air conditioning device" "the sky is a little dark" "turn on the bedroom lighting device".
[0098] Step S504, the number of each information semantics is counted to determine the number of sub-information corresponding to each information semantics.
[0099] Specifically, after determining the information semantics of each sub-information, the intelligent device will sort and count these information semantics. The specific method is to count the number of each different information semantics appearing in all sub-information, that is, to determine the number of sub-information corresponding to each information semantics. For example, after step S702, the information semantics "turn on the air conditioner" "turn on the light" "turn on the bedroom light" is obtained.
[0100] Step S506, in the case where there is a first information semantics in the information semantics, the number of sub-information corresponding to which satisfies the preset number condition, the first information semantics is converted into a control instruction matched with the voice control information.
[0101] The first information semantics is the information semantics in the information semantics, the number of sub-information corresponding to which satisfies the preset number condition.
[0102] Specifically, after completing the quantity statistics of the sub-information corresponding to each information semantic, the smart device checks the quantity conditions. If in each information semantic, there is a sub-information corresponding to the information semantic quantity that meets the preset quantity condition, that is, reaches or exceeds the pre-set threshold, then the information semantic that meets the condition is determined as the first information semantic. The first information semantic represents the relatively concentrated operation intention of the user. Then, the smart device converts the first information semantic into a control instruction matched with the whole voice control information according to the pre-set instruction conversion rule in the system, so as to drive the smart device to perform the corresponding operation.
[0103] In the embodiment, when the quantity of the sub-information corresponding to the information content cannot meet the preset condition, further analysis and statistics are performed from the information semantic level, which can more accurately capture the main operation intention of the user, avoid instruction determination errors caused by inaccurate information content analysis, improve the processing ability and accuracy of the smart device for complex voice control information, and enhance the stability and reliability of the user interaction with the smart device.
[0104] In one embodiment, as shown in Figure 6 The smart device control method further includes:
[0105] In step S602, in the case that each sub-formant is not matched with the first formant, the semantic analysis of the sub-information is performed for each sub-information to determine the information semantic of the sub-information.
[0106] Specifically, in the case that each sub-formant is not matched with the first formant, after receiving the voice control information containing multiple sub-information, the semantic analysis algorithm in the natural language processing technology is used to analyze each sub-information. By analyzing the vocabulary, syntax structure and semantic relationship of the sub-information, the information semantic of each sub-information is accurately determined, that is, it is clear what operation each sub-information specifically wants to implement.
[0107] In step S604, the sub-information weight of the sub-information is determined based on the information volume of the sub-information.
[0108] The information volume is the sound intensity of the sub-information when the voice is input. The sub-information weight is a value determined according to the information volume of the sub-information, which is used to measure the relative importance of the sub-information in the whole voice control information.
[0109] Specifically, after determining the information semantics of each sub-information, the server obtains the information volume of each sub-information when the voice input. According to the preset corresponding relationship rule between the volume and the weight, the information volume is converted into the sub-information weight. Generally speaking, the larger the volume is, the higher the sub-information weight can be, which means that the sub-information is more important in the overall voice control information. For example, if the preset rule is that the weight increases by a certain value when the volume increases by a certain decibel, the sub-information weight of "adjusting the brightness of the living room lamp" is higher after conversion because the volume of the sub-information is larger.
[0110] In step S606, for each information semantics, the semantic proportion of the information semantics is determined according to the sub-information weights of the sub-information in the information semantics.
[0111] The semantic proportion is a value reflecting the importance of the information semantics in the overall voice control information, which is obtained by comprehensively considering the sub-information weights of all sub-information containing the information semantics.
[0112] Specifically, for each information semantics, the server finds all sub-information containing the information semantics. Then, the sub-information weights of these sub-information are summarized and calculated to determine the semantic proportion of the information semantics according to a specific calculation method (such as summation, weighted average, etc.). The semantic proportion reflects the importance of the information semantics in the overall voice control information. For example, there are two sub-information "adjusting the brightness of the living room lamp" and "adjusting the brightness of the living room lamp but adjusting the volume is larger" under the information semantics "adjusting the brightness of the living room lamp", and the semantic proportion of the information semantics is calculated according to the sub-information weights.
[0113] In step S608, if the semantic proportion of the second information semantics is greater than the preset proportion condition, the second information semantics is converted into a control instruction matching the voice control information.
[0114] The second information semantics is the information semantics whose semantic proportion is greater than the preset proportion condition among the information semantics. The preset proportion condition is a preset proportion threshold value for determining whether the semantic proportion of the information semantics is large enough to determine whether it is the main operation intention.
[0115] Specifically, the server compares the semantic proportion of each information semantic with a preset proportion condition. If the semantic proportion of a certain information semantic (i.e., a second information semantic) in each information semantic is greater than the preset proportion condition, it indicates that the operation intention represented by the information semantic occupies a dominant position in the overall voice control information. At this time, the smart device converts the second information semantic into a control instruction matched with the voice control information according to the preset instruction conversion rule in the system, so as to drive the smart device to perform the corresponding operation. For example, the preset proportion condition is 0.6, and the semantic proportion of the "adjust the brightness of the living room lamp" information semantic is 0.7, which is greater than the condition, so it is converted into a control instruction for adjusting the brightness of the living room lamp.
[0116] The embodiment can more comprehensively consider the importance of different sub-information in the voice control information by determining the sub-information weight by combining the information volume of the sub-information and calculating the semantic proportion of the information semantic, accurately grasp the main operation intention of the user from two dimensions of volume and semantic, improve the processing capability and accuracy of the smart device for complex voice control information, and enhance the stability and reliability of the user interaction with the smart device.
[0117] In a specific embodiment, a smart device control method in an actual application scenario is also provided. After obtaining voice wake-up information received by a smart device in an environment, a first formant A1 corresponding to the voice wake-up information is recorded. If a second formant A2 is different next time, the voice control information corresponding to the second formant A2 is stored. If the same first formant A1 appears within T (10 seconds) time, the control instruction of the user who issues the first formant A1 is executed. The voice control information includes multiple sub-information.
[0118] If the same first formant A1 does not appear again within T time, the multiple stored sub-information is analyzed. Based on the information content of each sub-information, the number of each information content is counted, and the number of sub-information corresponding to each information content is determined. If the maximum number in the number of sub-information satisfies a preset number condition (the proportion is greater than P (60%-80%)), the information content corresponding to the maximum number is subjected to semantic analysis, a control semantic is determined, and the control semantic is converted into a control instruction matched with the voice control information.
[0119] If the proportion is less than P, the number of each information semantic is counted, and the number of sub-information corresponding to each information semantic is determined. In each information semantic, if there is a first information semantic whose corresponding sub-information number satisfies the preset number condition, the first information semantic is converted into a control instruction matched with the voice control information. If the proportion is still less than P, all stored information is cleared, and no control instruction is executed.
[0120] If greater than T time, no voice control information appears, the first formant A1 is cleared, and no control instruction is executed.
[0121] In one specific embodiment, as Figure 7 shown, an intelligent device control method is also provided, comprising:
[0122] Step S701, obtaining voice wake-up information received by the intelligent device in an environment;
[0123] Step S702, determining a receiving time of the voice wake-up information and a first formant corresponding to the voice wake-up information;
[0124] Step S703, obtaining voice control information received by the intelligent device in the environment within a preset time after the receiving time;
[0125] Step S704, determining a second formant corresponding to the voice control information; the voice control information comprises a plurality of sub-information; the second formant comprises a sub-formant of each sub-information;
[0126] Step S705, in a case where a target formant matching the first formant exists in each sub-formant, extracting a semantic feature vector of target sub-information corresponding to the target formant;
[0127] Step S706, performing similarity matching on the semantic feature vector and a candidate instruction template in a preset instruction template library, and taking a target instruction template satisfying a similarity condition as a control semantic of the target sub-information;
[0128] Step S707, converting the control semantic into a control instruction matching the voice control information;
[0129] Step S708, in a case where each sub-formant does not match the first formant, respectively performing quantity statistics on information content of each sub-information based on the information content, and determining a sub-information quantity corresponding to each information content;
[0130] Step S709, if a maximum quantity in the sub-information quantity satisfies a preset quantity condition, performing semantic analysis on information content corresponding to the maximum quantity to determine a control semantic;
[0131] Step S710, converting the control semantic into a control instruction matching the voice control information;
[0132] Step S711, if the maximum quantity in the sub-information quantity does not satisfy the preset quantity condition, performing semantic analysis on each sub-information to determine an information semantic of the sub-information;
[0133] Step S712, respectively, the number of information semantics is statistically determined, and the number of sub-information corresponding to each information semantics is determined;
[0134] Step S713, in each information semantics, there is a first information semantics corresponding to the number of sub-information that meets the preset number condition, the first information semantics is converted into a control instruction matched with the voice control information;
[0135] The control instruction is used to drive the smart device to perform the corresponding control operation.
[0136] Optionally, in the case that each sub-resonance peak does not match the first resonance peak, the server can also perform semantic analysis on the sub-information for each sub-information, determine the information semantics of the sub-information; based on the information volume of the sub-information, determine the sub-information weight of the sub-information; for each information semantics, according to the sub-information weight of the sub-information in the information semantics, determine the semantic proportion of the information semantics; in the case that the semantic proportion of the second information semantics is greater than the preset proportion condition, the second information semantics is converted into a control instruction matched with the voice control information.
[0137] It should be understood that, although each step in the flowchart involved in the above embodiments is displayed in sequence according to the arrow, these steps are not necessarily executed in sequence according to the arrow. Unless otherwise specified herein, the execution of these steps is not strictly limited in sequence, and these steps can be executed in other sequences. Moreover, at least part of the steps in the flowchart involved in the above embodiments can include multiple steps or multiple stages, which are not necessarily executed at the same time, but can be executed at different times, and the execution sequence of these steps or stages is not necessarily sequential, but can be executed alternately or alternately with at least part of other steps or steps or stages in other steps.
[0138] Based on the same inventive concept, the embodiments of the present application also provide a smart device control device for implementing the above-mentioned smart device control method. The implementation scheme of the problem solving provided by the device is similar to the implementation scheme described in the above method, so the specific limitations in one or more smart device control device embodiments provided below can refer to the limitations of the smart device control method in the above text, and will not be repeated here.
[0139] In one embodiment, as shown in Figure 8 A smart device control device 800 is provided, which includes a voice wake-up information acquisition module 802, a first resonance peak determination module 804, a voice control information acquisition module 806, a second resonance peak determination module 808, a number statistics module 810, a control semantic determination module 812, and a first control instruction determination module 814.
[0140] The voice wake-up information acquisition module 802 is configured to acquire voice wake-up information received by the smart device in an environment;
[0141] The first formant determination module 804 is configured to determine a receiving time of the voice wake-up information and a first formant corresponding to the voice wake-up information;
[0142] The voice control information acquisition module 806 is configured to acquire voice control information received by the smart device in the environment within a preset time after the receiving time;
[0143] The second formant determination module 808 is configured to determine a second formant corresponding to the voice control information; the voice control information includes a plurality of sub-information; and the second formant includes a sub-formant of each of the sub-information;
[0144] The quantity statistics module 810 is configured to, in a case where each of the sub-formants does not match the first formant, respectively perform quantity statistics on information content of each of the sub-information based on the information content, and determine a quantity of each of the sub-information corresponding to the information content;
[0145] The control semantic determination module 812 is configured to, if a maximum quantity in the quantities of the sub-information satisfies a preset quantity condition, perform semantic analysis on information content corresponding to the maximum quantity, and determine a control semantic;
[0146] The first control instruction determination module 814 is configured to convert the control semantic into a control instruction matching the voice control information; and the control instruction is used to drive the smart device to perform a corresponding control operation.
[0147] In an embodiment, the smart device control apparatus 800 further includes a second control instruction determination module, including:
[0148] The control semantic determination unit is configured to, in a case where there is a target formant matching the first formant in the sub-formants, perform semantic analysis on target sub-information corresponding to the target formant, and determine a control semantic of the target sub-information;
[0149] The conversion unit is configured to convert the control semantic into a control instruction matching the voice control information.
[0150] In an embodiment, the control semantic determination unit is specifically configured to:
[0151] extract a semantic feature vector of the target sub-information corresponding to the target formant;
[0152] perform similarity matching on the semantic feature vector and a candidate instruction template in a preset instruction template library, and take a target instruction template with a similarity satisfying a similarity condition as the control semantic of the target sub-information.
[0153] In an embodiment, the smart device control apparatus 800 further comprises a third control instruction determination module, specifically configured to: if the maximum number of the numbers of sub-information does not satisfy the preset number condition, perform semantic analysis on the sub-information for each sub-information, and determine the information semantics of the sub-information;
[0154] count the numbers of the information semantics respectively, and determine the numbers of sub-information corresponding to the information semantics respectively;
[0155] In the case that the first information semantics corresponding to the number of sub-information satisfies the preset number condition, the first information semantics is converted into the control instruction matched with the voice control information.
[0156] In an embodiment, the smart device control apparatus 800 further comprises a fourth control instruction determination module, specifically configured to:
[0157] In the case that each of the sub-formants does not match the first formant, the semantic analysis is performed on the sub-information for each sub-information, and the information semantics of the sub-information is determined;
[0158] Based on the information volume of the sub-information, the sub-information weight of the sub-information is determined;
[0159] For each information semantics, the semantic proportion of the information semantics is determined according to the sub-information weight of the sub-information in the information semantics;
[0160] In the case that the semantic proportion of the second information semantics is greater than the preset proportion condition, the second information semantics is converted into the control instruction matched with the voice control information.
[0161] The above-mentioned modules in the smart device control apparatus can be all or part realized by software, hardware and combination thereof. The above-mentioned modules can be embedded in or independent of the processor in the computer device in hardware form, or stored in the memory in the computer device in software form, so as to be called and executed by the processor to perform the operations corresponding to the above-mentioned modules.
[0162] In an embodiment, a computer device is provided, which can be a terminal, and the internal structure diagram thereof can be as shown in Figure 9The computer device includes a processor, a memory, an input / output interface, a communication interface, a display unit and an input device. The processor, the memory and the input / output interface are connected through a system bus, and the communication interface, the display unit and the input device are connected to the system bus through the input / output interface. The processor of the computer device is configured to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for running the operating system and the computer program in the non-volatile storage medium. The input / output interface of the computer device is configured to exchange information between the processor and external devices. The communication interface of the computer device is configured to perform wired or wireless communication with external terminals. The wireless communication can be achieved through WIFI, mobile cellular network, NFC (Near Field Communication) or other technologies. The computer program is executed by the processor to implement an intelligent device control method. The display unit of the computer device is configured to form a visually visible picture, which can be a display screen, a projection device or a virtual reality imaging device. The display screen can be a liquid crystal display screen or an electronic ink display screen. The input device of the computer device can be a touch layer overlaid on the display screen, or a key, trackball or touchpad arranged on the shell of the computer device, or an external keyboard, touchpad or mouse, etc.
[0163] Those skilled in the art can understand that Figure 9 The structure shown in the figure is only a block diagram of part of the structure related to the scheme of the present application, and does not constitute a limitation on the computer device to which the scheme of the present application is applied. The specific computer device can include more or fewer components than those shown in the figure, or combine certain components, or have a different component arrangement.
[0164] In one embodiment, a computer device is provided, including a memory and a processor, the memory storing a computer program, and the processor executing the computer program to implement the following steps:
[0165] Obtaining voice wake-up information received by the intelligent device in the environment;
[0166] Determining the receiving time of the voice wake-up information and the first formant corresponding to the voice wake-up information;
[0167] Within a preset time after the receiving time, obtaining voice control information received by the intelligent device in the environment;
[0168] Determining the second formant corresponding to the voice control information; the voice control information includes a plurality of sub-information; the second formant includes a sub-formant of each sub-information;
[0169] In a case where each sub-resonance peak is not matched with the first resonance peak, based on information content of each sub-information, the information content is counted respectively to determine a number of sub-information corresponding to each information content;
[0170] If the maximum number of the sub-information numbers satisfies a preset number condition, the information content corresponding to the maximum number is subjected to semantic analysis to determine a control semantic;
[0171] The control semantic is converted into a control instruction matched with the voice control information; the control instruction is used to drive the smart device to perform a corresponding control operation.
[0172] In one embodiment, the processor, when executing the computer program, further implements the following steps:
[0173] In a case where a target resonance peak matched with the first resonance peak exists in each sub-resonance peak, the target sub-information corresponding to the target resonance peak is subjected to semantic analysis to determine a control semantic of the target sub-information;
[0174] The control semantic is converted into a control instruction matched with the voice control information.
[0175] In one embodiment, the processor, when executing the computer program, further implements the following steps:
[0176] A semantic feature vector of the target sub-information corresponding to the target resonance peak is extracted;
[0177] The semantic feature vector is subjected to similarity matching with a candidate instruction template in a preset instruction template library, and a target instruction template satisfying a similarity condition is taken as the control semantic of the target sub-information.
[0178] In one embodiment, the processor, when executing the computer program, further implements the following steps:
[0179] If the maximum number of the sub-information numbers does not satisfy the preset number condition, each sub-information is subjected to semantic analysis to determine an information semantic of the sub-information;
[0180] Each information semantic is counted respectively to determine a number of sub-information corresponding to each information semantic;
[0181] In a case where, in each information semantic, a first information semantic corresponding to a sub-information number satisfying a preset number condition exists, the first information semantic is converted into a control instruction matched with the voice control information.
[0182] In one embodiment, the processor, when executing the computer program, further implements the following steps:
[0183] In a case where each of the sub-resonance peaks does not match the first resonance peak, the semantic analysis is performed on the sub-information for each of the sub-information to determine information semantics of the sub-information;
[0184] The sub-information weight of the sub-information is determined based on the information volume of the sub-information;
[0185] The semantic proportion of the information semantics is determined according to the sub-information weight of the sub-information in the information semantics for each of the information semantics;
[0186] In a case where the semantic proportion of the second information semantics is greater than a preset proportion, the second information semantics is converted into a control instruction matching the voice control information.
[0187] In an embodiment, a computer readable storage medium is provided, and a computer program is stored on the computer readable storage medium. The computer program is executed by a processor to implement the following steps:
[0188] Voice wake-up information received by the smart device in an environment is acquired;
[0189] The receiving time of the voice wake-up information and the first resonance peak corresponding to the voice wake-up information are determined;
[0190] Voice control information received by the smart device in the environment is acquired within a preset time after the receiving time;
[0191] The second resonance peak corresponding to the voice control information is determined; the voice control information includes a plurality of sub-information; and the second resonance peak includes a sub-resonance peak of each of the sub-information;
[0192] In a case where each of the sub-resonance peaks does not match the first resonance peak, the quantity of each information content is respectively counted based on the information content of each of the sub-information to determine the sub-information quantity corresponding to each of the information content;
[0193] If the maximum quantity in the sub-information quantity satisfies a preset quantity condition, the information content corresponding to the maximum quantity is subjected to semantic analysis to determine a control semantics;
[0194] The control semantics is converted into a control instruction matching the voice control information; and the control instruction is used to drive the smart device to perform a corresponding control operation.
[0195] In an embodiment, the computer program is executed by the processor to further implement the following steps:
[0196] In a case where a target resonance peak matching the first resonance peak exists in the sub-resonance peaks, a target sub-information corresponding to the target resonance peak is subjected to semantic analysis to determine a control semantics of the target sub-information;
[0197] The control semantics is converted into a control instruction matching the voice control information.
[0198] In one embodiment, the computer program, when executed by the processor, further implements the following steps:
[0199] extracting a semantic feature vector of the target sub-information corresponding to the target formant;
[0200] performing similarity matching on the semantic feature vector and a candidate instruction template in a preset instruction template library, and taking a target instruction template satisfying a similarity condition as a control semantic of the target sub-information.
[0201] In one embodiment, the computer program, when executed by the processor, further implements the following steps:
[0202] If the maximum number of the sub-information quantities does not satisfy the preset quantity condition, performing semantic analysis on each sub-information to determine an information semantic of the sub-information.
[0203] performing quantity statistics on each information semantic to determine a sub-information quantity corresponding to each information semantic;
[0204] In a case where, among the information semantics, there is a first information semantic whose corresponding sub-information quantity satisfies the preset quantity condition, converting the first information semantic into a control instruction matching the voice control information.
[0205] In one embodiment, the computer program, when executed by the processor, further implements the following steps:
[0206] In a case where each sub-formant does not match the first formant, performing semantic analysis on each sub-information to determine an information semantic of the sub-information.
[0207] determining a sub-information weight of the sub-information based on an information volume of the sub-information;
[0208] For each information semantic, determining a semantic proportion of the information semantic according to the sub-information weight of the sub-information in the information semantic.
[0209] In a case where the semantic proportion of the second information semantic is greater than a preset proportion condition, converting the second information semantic into a control instruction matching the voice control information.
[0210] In one embodiment, a computer program product is provided, comprising a computer program which, when executed by a processor, implements the following steps:
[0211] obtaining voice wake-up information received by the intelligent device in an environment;
[0212] determining a receiving time of the voice wake-up information and a first formant corresponding to the voice wake-up information;
[0213] acquire voice control information received by the smart device in the environment within a preset time after the receiving time;
[0214] determine a second formant corresponding to the voice control information; the voice control information comprises a plurality of sub-information; the second formant comprises a sub-formant of each sub-information;
[0215] in a case where each sub-formant does not match the first formant, respectively count quantities of information contents of each sub-information based on the information contents, and determine a sub-information quantity corresponding to each information content;
[0216] if a maximum quantity in the sub-information quantities satisfies a preset quantity condition, perform semantic analysis on an information content corresponding to the maximum quantity to determine a control semantic;
[0217] convert the control semantic into a control instruction matching the voice control information; the control instruction is used to drive the smart device to perform a corresponding control operation.
[0218] In one embodiment, the computer program, when executed by the processor, further implements the following steps:
[0219] in a case where a target formant matching the first formant exists in the sub-formants, perform semantic analysis on a target sub-information corresponding to the target formant to determine a control semantic of the target sub-information;
[0220] convert the control semantic into a control instruction matching the voice control information.
[0221] In one embodiment, the computer program, when executed by the processor, further implements the following steps:
[0222] extract a semantic feature vector of the target sub-information corresponding to the target formant;
[0223] perform similarity matching on the semantic feature vector and a candidate instruction template in a preset instruction template library, and take a target instruction template satisfying a similarity condition as the control semantic of the target sub-information.
[0224] In one embodiment, the computer program, when executed by the processor, further implements the following steps:
[0225] if the maximum quantity in the sub-information quantities does not satisfy the preset quantity condition, perform semantic analysis on each sub-information to determine an information semantic of the sub-information;
[0226] respectively count quantities of the information semantics to determine a sub-information quantity corresponding to each information semantic;
[0227] In the case that the number of sub-information corresponding to each information semantic satisfies the preset number condition, the first information semantic is converted into a control instruction matching the voice control information.
[0228] In one embodiment, the computer program, when executed by the processor, further implements the following steps:
[0229] In the case that each sub-resonance peak does not match the first resonance peak, for each sub-information, the semantic analysis is performed on the sub-information to determine the information semantic of the sub-information;
[0230] Based on the information volume of the sub-information, the sub-information weight of the sub-information is determined;
[0231] For each information semantic, the semantic proportion of the information semantic is determined according to the sub-information weight of the sub-information in the information semantic;
[0232] In the case that the semantic proportion of the second information semantic is greater than the preset proportion condition, the second information semantic is converted into a control instruction matching the voice control information.
[0233] It should be noted that the user information (including but not limited to user equipment information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in the present application are all information and data authorized by the user or authorized by all parties, and the collection, use and processing of related data need to comply with relevant national and regional laws, regulations and standards.
[0234] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer readable storage medium, and when the computer program is executed, the processes of the above-mentioned embodiments of the methods can be included. Any reference to memory, database or other medium used in the embodiments provided in the present application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical storage, high-density embedded non-volatile memory, resistive memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. As an illustration but not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc. The database involved in the embodiments provided in the present application can include at least one of a relational database and a non-relational database. The non-relational database can include a distributed database based on a block chain, etc., without being limited thereto. The processor involved in the embodiments provided in the present application can be a general-purpose processor, a central processing unit, a graphics processing unit, a digital signal processor, a programmable logic device, a data processing logic device based on quantum computing, etc., without being limited thereto.
[0235] Any combination of the technical features of the above embodiments can be made. In order to make the description simple, all possible combinations of the technical features in the above embodiments are not described, however, as long as the combination of the technical features does not exist, it should be considered as the scope of the present application.
[0236] The above embodiments only express several implementation manners of the present application, and the description is more specific and detailed, but it should not be understood as a limitation on the scope of the patent of the present application. It should be pointed out that for ordinary skilled in the art, without departing from the concept of the present application, a number of modifications and improvements can be made, which are within the scope of protection of the present application. Therefore, the protection scope of the present application should be subject to the appended claims.
Claims
1. A method for controlling an intelligent device, characterized in that, The method includes: Acquire voice wake-up information received by smart devices in their environment; Determine the reception time of the voice wake-up information and the first formant corresponding to the voice wake-up information; Within a preset time after the receiving time, acquire the voice control information received by the smart device in the environment; Determine the second formant corresponding to the voice control information; the voice control information includes multiple sub-information; the second formant includes the sub-formant of each of the sub-information. When none of the sub-resonance peaks match the first resonance peak, the number of sub-information corresponding to each sub-information is determined by counting the number of each sub-information content based on the information content of each sub-information. If the maximum quantity among the sub-information quantities satisfies a preset quantity condition, then semantic analysis is performed on the information content corresponding to the maximum quantity to determine the control semantics; The control semantics are converted into control commands that match the voice control information; the control commands are used to drive the smart device to perform corresponding control operations.
2. The method according to claim 1, characterized in that, The method further includes: If a target resonance peak matching the first resonance peak exists among the sub-resonance peaks, semantic analysis is performed on the target sub-information corresponding to the target resonance peak to determine the control semantics of the target sub-information. The control semantics are converted into control commands that match the voice control information.
3. The method according to claim 2, characterized in that, The step of performing semantic analysis on the target sub-information corresponding to the target resonance peak to determine the control semantics of the target sub-information includes: Extract the semantic feature vector of the target sub-information corresponding to the target resonance peak; The semantic feature vector is matched with the candidate instruction templates in the preset instruction template library for similarity. The target instruction template whose similarity meets the similarity condition is used as the control semantics of the target sub-information.
4. The method according to claim 1, characterized in that, The method further includes: If the maximum number of each of the sub-information quantities does not meet the preset quantity condition, then for each of the sub-information quantities, semantic analysis is performed on the sub-information to determine the information semantics of the sub-information. Quantitative statistics are performed on each of the aforementioned information semantics to determine the number of sub-information corresponding to each of the aforementioned information semantics; In each of the aforementioned information semantics, if there exists a first information semantic with a corresponding number of sub-information that satisfies the preset quantity condition, the first information semantic is converted into a control command that matches the voice control information.
5. A smart device control device, characterized in that, The device includes: The voice wake-up information acquisition module is used to acquire the voice wake-up information received by the smart device in its environment. The first formant determination module is used to determine the reception time of the voice wake-up information and the first formant corresponding to the voice wake-up information; The voice control information acquisition module is used to acquire the voice control information received by the smart device in the environment within a preset time after the reception time. The second formant determination module is used to determine the second formant corresponding to the voice control information; the voice control information includes multiple sub-information; the second formant includes the sub-formant of each of the sub-information. The quantity statistics module is used to determine the quantity of each sub-information based on the information content of each sub-information when none of the sub-resonance peaks match the first resonance peak. The control semantic determination module is used to perform semantic analysis on the information content corresponding to the maximum quantity if the maximum quantity among the quantities of each sub-information satisfies a preset quantity condition, and to determine the control semantics. The first control instruction determination module is used to convert the control semantics into control instructions that match the voice control information; the control instructions are used to drive the smart device to perform corresponding control operations.
6. The apparatus according to claim 5, characterized in that, The device further includes a second control command determination module, comprising: The control semantic determination unit is used to perform semantic analysis on the target sub-information corresponding to the target resonance peak and determine the control semantics of the target sub-information when there is a target resonance peak that matches the first resonance peak among the sub-resonance peaks. A conversion unit is used to convert the control semantics into control commands that match the voice control information.
7. The apparatus according to claim 6, characterized in that, The control semantic determination unit is specifically used for: Extract the semantic feature vector of the target sub-information corresponding to the target resonance peak; The semantic feature vector is matched with the candidate instruction templates in the preset instruction template library for similarity. The target instruction template whose similarity meets the similarity condition is used as the control semantics of the target sub-information.
8. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 4.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 4.
10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 4.
Citation Information
Patent Citations
Intelligent device awakening feedback method and intelligent device
CN110673821A
Speech recognition method, device, system and equipment and computer readable storage medium
CN111210829A