Elevator control method, device, electronic equipment, storage medium and product

By combining the acoustic model and language model in the elevator control method, the first and second control parameters of the voice signal are used to solve the problem of misrecognition of voice-controlled elevators, achieving higher accuracy and safety.

CN114333821BActive Publication Date: 2025-09-30SOUNDAI TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202111657516.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-30
Publication Date
2025-09-30
Estimated Expiration
2041-12-30

AI Technical Summary

Technical Problem

Existing voice-controlled elevator methods have limited acoustic models due to the storage capacity of storage devices, resulting in a limited corpus. This makes it easy for similar pronunciations, such as "airflow" and "seventh floor," to be misrecognized, leading to a high miscontrol rate.

Method used

By obtaining the first control parameter and the second control parameter of the voice signal, the first control parameter represents the probability that the voice signal is a command word, and the second control parameter represents the probability that the text information corresponding to the voice signal matches the text information corresponding to the command word, the two are combined to control the elevator, using a combined recognition technology of acoustic model and language model.

Benefits of technology

The accuracy of voice-controlled elevators is improved, the error control rate is reduced, and the accuracy and safety of elevator responses are ensured.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114333821B_ABST
    Figure CN114333821B_ABST
Patent Text Reader

Abstract

The present application provides an elevator control method, device, electronic device, storage medium, and product, belonging to the field of speech recognition technology. The method includes: obtaining a speech signal, the speech signal being used to control the elevator; determining a first control parameter corresponding to the speech signal, the first control parameter being used to indicate the probability that the speech signal is a command word; determining a second control parameter corresponding to the speech signal, the second control parameter being used to indicate the probability that the text information corresponding to the speech signal matches the text information corresponding to the command word; and controlling the elevator based on the first control parameter and the second control parameter. The method controls the elevator based on the first control parameter and the second control parameter, achieving elevator control based on two recognition results of the speech signal, thereby improving the accuracy of voice-based elevator control.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of speech recognition technology, and in particular to an elevator control method, device, electronic device, storage medium and product. Background Art

[0002] In scenarios like elevators, which are frequently used, have a complex user base, and are densely populated, voice control eliminates the risk of virus infection from touching elevator buttons, offering advantages such as hygiene, safety, and efficiency. Therefore, voice control of elevators has become a research hotspot in this field.

[0003] In related art, voice control of elevators is typically achieved through acoustic models. These models are trained based on command words used to wake up the elevator; for example, command words include "go to the first floor" or "go to the seventh floor." When controlling the elevator via voice, the acoustic model identifies the probability that the voice signal matches the command word. If the probability exceeds a preset threshold for the command word, the elevator is controlled based on the command word.

[0004] Because the current acoustic model is limited by the storage capacity of storage devices, the corpus storage is not very rich. Similar pronunciations, such as "airflow" and "seventh floor", are easily misrecognized, resulting in a high error control rate for elevators. Summary of the Invention

[0005] The embodiments of the present application provide an elevator control method, device, electronic device, storage medium, and product that can improve the accuracy of voice-based elevator control. The technical solution is as follows:

[0006] In one aspect, a method for controlling an elevator is provided, the method comprising:

[0007] Acquiring a voice signal, where the voice signal is used to control the elevator;

[0008] determining a first control parameter corresponding to the voice signal, where the first control parameter is used to indicate a probability that the voice signal is a command word;

[0009] determining a second control parameter corresponding to the voice signal, where the second control parameter is used to represent a probability that the text information corresponding to the voice signal matches the text information corresponding to the command word;

[0010] The elevator is controlled based on the first control parameter and the second control parameter.

[0011] In one implementation, the first control parameter includes first control sub-parameters of multiple groups of phoneme sequences corresponding to the speech signal, and the second control parameter includes second control sub-parameters of multiple text information corresponding to the multiple groups of phoneme sequences, each group of phoneme sequences corresponding to one text information;

[0012] The controlling the elevator based on the first control parameter and the second control parameter includes:

[0013] determining a target phoneme sequence from the multiple groups of phoneme sequences based on the first control sub-parameters and the second control sub-parameters respectively corresponding to the multiple groups of phoneme sequences;

[0014] If the first control sub-parameter corresponding to the target phoneme sequence meets the target condition, the elevator is controlled to make an elevator-taking response corresponding to the target text information, where the target text information is text information corresponding to the target phoneme sequence.

[0015] In one implementation, the speech signal includes multiple groups of audio frames, each group of audio frames corresponds to a highest probability of a phoneme, and the first control sub-parameter of each group of phoneme sequences is the accumulated value of the probability of each phoneme included in the phoneme sequence;

[0016] The process of determining whether the first control sub-parameter corresponding to the target phoneme sequence meets the target condition includes:

[0017] Determining a ratio of the first control subparameter to a target sum to obtain a first confidence level of the target phoneme sequence, where the target sum is a sum of a plurality of highest probabilities corresponding to the plurality of groups of audio frames;

[0018] If the first confidence is greater than a first threshold corresponding to the target phoneme sequence, it is determined that the first control sub-parameter corresponding to the target phoneme sequence meets the target condition.

[0019] In one implementation, the speech signal includes multiple groups of audio frames, each group of audio frames corresponds to a highest probability of a phoneme, and the first control sub-parameter of each group of phoneme sequences is the accumulated value of the probability of each phoneme included in the phoneme sequence;

[0020] The process of determining whether the first control sub-parameter corresponding to the target phoneme sequence meets the target condition includes:

[0021] For each phoneme in the target phoneme sequence, determining a ratio of a probability of the phoneme to a target probability to obtain a second confidence level of the phoneme, wherein the target probability is a maximum probability of the audio frame corresponding to the phoneme;

[0022] If the second confidence of each phoneme is greater than its corresponding second threshold, it is determined that the first control sub-parameter corresponding to the target phoneme sequence meets the target condition.

[0023] In one implementation, determining a target phoneme sequence from the multiple groups of phoneme sequences based on the first control sub-parameters and the second control sub-parameters respectively corresponding to the multiple groups of phoneme sequences includes:

[0024] For each group of phoneme sequences, determining a total parameter of the phoneme sequence based on a first control sub-parameter and a second control sub-parameter corresponding to the phoneme sequence;

[0025] Determine the phoneme sequence with the largest total parameter from the multiple groups of phoneme sequences;

[0026] If the text information corresponding to the selected phoneme sequence matches the preset command word, the selected phoneme sequence is determined to be the target phoneme sequence.

[0027] In one implementation, determining the total parameter of the phoneme sequence based on the first control sub-parameter and the second control sub-parameter corresponding to the phoneme sequence includes:

[0028] The first control sub-parameter and the second control sub-parameter are weighted and summed to obtain the total parameter of the phoneme sequence.

[0029] In one implementation, determining the first control parameter corresponding to the voice signal includes:

[0030] The speech signal is input into an acoustic model, and a first control parameter corresponding to the speech signal is output. The acoustic model is used to determine the first control parameter of the speech signal.

[0031] In one implementation, determining the second control parameter corresponding to the voice signal includes:

[0032] The speech signal is input into a language model, and a second control parameter corresponding to the speech signal is output. The language model is used to determine the second control parameter of the speech signal.

[0033] In one implementation, the language model training process includes:

[0034] Acquire a plurality of first sample information and a plurality of second sample information, wherein the first sample information is text information containing a command word, and the second sample information is text information not containing the command word, wherein the command word is used to control the elevator;

[0035] Based on the multiple first sample information and the multiple second sample information, model training is performed to obtain the language model.

[0036] In another aspect, an elevator control device is provided, the device comprising:

[0037] A first acquisition module is used to acquire a voice signal, where the voice signal is used to control the elevator;

[0038] a first determining module, configured to determine a first control parameter corresponding to the voice signal, wherein the first control parameter is used to indicate a probability that the voice signal is a command word;

[0039] a second determining module, configured to determine a second control parameter corresponding to the voice signal, wherein the second control parameter is used to indicate a probability that the text information corresponding to the voice signal matches the text information corresponding to the command word;

[0040] A control module is used to control the elevator based on the first control parameter and the second control parameter.

[0041] In one implementation, the first control parameter includes first control sub-parameters of multiple groups of phoneme sequences corresponding to the speech signal, and the second control parameter includes second control sub-parameters of multiple text information corresponding to the multiple groups of phoneme sequences, each group of phoneme sequences corresponding to one text information;

[0042] The control module includes:

[0043] A determining unit, configured to determine a target phoneme sequence from the multiple groups of phoneme sequences based on the first control sub-parameters and the second control sub-parameters respectively corresponding to the multiple groups of phoneme sequences;

[0044] The control unit is configured to control the elevator to make an elevator-taking response corresponding to target text information if the first control sub-parameter corresponding to the target phoneme sequence meets a target condition, where the target text information is text information corresponding to the target phoneme sequence.

[0045] In one implementation, the speech signal includes multiple groups of audio frames, each group of audio frames corresponds to a highest probability of a phoneme, and the first control sub-parameter of each group of phoneme sequences is the accumulated value of the probability of each phoneme included in the phoneme sequence;

[0046] The device further comprises:

[0047] a third determination module, configured to determine a ratio of the first control sub-parameter to a target sum to obtain a first confidence level of the target phoneme sequence, wherein the target sum is a sum of a plurality of highest probabilities corresponding to the plurality of groups of audio frames;

[0048] The fourth determination module is configured to determine whether the first control sub-parameter corresponding to the target phoneme sequence satisfies the target condition if the first confidence is greater than a first threshold corresponding to the target phoneme sequence.

[0049] In one implementation, the speech signal includes multiple groups of audio frames, each group of audio frames corresponds to a phoneme with the highest probability, and the first control sub-parameter of each group of phoneme sequences is the accumulated value of the probability of each phoneme included in the phoneme sequence; the apparatus further includes:

[0050] a fifth determination module, configured to determine, for each phoneme in the target phoneme sequence, a ratio of a probability of the phoneme to a target probability to obtain a second confidence level of the phoneme, wherein the target probability is a maximum probability of the audio frame corresponding to the phoneme;

[0051] The sixth determination module is configured to determine whether the first control sub-parameter corresponding to the target phoneme sequence satisfies the target condition if the second confidence of each phoneme is greater than its corresponding second threshold.

[0052] In one implementation, the determining unit includes:

[0053] a first determining subunit, configured to determine, for each group of phoneme sequences, a total parameter of the phoneme sequence based on a first control subparameter and a second control subparameter corresponding to the phoneme sequence;

[0054] A second determining subunit is configured to determine a phoneme sequence having the largest total parameter from the plurality of phoneme sequences;

[0055] The third determining subunit is configured to determine that the selected phoneme sequence is the target phoneme sequence if the text information corresponding to the selected phoneme sequence matches a preset command word.

[0056] In one implementation, the first determining subunit is configured to:

[0057] The first control sub-parameter and the second control sub-parameter are weighted and summed to obtain the total parameter of the phoneme sequence.

[0058] In one implementation, the first determination module is configured to input the speech signal into an acoustic model and output a first control parameter corresponding to the speech signal, wherein the acoustic model is used to determine the first control parameter of the speech signal.

[0059] In one implementation, the second determination module is configured to input the speech signal into a language model and output a second control parameter corresponding to the speech signal, wherein the language model is used to determine the second control parameter of the speech signal.

[0060] In one implementation, the apparatus further includes:

[0061] a second acquisition module, configured to acquire a plurality of first sample information and a plurality of second sample information, wherein the first sample information is text information containing a command word, and the second sample information is text information not containing the command word, wherein the command word is used to control the elevator;

[0062] A training module is used to perform model training based on the multiple first sample information and the multiple second sample information to obtain the language model.

[0063] On the other hand, an electronic device is provided, comprising one or more processors and one or more memories, wherein the one or more memories store at least one program code, and the at least one program code is loaded and executed by the one or more processors to implement the elevator control method described in any of the above implementations.

[0064] On the other hand, a computer-readable storage medium is provided, in which at least one program code is stored. The at least one program code is loaded and executed by a processor to implement the elevator control method described in any of the above implementations.

[0065] On the other hand, a computer program product is provided, which includes computer program code. The computer program code is stored in a computer-readable storage medium. A processor of an electronic device reads the computer program code from the computer-readable storage medium, and the processor executes the computer program code, so that the electronic device performs the elevator control method described in any of the above implementations.

[0066] The beneficial effects of the technical solutions provided by the embodiments of the present application include at least:

[0067] An embodiment of the present application provides an elevator control method. Since the method controls the elevator based on a first control parameter and a second control parameter, and since the first control parameter is used to represent the probability that the voice signal is a naming word, and the second control parameter is used to represent the probability that the text information corresponding to the voice signal matches the text information corresponding to the command word, the elevator is controlled based on the first control parameter and the second control parameter, thereby achieving elevator control based on two recognition results of the voice signal, thereby improving the accuracy of voice-based elevator control. BRIEF DESCRIPTION OF THE DRAWINGS

[0068] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0069] Figure 1 This is a schematic diagram of an implementation environment provided by an embodiment of the present application;

[0070] Figure 2 This is a flow chart of an elevator control method provided by an embodiment of the present application;

[0071] Figure 3 This is a flow chart of an elevator control method provided by an embodiment of the present application;

[0072] Figure 4 This is a block diagram of an elevator control device provided by an embodiment of the present application;

[0073] Figure 5 This is a block diagram of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0074] In order to make the objectives, technical solutions and advantages of this application clearer, the implementation methods of this application will be further described in detail below with reference to the accompanying drawings.

[0075] The terms "first," "second," "third," and "fourth," etc. in the specification and claims of this application and the accompanying drawings are used to distinguish different objects, not to describe a specific order. In addition, the terms "including" and "having," and any variations thereof, are intended to cover non-exclusive inclusions. For example, a process, method, system, product, or apparatus comprising a series of steps or elements is not limited to the listed steps or elements, but may optionally include steps or elements not listed, or may optionally include other steps or elements inherent to the process, method, product, or apparatus.

[0076] Figure 1 The present application provides an implementation environment for an elevator control method, comprising an elevator 10, an electronic device 20, and a sound pickup device 30. In some embodiments, the electronic device 20 and the sound pickup device 30 are installed inside the elevator 10. The sound pickup device 30 is configured to capture voice signals from passengers and transmit the voice signals to the electronic device 20. The electronic device 20 is configured to recognize command words in the voice signals and control the elevator 10 based on the command words, thereby implementing the application of the method in the context of voice-based elevator control.

[0077] Figure 2 An elevator control method provided in an embodiment of the present application includes:

[0078] 201. Acquire a voice signal, where the voice signal is used to control the elevator.

[0079] 202. Determine a first control parameter corresponding to the speech signal, where the first control parameter is used to indicate a probability that the speech signal is a command word.

[0080] 203. Determine a second control parameter corresponding to the voice signal, where the second control parameter is used to indicate a probability of matching between text information corresponding to the voice signal and text information corresponding to the command word.

[0081] 204. Control the elevator based on the first control parameter and the second control parameter.

[0082] In one implementation, the first control parameter includes first control sub-parameters of multiple groups of phoneme sequences corresponding to the speech signal, and the second control parameter includes second control sub-parameters of multiple text messages corresponding to the multiple groups of phoneme sequences, each group of phoneme sequences corresponding to one text message;

[0083] Controlling the elevator based on the first control parameter and the second control parameter includes:

[0084] determining a target phoneme sequence from the plurality of phoneme sequences based on the first control sub-parameters and the second control sub-parameters corresponding to the plurality of phoneme sequences respectively;

[0085] If the first control sub-parameter corresponding to the target phoneme sequence meets the target condition, the elevator is controlled to make an elevator ride response corresponding to the target text information, where the target text information is text information corresponding to the target phoneme sequence.

[0086] In one implementation, the speech signal includes multiple groups of audio frames, each group of audio frames corresponds to a highest probability of a phoneme, and the first control sub-parameter of each group of phoneme sequences is the accumulated value of the probability of each phoneme included in the phoneme sequence;

[0087] The process of determining whether the first control sub-parameter corresponding to the target phoneme sequence meets the target condition includes:

[0088] Determining a ratio of the first control subparameter to a target sum to obtain a first confidence level of the target phoneme sequence, where the target sum is a sum of multiple highest probabilities corresponding to multiple groups of audio frames;

[0089] If the first confidence is greater than a first threshold corresponding to the target phoneme sequence, it is determined that the first control sub-parameter corresponding to the target phoneme sequence meets the target condition.

[0090] In one implementation, the speech signal includes multiple groups of audio frames, each group of audio frames corresponds to a highest probability of a phoneme, and the first control sub-parameter of each group of phoneme sequences is the accumulated value of the probability of each phoneme included in the phoneme sequence;

[0091] The process of determining whether the first control sub-parameter corresponding to the target phoneme sequence meets the target condition includes:

[0092] For each phoneme in the target phoneme sequence, determining a ratio of the probability of the phoneme to the target probability to obtain a second confidence level of the phoneme, where the target probability is the highest probability of the audio frame corresponding to the phoneme;

[0093] If the second confidence of each phoneme is greater than its corresponding second threshold, it is determined that the first control sub-parameter corresponding to the target phoneme sequence meets the target condition.

[0094] In one implementation, determining a target phoneme sequence from a plurality of phoneme sequences based on first control sub-parameters and second control sub-parameters corresponding to the plurality of phoneme sequences respectively includes:

[0095] For each group of phoneme sequences, determining a total parameter of the phoneme sequence based on the first control sub-parameter and the second control sub-parameter corresponding to the phoneme sequence;

[0096] Determine the phoneme sequence with the largest total parameter from multiple groups of phoneme sequences;

[0097] If the text information corresponding to the selected phoneme sequence matches the preset command word, the selected phoneme sequence is determined to be the target phoneme sequence.

[0098] In one implementation, determining the total parameter of the phoneme sequence based on the first control sub-parameter and the second control sub-parameter corresponding to the phoneme sequence includes:

[0099] The first control sub-parameter and the second control sub-parameter are weighted and summed to obtain the total parameter of the phoneme sequence.

[0100] In one implementation, determining a first control parameter corresponding to a speech signal includes:

[0101] The speech signal is input into the acoustic model, and the first control parameter corresponding to the speech signal is output. The acoustic model is used to determine the first control parameter of the speech signal.

[0102] In one implementation, determining the second control parameter corresponding to the speech signal includes:

[0103] The speech signal is input into the language model, and the second control parameter corresponding to the speech signal is output. The language model is used to determine the second control parameter of the speech signal.

[0104] In one implementation, the language model training process includes:

[0105] Acquire a plurality of first sample information and a plurality of second sample information, wherein the first sample information is text information containing a command word, and the second sample information is text information not containing a command word, where the command word is used to control the elevator;

[0106] Based on the plurality of first sample information and the plurality of second sample information, model training is performed to obtain a language model.

[0107] An embodiment of the present application provides an elevator control method. Since the method controls the elevator based on a first control parameter and a second control parameter, and since the first control parameter is used to represent the probability that the voice signal is a naming word, and the second control parameter is used to represent the probability that the text information corresponding to the voice signal matches the text information corresponding to the command word, the elevator is controlled based on the first control parameter and the second control parameter, thereby achieving elevator control based on two recognition results of the voice signal, thereby improving the accuracy of voice-based elevator control.

[0108] Figure 3 An elevator control method provided in an embodiment of the present application includes:

[0109] 301. The electronic device obtains a voice signal, and the voice signal is used to control the elevator.

[0110] In one implementation, the electronic device captures a voice signal from a passenger via a sound pickup device. Optionally, the electronic device is configured to recognize not only the passenger's voice signal after entering the elevator, but also the passenger's voice signal while waiting for the elevator, such as "going up" and "going down."

[0111] In one implementation, the electronic device detects passengers in the waiting area or elevator through an infrared detection device. When the passengers are detected, the electronic device obtains the voice signal through the sound pickup device, thus avoiding the waste of resources caused by constantly obtaining voice signals.

[0112] 302. The electronic device determines a first control parameter corresponding to the voice signal, where the first control parameter is used to indicate a probability that the voice signal is a command word.

[0113] The first control parameter includes first control sub-parameters for multiple phoneme sequences corresponding to the speech signal. The first control sub-parameter for each phoneme sequence represents the probability that the phoneme sequence is a command word. The first control sub-parameter for each phoneme sequence is the cumulative value of the probabilities of each phoneme in the phoneme sequence. A phoneme is the smallest unit of speech, defined based on the natural properties of speech, and is the smallest unit or smallest speech segment that constitutes speech.

[0114] In one implementation, the electronic device inputs a speech signal into an acoustic model and outputs a first control parameter corresponding to the speech signal, and the acoustic model is used to determine the first control parameter of the speech signal. The electronic device inputs a speech signal into the acoustic model and outputs multiple phoneme sequences corresponding to the speech signal and the first control sub-parameters for each phoneme sequence.

[0115] It should be noted that because the speech signal input to the acoustic model may cause errors due to various reasons, the acoustic model outputs multiple phoneme sequences. For example, for the input speech signal "seventh floor", the acoustic model outputs phoneme sequences including "qi lou", "qi liu", and "qi you". Therefore, based on the acoustic model, at least one phoneme sequence corresponding to the speech signal and the first control sub-parameter of each phoneme sequence can be determined.

[0116] In some embodiments, the training process of the acoustic model includes: for each phoneme, obtaining multiple third sample information and multiple fourth sample information of the phoneme. The third sample information includes an audio frame corresponding to the phoneme, which is a positive sample. The fourth sample information includes an audio frame that does not correspond to the phoneme, which is a negative sample. For each command word, obtaining multiple fifth sample information and multiple sixth sample information of the command word. The fifth sample information includes a phoneme sequence corresponding to the command word, which is a positive sample; the sixth sample information includes a phoneme sequence that does not correspond to the command word, which is a negative sample. The electronic device performs model training based on the multiple third sample information, the multiple fourth sample information, the multiple fifth sample information, and the multiple sixth sample information to obtain an acoustic model.

[0117] In an embodiment of the present application, since the acoustic model can recognize phonemes, the probability that multiple groups of phoneme sequences corresponding to the speech signal are the phoneme sequences of command words can be obtained through the acoustic model; and since phonemes are the smallest pronunciation units, the recognition of phonemes through the acoustic model realizes the bottom-level recognition of the speech signal, and then when the second control sub-parameters of the text information corresponding to the phoneme sequence are subsequently determined, the accuracy of the determined second control sub-parameters can be improved.

[0118] 303. The electronic device determines a second control parameter corresponding to the voice signal, where the second control parameter is used to represent a probability of matching the text information corresponding to the voice signal with the text information corresponding to the command word.

[0119] The second control parameter includes second control sub-parameters of multiple text information corresponding to multiple groups of phoneme sequences, and each group of phoneme sequences corresponds to one text information.

[0120] In some embodiments, the electronic device inputs a speech signal into a language model and outputs a second control parameter corresponding to the speech signal, and the language model is used to determine the second control parameter of the speech signal.

[0121] In one implementation, the electronic device inputs multiple phoneme sequences corresponding to the speech signal into a language model, and outputs multiple text messages and a second control sub-parameter corresponding to each piece of text information. Alternatively, the electronic device inputs multiple phoneme sequences output by the acoustic model into the language model, and outputs multiple text messages and a second control sub-parameter corresponding to each piece of text information.

[0122] In another implementation, the electronic device can also obtain a phoneme sequence through a language model. After the electronic device inputs the speech signal into the language model, multiple groups of phoneme sequences corresponding to the speech signal are obtained, and then the text information corresponding to each group of phoneme sequences and the second control subparameter corresponding to each text information are obtained.

[0123] Among them, each group of phoneme sequences can also correspond to multiple text information. For example, the phoneme sequence "qi lou" can correspond to outputs such as "seven floors" and "ventilator shaft"; the language model assigns a higher probability to the text information that conforms to the semantics and matches the command word, and assigns a lower probability to the text information that does not conform to the semantics or does not match the command word.

[0124] In the embodiments of the present application, since the language model is a word-level recognition model, the language model can determine whether a sentence conforms to human language habits, that is, whether it conforms to the semantic logic of human speech, and can also determine whether a sentence is a sentence expressed by general text, so that the language model can assign probabilities based on the semantics of the text information; in this way, by determining the second control parameter of the speech signal through the language model, the accuracy of the probability of the determined text information matching the text information of the command word can be effectively improved.

[0125] In some embodiments, the training process of the language model includes steps (1)-(2):

[0126] (1) The electronic device obtains multiple first sample information and multiple second sample information. The first sample information is text information containing a command word, and the second sample information is text information not containing a command word. The command word is used to control the elevator.

[0127] Optionally, the first sample information is text information such as "go to the first floor", "I want to go to the tenth floor", "the seventh floor", etc. The second sample information is general text information such as "news", "chat", "story", etc.

[0128] (2) The electronic device performs model training based on multiple first sample information and multiple second sample information to obtain a language model.

[0129] In the embodiments of the present application, the language model is obtained through model training with text information containing a command word and text information not containing a command word, making the corpus of the language model rich, and thus improving the accuracy of the second control parameter of the determined speech signal.

[0130] 304. The electronic device determines a target phoneme sequence from multiple groups of phoneme sequences based on the first control subparameters and the second control subparameters corresponding to the multiple groups of phoneme sequences.

[0131] In one implementation, this step includes the following steps (1)-(3)

[0132] (1) For each group of phoneme sequences, the electronic device determines the total parameter of the phoneme sequence based on the first control sub-parameter and the second control sub-parameter corresponding to the phoneme sequence.

[0133] In one implementation, the electronic device performs weighted summation on the first control sub-parameter and the second control sub-parameter to obtain the total parameter of the phoneme sequence.

[0134] The electronic device determines the first control sub-parameter and the first weight and the second weight of the second control sub-parameter respectively, and based on the first weight and the second weight, performs weighted summation of the first control sub-parameter and the second control sub-parameter to obtain the total parameter of the phoneme sequence.

[0135] It should be noted that the first weight and the second weight can be set and changed as needed, and are not specifically limited in the embodiments of the present application. Optionally, if the first weight is 1, the electronic device determines the total parameter of the phoneme sequence by the following formula 1.

[0136] Formula 1: Total parameter = first control sub-parameter + second weight * second control sub-parameter

[0137] In one implementation, the second weight is determined based on the type of language model. Optionally, the language model includes a general language model and an elevator language model, with the elevator language model being a language model specifically designed to recognize elevator phrases. If the language model in the embodiment of this application is a general language model, the second weight is optionally 0.5 or 0.6; if the language model in the embodiment of this application is an elevator language model, the second weight is optionally 0.8 or 0.9.

[0138] In an embodiment of the present application, the total parameter is determined by taking the weighted sum of the first control parameter and the second control parameter, which fully considers the importance of each control sub-parameter to the total parameter; and the total parameter is determined based on the weighted sum of the first control sub-parameter and the second control sub-parameter, so that the total parameter combines the two recognition results of the speech signal, and the total parameter determined based on the two recognition results is more comprehensive and accurate.

[0139] (2) The electronic device determines the phoneme sequence with the largest total parameter from multiple groups of phoneme sequences.

[0140] It should be noted that the electronic device realizes the application of the Viterbi algorithm in the embodiment of the present application by determining the phoneme sequence with the largest total parameter. The Viterbi algorithm is an algorithm that selects the optimal path from multiple paths, so that the phoneme sequence with the largest total parameter determined in the embodiment of the present application is the optimal recognition result, thereby improving the accuracy of the determined phoneme sequence.

[0141] (3) If the text information corresponding to the selected phoneme sequence matches the preset command word, the electronic device determines that the selected phoneme sequence is the target phoneme sequence.

[0142] In one implementation, if the first control sub-parameter of the selected phoneme sequence is greater than a preset sub-parameter threshold, the electronic device determines that the text information corresponding to the phoneme sequence matches the preset command word. In another implementation, a similarity calculation is performed between the text information corresponding to the selected phoneme sequence and the preset command word. If the similarity between the text information and the preset command word is greater than a preset similarity threshold, the electronic device determines that the text information corresponding to the phoneme sequence matches the preset command word.

[0143] In an embodiment of the present application, the optimal recognition result among multiple groups of phoneme sequences is determined through the Viterbi algorithm; and the selected phoneme sequence is also matched with a preset command word, and the target phoneme sequence is determined only when the text information corresponding to the selected phoneme sequence matches the preset command word, thereby improving the accuracy of the determined target phoneme sequence.

[0144] 305. The electronic device determines whether the first control sub-parameter corresponding to the target phoneme sequence meets the target condition.

[0145] It should be noted that a speech signal includes multiple groups of audio frames, each of which corresponds to the highest probability of a phoneme. Typically, a speech signal includes multiple words, each of which includes multiple phonemes, each of which corresponds to multiple audio frames, and each of which corresponds to a group of audio frames. Thus, multiple groups of audio frames constitute a complete speech signal.

[0146] In some embodiments, the acoustic model determines the first control parameter using a decoding graph, where the decoding graph includes multiple decoding paths, each decoding path corresponding to a group of phoneme sequences. After the electronic device inputs a speech signal into the acoustic model, it decodes the signal to obtain multiple phoneme sequences and the first control sub-parameter for each group of phoneme sequences.

[0147] In one implementation, the electronic device defines multiple token structures in a decoding graph of the acoustic model. The token structures are used to record historical path information for each decoding path, including the probability of each phoneme corresponding to each node in the decoding path. For each decoding path, the probability of each phoneme at each node in the decoding path is recorded by a token. In some embodiments, the token of the last node in any decoding path is also used to store the accumulated value of the probability of the decoding path to obtain the first control sub-parameter.

[0148] In another implementation, each decoding path corresponds to a token, which is used to record the probabilities of the phonemes of all nodes in the decoding path. In this way, all historical path information of the decoding path can be read or traced back through a token, and the accumulated value of the probability of the decoding path can be stored to obtain the first control sub-parameter of the decoding path. It should be noted that multiple tokens can also record multiple highest probabilities corresponding to each group of audio frames; in this way, when tracing back the historical decoding path through the token, not only the first control sub-parameter of the optimal path can be obtained, but also the highest probability in the decoding path corresponding to each group of audio frames can be obtained.

[0149] The electronic device determines whether the first control sub-parameter corresponding to the target phoneme sequence meets the target condition, including the following two implementation methods:

[0150] In one implementation, the electronic device determines the ratio of the first control sub-parameter to the target sum to obtain a first confidence level of the target phoneme sequence, where the target sum is the sum of multiple highest probabilities corresponding to multiple groups of audio frames. If the first confidence level is greater than a first threshold corresponding to the target phoneme sequence, the electronic device determines that the first control sub-parameter corresponding to the target phoneme sequence meets the target condition.

[0151] Among them, the first threshold is set in advance, and the size of the first threshold can be set and changed as needed. In the embodiment of the present application, there is no specific limitation on this; optionally, the first threshold is 0.8 or 0.9.

[0152] In an embodiment of the present application, the accuracy of the determined target phoneme sequence is further improved by further comparing the first control sub-parameters of the target phoneme sequence based on a first threshold. For example, for voice signals indicating the 4th and 10th floors, the passenger's voice signal indicates the 4th floor, but due to an accent or other errors, the combined recognition result of the acoustic model and language model indicates the 10th floor. Without further comparison using the first threshold, this could result in misidentification and, in turn, incorrect elevator control. In an embodiment of the present application, further comparison of the first control sub-parameters of the target phoneme sequence using the first threshold further improves recognition accuracy.

[0153] For example, the first threshold is 0.9. Only when the first confidence of the target phoneme sequence is greater than 0.9, it is determined that the identified target phoneme sequence is accurate and the elevator is controlled to respond to the elevator. When the first confidence of the target phoneme sequence is not greater than 0.9, it means that the voice signal is recognized relatively vaguely between the 4th and 10th floors, and it is impossible to accurately judge whether the identified target phoneme sequence is accurate, and thus the elevator will not be controlled to respond to the elevator. After the person taking the elevator notices it, a second clearer voice signal will be sent to control the elevator. It can be seen that the method provided by the embodiment of the present application can effectively improve the recognition accuracy, thereby reducing the error control rate of the elevator.

[0154] Optionally, the electronic device obtains a first confidence of the target phoneme sequence based on the first control sub-parameter and the target sum by using the following formula 2.

[0155] Formula 2: First confidence = first control sub-parameter / target sum

[0156] The target sum refers to the sum of multiple highest probabilities corresponding to multiple groups of audio frames.

[0157] In this implementation, the first confidence level of the target phoneme sequence is determined by the sum of the multiple highest probabilities corresponding to the first control sub-parameter and multiple groups of audio frames, so that the value of the first confidence level is more consistent with the probability value of the current speech signal, making the first confidence level more targeted, and then determining whether the target condition is met based on the comparison result between the first confidence level and the target threshold, which can improve the accuracy of determining whether the first control sub-parameter meets the target condition.

[0158] In another implementation, the electronic device determines, for each phoneme in the target phoneme sequence, the ratio of the probability of the phoneme to the target probability, and obtains a second confidence level of the phoneme, where the target probability is the highest probability of the audio frame corresponding to the phoneme; if the second confidence level of each phoneme is greater than its corresponding second threshold, the electronic device determines that the first control sub-parameter corresponding to the target phoneme sequence meets the target condition.

[0159] It should be noted that the first thresholds for multiple phonemes are not the same, and the second threshold for each phoneme can be set and modified as needed, which is not specifically limited in the embodiments of the present application. For example, a higher second threshold is set for phonemes corresponding to "si" and "shi", "lou" and "liu", and "n" and "l" in speech signals that are easily misrecognized to reduce the misrecognition rate.

[0160] In an embodiment of the present application, by making the second confidence of each phoneme in the target phoneme sequence greater than its corresponding second threshold, when determining that the first control sub-parameter corresponding to the target phoneme sequence meets the target condition, each phoneme in the target phoneme sequence meets the condition and has high accuracy, thereby improving the accuracy of determining the target phoneme sequence that meets the target condition, and then when subsequently controlling the elevator to make an elevator response corresponding to the target text information corresponding to the target phoneme sequence, the accuracy of the elevator response can be improved.

[0161] 306. If the first control sub-parameter corresponding to the target phoneme sequence meets the target condition, the electronic device controls the elevator to make an elevator ride response corresponding to the target text information, where the target text information is text information corresponding to the target phoneme sequence.

[0162] It should be noted that if there are multiple text information corresponding to the target phoneme sequence, the target text information is the text information corresponding to the target phoneme sequence with the largest second control sub-parameter, that is, the second control sub-parameter with the largest total parameter is determined.

[0163] In one implementation, the electronic device sends a control instruction to the elevator control panel, where the control instruction carries a command word that matches the target text information, and the control panel makes an elevator-taking response corresponding to the target text information based on the command word.

[0164] In some embodiments, if the first control sub-parameter corresponding to the target phoneme sequence does not meet the target condition, the determined target phoneme sequence is considered inaccurate and is a misrecognition, and the step of issuing a control instruction to the elevator control panel will not be executed.

[0165] In an embodiment of the present application, the electronic device controls the elevator to respond to the elevator request only when the first control sub-parameter corresponding to the target phoneme sequence meets the target condition. This avoids erroneous control caused by controlling the elevator directly based on the text information corresponding to the target phoneme sequence, thereby improving the accuracy of the elevator response. It should be noted that compared with the prior art, the embodiment of the present application provides a reliable reference standard for the target threshold, reducing erroneous control of elevators based on voice control. In some embodiments, the elevator control method provided by this application can reduce the erroneous control rate of elevators by 30%.

[0166] An embodiment of the present application provides an elevator control method. Since the method controls the elevator based on a first control parameter and a second control parameter, and since the first control parameter is used to represent the probability that the voice signal is a naming word, and the second control parameter is used to represent the probability that the text information corresponding to the voice signal matches the text information corresponding to the command word, the elevator is controlled based on the first control parameter and the second control parameter, thereby achieving elevator control based on two recognition results of the voice signal, thereby improving the accuracy of voice-based elevator control.

[0167] The present application also provides an elevator control device. Figure 4 , the device comprises:

[0168] A first acquisition module 401 is used to acquire a voice signal, which is used to control the elevator;

[0169] A first determination module 402 is configured to determine a first control parameter corresponding to the speech signal, the first control parameter being used to indicate a probability that the speech signal is a command word;

[0170] A second determination module 403 is configured to determine a second control parameter corresponding to the voice signal, where the second control parameter represents a probability that the text information corresponding to the voice signal matches the text information corresponding to the command word;

[0171] The control module 404 is configured to control the elevator based on the first control parameter and the second control parameter.

[0172] In one implementation, the first control parameter includes first control sub-parameters of multiple groups of phoneme sequences corresponding to the speech signal, and the second control parameter includes second control sub-parameters of multiple text messages corresponding to the multiple groups of phoneme sequences, each group of phoneme sequences corresponding to one text message;

[0173] The control module 404 includes:

[0174] A determining unit, configured to determine a target phoneme sequence from the multiple phoneme sequences based on the first control sub-parameters and the second control sub-parameters respectively corresponding to the multiple phoneme sequences;

[0175] The control unit is used to control the elevator to make an elevator-taking response corresponding to the target text information if the first control sub-parameter corresponding to the target phoneme sequence meets the target condition, where the target text information is the text information corresponding to the target phoneme sequence.

[0176] In one implementation, the speech signal includes multiple groups of audio frames, each group of audio frames corresponds to a phoneme with the highest probability, and the first control sub-parameter of each group of phoneme sequences is the accumulated value of the probability of each phoneme included in the phoneme sequence; the apparatus further includes:

[0177] a third determination module, configured to determine a ratio of the first control subparameter to a target sum to obtain a first confidence level of the target phoneme sequence, wherein the target sum is a sum of multiple highest probabilities corresponding to multiple groups of audio frames;

[0178] The fourth determination module is configured to determine whether the first control sub-parameter corresponding to the target phoneme sequence meets the target condition if the first confidence level is greater than a first threshold corresponding to the target phoneme sequence.

[0179] In one implementation, the speech signal includes multiple groups of audio frames, each group of audio frames corresponds to a phoneme with the highest probability, and the first control sub-parameter of each group of phoneme sequences is the accumulated value of the probability of each phoneme included in the phoneme sequence; the apparatus further includes:

[0180] a fifth determination module, configured to determine, for each phoneme in the target phoneme sequence, a ratio of the probability of the phoneme to the target probability, to obtain a second confidence level of the phoneme, wherein the target probability is a maximum probability of the audio frame corresponding to the phoneme;

[0181] The sixth determination module is configured to determine whether the first control sub-parameter corresponding to the target phoneme sequence meets the target condition if the second confidence of each phoneme is greater than its corresponding second threshold.

[0182] In one implementation, the determining unit includes:

[0183] a first determining subunit, configured to determine, for each group of phoneme sequences, a total parameter of the phoneme sequence based on the first control subparameter and the second control subparameter corresponding to the phoneme sequence;

[0184] A second determining subunit is used to determine a phoneme sequence with the largest total parameter from multiple groups of phoneme sequences;

[0185] The third determining subunit is configured to determine that the selected phoneme sequence is a target phoneme sequence if the text information corresponding to the selected phoneme sequence matches a preset command word.

[0186] In one implementation, the first determining subunit is configured to:

[0187] The first control sub-parameter and the second control sub-parameter are weighted and summed to obtain the total parameter of the phoneme sequence.

[0188] In one implementation, the first determination module 402 is configured to input a speech signal into an acoustic model and output a first control parameter corresponding to the speech signal, and the acoustic model is configured to determine the first control parameter of the speech signal.

[0189] In one implementation, the second determination module 403 is configured to input a speech signal into a language model and output a second control parameter corresponding to the speech signal, and the language model is used to determine the second control parameter of the speech signal.

[0190] In one implementation, the apparatus further includes:

[0191] a second acquisition module, configured to acquire a plurality of first sample information and a plurality of second sample information, wherein the first sample information is text information containing a command word, and the second sample information is text information not containing a command word, and the command word is used to control the elevator;

[0192] The training module is used to perform model training based on multiple first sample information and multiple second sample information to obtain a language model.

[0193] Figure 5 The following is a block diagram of an electronic device 500 according to an exemplary embodiment of the present application. The electronic device 500 may be a portable mobile electronic device, such as a smartphone, a tablet computer, an MP3 player (Moving Picture Experts Group Audio Layer III), an MP4 player (Moving Picture Experts Group Audio Layer IV), a laptop computer, or a desktop computer. The electronic device 500 may also be referred to as a user device, a portable electronic device, a laptop electronic device, a desktop electronic device, or other similar names.

[0194] Typically, the electronic device 500 includes a processor 501 and a memory 502 .

[0195] The processor 501 may include one or more processing cores, such as a 4-core processor, an 8-core processor, etc. The processor 501 may be implemented in at least one hardware form of DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), or PLA (Programmable Logic Array). The processor 501 may also include a main processor and a coprocessor. The main processor is a processor for processing data in the awake state, also known as a CPU (Central Processing Unit); the coprocessor is a low-power processor for processing data in the standby state. In some embodiments, the processor 501 may be integrated with a GPU (Graphics Processing Unit), which is responsible for rendering and drawing the content to be displayed on the display screen. In some embodiments, the processor 501 may also include an AI (Artificial Intelligence) processor, which is used to process computing operations related to machine learning.

[0196] The memory 502 may include one or more computer-readable storage media, which may be non-transitory. The memory 502 may also include high-speed random access memory and non-volatile memory, such as one or more disk storage devices and flash memory storage devices. In some embodiments, the non-transitory computer-readable storage medium in the memory 502 is used to store at least one program code, which is executed by the processor 501 to implement the elevator control method provided in the method embodiment of the present application.

[0197] In some embodiments, electronic device 500 may optionally include a peripheral device interface 503 and at least one peripheral device. Processor 501, memory 502, and peripheral device interface 503 may be connected via a bus or signal lines. Each peripheral device may be connected to peripheral device interface 503 via a bus, signal lines, or circuit boards. Specifically, the peripheral device may include at least one of a radio frequency circuit 504, a display screen 505, a camera assembly 506, an audio circuit 507, a positioning assembly 508, and a power supply 509.

[0198] The peripheral device interface 503 can be used to connect at least one I / O (Input / Output)-related peripheral device to the processor 501 and the memory 502. In some embodiments, the processor 501, the memory 502, and the peripheral device interface 503 are integrated on the same chip or circuit board; in some other embodiments, any one or two of the processor 501, the memory 502, and the peripheral device interface 503 can be implemented on separate chips or circuit boards, which is not limited in this embodiment.

[0199] The radio frequency circuit 504 is used to receive and transmit RF (Radio Frequency) signals, also known as electromagnetic signals. The radio frequency circuit 504 communicates with communication networks and other communication devices via electromagnetic signals. The radio frequency circuit 504 converts electrical signals into electromagnetic signals for transmission, or converts received electromagnetic signals into electrical signals. Optionally, the radio frequency circuit 504 includes: an antenna system, an RF transceiver, one or more amplifiers, a tuner, an oscillator, a digital signal processor, a codec chipset, a user identity module card, etc. The radio frequency circuit 504 can communicate with other electronic devices via at least one wireless communication protocol. Such wireless communication protocols include, but are not limited to, the World Wide Web, metropolitan area networks, intranets, various generations of mobile communication networks (2G, 3G, 4G, and 5G), wireless local area networks, and / or WiFi (Wireless Fidelity) networks. In some embodiments, the radio frequency circuit 504 may also include circuits related to NFC (Near Field Communication), which is not limited in this application.

[0200] The display screen 505 is used to display a UI (User Interface). The UI may include graphics, text, icons, videos, or any combination thereof. When the display screen 505 is a touch screen display, the display screen 505 is also capable of collecting touch signals on or above the surface of the display screen 505. The touch signals can be input as control signals to the processor 501 for processing. In this case, the display screen 505 can also be used to provide virtual buttons and / or a virtual keyboard, also known as soft buttons and / or a soft keyboard. In some embodiments, there can be one display screen 505, disposed on the front panel of the electronic device 500; in other embodiments, there can be at least two display screens 505, disposed on different surfaces of the electronic device 500 or in a foldable design; in other embodiments, the display screen 505 can be a flexible display, disposed on a curved or foldable surface of the electronic device 500. The display screen 505 can even be configured as a non-rectangular irregular shape, i.e., a special-shaped screen. The display screen 505 can be made of materials such as LCD (Liquid Crystal Display) and OLED (Organic Light-Emitting Diode).

[0201] The camera assembly 506 is used to capture images or videos. Optionally, the camera assembly 506 includes a front camera and a rear camera. Typically, the front camera is arranged on the front panel of the electronic device, and the rear camera is arranged on the back of the electronic device. In some embodiments, there are at least two rear cameras, which are any one of a main camera, a depth of field camera, a wide-angle camera, and a telephoto camera, so as to realize the fusion of the main camera and the depth of field camera to realize the background blur function, the fusion of the main camera and the wide-angle camera to realize panoramic shooting and VR (Virtual Reality) shooting function or other fusion shooting functions. In some embodiments, the camera assembly 506 may also include a flash. The flash can be a single-color temperature flash or a dual-color temperature flash. A dual-color temperature flash refers to a combination of a warm light flash and a cold light flash, which can be used for light compensation at different color temperatures.

[0202] The audio circuit 507 may include a microphone and a speaker. The microphone is used to collect sound waves from the user and the environment, and convert the sound waves into electrical signals to be input into the processor 501 for processing, or input into the radio frequency circuit 504 to achieve voice communication. For the purpose of stereo sound collection or noise reduction, there can be multiple microphones, which are respectively arranged in different parts of the electronic device 500. The microphone can also be an array microphone or an omnidirectional collection microphone. The speaker is used to convert the electrical signals from the processor 501 or the radio frequency circuit 504 into sound waves. The speaker can be a traditional thin film speaker or a piezoelectric ceramic speaker. When the speaker is a piezoelectric ceramic speaker, it can not only convert the electrical signals into sound waves audible to humans, but also convert the electrical signals into sound waves inaudible to humans for purposes such as ranging. In some embodiments, the audio circuit 507 may also include a headphone jack.

[0203] The positioning component 508 is used to locate the current geographic location of the electronic device 500 to implement navigation or LBS (Location Based Service). The positioning component 508 can be a positioning component based on the US GPS (Global Positioning System), China's Beidou system, or Russia's Galileo system.

[0204] Power supply 509 is used to power the various components in electronic device 500. Power supply 509 can be AC ​​power, DC power, a disposable battery, or a rechargeable battery. When power supply 509 includes a rechargeable battery, the rechargeable battery can be a wired rechargeable battery or a wireless rechargeable battery. A wired rechargeable battery is a battery that is charged via a wired line, while a wireless rechargeable battery is a battery that is charged via a wireless coil. The rechargeable battery can also be used to support fast charging technology.

[0205] In some embodiments, the electronic device 500 further includes one or more sensors 510 , including but not limited to: an acceleration sensor 511 , a gyroscope sensor 512 , a pressure sensor 513 , a fingerprint sensor 514 , an optical sensor 515 , and a proximity sensor 516 .

[0206] The accelerometer 511 can detect the magnitude of acceleration along the three coordinate axes of the coordinate system established by the electronic device 500. For example, the accelerometer 511 can be used to detect the components of gravity acceleration along the three coordinate axes. The processor 501 can control the display screen 505 to display the user interface in a landscape or portrait view based on the gravity acceleration signal collected by the accelerometer 511. The accelerometer 511 can also be used to collect game or user motion data.

[0207] The gyroscope sensor 512 can detect the orientation and rotation angle of the electronic device 500. It can work in conjunction with the accelerometer 511 to collect the user's 3D movements of the electronic device 500. Based on the data collected by the gyroscope sensor 512, the processor 501 can implement the following functions: motion sensing (for example, changing the UI based on the user's tilt operation), image stabilization during shooting, game control, and inertial navigation.

[0208] The pressure sensor 513 can be set on the side frame of the electronic device 500 and / or the lower layer of the display screen 505. When the pressure sensor 513 is set on the side frame of the electronic device 500, it can detect the user's grip signal of the electronic device 500, and the processor 501 performs left and right hand recognition or shortcut operations based on the grip signal collected by the pressure sensor 513. When the pressure sensor 513 is set on the lower layer of the display screen 505, the processor 501 controls the operable controls on the UI interface based on the user's pressure operation on the display screen 505. The operable controls include at least one of a button control, a scroll bar control, an icon control, and a menu control.

[0209] The fingerprint sensor 514 is used to collect the user's fingerprint, and the processor 501 identifies the user's identity based on the fingerprint collected by the fingerprint sensor 514, or the fingerprint sensor 514 identifies the user's identity based on the collected fingerprint. When the user's identity is identified as a trusted identity, the processor 501 authorizes the user to perform relevant sensitive operations, which include unlocking the screen, viewing encrypted information, downloading software, making payments, and changing settings. The fingerprint sensor 514 can be set on the front, back, or side of the electronic device 500. When a physical button or manufacturer logo is set on the electronic device 500, the fingerprint sensor 514 can be integrated with the physical button or manufacturer logo.

[0210] The optical sensor 515 is used to detect ambient light intensity. In one embodiment, the processor 501 can control the display brightness of the display screen 505 based on the ambient light intensity detected by the optical sensor 515. Specifically, when the ambient light intensity is high, the display brightness of the display screen 505 is increased; when the ambient light intensity is low, the display brightness of the display screen 505 is decreased. In another embodiment, the processor 501 can also dynamically adjust the shooting parameters of the camera assembly 506 based on the ambient light intensity detected by the optical sensor 515.

[0211] Proximity sensor 516, also known as a distance sensor, is typically located on the front panel of electronic device 500. Proximity sensor 516 is used to detect the distance between the user and the front of electronic device 500. In one embodiment, when proximity sensor 516 detects that the distance between the user and the front of electronic device 500 is gradually decreasing, processor 501 controls display screen 505 to switch from the screen-on state to the screen-off state. When proximity sensor 516 detects that the distance between the user and the front of electronic device 500 is gradually increasing, processor 501 controls display screen 505 to switch from the screen-off state to the screen-on state.

[0212] Those skilled in the art will understand that Figure 5 The structure shown in the figure does not constitute a limitation on the electronic device 500, and the electronic device 500 may include more or fewer components than shown in the figure, or combine certain components, or adopt a different component arrangement.

[0213] An embodiment of the present application further provides a computer-readable storage medium, in which at least one program code is stored. The at least one program code is loaded and executed by a processor to implement the elevator control method of any of the above implementations.

[0214] An embodiment of the present application further provides a computer program product, which includes computer program code. The computer program code is stored in a computer-readable storage medium. A processor of an electronic device reads the computer program code from the computer-readable storage medium, and the processor executes the computer program code, so that the electronic device performs the elevator control method of any of the above implementations.

[0215] In some embodiments, the computer program product involved in the embodiments of the present application can be deployed and executed on an electronic device, or on multiple electronic devices located at one location, or on multiple electronic devices distributed at multiple locations and interconnected through a communication network. Multiple electronic devices distributed at multiple locations and interconnected through a communication network can constitute a blockchain system.

[0216] An embodiment of the present application provides an elevator control method. Since the method controls the elevator based on a first control parameter and a second control parameter, and since the first control parameter is used to represent the probability that the voice signal is a naming word, and the second control parameter is used to represent the probability that the text information corresponding to the voice signal matches the text information corresponding to the command word, the elevator is controlled based on the first control parameter and the second control parameter, thereby achieving elevator control based on two recognition results of the voice signal, thereby improving the accuracy of voice-based elevator control.

[0217] The above are only optional embodiments of the present application and are not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application should be included in the scope of protection of the present application.

Claims

1. An elevator control method, characterized in that: The method comprises: Acquiring a voice signal, where the voice signal is used to control the elevator; Determining a first control parameter corresponding to the speech signal, the first control parameter being used to indicate a probability that the speech signal is a command word, the first control parameter comprising first control sub-parameters of multiple groups of phoneme sequences corresponding to the speech signal; determining a second control parameter corresponding to the voice signal, the second control parameter being used to represent a probability that the text information corresponding to the voice signal matches the text information corresponding to the command word, the second control parameter comprising second control sub-parameters of a plurality of text information corresponding to the plurality of groups of phoneme sequences, each group of phoneme sequences corresponding to a piece of text information; determining a target phoneme sequence from the multiple groups of phoneme sequences based on the first control sub-parameters and the second control sub-parameters respectively corresponding to the multiple groups of phoneme sequences; If the first control sub-parameter corresponding to the target phoneme sequence meets the target condition, the elevator is controlled to make an elevator-taking response corresponding to the target text information, where the target text information is the text information corresponding to the target phoneme sequence.

2. The method according to claim 1, characterized in that The speech signal includes multiple groups of audio frames, each group of audio frames corresponds to the highest probability of a phoneme, and the first control sub-parameter of each group of phoneme sequences is the accumulated value of the probability of each phoneme included in the phoneme sequence; The process of determining whether the first control sub-parameter corresponding to the target phoneme sequence meets the target condition includes: Determining a ratio of the first control subparameter to a target sum to obtain a first confidence level of the target phoneme sequence, where the target sum is a sum of a plurality of highest probabilities corresponding to the plurality of groups of audio frames; If the first confidence is greater than a first threshold corresponding to the target phoneme sequence, it is determined that the first control sub-parameter corresponding to the target phoneme sequence meets the target condition.

3. The method according to claim 1, characterized in that The speech signal includes multiple groups of audio frames, each group of audio frames corresponds to the highest probability of a phoneme, and the first control sub-parameter of each group of phoneme sequences is the accumulated value of the probability of each phoneme included in the phoneme sequence; The process of determining whether the first control sub-parameter corresponding to the target phoneme sequence meets the target condition includes: For each phoneme in the target phoneme sequence, determining a ratio of a probability of the phoneme to a target probability to obtain a second confidence level of the phoneme, wherein the target probability is a maximum probability of the audio frame corresponding to the phoneme; If the second confidence of each phoneme is greater than its corresponding second threshold, it is determined that the first control sub-parameter corresponding to the target phoneme sequence meets the target condition.

4. The method according to claim 1, wherein The determining of a target phoneme sequence from the multiple groups of phoneme sequences based on the first control sub-parameter and the second control sub-parameter respectively corresponding to the multiple groups of phoneme sequences includes: For each group of phoneme sequences, determining a total parameter of the phoneme sequence based on a first control sub-parameter and a second control sub-parameter corresponding to the phoneme sequence; Determine the phoneme sequence with the largest total parameter from the multiple groups of phoneme sequences; If the text information corresponding to the selected phoneme sequence matches the preset command word, the selected phoneme sequence is determined to be the target phoneme sequence.

5. The method according to claim 4, characterized in that The determining of the total parameter of the phoneme sequence based on the first control sub-parameter and the second control sub-parameter corresponding to the phoneme sequence includes: The first control sub-parameter and the second control sub-parameter are weighted and summed to obtain the total parameter of the phoneme sequence.

6. The method according to claim 1, characterized in that The determining the first control parameter corresponding to the voice signal includes: The speech signal is input into an acoustic model, and a first control parameter corresponding to the speech signal is output. The acoustic model is used to determine the first control parameter of the speech signal.

7. The method according to claim 1, characterized in that The determining the second control parameter corresponding to the voice signal includes: The speech signal is input into a language model, and a second control parameter corresponding to the speech signal is output. The language model is used to determine the second control parameter of the speech signal.

8. The method according to claim 7, characterized in that The training process of the language model includes: Acquire a plurality of first sample information and a plurality of second sample information, wherein the first sample information is text information containing a command word, and the second sample information is text information not containing the command word, wherein the command word is used to control the elevator; Based on the multiple first sample information and the multiple second sample information, model training is performed to obtain the language model.

9. An elevator control device, characterized in that: The device comprises: A first acquisition module is used to acquire a voice signal, where the voice signal is used to control the elevator; a first determining module, configured to determine a first control parameter corresponding to the speech signal, the first control parameter being used to indicate a probability that the speech signal is a command word, the first control parameter comprising first control sub-parameters of multiple groups of phoneme sequences corresponding to the speech signal; a second determination module, configured to determine a second control parameter corresponding to the voice signal, the second control parameter being used to represent a probability of matching between the text information corresponding to the voice signal and the text information corresponding to the command word, the second control parameter comprising second control sub-parameters of the plurality of text information corresponding to the plurality of groups of phoneme sequences, each group of phoneme sequences corresponding to one piece of text information; A control module is configured to determine a target phoneme sequence from the multiple phoneme sequences based on first and second control sub-parameters corresponding to the multiple phoneme sequences, respectively; and control the elevator to make an elevator-taking response corresponding to target text information if the first control sub-parameter corresponding to the target phoneme sequence satisfies a target condition, where the target text information is text information corresponding to the target phoneme sequence.

10. An electronic device, characterized in that: The electronic device includes one or more processors and one or more memories, wherein at least one program code is stored in the one or more memories, and the at least one program code is loaded and executed by the one or more processors to implement the elevator control method according to any one of claims 1 to 8.

11. A computer-readable storage medium, characterized in that The storage medium stores at least one program code, and the at least one program code is loaded and executed by the processor to implement the elevator control method according to any one of claims 1 to 8.

12. A computer program product, characterized in that The computer program product includes a computer program code, which is stored in a computer-readable storage medium. A processor of an electronic device reads the computer program code from the computer-readable storage medium, and the processor executes the computer program code, so that the electronic device performs the elevator control method according to any one of claims 1 to 8.