A method, apparatus, set-top box and medium for determining operation instructions
By combining the recognition probability evaluation method of video and voice data streams, the problems of high complexity and insufficient robustness of algorithms in the prior art are solved, and efficient and accurate identification of operation instructions are achieved.
Patent Information
- Application Number
- CN202210345825.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-03-31
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2042-03-31
AI Technical Summary
In the prior art, algorithms based on hidden Markov models require a large number of training samples and storage volumes when identifying operation instructions, resulting in a long training process, and the recognition algorithm is highly complex and has insufficient robustness.
By acquiring coordinate pairs in the video data stream and audio data in the voice data stream, the target gesture recognition probability and speech recognition probability are evaluated respectively by using the reference gesture library and the voice library, and the operation instructions are determined in combination with the two.
It reduces the complexity of the algorithm, improves the robustness of the recognition algorithm, and ensures the correctness of the operation instructions.
Smart Images

Figure CN115047967B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of video and speech recognition, and particularly to a method and device for determining an operation instruction, a set-top box, and a medium. Background Art
[0002] In recent years, intelligent navigation has become a hot topic in the field of artificial intelligence. Intelligent navigation mainly uses speech or gesture recognition algorithms to recognize operation instructions, and can control the actions of target devices (such as mobile phones, televisions, etc.) within a certain distance according to the operation instructions, without the need for devices such as remote controls.
[0003] The algorithms adopted by existing technical solutions are mainly algorithms based on the Hidden Markov Model (HMM). The HMM algorithm uses a state sequence to describe the temporal logic of the observation vector, and represents the spatial distribution of the observation vector sequence through a multivariate mixture Gaussian distribution. It requires a large number of training samples and storage, so the training process takes a long time. Summary of the Invention
[0004] The present invention provides a method and device for determining an operation instruction, a set-top box, and a medium, so as to improve the robustness of the recognition algorithm and reduce the complexity of the algorithm.
[0005] According to one aspect of the present invention, there is provided a method for determining an operation instruction, which is applied to a set-top box and includes:
[0006] Obtaining a coordinate pair in a video data stream, where the coordinate pair is formed by the coordinates of two joint points on the same arm in a frame of image;
[0007] Obtaining at least one audio data in a voice data stream;
[0008] Evaluating the coordinate pair according to a reference gesture library to obtain a target gesture recognition probability of the coordinate pair, where the target gesture recognition probability is the recognition probability corresponding to a target gesture, and the reference gesture library includes at least one standard gesture instruction;
[0009] Evaluating each audio data according to a reference voice library to obtain a target voice recognition probability of each audio data, where the target voice recognition probability is the recognition probability corresponding to a target voice, and the reference voice library includes at least one standard voice instruction; determining an operation instruction based on the target gesture recognition probability and the target voice recognition probability.
[0010] According to another aspect of the present invention, there is provided an apparatus for determining an operation instruction, including:
[0011] A first obtaining module, configured to obtain a coordinate pair in a video data stream, where the coordinate pair is formed by the coordinates of two joint points on the same arm in a frame of image;
[0012] A second acquisition module, configured to acquire at least one audio data from the voice data stream;
[0013] A first evaluation module, configured to evaluate the coordinate pair according to a reference gesture library to obtain a target gesture recognition probability of the coordinate pair, where the target gesture recognition probability is the recognition probability corresponding to a target gesture, and the reference gesture library includes at least one standard gesture instruction;
[0014] A second evaluation module, configured to evaluate each of the audio data according to a reference voice library to obtain a target voice recognition probability of each of the audio data, where the target voice recognition probability is the recognition probability corresponding to a target voice, and the reference voice library includes at least one standard voice instruction;
[0015] A determination module, configured to determine an operation instruction based on the target gesture recognition probability and the target voice recognition probability.
[0016] According to another aspect of the present invention, there is provided a set-top box, where the set-top box includes:
[0017] A camera;
[0018] A microphone;
[0019] A controller, communicatively connected to the camera and the microphone respectively, where the controller includes:
[0020] At least one processor; and
[0021] A memory communicatively connected to the at least one processor; wherein,
[0022] The memory stores a computer program executable by the at least one processor, and when the computer program is executed by the at least one processor, the at least one processor is enabled to execute the operation instruction determination method according to any embodiment of the present invention.
[0023] According to another aspect of the present invention, there is provided a computer-readable storage medium, where the computer-readable storage medium stores computer instructions, and when the computer instructions are executed by a processor, the operation instruction determination method according to any embodiment of the present invention is implemented.
[0024] An embodiment of the present invention provides a method, apparatus, set-top box, and medium for determining an operation instruction. The method is applied to a set-top box. By evaluating the coordinate pairs in the acquired video data stream according to a reference gesture library, the target gesture recognition probability is accurately obtained. By evaluating the audio data in the acquired voice data stream according to a reference voice library, the target voice recognition probability is accurately obtained, reducing the complexity of the algorithm. At the same time, the operation instruction is determined by calibrating the target gesture recognition probability and the target voice recognition probability with each other, improving the robustness of the algorithm, and thus ensuring the correctness of the operation instruction.
[0025] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present invention, nor is it used to limit the scope of the present invention. Other features of the present invention will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0027] Figure 1 is a flowchart of a method for determining an operation instruction according to Embodiment 1 of the present invention;
[0028] Figure 2 is a flowchart of a method for determining an operation instruction according to Embodiment 2 of the present invention;
[0029] Figure 3 is a schematic structural diagram of an apparatus for determining an operation instruction according to Embodiment 3 of the present invention;
[0030] Figure 4 is a schematic structural diagram of a set-top box for implementing the method for determining an operation instruction according to Embodiment 1 of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0031] In order to enable those skilled in the art to better understand the solutions of the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only some of the embodiments of the present invention, rather than all of them. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0032] It should be noted that the terms "first", "second", etc. in the specification, claims and above-mentioned drawings of the present invention are used to distinguish similar objects, and do not necessarily have to be used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of the present invention described here can be implemented in an order other than those illustrated or described here. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device comprising a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0033] Embodiment 1
[0034] Figure 1 is a flowchart of a method for determining an operation instruction provided according to Embodiment 1 of the present invention. This embodiment is applicable to the situation of determining an operation instruction. This method can be executed by an operation instruction determining device, which can be implemented in the form of hardware and / or software, and the operation instruction determining device can be configured in a set-top box. As Figure 1 shown, the method includes:
[0035] S110. Obtain coordinate pairs in the video data stream, where the coordinate pairs are formed by the coordinates of two joint points on the same arm in one frame of the image.
[0036] Among them, the video data stream can be regarded as the video data within a specific duration obtained by the camera in the set-top box. The video data within a specific duration can be the video data at the current moment and within a period of time before the current moment. This embodiment does not limit the specific duration, which can be set by the system or relevant personnel. Among them, the video data stream can contain the gesture action information of the user, such as the coordinates of the joint points, etc.
[0037] The coordinate pair can be regarded as a coordinate pair formed by the coordinates of two joint points on the same arm in one frame of the image, and can be used to obtain the target gesture recognition probability. The specific positions of the joint points are not limited. For example, they can be fingers, elbows, wrists, etc.
[0038] In this embodiment, the coordinate pairs in the video data stream can be obtained for subsequent steps. This embodiment does not limit the specific obtaining means. Exemplarily, the image at the current moment can be obtained, and then the coordinate pairs in the image can be obtained.
[0039] S120. Obtain at least one audio data in the voice data stream.
[0040] Among them, the voice data stream can be regarded as the audio data within a specific duration obtained by the microphone in the set-top box. The audio data within the specific duration can be the audio data at the current moment and within a certain duration before the current moment. This step does not limit the type and number of audio data. For example, the audio data can be phrases, words, etc.
[0041] Specifically, at least one audio data in the voice data stream can be obtained for subsequent steps. This embodiment does not limit the specific obtaining means. Exemplarily, the voice data stream at the current moment and within 2 minutes before the current moment can be obtained, and then each audio data in the voice data stream can be obtained.
[0042] S130. Evaluate the coordinate pair according to the reference gesture library to obtain the target gesture recognition probability of the coordinate pair. The target gesture recognition probability is the recognition probability corresponding to the target gesture. The reference gesture library includes at least one standard gesture instruction.
[0043] The reference gesture library can be understood as a gesture library for reference when evaluating the coordinate pair. The reference gesture library can include at least one standard gesture instruction. The standard gesture instruction can be the instruction corresponding to the standard gesture, such as instructions like click, long press, etc. Among them, the standard gesture can be a gesture preset by the system or relevant personnel to represent the standard gesture instruction. It can be understood that each standard gesture instruction corresponds to a standard gesture. When it is recognized that the coordinate pair corresponds to a certain standard gesture, the standard gesture instruction corresponding to this standard gesture can be executed.
[0044] The target gesture recognition probability can be regarded as the probability that the gesture indicated by the coordinate pair is the target gesture. The target gesture is a gesture similar to or the same as one of the standard gestures.
[0045] In this embodiment, the coordinate pair can be evaluated according to the standard gestures represented by each standard gesture instruction to obtain the target gesture recognition probability of the coordinate pair. This embodiment does not limit the evaluation means, as long as the target gesture recognition probability can be obtained. Exemplarily, the minimum included angle between the vector formed by the coordinate pair and each standard vector corresponding to each standard gesture can be calculated, and then the respective historical minimum included angles corresponding to the first set number of coordinate pairs before the moment where the coordinate pair is located can be calculated; finally, the target gesture recognition probability can be calculated according to the minimum included angle and each historical minimum included angle. This embodiment does not make a limit on this.
[0046] S140. Evaluate each of the audio data according to the reference voice library to obtain the target voice recognition probability of each of the audio data. The target voice recognition probability is the recognition probability corresponding to the target voice. The reference voice library includes at least one standard voice instruction.
[0047] The reference speech library can be understood as a speech library for reference when evaluating each audio data. The reference speech library can include at least one standard speech instruction. The standard speech instruction can be an instruction corresponding to the standard speech, such as instructions like click and long press. Among them, the standard speech can be a gesture preset by the system or relevant personnel to represent the standard speech instruction. The target speech recognition probability can be considered as the probability that the speech indicated by each audio data is the target speech, and the target speech is one of the standard speeches among each standard speech.
[0048] In this embodiment, each audio data can be evaluated according to the standard speech represented by each standard speech instruction to obtain the target speech recognition probability of each audio data. This embodiment does not limit the means of evaluation, as long as the target speech recognition probability can be obtained. Exemplarily, first, each audio data in the speech data stream can be denoised, echo-canceled, etc. to improve the audio quality, and then the processed audio data can be matched with the reference speech library to obtain the target speech recognition probability. This embodiment does not make a limit on this.
[0049] S150. Determine an operation instruction based on the target gesture recognition probability and the target speech recognition probability.
[0050] After obtaining the target gesture recognition probability and the target speech recognition probability, the operation instruction can be determined based on the target gesture recognition probability and the target speech recognition probability. This step does not limit the method for determining the operation instruction. Exemplarily, the magnitudes of the target gesture recognition probability and the target speech recognition probability can be compared, and the operation instruction can be determined according to the larger one of the comparison results; it is also possible to determine an operation probability based on the target gesture recognition probability and the target speech recognition probability, and then determine the operation instruction according to the operation probability.
[0051] A method for determining an operation instruction provided in the first embodiment of the present invention includes: obtaining coordinate pairs in a video data stream, where the coordinate pairs are formed by the coordinates of two joint points on the same arm in a frame of image; obtaining at least one audio data in a voice data stream; evaluating the coordinate pairs according to a reference gesture library to obtain the target gesture recognition probability of the coordinate pairs, where the target gesture recognition probability is the recognition probability corresponding to a target gesture, and the reference gesture library includes at least one standard gesture instruction; evaluating each audio data according to a reference voice library to obtain the target voice recognition probability of each audio data, where the target voice recognition probability is the recognition probability corresponding to a target voice, and the reference voice library includes at least one standard voice instruction; determining an operation instruction based on the target gesture recognition probability and the target voice recognition probability. Using this method, the target gesture recognition probability is accurately obtained by evaluating the coordinate pairs in the obtained video data stream according to the reference gesture library, and the target voice recognition probability is accurately obtained by evaluating the audio data in the obtained voice data stream according to the reference voice library, reducing the complexity of the algorithm; at the same time, by determining the operation instruction in a way that the target gesture recognition probability and the target voice recognition probability are mutually calibrated, the robustness of the algorithm is improved, thus ensuring the correctness of the operation instruction.
[0052] In one embodiment, different gesture instructions correspond to different coordinate pairs.
[0053] In this embodiment, different gesture instructions may correspond to coordinate pairs formed at different joint points. For example, when the gesture instruction is a click instruction, the coordinate pair may be formed by the joint points of the fingertip and the wrist; when the gesture instruction is a rightward instruction, the coordinate pair may be formed by the joint points of the elbow and the wrist.
[0054] It can be understood that the process of evaluating the coordinate pairs can also be understood as a process of matching the gesture instruction corresponding to the coordinate pairs with each standard gesture instruction in the reference gesture library to obtain the target gesture recognition probability.
[0055] In one embodiment, the determining an operation instruction based on the target gesture recognition probability and the target voice recognition probability includes:
[0056] Determining an operation probability based on the target gesture recognition probability and the target voice recognition probability;
[0057] Determining an operation instruction according to the operation probability.
[0058] Among them, the operation probability can be considered as the probability jointly obtained by gesture recognition and voice recognition for determining an operation instruction.
[0059] In this embodiment, first, an operation probability can be determined based on a target gesture recognition probability and a target speech recognition probability, and then an operation instruction can be determined according to the operation probability. This embodiment does not expand on the specific steps for determining the operation probability. Exemplarily, when determining the operation probability, the target gesture recognition probability and the target speech recognition probability can be compared with a first set value. When the target gesture recognition probability is greater than the first set value or the target speech recognition probability is greater than the first set value, the larger value of the target gesture recognition probability and the target speech recognition probability can be taken as the operation probability; when both the target gesture recognition probability and the target speech recognition probability are less than the first set value, a weighted process is performed on the target gesture recognition probability and the target speech recognition probability to determine the operation probability, where the specific value of the first set value is not limited and can be determined by an empirical value.
[0060] In one embodiment, determining the operation probability based on the target gesture recognition probability and the target speech recognition probability includes:
[0061] If the target gesture recognition probability is equal to zero, the target speech recognition probability is determined as the operation probability;
[0062] If the target speech recognition probability is equal to zero, the target gesture recognition probability is determined as the operation probability;
[0063] Otherwise, the weighted value of the target gesture recognition probability, the historical target gesture recognition probabilities of a second set number, the target speech recognition probability, and the historical speech gesture recognition probabilities of a third set number is determined as the operation probability.
[0064] After obtaining the target gesture recognition probability and the target speech recognition probability, the operation probability can be determined comprehensively. Specifically, when the target gesture recognition probability is equal to zero, it can be considered that the accuracy of gesture recognition is not high. At this time, according to the target speech recognition probability obtained by speech recognition, the target speech recognition probability is determined as the operation probability; when the target speech recognition probability is equal to zero, it can be considered that the accuracy of speech recognition is not high. At this time, according to the target gesture recognition probability obtained by gesture recognition, the target gesture recognition probability is determined as the operation probability; when both the target gesture recognition probability and the target speech recognition probability are not equal to zero, it indicates that the operation probability can be determined by comprehensively considering gesture recognition and speech recognition. Here, the weighted value of the target gesture recognition probability, the historical target gesture recognition probabilities of a second set number, the target speech recognition probability, and the historical speech gesture recognition probabilities of a third set number is determined as the operation probability. This embodiment does not limit the specific determination method. Among them, the second set number and the third set number can be limited by relevant personnel, and they can be the same or different.
[0065] Exemplarily, the operation probability can be determined according to the following formula:
[0066]
[0067] Among them, P gr is the target gesture recognition probability, and P vr is the target speech recognition probability, and P 1gr ,..., P (n-1)gr are the recognition probabilities of n-1 historical target gestures, and P 1vr ,..., P (n-1)vr are the recognition probabilities of n-1 historical target speeches, and P ngr = P gr , P nvr = P vr , the second set quantity and the third set quantity are n, and w1,..., w n are the weight values of the recognition probabilities of each target gesture or each target speech.
[0068] Embodiment 2
[0069] Figure 2 is a flowchart of a method for determining an operation instruction according to Embodiment 2 of the present invention. Embodiment 2 is optimized on the basis of the above embodiments.
[0070] In this embodiment, further specifying the target gesture recognition probability of the coordinate pair obtained by evaluating the coordinate pair according to the reference gesture library as: determining the minimum included angle corresponding to the coordinate pair according to the reference gesture library, where the minimum included angle is the minimum value among the included angles of each vector, and the included angle of each vector is the included angle between the feature vector corresponding to the coordinate pair and the standard vector corresponding to each standard gesture instruction in the reference gesture library; determining the target gesture probability of the coordinate pair according to the average value of each of the vector included angles and the minimum variance of the minimum included angle, where the minimum variance of the minimum included angle is determined based on the minimum included angle and the first set quantity of historical minimum included angles; determining the target gesture recognition probability of the coordinate pair based on the target gesture probability, the minimum included angle, and a set threshold.
[0071] As Figure 2 shown, the method includes:
[0072] S210. Obtain a coordinate pair in the video data stream, where the coordinate pair is formed by the coordinates of two joint points on the same arm in one frame of image.
[0073] S220. Obtain at least one audio data in the voice data stream.
[0074] S230. Determine the minimum included angle corresponding to the coordinate pair according to the reference gesture library. The minimum included angle is the minimum value among the included angles of vectors, and the included angles of vectors are the included angles between the feature vectors corresponding to the coordinate pair and the standard vectors corresponding to each standard gesture command in the reference gesture library.
[0075] Among them, the minimum included angle can be the minimum value among the included angles of vectors. The included angles of vectors can be considered as the included angles between the feature vectors corresponding to the coordinate pair and the standard vectors corresponding to each standard gesture command in the reference gesture library. It can be understood that the reference gesture library includes at least one standard gesture command, and each standard gesture command can be represented in the form of a spatial vector, that is, each standard gesture command corresponds to a standard vector.
[0076] After obtaining the coordinate pair in the video data stream, the included angles of vectors between the feature vector corresponding to the coordinate pair and the standard vectors corresponding to each standard gesture command in the reference gesture library can be determined, and then the minimum value among the included angles of vectors is determined as the minimum included angle corresponding to the coordinate pair.
[0077] For example, the reference gesture library contains 9 standard gesture commands, corresponding to 9 standard vectors respectively. Set the 9 standard vectors as a set Y, that is, Y = {M i |i = [1, 9]}. Select the right elbow (N qr ) and the right wrist (N er ) to form a coordinate pair. Let the coordinates of N qr and N er be (x1, y1, z1) and (x2, y2, z2) respectively. Then the feature vector M qe corresponding to the coordinate pair = (x2 - x1, y2 - y1, z2 - z1). Finally, the included angles of vectors can be calculated according to the feature vector M qe and Y = {M i |i = [1, 9]}, and the minimum value among the included angles of vectors is determined as the minimum included angle.
[0078] S240. Determine the target gesture probability of the coordinate pair according to the average value of the included angles of vectors and the minimum variance of the minimum included angle. The minimum variance of the minimum included angle is determined based on the minimum included angle and the minimum included angles of the first set number of historical data.
[0079] The minimum variance of the minimum included angle can refer to the variance corresponding to the minimum included angle, and can be determined based on the minimum included angle and the minimum included angles of the first set number of historical data. Among them, the historical minimum included angle can be considered as the minimum included angle corresponding to the coordinate pair before the moment of the coordinate pair. The first set number can be set by relevant personnel, and this embodiment does not limit this.
[0080] Specifically, after determining the minimum angle corresponding to the coordinate pair, the target gesture probability of the coordinate pair can be determined according to the average value of the angles between vectors and the minimum variance of the minimum angle. The specific determination steps are not limited. For example, first, the average value ave(θ r ) of the angles between vectors corresponding to the coordinate pair can be calculated; then, take 9 historical minimum angles θ t,min , t ∈ [1, 9]. Based on the minimum angle θ min and the 9 (i.e., the first set number) historical minimum angles, the minimum variance can be obtained. Finally, according to the calculated average value ave(θ r ) and the minimum variance , the target gesture probability corresponding to the coordinate pair is determined
[0081] S250. Determine the target gesture recognition probability of the coordinate pair based on the target gesture probability, the minimum angle, and a set threshold.
[0082] The set threshold can be understood as the critical value of the minimum angle. The specific value is not limited and can be set by relevant personnel.
[0083] In this embodiment, the target gesture recognition probability of the coordinate pair can be determined based on the target gesture probability, the minimum angle, and the set threshold. For example, the magnitudes of the minimum angle and the set threshold can be compared. When the minimum angle is less than the set threshold, the target gesture probability is determined as the target gesture recognition probability; when the minimum angle is greater than the set threshold, the target gesture recognition probability is set to zero. This embodiment does not make any limitations in this regard.
[0084] In one embodiment, the determining the target gesture recognition probability of the coordinate pair based on the target gesture probability, the minimum angle, and the set threshold includes:
[0085] If the target gesture probability is less than zero and the minimum angle is greater than the set threshold, then the target gesture recognition probability is equal to zero; otherwise, the target gesture recognition probability is equal to the target gesture probability.
[0086] In this step, when the target gesture probability is less than zero and the minimum angle is greater than the set threshold, it can be considered that the accuracy rate of gesture recognition is very low, and the target gesture recognition probability is determined to be equal to zero; otherwise, it can be considered that the target gesture probability is the target gesture recognition probability, that is, the target gesture recognition probability is equal to the target gesture probability. Exemplarily, when the target gesture probability is P ar , the minimum angle is θ min , and T θ is the set threshold, the formula can be used to calculate the target gesture recognition probability.
[0087] S260. Evaluate each of the audio data according to a reference audio library to obtain the target speech recognition probability of each of the audio data.
[0088] S270. Determine an operation instruction based on the target gesture recognition probability and the target speech recognition probability.
[0089] An operation instruction determination method provided in the second embodiment of the present invention can improve the accuracy of the target gesture recognition probability by first determining the target gesture probability of a coordinate pair and then determining the target gesture recognition probability, thereby making the operation instruction more accurate. At the same time, by comprehensively determining the target gesture probability based on the minimum included angle corresponding to the coordinate pair and the historical minimum included angle, the accuracy of the target gesture probability can be further improved by combining the gesture situation reflected by the coordinate pairs at historical moments.
[0090] In one embodiment, the determining the minimum included angle corresponding to the coordinate pair according to the reference gesture library includes:
[0091] Determine a feature vector based on the coordinate pair, where the starting point of the feature vector is the first joint point and the ending point of the feature vector is the second joint point, and the first joint point and the second joint point are determined based on the determined target gesture.
[0092] Determine the minimum included angle between the feature vector and each standard vector corresponding to each standard gesture instruction in the reference gesture library according to the feature vector and the standard vectors.
[0093] In this embodiment, first, a feature vector can be determined based on the coordinate pair. The starting point of the feature vector can be the first joint point, and the ending point of the feature vector can be the second joint point. Among them, the first joint point and the second joint point are joint points on the same arm in a frame of image, and the first joint point and the second joint point can be determined based on the determined target gesture. For example, when the target gesture is a rightward gesture, the first joint point can be the right elbow and the second joint point can be the right wrist. The first joint point and the second joint point are only used to distinguish different joint points, and this embodiment does not limit this.
[0094] Subsequently, the minimum included angle between the determined feature vector and each standard vector corresponding to each standard gesture instruction in the reference gesture library can be determined.
[0095] The following gives an exemplary description of an operation instruction determination method provided by an embodiment of the present invention.
[0096] First, obtain gesture instructions and audio:
[0097] Twenty bone joint points in a 3D coordinate system and their three-dimensional coordinate systems can be obtained through the video data stream of the PTZ high-definition camera in a 4K smart set-top box (i.e., obtain the coordinate pairs in the video data stream).
[0098] At the bottom of the 4K smart set-top box device, there is a microphone array composed of four independent digital microphones, which can collect voice commands even when far away from the microphone (i.e., obtain at least one audio data in the voice data stream).
[0099] Then, calculate the probability of gesture recognition:
[0100] Represent the nine standard gesture commands in the reference gesture library in the form of spatial vectors, that is, each standard gesture can be represented by a three-dimensional standard direction vector. Set the nine standard vectors as set Y, that is, Y = {M i |i = [1, 9]}.
[0101] Select the right elbow (N qr ) and the right wrist (N er ) as the two joint points for gesture recognition to form a coordinate pair. Use the feature vector M qe starting from the right elbow and ending at the right wrist to recognize various gesture commands. For example, if the coordinates of N qr and N er are (x1, y1, z1) and (x2, y2, z2) respectively, then M qe = (x2 - x1, y2 - y1, z2 - z1).
[0102] Use the obtained M qe to calculate the included angles with the nine standard vectors in set Y, and find the minimum included angle θ min , that is
[0103] Assume that the minimum angle between the current gesture vector and the feature vector is θ min , and the window size is 10 frames. From the 1st frame to the 100th frame, it can be divided into 100 time windows. It is necessary to calculate the minimum variance of the gesture vectors in each window Then the target gesture probability where, ave(θ r ) is the average value of the included angles of each vector, is the variance of the current gesture duration, and θ t,min is the nine (i.e., the first set number) historical minimum included angles.
[0104] Through the formula the correct gesture recognition probability P gr (i.e., the target gesture recognition probability) can be calculated, where, T θ is a fixed threshold (i.e., the set threshold).
[0105] Subsequently, calculate the probability of speech recognition:
[0106] A voice recognition engine is established to analyze and search from specific grammar objects. The grammar object (i.e., the reference voice library) consists of a series of words and phrases (i.e., standard voice instructions). In this embodiment, 9 instructions can be recognized for interactive operations of remote medical consultations in the medical consultation App based on the set-top box. The 9 instructions are as follows:
[0107] Grammar = {"click", "longClick", "left", "right", "northeast", "southeast", "south west", "northwest", "touch"}.
[0108] At least one audio data is obtained from the voice data stream of the microphone, and the audio quality is improved through noise reduction, automatic gain control, and echo cancellation. The voice recognition engine receives the processed audio data to match the grammar library and parse the text result. The parsed result is matched with the words in the grammar object, and the matching probability of each word is calculated, and the maximum matching probability P vr (i.e., the target voice recognition probability) is taken out.
[0109] Finally, the generation of the operation instruction: The operation instruction can be jointly obtained by gesture recognition and voice recognition. Taking the right direction as an example, the probability P of the right operation r is calculated by the formula:
[0110]
[0111] As shown in the above formula, the operation probability P r is calculated from the probability P of right hand gesture recognition gr (i.e., the target gesture recognition probability) and the probability P of right voice recognition vr (the target voice recognition probability). If the voice recognition is unreliable (i.e., P vr = 0), then only gesture recognition can be relied on, and vice versa. If both are reliable, the weighted average of the gesture and voice recognition probabilities is taken for calculation to obtain the correct operation instruction.
[0112] It should be noted that the operation instruction determination method provided in the embodiment of the present invention can be used in the medical consultation APP built in the set-top box. The operation instruction can be determined through the video data stream and voice data stream obtained by the set-top box, and the operation instruction is executed in the medical consultation APP.
[0113] Since human-computer interaction is a hot topic in the field of artificial intelligence, intelligent navigation, as one of the important applications of human-computer interaction, controls the action of the target device through voice or gesture information. The main advantage of intelligent navigation is that it can control the target device within a certain distance without any remote control device. Therefore, the human-computer interaction navigation algorithm can be integrated into the intelligent audio live broadcast function of the consultation APP. Therefore, this embodiment proposes a navigation algorithm that combines gesture recognition and voice recognition and integrates it into the consultation APP. The main steps are as follows:
[0114] Integrate the navigation algorithm SO library into the consultation app in the form of SDK.
[0115] A reference model (ie, a reference voice library and a reference gesture library) is established through 9 gesture instructions and 9 language commands.
[0116] Real-time video and audio information (i.e., video data stream and voice data stream) is extracted through the 4K smart set-top box camera and microphone.
[0117] The matching degree of the current gesture and voice information is evaluated through the reference model, and the operation instructions of the medical consultation app are derived.
[0118] The specific steps for integration in the consultation APP can be: first configure the C++11 compilation environment, then write C++ to implement the interface and pass the interface to the Android side, and finally package the algorithm into a dynamic library (so library) in the configured C++ compilation environment. On the Android side, integrate the packaged dynamic library into the Android system, and then use NDK-build to package and compile again. The Android side directly transmits the audio and video stream data to the algorithm through the interface exposed on the C++ side. At this point, the entire algorithm has been integrated into the consultation app.
[0119] Embodiment 3
[0120] Figure 3 is a schematic diagram of the structure of an operation instruction determination device provided according to Embodiment 3 of the present invention. Figure 3 As shown, the device comprises:
[0121] A first acquisition module 310 is used to acquire a coordinate pair in a video data stream, wherein the coordinate pair is formed by the coordinates of two joint points on the same arm in a frame of image;
[0122] A second acquisition module 320, configured to acquire at least one audio data in the voice data stream;
[0123] The first evaluation module 330 is configured to evaluate each of the coordinate pairs according to a reference gesture library to obtain a target gesture recognition probability for each of the coordinate pairs, where the target gesture recognition probability is the recognition probability corresponding to a target gesture, and the reference gesture library includes at least one standard gesture instruction;
[0124] The second evaluation module 340 is configured to evaluate each of the audio data according to a reference speech library to obtain a target speech recognition probability for each of the audio data, where the target speech recognition probability is the recognition probability corresponding to a target speech, and the reference speech library includes at least one standard speech instruction;
[0125] The determination module 350 is configured to determine an operation instruction based on the target gesture recognition probability and the target speech recognition probability.
[0126] An operation instruction determination device provided in Embodiment 3 of the present invention obtains coordinate pairs in a video data stream through a first acquisition module 310, where the coordinate pairs are formed by the coordinates of two joint points on the same arm in one frame of an image; obtains at least one audio data in a speech data stream through a second acquisition module 320; evaluates each of the coordinate pairs according to a reference gesture library through a first evaluation 330 to obtain a target gesture recognition probability for each of the coordinate pairs, where the target gesture recognition probability is the recognition probability corresponding to a target gesture, and the reference gesture library includes at least one standard gesture instruction; evaluates each of the audio data according to a reference speech library through a second evaluation 340 to obtain a target speech recognition probability for each of the audio data, where the target speech recognition probability is the recognition probability corresponding to a target speech, and the reference speech library includes at least one standard speech instruction; determines an operation instruction based on the target gesture recognition probability and the target speech recognition probability through a determination module 350. Using this device, the target gesture recognition probability is accurately obtained by evaluating the coordinate pairs in the acquired video data stream according to the reference gesture library, and the target speech recognition probability is accurately obtained by evaluating the audio data in the acquired speech data stream according to the reference speech library, reducing the complexity of the algorithm; at the same time, by determining the operation instruction in a mutually calibrated manner based on the target gesture recognition probability and the target speech recognition probability, the robustness of the algorithm is improved, thereby ensuring the correctness of the operation instruction.
[0127] Optionally, the first evaluation module 330 includes:
[0128] The first determination unit is configured to, for each coordinate pair, determine a minimum included angle corresponding to the coordinate pair according to the reference gesture library, where the minimum included angle is the minimum value among the included angles of vectors, and the included angles of vectors are the included angles between the feature vectors corresponding to the coordinate pair and the standard vectors corresponding to the standard gesture instructions in the reference gesture library;
[0129] A second determination unit, configured to determine, for each coordinate pair, a target gesture probability of the coordinate pair according to an average value of the vector angles and a minimum variance of the minimum angle, where the minimum variance of the minimum angle is determined based on the minimum angle and a first set number of historical minimum angles;
[0130] A third determination unit, configured to determine a target gesture recognition probability of the coordinate pair based on the target gesture probability, the minimum angle, and a set threshold.
[0131] Optionally, the first determination unit is specifically configured to:
[0132] Determine a feature vector based on the coordinate pair, where a starting point of the feature vector is a first joint point, and an end point of the feature vector is a second joint point, and the first joint point and the second joint point are determined based on the determined target gesture;
[0133] Determine a minimum angle between the feature vector and each standard vector corresponding to each standard gesture command in a reference gesture library according to the feature vector and the standard vectors.
[0134] Optionally, the third determination unit is specifically configured to:
[0135] If the target gesture probability is less than zero and the minimum angle is greater than the set threshold, then the target gesture recognition probability is equal to zero; otherwise, the target gesture recognition probability is equal to the target gesture probability.
[0136] Optionally, different gesture commands correspond to different coordinate pairs.
[0137] Optionally, the determination module 350 includes:
[0138] An operation probability determination unit, configured to determine an operation probability based on the target gesture recognition probability and the target speech recognition probability;
[0139] An operation instruction determination unit, configured to determine an operation instruction according to the operation probability.
[0140] Optionally, the operation probability determination unit is specifically configured to:
[0141] If the target gesture recognition probability is equal to zero, then determine the target speech recognition probability as the operation probability;
[0142] If the target speech recognition probability is equal to zero, then determine the target gesture recognition probability as the operation probability;
[0143] Otherwise, determine a weighted value of the target gesture recognition probability, a second set number of historical target gesture recognition probabilities, the target speech recognition probability, and a third set number of historical speech gesture recognition probabilities as the operation probability.
[0144] The operation instruction determination device provided by the embodiments of the present invention can execute the operation instruction determination method provided by any embodiment of the present invention, and has the corresponding functional modules and beneficial effects for executing the method.
[0145] Embodiment 4
[0146] Figure 4 is a schematic structural diagram of a set-top box for implementing the operation instruction determination method of Embodiment 1 of the present invention. As Figure 4 shown, the set-top box provided in Embodiment 4 of the present invention includes: a camera 1; a microphone 2; and a controller 3, which is communicatively connected to the camera 1 and the microphone 2 respectively.
[0147] The controller 3 includes: at least one processor 31; and a storage device 32 communicatively connected to the at least one processor 31; the processor 31 in the controller 3 can be one or more, Figure 4 taking one processor 31 as an example; the storage device 32 is used to store one or more programs; the one or more programs are executed by the one or more processors 31, so that the one or more processors 31 implement the operation instruction determination method described in any one of the embodiments of the present invention.
[0148] The processor 31 and the storage device 32 in the set-top box can be connected by a bus or other means, Figure 4 taking the connection by bus as an example.
[0149] The storage device 32 in the set-top box, as a computer-readable storage medium, can be used to store one or more programs, and the programs can be software programs, computer-executable programs, and modules, such as the program instructions / modules corresponding to the operation instruction determination methods provided in Embodiment 1 or Embodiment 2 of the present invention (for example, the modules in the operation instruction determination device shown in Figure 3 include: a first acquisition module 310, a second acquisition module 320, a first evaluation module 330, a second evaluation module 340, and a determination module 350). The processor 31 executes various functional applications and data processing of the electronic device by running the software programs, instructions, and modules stored in the storage device 32, that is, implements the operation instruction determination method in the above method embodiments.
[0150] The storage device 32 may include a program storage area and a data storage area. The program storage area may store an operating system and application programs required for at least one function. The data storage area may store data created according to the use of the electronic device and the like. In addition, the storage device 32 may include a high-speed random access memory and may also include a non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other non-volatile solid-state storage devices. In some instances, the storage device 32 may further include a memory remotely provided with respect to the processor 31, and these remote memories may be connected to the device through a network. Examples of the above network include but are not limited to the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0151] Moreover, when one or more programs included in the above controller 3 are executed by the one or more processors 31, the programs perform the following operations:
[0152] Obtain coordinate pairs in the video data stream, where the coordinate pairs are formed by the coordinates of two joint points on the same arm in one frame of the image;
[0153] Obtain at least one audio data in the voice data stream;
[0154] Evaluate each of the coordinate pairs according to a reference gesture library to obtain the target gesture recognition probability of each of the coordinate pairs. The target gesture recognition probability is the recognition probability corresponding to the target gesture, and the reference gesture library includes at least one standard gesture instruction;
[0155] Evaluate each of the audio data according to a reference voice library to obtain the target voice recognition probability of each of the audio data. The target voice recognition probability is the recognition probability corresponding to the target voice, and the reference voice library includes at least one standard voice instruction;
[0156] Determine an operation instruction based on the target gesture recognition probability and the target voice recognition probability.
[0157] Embodiment Five
[0158] Embodiment Five of the present invention provides a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, it is used to execute a method for determining an operation instruction. The method includes:
[0159] Obtain coordinate pairs in the video data stream, where the coordinate pairs are formed by the coordinates of two joint points on the same arm in one frame of the image;
[0160] Obtain at least one audio data in the voice data stream;
[0161] Evaluating each of the coordinate pairs according to a reference gesture library to obtain a target gesture recognition probability for each of the coordinate pairs, where the target gesture recognition probability is the recognition probability corresponding to a target gesture, and the reference gesture library includes at least one standard gesture instruction;
[0162] Evaluating each of the audio data according to a reference voice library to obtain a target voice recognition probability for each of the audio data, where the target voice recognition probability is the recognition probability corresponding to a target voice, and the reference voice library includes at least one standard voice instruction;
[0163] Determining an operation instruction based on the target gesture recognition probability and the target voice recognition probability.
[0164] Optionally, when executed by a processor, the program can also be used to execute the operation instruction determination method provided in any embodiment of the present invention.
[0165] The computer storage medium of the embodiments of the present invention can adopt any combination of one or more computer-readable media. The computer-readable medium can be a computer-readable signal medium or a computer-readable storage medium. The computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (non-exhaustive list) of the computer-readable storage medium include: an electrical connection having one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM), a flash memory, an optical fiber, a portable CD-ROM, an optical storage device, a magnetic storage device, or any suitable combination of the above. The computer-readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or in combination with an instruction execution system, apparatus, or device.
[0166] The computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal can take various forms, including but not limited to: electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium can also be any computer-readable medium other than the computer-readable storage medium, and this computer-readable medium can send, propagate, or transmit a program for use by or in combination with an instruction execution system, apparatus, or device.
[0167] The program code contained on a computer-readable medium can be transmitted using any suitable medium, including but not limited to: wireless, wire, optical fiber cable, radio frequency (RF), etc., or any suitable combination of the above.
[0168] The computer program code for performing the operations of the present invention can be written in one or more programming languages or combinations thereof. The programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computer (for example, by using an Internet service provider to connect through the Internet).
[0169] Note that the above is only the preferred embodiment of the present invention and the technical principles applied. Those skilled in the art will understand that the present invention is not limited to the specific embodiments described herein. Various obvious changes, re-adjustments, and substitutions can be made by those skilled in the art without departing from the protection scope of the present invention. Therefore, although the present invention has been described in more detail through the above embodiments, the present invention is not limited to the above embodiments. Without departing from the concept of the present invention, it can also include more other equivalent embodiments, and the scope of the present invention is determined by the scope of the appended claims.
Claims
1. A method for determining an operation instruction, characterized in that, Applied to a set-top box, the method includes: Obtain coordinate pairs in a video data stream, where the coordinate pairs are formed by the coordinates of two joint points on the same arm in a frame of an image; Obtain at least one audio data in a voice data stream; Evaluate the coordinate pairs according to a reference gesture library to obtain the target gesture recognition probability of the coordinate pairs, including: Determine the minimum included angle corresponding to the coordinate pairs according to the reference gesture library, where the minimum included angle is the minimum value among the included angles of vectors, and each included angle of vectors is the included angle between the feature vector corresponding to the coordinate pairs and the standard vectors corresponding to each standard gesture instruction in the reference gesture library; Determine the target gesture probability of the coordinate pairs according to the average value of each of the included angles of vectors and the minimum variance of the minimum included angle, where the minimum variance of the minimum included angle is determined based on the minimum included angle and the first set number of historical minimum included angles; Based on the target gesture probability, the minimum included angle, and a set threshold, determine the target gesture recognition probability of the coordinate pairs; The target gesture recognition probability is the recognition probability corresponding to the target gesture, and the reference gesture library includes at least one standard gesture instruction; Evaluate each of the audio data according to a reference voice library to obtain the target voice recognition probability of each of the audio data, where the target voice recognition probability is the recognition probability corresponding to the target voice, and the reference voice library includes at least one standard voice instruction; Determine an operation instruction based on the target gesture recognition probability and the target voice recognition probability.
2. The method according to claim 1, wherein The determining the minimum included angle corresponding to the coordinate pairs according to the reference gesture library includes: Determine a feature vector based on the coordinate pairs, where the starting point of the feature vector is the first joint point, and the ending point of the feature vector is the second joint point, and the first joint point and the second joint point are determined based on the determined target gesture; Determine the minimum included angle between the feature vector and each of the standard vectors corresponding to the standard gesture instructions in the reference gesture library according to the feature vector and the standard vectors.
3. The method according to claim 1, wherein The determining the target gesture recognition probability of the coordinate pairs based on the target gesture probability, the minimum included angle, and a set threshold includes: If the target gesture probability is less than zero and the minimum included angle is greater than the set threshold, then the target gesture recognition probability is equal to zero; otherwise, the target gesture recognition probability is equal to the target gesture probability.
4. The method according to claim 1, wherein Different gesture instructions correspond to different coordinate pairs.
5. The method according to claim 1, characterized in that, The determining an operation instruction based on the target gesture recognition probability and the target voice recognition probability includes: Determine an operation probability based on the target gesture recognition probability and the target voice recognition probability; Determine an operation instruction according to the operation probability.
6. The method according to claim 5, characterized in that The determining an operation probability based on the target gesture recognition probability and the target voice recognition probability includes: If the target gesture recognition probability is equal to zero, then determine the target voice recognition probability as the operation probability; If the target voice recognition probability is equal to zero, then determine the target gesture recognition probability as the operation probability; Otherwise, determine the weighted value of the target gesture recognition probability, the second set number of historical target gesture recognition probabilities, the target voice recognition probability, and the third set number of historical voice recognition probabilities as the operation probability.
7. An operation instruction determination device, characterized in that, Includes: A first acquisition module, configured to acquire coordinate pairs in a video data stream, where the coordinate pairs are formed by coordinates of two joint points on the same arm in a frame of image; A second acquisition module, configured to acquire at least one audio data in an audio data stream; A first evaluation module, configured to evaluate the coordinate pairs according to a reference gesture library to obtain a target gesture recognition probability of the coordinate pairs, where the target gesture recognition probability is a recognition probability corresponding to a target gesture, and the reference gesture library includes at least one standard gesture instruction; A second evaluation module, configured to evaluate each of the audio data according to a reference audio library to obtain a target audio recognition probability of each of the audio data, where the target audio recognition probability is a recognition probability corresponding to a target audio, and the reference audio library includes at least one standard audio instruction; A determination module, configured to determine an operation instruction based on the target gesture recognition probability and the target audio recognition probability; The first evaluation module further includes a first determination unit, a second determination unit, and a third determination unit; The first determination unit is configured to determine a minimum included angle corresponding to the coordinate pairs according to the reference gesture library, where the minimum included angle is the minimum value among included angles of vectors, and the included angles of vectors are included angles between feature vectors corresponding to the coordinate pairs and standard vectors corresponding to respective standard gesture instructions in the reference gesture library; The second determination unit is configured to determine a target gesture probability of the coordinate pairs according to an average value of the included angles of the vectors and a minimum variance of the minimum included angle, where the minimum variance of the minimum included angle is determined based on the minimum included angle and a first set number of historical minimum included angles; The third determination unit is configured to determine a target gesture recognition probability of the coordinate pairs based on the target gesture probability, the minimum included angle, and a set threshold.
8. A set-top box, characterized in that, The set-top box includes: A camera; A microphone; A controller, communicatively connected to the camera and the microphone respectively, where the controller includes: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program executable by the at least one processor, and when the computer program is executed by the at least one processor, the at least one processor is enabled to execute the operation instruction determination method according to any one of claims 1-6.
9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions, and when the computer instructions are executed by a processor, the operation instruction determination method according to any one of claims 1-6 is implemented.
Citation Information
Patent Citations
Emotion interaction based Peking Opera teaching system
CN106020440A
Man-machine interaction method, device and terminal
CN108986801A
3D gesture recognition method, device and system
CN112507924A