Method, device and electronic device for controlling a robotic arm based on voice information

By extracting features from the ambient sound information and determining the action instruction set, the problem of poor interactivity of the existing robot arm control methods is solved, and efficient interaction and dynamic response between the robot arm and the sound information is achieved.

CN119772886BActive Publication Date: 2025-06-10BEIJING APAILANG CREATIVITY TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411966041.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-30
Publication Date
2025-06-10
Estimated Expiration
2044-12-30

AI Technical Summary

Technical Problem

The existing robotic arm control method relies on pre-set fixed action instructions, resulting in poor interactivity and difficult to meet the control needs of different users.

Method used

By obtaining target sound information from the environmental sound information of the environment in which the robot arm is located, feature extraction is performed, the target performance data and action command set of the robot arm are determined, and the movement of the robot arm is controlled based on these instructions.

Benefits of technology

It realizes good interaction between the robotic arm and the sound information, improves the interactivity and flexibility in the robotic arm control process, and can adjust and respond to actions in real time according to the target sound information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119772886B_ABST
    Figure CN119772886B_ABST
Patent Text Reader

Abstract

The present application provides a robotic arm control method, device and electronic device based on sound information, relating to the technical field of robot control. In the present application, target sound information is obtained from the environmental sound information of the environment where the robotic arm is located; the target sound information includes motion control keywords for controlling the movement of the robotic arm; feature extraction is performed on the target sound information to obtain a first feature set and a second feature set; wherein, the first feature set includes multiple physical features of the target sound information, and the second feature set includes multiple semantic features of the target sound information; based on the first feature set and the second feature set, the target performance data of the robotic arm and the action instruction set associated with the target performance data are determined; the robotic arm is controlled based on the action instruction set to display the target performance data. It can be seen that according to the target sound information, real-time adjustment and dynamic response of the robotic arm actions can be achieved, thereby increasing the interactivity and flexibility in the process of robotic arm control.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of robot control technology, and particularly to a robotic arm control method, device, and electronic device based on voice information. Background Art

[0002] With the continuous development of artificial intelligence (AI) technology and robotics technology, robotic arms, which are common basic forms of robots, have been widely used in many fields of daily life. In related application scenarios, a robotic arm usually completes a series of actions corresponding to a preset action instruction.

[0003] Therefore, with the above-mentioned robotic arm control method, since the preset action instruction is a fixed control instruction, the series of actions corresponding to the action instruction completed by the robotic arm will also be fixed, resulting in poor interactivity and difficulty in meeting the robotic arm control requirements of different users.

[0004] In view of this, how to achieve good interaction between a robotic arm and voice information during the robotic arm control process is an urgent problem to be solved currently. Summary of the Invention

[0005] Embodiments of this application provide a robotic arm control method, device, and electronic device based on voice information, so as to achieve good interaction between a robotic arm and voice information during the robotic arm control process.

[0006] In a first aspect, embodiments of this application provide a robotic arm control method based on voice information, and the method includes:

[0007] Obtain target voice information from the environmental voice information of the environment where the robotic arm is located; the target voice information includes motion control keywords for controlling the movement of the robotic arm;

[0008] Extract features from the target voice information to obtain a first feature set and a second feature set; wherein, the first feature set includes multiple physical features of the target voice information, and the second feature set includes multiple semantic features of the target voice information;

[0009] Based on the first feature set and the second feature set, determine the target performance data of the robotic arm and the action instruction set associated with the target performance data;

[0010] Control the robotic arm to display the target performance data based on the action instruction set.

[0011] In an optional embodiment, obtaining target voice information from the environmental voice information of the environment where the robotic arm is located includes:

[0012] Perform noise reduction processing on the environmental sound information, filter out the environmental noise included in the environmental sound information, and obtain the environmental sound information after noise reduction processing;

[0013] Perform normalization processing on the environmental sound information after noise reduction processing to obtain the target sound information.

[0014] In an alternative embodiment, perform feature extraction on the target sound information to obtain a first feature set and a second feature set, including:

[0015] Perform sound feature extraction on the target sound information based on the sound feature extraction methods respectively set for multiple physical features to obtain the first feature set; and,

[0016] Perform semantic analysis on the target sound information to obtain the second feature set.

[0017] In an alternative embodiment, the first feature set includes any one or a combination of pitch feature, volume feature, timbre feature, speech rate feature, duration feature, and spatial position feature; wherein, the duration feature is used to represent the duration of the target sound information, and the spatial position feature is used to represent the sound source position coordinates corresponding to the target sound information.

[0018] In an alternative embodiment, based on the first feature set and the second feature set, determine the target performance data of the robotic arm and the action instruction set associated with the target performance data, including:

[0019] Determine the target performance data that matches the target semantic feature representing the motion control keyword in the second feature set from the preset candidate performance data set;

[0020] Based on the mapping relationship between the preset features and action parameters, and the multiple physical features included in the first feature set and at least one semantic feature other than the target semantic feature in the second feature set, determine the subset of action parameters corresponding to each action in the target performance data;

[0021] Generate an action instruction set based on the obtained multiple subsets of action parameters.

[0022] In an alternative embodiment, generating an action instruction set based on the obtained multiple subsets of action parameters includes:

[0023] For multiple subsets of action parameters, perform the following operations respectively:

[0024] Based on the conversion relationship between the preset action parameters and action instructions, generate the first action instruction corresponding to the first subset of action parameters; wherein, the first subset of action parameters is any one of the multiple subsets of action parameters;

[0025] Save the first action instruction to the action instruction set.

[0026] In an alternative embodiment, controlling the robotic arm to display target performance data based on an action instruction set includes:

[0027] Controlling the robotic arm to display multiple actions included in the target performance data according to the instruction execution order of multiple action instructions included in the action instruction set;

[0028] Wherein, after each action is displayed by the robotic arm, the following operations are performed:

[0029] Obtaining feedback information of the action displayed by the robotic arm at the current moment; the feedback information is used to indicate the completion of the action, as well as the first feature subset and the second feature subset at the next moment adjacent to the current moment;

[0030] Modifying the mapping relationship between the preset features and action parameters and the conversion relationship between the preset action parameters and action instructions based on the feedback information to obtain a modified mapping relationship and a modified conversion relationship;

[0031] Adjusting the motion instruction at the next moment based on the modified mapping relationship and the modified conversion relationship.

[0032] In an alternative embodiment, modifying the mapping relationship between the preset features and action parameters and the conversion relationship between the preset action parameters and action instructions based on the feedback information includes:

[0033] Determining a first modification factor for the mapping relationship and a second modification factor for the conversion relationship based on the feedback information;

[0034] Modifying the mapping relationship based on the first modification factor and modifying the conversion relationship based on the second modification factor.

[0035] In a second aspect, an embodiment of the present application further provides a robotic arm control device based on sound information, and the device includes:

[0036] An information acquisition module, configured to acquire target sound information from the ambient sound information of the environment where the robotic arm is located; the target sound information includes motion control keywords for controlling the movement of the robotic arm;

[0037] A feature extraction module, configured to extract features from the target sound information to obtain a first feature set and a second feature set; wherein, the first feature set includes multiple physical features of the target sound information, and the second feature set includes multiple semantic features of the target sound information;

[0038] An instruction determination module, configured to determine the target performance data of the robotic arm and an action instruction set associated with the target performance data based on the first feature set and the second feature set;

[0039] The instruction control module is used to control the robotic arm to display target performance data based on an action instruction set.

[0040] In an alternative embodiment, when obtaining target sound information from the environmental sound information of the environment where the robotic arm is located, the information acquisition module specifically is used for:

[0041] Perform noise reduction processing on the environmental sound information to filter out the environmental noise included in the environmental sound information, and obtain the environmental sound information after noise reduction processing;

[0042] Perform normalization processing on the environmental sound information after noise reduction processing to obtain the target sound information.

[0043] In an alternative embodiment, when performing feature extraction on the target sound information to obtain a first feature set and a second feature set, the feature extraction module specifically is used for:

[0044] Perform sound feature extraction on the target sound information based on the sound feature extraction methods respectively set for multiple physical features to obtain a first feature set; and,

[0045] Perform semantic analysis on the target sound information to obtain a second feature set.

[0046] In an alternative embodiment, when determining the target performance data of the robotic arm and the action instruction set associated with the target performance data based on the first feature set and the second feature set, the instruction determination module specifically is used for:

[0047] Determine the target performance data that matches the target semantic feature representing the motion control keyword in the second feature set from a preset candidate performance data set;

[0048] Based on the mapping relationship between the preset features and action parameters, and multiple physical features included in the first feature set and at least one semantic feature other than the target semantic feature in the second feature set, determine the subset of action parameters corresponding to each action in the target performance data;

[0049] Generate an action instruction set based on the obtained multiple subsets of action parameters.

[0050] In an alternative embodiment, when generating an action instruction set based on the obtained multiple subsets of action parameters, the instruction determination module specifically is used for:

[0051] For multiple subsets of action parameters, respectively perform the following operations:

[0052] Based on the conversion relationship between the preset action parameters and action instructions, generate the first action instruction corresponding to the first subset of action parameters; wherein, the first subset of action parameters is any one of the multiple subsets of action parameters;

[0053] Save the first action instruction to the action instruction set.

[0054] In an alternative embodiment, when controlling the robotic arm to display target performance data based on the action instruction set, the instruction control module is specifically configured to:

[0055] Control the robotic arm to display multiple actions included in the target performance data according to the instruction execution order of the multiple action instructions included in the action instruction set;

[0056] Wherein, after each action is displayed by the robotic arm, the following operations are performed:

[0057] Obtain feedback information on the action currently displayed by the robotic arm; the feedback information is used to indicate the completion status of the action, as well as the first feature subset and the second feature subset at the next moment adjacent to the current moment;

[0058] Modify the preset mapping relationship between features and action parameters and the preset conversion relationship between action parameters and action instructions based on the feedback information to obtain a modified mapping relationship and a modified conversion relationship;

[0059] Adjust the motion instruction for the next moment based on the modified mapping relationship and the modified conversion relationship.

[0060] In an alternative embodiment, when modifying the preset mapping relationship between features and action parameters and the preset conversion relationship between action parameters and action instructions based on the feedback information, the instruction control module is specifically configured to:

[0061] Determine a first modification factor for the mapping relationship and a second modification factor for the conversion relationship based on the feedback information;

[0062] Modify the mapping relationship based on the first modification factor and modify the conversion relationship based on the second modification factor.

[0063] In a third aspect, an embodiment of the present application further provides an electronic device, including:

[0064] A processor; and

[0065] A memory storing a program,

[0066] wherein the program includes instructions that, when executed by the processor, cause the processor to execute the robotic arm control method based on voice information as described in the first aspect.

[0067] In a fourth aspect, an embodiment of the present application further provides a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause a computer to execute the robotic arm control method based on voice information as described in the first aspect.

[0068] In a fifth aspect, the present application provides a computer program product which, when called by a computer, causes the computer to execute the steps of the robotic arm control method based on voice information as described in the first aspect.

[0069] The beneficial effects of the present application are as follows:

[0070] In the robotic arm control method based on voice information provided in the embodiments of the present application, after obtaining the ambient voice information of the environment where the robotic arm is located, the target voice information included in the ambient voice information can be subjected to feature extraction to obtain a first feature set including a plurality of physical features and a second feature set including a plurality of semantic features. Then, based on the first feature set and the second feature set, the target performance data of the robotic arm and the action instruction set associated with the target performance data are determined, and further, the robotic arm is controlled based on the action instruction set to display the target performance data. Thus, it can be seen that the real-time adjustment and dynamic response of the robotic arm's actions can be achieved according to the target voice information, thereby increasing the interactivity and flexibility in the process of robotic arm control.

[0071] In addition, other features and advantages of the present application will be described in the subsequent specification, and part of them will become obvious from the specification, or can be understood by implementing the present application. The objectives and other advantages of the present application can be realized and obtained through the structures specifically pointed out in the written specification, claims, and drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0072] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings described here are used to provide a further understanding of the present application and constitute a part of the present application, without unduly limiting the present application. In the drawings:

[0073] Figure 1 is a schematic flowchart of an implementation process of a robotic arm control method based on voice information applicable to the embodiments of the present application;

[0074] Figure 2 is a schematic logical diagram of feature extraction for target voice information provided in the embodiments of the present application;

[0075] Figure 3 is a schematic flowchart of an implementation process of a method for determining the target performance data of a robotic arm and its corresponding action instruction set provided in the embodiments of the present application;

[0076] Figure 4 is a schematic structural diagram of a robotic arm control device based on voice information provided in the embodiments of the present application;

[0077] Figure 5A schematic structural diagram of an electronic device provided by an embodiment of the present application. Detailed implementation manners

[0078] The embodiments of the present application will be described in more detail with reference to the accompanying drawings. Although some embodiments of the present application are shown in the drawings, it should be understood that the present application can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. On the contrary, these embodiments are provided to more thoroughly and completely understand the present application. It should be understood that the drawings and embodiments of the present application are only for exemplary purposes and are not used to limit the protection scope of the present application.

[0079] It should be understood that the steps described in the method embodiments of the present application can be executed in different orders and / or in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present application is not limited in this regard.

[0080] The term "including" and its variants used herein are open-ended, that is, "including but not limited to". The term "based on" is "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". The relevant definitions of other terms will be given in the following description. It should be noted that the concepts such as "first" and "second" mentioned in the present application are only used to distinguish different devices, modules or units, and are not used to limit the order or interdependence relationship of the functions performed by these devices, modules or units.

[0081] It should be noted that the modifications of "one" and "multiple" mentioned in the present application are illustrative rather than restrictive. Those skilled in the art should understand that unless clearly stated otherwise in the context, it should be understood as "one or more".

[0082] The names of the messages or information exchanged between multiple devices in the embodiments of the present application are only for illustrative purposes and are not used to limit the scope of these messages or information.

[0083] First, a brief introduction to the design concept of the embodiments of the present application is given below:

[0084] With the continuous development of AI technology and robotics, robots are gradually moving from traditional industrial application fields to many more life - related fields such as service and entertainment. Among them, robot performances (such as robot dances), as a novel form of performance, have attracted more and more attention. In related technologies, the implementation of robot dances mainly relies on pre - compiled program instructions to control the robot to complete a series of fixed actions. Although this method can achieve robot dances, it lacks interactivity and cannot adjust dance movements in real - time according to changes in the external environment, making it difficult to bring a better viewing experience to the audience.

[0085] In view of this, how to achieve good interaction between the robotic arm and sound information during the control process of the robotic arm, which is a common basic form of robots, that is, to improve interactivity, is an urgent problem to be solved currently. Therefore, in order to improve the interactivity during the control process of the robotic arm, the embodiments of this application propose a robotic - arm control method based on sound information, which may specifically include: First, obtain target sound information from the environmental sound information of the environment where the robotic arm is located; the target sound information includes motion - control keywords for controlling the movement of the robotic arm. Then, perform feature extraction on the target sound information to obtain a first feature set and a second feature set; among them, the first feature set may include multiple physical features of the target sound information, and the second feature set may include multiple semantic features of the target sound information. Further, based on the first feature set and the second feature set, determine the target performance data of the robotic arm and the action instruction set associated with the target performance data. Finally, control the robotic arm to display the target performance data based on the action instruction set.

[0086] It should be noted that the execution subject of the robotic - arm control method based on sound information provided by the embodiments of this application may be one or more communication devices, and the embodiments of this application do not make any limitations in this regard. Among them, the communication device may be a terminal device or a server that can control the robotic arm, and of course, it may also be other communication devices, and the embodiments of this application do not make any limitations in this regard.

[0087] A terminal device is a device that can provide voice and / or data connectivity to users and can be a device supporting wired and / or wireless connection methods. Exemplarily, the terminal device may include, but is not limited to: mobile phones, tablet computers, laptop computers, handheld computers, mobile internet devices (MIDs), wearable devices, virtual reality (VR) devices, augmented reality (AR) devices, wireless terminal devices in industrial control, wireless terminal devices in unmanned driving, wireless terminal devices in smart grids, wireless terminal devices in transportation safety, wireless terminal devices in smart cities, or wireless terminal devices in smart homes, etc. In addition, relevant clients may be installed on the terminal device, and the client may be software, for example, application programs (APPs), browsers, short video software, etc., or may also be web pages, applets, etc.

[0088] The server can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or can also be a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery network (CDN), as well as big data and artificial intelligence platforms.

[0089] The following describes the robotic arm control method based on voice information provided by the exemplary embodiments of the present application with reference to the accompanying drawings. Refer to Figure 1 As shown, it is a schematic diagram of the implementation process of a robotic arm control method based on voice information provided by an embodiment of the present application. Taking the server as an example for the execution subject, the specific implementation process of this method is as follows:

[0090] S101: Obtain target voice information from the environmental voice information of the environment where the robotic arm is located that is collected.

[0091] Among them, the above-mentioned target voice information may include motion control keywords for controlling the movement of the robotic arm. Taking the target voice information as "Spring has arrived" as an example, among them, "Spring" can be a motion control keyword for subsequent control of the movement of the robotic arm. After the server obtains the motion control keywords included in the target voice information, it can determine the performance data to be subsequently displayed by the robotic arm. For example, the robotic arm will display a dance program or a drama interface related to "Spring", etc.

[0092] For example, when performing step S101, the server can collect the environmental sound information of the environment where the robotic arm is located through the sound collection module on the terminal device, so as to receive the collected environmental sound information of the environment where the robotic arm is located uploaded by the terminal device. The aforementioned sound collection module can be one or more voice collection devices (such as a microphone array) for sound data collection. The aforementioned environmental sound information can include all sounds within a set range where the robotic arm is located (such as an area with the robotic arm as the center and a radius of 5 m), such as music, the voice of the target object, and other background sounds, etc.

[0093] In addition, the aforementioned target sound information can be voice information such as the voice information of the target object. Of course, it can also be non-voice information including music or knocking sounds, etc. This application does not make specific limitations in this regard. In this way, after obtaining the voice information and / or non-voice information from the environmental sound information, the information related to the motion control of the robotic arm can be obtained by parsing the voice information and / or non-voice information.

[0094] In an optional implementation manner, after the server obtains the environmental sound information of the environment where the robotic arm is located, it can perform noise reduction processing on the environmental sound information, filter out the environmental noise included in the environmental sound information, obtain the environmental sound information after noise reduction processing, and then perform normalization processing on the environmental sound information after noise reduction processing to obtain the target sound information.

[0095] In this way, the environmental noise included in the collected original sound information (i.e., environmental sound information) is removed, and the more accurate sound information containing the target sound information is retained, ensuring the accuracy of the subsequent motion control of the robotic arm based on the obtained target sound information. Moreover, the target sound information obtained by performing normalization processing on the environmental sound information after noise reduction processing can improve the efficiency of subsequent feature extraction of the target sound information.

[0096] Exemplarily, the server can filter out the environmental noise included in the environmental sound information through a set filtering method (such as Kalman filtering, wavelet filtering, etc.) to obtain the environmental sound information after noise reduction processing. Among them, the environmental sound information after noise reduction processing includes the target sound information. The aforementioned normalization processing can normalize the signal amplitude in the environmental sound information after noise reduction processing to a certain range, such as within [0, 1]. This application does not make specific limitations in this regard.

[0097] S102: Extract features from the target sound information to obtain a first feature set and a second feature set.

[0098] The above-mentioned first feature set may include multiple physical features of the target sound information, and the above-mentioned second feature set may include multiple semantic features of the target sound information. Optionally, the above-mentioned first feature set may include any one or a combination of pitch feature, volume feature, timbre feature, speech rate feature, duration feature, and spatial position feature corresponding to the target sound information. Among them, the aforementioned duration feature may be used to represent the duration of the target sound information, and the aforementioned spatial position feature may be used to represent the sound source position coordinates corresponding to the target sound information.

[0099] In an alternative implementation, referring to Figure 2 As shown, when executing step S102, after obtaining the target sound information, the server may perform sound feature extraction on the target sound information based on the sound feature extraction methods respectively set for multiple physical features to obtain the first feature set; and perform semantic analysis on the target sound information to obtain the second feature set.

[0100] Exemplarily, the server may adopt time-frequency transformation methods such as fast fourier transform (FFT) (i.e., the sound feature extraction method corresponding to the pitch feature) to convert the time-domain signal into a frequency-domain signal, so as to extract the spectral information of the target sound information, and further obtain pitch features such as the main frequency and spectral range. Assuming that the speech signal of the above-mentioned target object is x(n), the aforementioned FFT transformation formula may be specifically expressed as follows:

[0101]

[0102] Among them, X(k) is the frequency-domain signal, x(n) is the time-domain signal, that is, the speech signal of the target object, and N is the length of the time-domain signal.

[0103] The server may represent the volume of the target object by calculating the root mean square (RMS) of the speech signal of the target object. Optionally, the calculation formula of the aforementioned RMS may be specifically expressed as follows:

[0104]

[0105] Among them, V is the root mean square of the speech signal (i.e., the volume feature), x(n) is the time-domain signal, that is, the speech signal of the target object, and N is the length of the time-domain signal.

[0106] The server can extract timbre features using mel-scale frequency cepstral coefficients (MFCC), map the speech signal to the mel frequency domain, and perform discrete cosine transform (DCT) to obtain MFCC. Optionally, the calculation formula for the mel frequency corresponding to the speech signal can be specifically expressed as follows:

[0107]

[0108] Where m is the mel frequency corresponding to the speech signal, f is the actual speech frequency of the speech signal, and a is derived from an empirical formula. The purpose is to maintain a certain linear relationship when mapping the frequency to the mel frequency domain and at the same time conform to the characteristics of the human ear's perception of frequency. a can be 2595. The frequency near b is a region where the human ear is relatively sensitive to frequency changes. Therefore, this value is commonly used as a reference in the calculation of mel frequency, and b can be 700.

[0109] The above-mentioned speech rate feature can be determined by the server by counting the number of words in the above-mentioned target voice information per unit time. The aforementioned unit time can be 0.5 seconds, 1 second, 2 seconds, etc., and the embodiments of the present application do not make specific limitations in this regard.

[0110] The above-mentioned duration feature can be that the server determines the start time and end time of the speech signal through endpoint detection of the speech signal, and calculates the duration of the speech signal. The calculation formula for the aforementioned duration can be specifically expressed as: Δt = t 2 -t 1 , where Δt is the duration, t 1 is the start time, and t 2 is the end time.

[0111] The above-mentioned spatial position feature can be determined by the server through the sound source localization technology set for the sound collection module. Taking the aforementioned sound collection module as a microphone array as an example. According to the time difference or phase difference of the speech signals received by different microphone units, the azimuth angle and elevation angle of the target object are calculated, that is, the spatial position feature of the target object.

[0112] It should be understood that when the server performs sound feature extraction on the target sound information based on the sound feature extraction methods respectively set for multiple physical features and obtains the first feature set, it can analyze parameters such as the fundamental frequency and energy of the speech signal, so as to determine physical features such as the pitch feature, volume feature, timbre feature, speech rate feature, duration feature, and spatial position feature corresponding to the target sound information.

[0113] For another example, the server may adopt natural language processing (NLP) technology to perform automatic speech recognition (ASR) on the speech signal, convert the target voice information into text information, and then perform semantic understanding on the foregoing text information to obtain the second feature set.

[0114] S103: Based on the first feature set and the second feature set, determine the target performance data of the robotic arm and the action instruction set associated with the target performance data.

[0115] In an alternative implementation manner, when executing step S103, the server may refer to the Figure 3 method shown to determine the target performance data of the robotic arm and the action instruction set associated with the target performance data. The specific process is as follows:

[0116] S301: From the preset candidate performance data set, determine the target performance data that matches the target semantic feature representing the motion control keyword in the second feature set.

[0117] Exemplarily, the above-mentioned candidate performance data set may include dance tracks for different festivals. Assuming that the target semantic feature representing the motion control keyword in the second feature set is "Spring Festival", then the server may determine the dance track related to the "Spring Festival" from the above-mentioned candidate performance data set. Optionally, if there are multiple dance tracks related to the "Spring Festival", the dance track with the highest popularity may be used as the target performance data.

[0118] S302: Based on the mapping relationship between the preset features and action parameters, and the multiple physical features included in the first feature set and at least one semantic feature other than the target semantic feature in the second feature set, determine the subset of action parameters corresponding to each action in the target performance data.

[0119] Among them, the above-mentioned preset mapping relationship between features and action parameters may specifically include: the mapping relationship between the pitch feature and the distal action of the robotic arm, the mapping relationship between the volume feature and the amplitude of the robotic arm action, the mapping relationship between the timbre feature and the type of robotic arm action, the mapping relationship between the speech rate feature and the speed of the robotic arm action, the mapping relationship between the duration feature and the duration of the robotic arm action, the mapping relationship between the spatial position feature and the direction of the robotic arm action, and the mapping relationship between the semantic feature and the combination of robotic arm actions.

[0120] Exemplarily, the mapping relationship between the above-mentioned pitch feature and the movement speed of the robotic arm can be: the higher the pitch, the higher the position reached by the distal end of the robotic arm; the lower the pitch, the lower the position reached by the distal end of the robotic arm. The mapping relationship between the above-mentioned volume feature and the movement amplitude of the robotic arm can be: the larger the volume, the larger the movement amplitude of the robotic arm; the smaller the volume, the smaller the movement amplitude of the robotic arm. The mapping relationship between the above-mentioned timbre feature and the movement type of the robotic arm can be: according to different timbre features, different robotic arm movement libraries are called to select the corresponding movement type. The mapping relationship between the above-mentioned speech rate (or music rhythm) feature and the distal movement of the robotic arm can be: the faster the speech rate (or music rhythm), the faster the movement speed of the robotic arm; the slower the speech rate (or music rhythm), the slower the movement speed of the robotic arm. The mapping relationship between the above-mentioned duration feature and the movement duration of the robotic arm can be: the longer the duration, the longer the movement duration of the robotic arm; the shorter the duration, the shorter the movement duration of the robotic arm. The mapping relationship between the above-mentioned spatial position feature and the movement direction of the robotic arm can be: according to the spatial position information of the sound source, the movement direction of the robotic arm is determined. The mapping relationship between the above-mentioned semantic feature and the movement combination of the robotic arm can be: according to the speech recognition result and semantic understanding (i.e., multiple semantic features included in the second feature set), the corresponding robotic arm movement combination library is called to generate the action sequence corresponding to the target performance data. For example, if the semantic feature is "forward", the robotic arm executes the forward movement combination.

[0121] Optionally, the mapping relationship between the above-mentioned preset physical feature or semantic feature and the robotic arm movement parameter can be represented by the mapping function R, and the embodiments of the present application do not make specific limitations on this. It should be understood that the specific mapping relationships corresponding to different types of target sound information (such as, speech information or non-speech information) are different. For example, if the target sound information is speech information or non-speech information containing semantics, the above-mentioned mapping relationships may include: the mapping relationships of the pitch feature, volume feature, timbre feature, speech rate feature, duration feature, spatial position feature, and semantic feature respectively.

[0122] For another example, if the target sound information is non-speech information without semantics, the above-mentioned mapping relationships may include: the mapping relationships of the volume feature, timbre feature, speech rate feature, duration feature, spatial position feature, and semantic feature respectively. The embodiments of the present application do not make specific limitations on this.

[0123] S303: Generate an action instruction set based on the obtained multiple subsets of action parameters.

[0124] In an alternative implementation, when performing step S303, after the server obtains the subset of action parameters corresponding to each action in the target performance data (i.e., obtains multiple subsets of action parameters), it can perform the following operations on any one of the aforementioned multiple subsets of action parameters, such as the first subset of action parameters: generate a first action instruction corresponding to the first subset of action parameters based on the preset conversion relationship between the action parameters and the action instructions; and save the first action instruction in the action instruction set. In this way, through the conversion relationship between the action parameters and the action instructions, the action instructions corresponding to each action in the target performance data can be quickly generated.

[0125] Exemplarily, the above-mentioned preset conversion relationship between the action parameters and the action instructions can be represented by a mapping function g, and the embodiments of the present application do not make specific limitations on this.

[0126] It should be understood that each action instruction included in the above action instruction set is generated by a basic action and its corresponding subset of action parameters. Among them, the basic actions may include, but are not limited to: raising the hand, turning around, swinging, etc. Each basic action has corresponding action parameters such as joint movement trajectories, speeds, amplitudes, etc. Optionally, the basic action set can be represented as {A k}, where k represents the serial number of the basic action, and the action parameter vector of each basic action A k can be represented as: P k = [P k1 , P k2 ,..., P kn , where P ki represents the i-th action parameter of the k-th basic action.

[0127] S104: Control the robotic arm to display the target performance data based on the action instruction set.

[0128] In an alternative implementation, after the server determines the target performance data of the robotic arm and the action instruction set associated with the target performance data based on the first feature set and the second feature set, it can control the robotic arm to display the multiple actions included in the target performance data according to the instruction execution order of the multiple action instructions included in the action instruction set.

[0129] To improve the real-time performance and accuracy of robotic arm control, after each action of the robotic arm is demonstrated, the server can also obtain the feedback information of the action demonstrated by the robotic arm at the current moment, so as to modify the mapping relationship between the preset features and action parameters and the conversion relationship between the preset action parameters and action instructions based on the foregoing feedback information, and obtain the modified mapping relationship and the modified conversion relationship; furthermore, adjust the motion instruction at the next moment based on the modified mapping relationship and the modified conversion relationship. Among them, the foregoing feedback information can be used to indicate the completion status of the action, as well as the first feature subset and the second feature subset at the next moment adjacent to the current moment.

[0130] Moreover, the server determines the first modification factor of the mapping relationship and the second modification factor of the conversion relationship based on the feedback information, so as to modify the mapping relationship based on the first modification factor to obtain the above-mentioned modified mapping relationship, and modify the conversion relationship based on the second modification factor to obtain the above-mentioned modified conversion relationship.

[0131] Exemplarily, assume that the above-mentioned feedback information is represented as R b , then the above-mentioned modified mapping relationship can be specifically represented as: R′ = R + ΔR(R b ), where R′ represents the mapping relationship corresponding to the next moment adjacent to the above-mentioned current moment (i.e., the modified mapping relationship), R represents the mapping relationship corresponding to the above-mentioned current moment, ΔR(R b ) represents the above-mentioned first modification factor, and ΔR() represents a function for modifying or adjusting the mapping relationship.

[0132] And, the above-mentioned modified conversion relationship can be specifically represented as: ′ = + Δ(R b ), where ′ represents the conversion relationship corresponding to the next moment adjacent to the above-mentioned current moment (i.e., the modified conversion relationship), g represents the conversion relationship corresponding to the current moment, Δ(R b ) represents the above-mentioned second modification factor, and Δ() represents a function for modifying or adjusting the conversion relationship.

[0133] Optionally, the above-mentioned feedback information R b may include: the action completion status R comp at the current moment, the sound change situation R sound , and other relevant information (such as, the image of the environment where the robotic arm is currently located, etc.). Among them, the sound change situation R sound can determine the first feature subset and the second feature subset at the next moment adjacent to the above-mentioned current moment.

[0134] It can be seen that based on the robotic arm control method based on voice information described in the above steps S101 to S104, by combining voice recognition technology with robotic arm control technology, the robotic arm can generate performance actions in real time according to the target voice information, thereby improving the interactivity and viewing pleasure when the robotic arm displays target performance data. Moreover, after collecting the environmental voice information of the environment where the robotic arm is located, it is possible to extract the features of the target voice information included in the environmental voice information, obtain a first feature set including multiple physical features and a second feature set including multiple semantic features, and thus determine the target performance data of the robotic arm and the action instruction set associated with the target performance data based on the first feature set and the second feature set, and then control the robotic arm to display the target performance data based on the action instruction set. It can be seen that according to the target voice information, the real-time adjustment and dynamic response of the robotic arm actions can be realized, thereby increasing the interactivity and flexibility in the process of robotic arm control.

[0135] Furthermore, based on the same technical concept, an embodiment of the present application provides a robotic arm control device based on voice information, and this robotic arm control device based on voice information is used to implement the above method flow of the embodiment of the present application. Exemplarily, referring to Figure 4 as shown, the robotic arm control device 400 based on voice information may include: an information acquisition module 401, a feature extraction module 402, an instruction determination module 403, and an instruction control module 404, where:

[0136] The information acquisition module 401 is used to obtain target voice information from the environmental voice information of the environment where the robotic arm is located; the target voice information includes motion control keywords for controlling the movement of the robotic arm;

[0137] The feature extraction module 402 is used to extract features from the target voice information to obtain a first feature set and a second feature set; wherein, the first feature set includes multiple physical features of the target voice information, and the second feature set includes multiple semantic features of the target voice information;

[0138] The instruction determination module 403 is used to determine the target performance data of the robotic arm and the action instruction set associated with the target performance data based on the first feature set and the second feature set;

[0139] The instruction control module 404 is used to control the robotic arm to display the target performance data based on the action instruction set.

[0140] In an optional embodiment, when obtaining the target voice information from the environmental voice information, the information acquisition module 401 is specifically used for:

[0141] Performing noise reduction processing on the environmental voice information to filter out the environmental noise included in the environmental voice information and obtain the environmental voice information after noise reduction processing;

[0142] Normalize the ambient sound information after noise reduction processing to obtain target sound information.

[0143] In an alternative embodiment, when extracting features from the target sound information to obtain a first feature set and a second feature set, the feature extraction module 402 is specifically configured to:

[0144] Extract sound features from the target sound information based on the sound feature extraction methods respectively set for multiple physical features to obtain a first feature set; and,

[0145] Perform semantic analysis on the target sound information to obtain a second feature set.

[0146] In an alternative embodiment, when determining the target performance data of the robotic arm and the action instruction set associated with the target performance data based on the first feature set and the second feature set, the instruction determination module 403 is specifically configured to:

[0147] Determine, from a preset candidate performance data set, the target performance data that matches the target semantic features representing motion control keywords in the second feature set;

[0148] Based on the preset mapping relationship between features and action parameters, and the multiple physical features included in the first feature set and at least one semantic feature other than the target semantic feature in the second feature set, determine the subset of action parameters corresponding to each action in the target performance data;

[0149] Generate an action instruction set based on the obtained multiple subsets of action parameters.

[0150] In an alternative embodiment, when generating an action instruction set based on the obtained multiple subsets of action parameters, the instruction determination module 403 is specifically configured to:

[0151] Perform the following operations respectively for multiple subsets of action parameters:

[0152] Generate a first action instruction corresponding to the first subset of action parameters based on the preset conversion relationship between action parameters and action instructions; wherein, the first subset of action parameters is any one of the multiple subsets of action parameters;

[0153] Save the first action instruction to the action instruction set.

[0154] In an alternative embodiment, when controlling the robotic arm to display the target performance data based on the action instruction set, the instruction control module 404 is specifically configured to:

[0155] Control the robotic arm to display the multiple actions included in the target performance data according to the instruction execution order of the multiple action instructions included in the action instruction set;

[0156] Among them, after each action demonstrated by the robotic arm, the following operations are performed:

[0157] Obtain the feedback information of the action demonstrated by the robotic arm at the current moment; the feedback information is used to indicate the completion status of the action, as well as the first feature subset and the second feature subset at the next moment adjacent to the current moment;

[0158] Modify the preset mapping relationship between features and action parameters and the preset conversion relationship between action parameters and action instructions based on the feedback information to obtain a modified mapping relationship and a modified conversion relationship;

[0159] Adjust the motion instruction at the next moment based on the modified mapping relationship and the modified conversion relationship.

[0160] In an alternative embodiment, when modifying the preset mapping relationship between features and action parameters and the preset conversion relationship between action parameters and action instructions based on the feedback information, the instruction control module 404 is specifically configured to:

[0161] Determine a first modification factor for the mapping relationship and a second modification factor for the conversion relationship based on the feedback information;

[0162] Modify the mapping relationship based on the first modification factor and modify the conversion relationship based on the second modification factor.

[0163] Based on the descriptions of the above method embodiments and apparatus embodiments, an exemplary embodiment of the present invention further provides an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor. The memory stores a computer program that can be executed by the at least one processor, and when the computer program is executed by the at least one processor, it is used to cause the electronic device to execute the method according to the embodiment of the present invention.

[0164] An embodiment of the present application further provides a non-transitory computer-readable storage medium storing a computer program, where the computer program, when executed by a processor of a computer, is used to cause the computer to execute the method according to the embodiment of the present application.

[0165] An embodiment of the present application further provides a computer program product, including a computer program, where the computer program, when executed by a processor of a computer, is used to cause the computer to execute the method according to the embodiment of the present application.

[0166] Refer to Figure 5As shown, the structural block diagram of an electronic device 500 that can be a server or a client of the present application will now be described. It is an example of a hardware device that can be applied to various aspects of the present application. The electronic device is intended to represent various forms of digital electronic computer devices, such as, laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, personal digital processors, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present application described and / or claimed herein.

[0167] As Figure 5 shown, the electronic device 500 includes a computing unit 501, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 502 or a computer program loaded from a storage unit 508 into a random access memory (RAM) 503. In the RAM 503, various programs and data required for the operation of the device 500 can also be stored. The computing unit 501, the ROM 502, and the RAM 503 are connected to each other through a bus 504. An input / output (I / O) interface 505 is also connected to the bus 504.

[0168] Multiple components in the electronic device 500 are connected to the I / O interface 505, including: an input unit 506, an output unit 507, a storage unit 508, and a communication unit 509. The input unit 506 can be any type of device that can input information into the electronic device 500. The input unit 506 can receive input digital or character information, and generate key signal inputs related to the user settings and / or function controls of the electronic device. The output unit 507 can be any type of device that can present information, and can include but is not limited to a display, a speaker, a video / audio output terminal, a vibrator, and / or a printer. The storage unit 508 can include but is not limited to magnetic disks, optical disks. The communication unit 509 allows the electronic device 500 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks, and can include but is not limited to a modem, a network card, an infrared communication device, a wireless communication transceiver, and / or a chipset, such as a Bluetooth device, a WiFi device, a worldwide interoperability for microwave access (WiMax) device, a cellular communication device, and / or the like.

[0169] The computing unit 501 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 501 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various AI computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 501 executes the various methods and processes described above.

[0170] For example, in some embodiments, the above-described robotic arm control method based on voice information can be implemented as a computer software program, which is tangibly included in a machine-readable medium, such as the storage unit 508. In some embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device 500 via the ROM 502 and / or the communication unit 509. In some embodiments, the computing unit 501 can be configured to execute the above-described robotic arm control method based on voice information in any other suitable manner (e.g., by means of firmware).

[0171] The program code for implementing the methods of the present application can be written in any combination of one or more programming languages. These program codes can be provided to the processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when the program codes are executed by the processor or controller, the functions / operations specified in the flowcharts and / or block diagrams are implemented. The program code can be executed entirely on the machine, partially on the machine, as an independent software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0172] In the context of this application, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM) or flash memory, an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0173] As used in this application, the terms "machine-readable medium" and "computer-readable medium" refer to any computer program product, apparatus, and / or device (e.g., a disk, optical disk, memory, programmable logic device (PLD)) used to provide machine instructions and / or data to a programmable processor, including a machine-readable medium that receives machine instructions as a machine-readable signal. The term "machine-readable signal" refers to any signal used to provide machine instructions and / or data to a programmable processor.

[0174] To provide for interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a cathode ray tube (CRT) or a liquid crystal display (LCD) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can also be used to provide for interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic, speech, or tactile input).

[0175] The systems and techniques described herein can be implemented in a computing system that includes backend components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes frontend components (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with an implementation of the systems and techniques described herein), or a computing system that includes any combination of such backend, middleware, or frontend components. The components of the system can be interconnected to each other by digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), and the Internet.

[0176] A computer system can include clients and servers. The clients and servers are generally remote from each other and typically interact through a communication network. The client - server relationship is created by computer programs that run on the respective computers and have a client - server relationship with each other.

[0177] Moreover, it should be understood that the above - disclosed is only the preferred embodiment of the present application, and of course, it cannot be used to limit the scope of the rights of the present invention. Therefore, equivalent changes made according to the claims of the present invention still fall within the scope covered by the present application.

Claims

1. A robotic arm control method based on sound information, characterized in that: include: Obtain target sound information from the collected environmental sound information of the environment in which the robot arm is located; The target sound information includes motion control keywords for controlling the motion of the robotic arm; Extracting features of the target sound information to obtain a first feature set and a second feature set; wherein the first feature set includes a plurality of physical features of the target sound information, and the second feature set includes a plurality of semantic features of the target sound information; Determining, from a preset candidate performance data set, target performance data that matches the target semantic feature representing the motion control keyword in the second feature set; Determine an action parameter subset corresponding to each action in the target performance data based on a mapping relationship between preset features and action parameters, the multiple physical features included in the first feature set, and at least one semantic feature in the second feature set except the target semantic feature; generating an action instruction set based on the obtained multiple action parameter subsets; The robotic arm is controlled to display the target performance data based on the action instruction set.

2. The method according to claim 1, characterized in that The step of obtaining target sound information from the collected ambient sound information of the environment in which the robot arm is located includes: Performing noise reduction processing on the ambient sound information, filtering out ambient noise included in the ambient sound information, and obtaining ambient sound information after noise reduction processing; The ambient sound information after the noise reduction process is normalized to obtain the target sound information.

3. The method according to claim 1, characterized in that The extracting features of the target sound information to obtain a first feature set and a second feature set includes: Based on the sound feature extraction methods respectively set for the multiple physical features, extracting the sound features of the target sound information to obtain the first feature set; and Perform semantic analysis on the target sound information to obtain the second feature set.

4. The method according to claim 1, 2 or 3, characterized in that: The first feature set includes any one or a combination of pitch features, volume features, timbre features, speaking speed features, duration features and spatial position features; wherein the duration feature is used to characterize the duration of the target sound information, and the spatial position feature is used to characterize the sound source position coordinates corresponding to the target sound information.

5. The method according to claim 1, 2 or 3, characterized in that: The step of generating an action instruction set based on the obtained plurality of action parameter subsets comprises: For the multiple action parameter subsets, the following operations are performed respectively: Based on the conversion relationship between the preset action parameters and the action instructions, generate a first action instruction corresponding to the first action parameter subset; wherein the first action parameter subset is any one of the multiple action parameter subsets; The first action instruction is saved in the action instruction set.

6. The method according to claim 1, 2 or 3, characterized in that: The controlling the mechanical arm to display the target performance data based on the action instruction set includes: Controlling the robot arm to display multiple actions included in the target performance data according to the instruction execution order of multiple action instructions included in the action instruction set; Wherein, after each action performed by the robotic arm, the following operations are performed: Acquire feedback information of the action performed by the robot arm at the current moment; the feedback information is used to indicate the completion status of the action, as well as the first feature subset and the second feature subset at the next moment adjacent to the current moment; Modifying the mapping relationship between the preset features and the action parameters and the conversion relationship between the preset action parameters and the action instructions based on the feedback information to obtain a modified mapping relationship and a modified conversion relationship; The motion instruction at the next moment is adjusted based on the modified mapping relationship and the modified conversion relationship.

7. The method according to claim 6, characterized in that The modifying, based on the feedback information, the mapping relationship between the preset features and the action parameters and the conversion relationship between the preset action parameters and the action instructions comprises: Determine a first modification factor of the mapping relationship and a second modification factor of the conversion relationship based on the feedback information; The mapping relationship is modified based on the first modification factor, and the conversion relationship is modified based on the second modification factor.

8. A robotic arm control device based on sound information, characterized in that: include: An information acquisition module is used to acquire target sound information from the collected environmental sound information of the environment in which the robot arm is located; The target sound information includes motion control keywords for controlling the motion of the robotic arm; A feature extraction module, configured to extract features from the target sound information to obtain a first feature set and a second feature set; wherein the first feature set includes a plurality of physical features of the target sound information, and the second feature set includes a plurality of semantic features of the target sound information; An instruction determination module is used to determine, from a preset set of candidate performance data, target performance data that matches the target semantic feature characterizing the motion control keyword in the second feature set; based on a mapping relationship between preset features and action parameters, as well as the multiple physical features included in the first feature set and at least one semantic feature in the second feature set other than the target semantic feature, determine an action parameter subset corresponding to each action in the target performance data; and generate an action instruction set based on the obtained multiple action parameter subsets; An instruction control module is used to control the robotic arm to display the target performance data based on the action instruction set.

9. An electronic device, comprising: processor; as well as Memory for storing programs, The program includes instructions, which, when executed by the processor, cause the processor to perform the method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Control system, method and device of mechanical arm equipment and storage medium

    CN113843814A

  • Scheduling voice interaction method and device based on semantic enhanced recognition and robot

    CN118116381A