Voice control device of a collaborative robot and collaborative robot system
Patent Information
- Application Number
- CN202610833774.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-10
- Publication Date
- 2026-08-21
AI Technical Summary
[0004]有鉴于此,本申请提供了一种协作机器人的语音控制装置及协作机器人系统,主要目的在于解决对协作机器人的控制效率较低的技术问题
[0014]根据本发明的第二个方面,提供了一种协作机器人系统,所述协作机器人系统包括协作机器人,以及如上述的协作机器人的语音控制装置;
Smart Images

Figure CN122606669A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of industrial control technology, and in particular to a voice control device and a collaborative robot system for collaborative robots. Background Technology
[0002] With the continuous improvement of automation and intelligence in the tobacco manufacturing industry, collaborative robots have been widely used in production processes such as tobacco processing, re-drying, tobacco pack handling, material transfer and fixed-point sampling. They can effectively improve production continuity, reduce manual labor intensity, and improve operational accuracy and production safety, and have become an important piece of equipment in intelligent tobacco manufacturing.
[0003] Currently, most collaborative robots are controlled via teach-in programming, button triggering, or panel input. During robot operation, operators need to issue motion control commands to the robot through physical interaction devices to drive it to perform the corresponding tasks. This type of control is cumbersome, slow in response, and easily limited by operating space and environmental conditions in industrial settings, resulting in low control efficiency for collaborative robots and failing to meet the fast, convenient, and real-time control requirements of tobacco production lines. Summary of the Invention
[0004] In view of this, this application provides a voice control device and a collaborative robot system, the main purpose of which is to solve the technical problem of low control efficiency of collaborative robots.
[0005] According to a first aspect of the present invention, a voice control device for a collaborative robot is provided for controlling the collaborative robot, the voice control device for the collaborative robot including a voice interaction unit and a central controller. The output of the voice interaction unit is connected to the interaction terminal of the central controller, and the output of the central controller is connected to the control terminal of the collaborative robot. The voice interaction unit is used to receive voice command information, convert the voice command information into a preset operation function message, and send the operation function message to the central controller; The central controller is used to receive the operation function message, obtain the preset action control instruction corresponding to the operation function message, and send the preset action control instruction to the collaborative robot so that the collaborative robot performs the operation action corresponding to the preset action control instruction.
[0006] In an optional embodiment, the voice interaction unit converts the voice command information into a preset job function message, including: the voice interaction unit converts the voice command information into command text and determines the semantic feature vector of the command text; obtains a preset message lookup table, wherein the message lookup table records multiple preset function messages and a corresponding semantic feature vector for each preset function message; calculates the cosine similarity between the semantic feature vector and each corresponding semantic feature vector to obtain the cosine similarity corresponding to each corresponding semantic feature vector; among all the cosine similarities corresponding to the corresponding semantic feature vectors, determines the target similarity with the largest value, and determines the preset function message corresponding to the target similarity as the job function message.
[0007] In an optional embodiment, the voice interaction unit determines the target similarity with the largest value among all the cosine similarities corresponding to the reference semantic feature vectors, and determines the preset function message corresponding to the target similarity as the job function message, including: determining the target similarity with the largest value among all the cosine similarities corresponding to the reference semantic feature vectors, and comparing the target similarity with a preset similarity threshold; when the target similarity is greater than the similarity threshold, determining the preset function message corresponding to the target similarity as the job function message.
[0008] In an optional embodiment, the voice interaction unit is further configured to receive a status control voice command, convert the status control voice command into a preset status function message, and send the status function message to the central controller; the central controller is further configured to receive the status function message and cause the central controller to enter the target operating state specified by the status function message.
[0009] In an optional embodiment, the target operating state includes an automatic control state; when the central controller enters the automatic control state, the central controller only controls the collaborative robot to perform the operation action after receiving the operation function message.
[0010] In an optional embodiment, the task function message corresponds to a specified operating state; the central controller is further configured to receive the task function message, determine the specified operating state corresponding to the task function message, and determine whether the target operating state currently in which the central controller is located is the same as the specified operating state. When the target operating state is the same as the specified operating state, the central controller acquires a preset action control instruction corresponding to the task function message and sends the preset action control instruction to the collaborative robot.
[0011] In an optional embodiment, the voice control device of the collaborative robot further includes a position sensor and an angle sensor. The position sensor is located at the end effector of the collaborative robot, and its output is connected to the receiver of the central controller. The position sensor is used to collect the position information of the end effector in real time and send the position information to the central controller. The central controller is also used to receive the position information of the end effector and determine whether the position information is within a preset safe position range for the end effector. If the position information is not within the safe position range, the central controller controls the voice interaction unit to issue a position exceedance voice alarm. The angle sensor is located at the robot joint of the collaborative robot, and its output is connected to the receiver of the central controller. The angle sensor is used to collect the rotation angle of the robot joint and send the rotation angle to the central controller. The central controller is also used to receive the rotation angle of the robot joint and determine whether the rotation angle is within a preset angle range for the robot joint. If the rotation angle is not within the angle range, the central controller controls the voice interaction unit to issue an angle exceedance voice alarm.
[0012] In an optional embodiment, when the collaborative robot performs a task, it sends the action identification information of the task to the central controller; the central controller is further configured to, upon receiving the action identification information, determine a preset voice prompt instruction corresponding to the action identification information, and send the preset voice prompt instruction to the voice interaction unit; the voice interaction unit is further configured to receive the preset voice prompt instruction and play pre-stored voice prompt information corresponding to the preset voice prompt instruction.
[0013] In an optional embodiment, the voice interaction unit receiving voice command information includes: the voice interaction unit receiving voice audio information and inputting the voice audio information into a pre-trained voice denoising model to obtain a voice mask output by the voice denoising model; fusing the voice mask with the voice audio information to obtain denoised voice audio information, and determining the denoised voice audio information as the voice command information.
[0014] According to a second aspect of the present invention, a collaborative robot system is provided, the collaborative robot system comprising a collaborative robot and a voice control device for the collaborative robot as described above; The central controller in the voice control device of the collaborative robot is connected to the control terminal of the collaborative robot and is used to control the collaborative robot to perform work actions.
[0015] This invention provides a voice control device and system for collaborative robots. Through a voice interaction unit working in conjunction with a central controller, it enables voice command control of the collaborative robot, eliminating the need for physical operations such as teach pendants, buttons, or panels. This significantly simplifies the operation process and effectively improves the control efficiency of the collaborative robot. Voice commands issued by personnel can be directly converted into work function messages and drive the robot to perform corresponding actions, resulting in faster response and more convenient operation. This makes the collaborative robot suitable for the high-paced, high-noise operating environment of tobacco production lines, better meeting the production needs of intelligent manufacturing.
[0016] The above description is only an overview of the technical solution of this application. In order to better understand the technical means of this application and to implement it in accordance with the contents of the specification, and to make the above and other objects, features and advantages of this application more obvious and understandable, the following are specific embodiments of this application. Attached Figure Description
[0017] The accompanying drawings, which are included to provide a further understanding of the invention and form part of this application, illustrate exemplary embodiments of the invention and, together with their description, serve to explain the invention and do not constitute an undue limitation thereof. In the drawings: Figure 1 A schematic diagram of the structure of a voice control device for a collaborative robot provided in an embodiment of the present invention is shown; Figure 2 A schematic diagram of the structure of a voice interaction unit provided in an embodiment of the present invention is shown. Detailed Implementation
[0018] The present invention will be described in detail below with reference to the accompanying drawings and embodiments. It should be noted that, unless otherwise specified, the embodiments and features described in the present application can be combined with each other.
[0019] With the continuous improvement of automation and intelligence in the tobacco manufacturing industry, collaborative robots have been widely used in production processes such as tobacco processing, re-drying, pack handling, material transfer, and fixed-point sampling. They effectively improve production continuity, reduce manual labor intensity, and enhance operational accuracy and production safety, becoming crucial equipment in intelligent tobacco manufacturing. Currently, most collaborative robots are controlled via teach-in programming, button triggering, or panel input. Operators need to issue motion control commands to the robot through physical interaction devices to drive it to perform the corresponding actions. This control method is cumbersome, slow in response, and easily limited by operating space and environmental conditions in industrial settings, resulting in low control efficiency and failing to meet the rapid, convenient, and real-time control requirements of tobacco production lines.
[0020] To address the above problems, in one embodiment, such as Figure 1 As shown, a voice control device for a collaborative robot is provided. Taking the control of a collaborative robot in a tobacco manufacturing scenario as an example, the voice control device for the collaborative robot includes a voice interaction unit 100 and a central controller 200. Here, the central controller 200 can be a programmable logic controller (PLC) that controls the collaborative robot 300.
[0021] Furthermore, the voice interaction unit 100 establishes a communication connection with the central controller 200 via RS485 communication or serial communication. The output terminal of the central controller 200 is connected to the control terminal of the collaborative robot 300, and a communication connection is established between the two based on the Modbus TCP communication protocol or the PROFINET communication protocol.
[0022] Here, as Figure 2 As shown, the voice interaction unit 100 includes a processing unit 110, a speaker 120, and a microphone 130. The processing unit 110 can be a computer device such as a microcontroller or digital signal processor. The speaker 120 and microphone 130 are respectively connected to the processing unit 110. The microphone 130 is used to collect audio information from the scene and send the audio information to the processing unit 110. The processing unit 110 can also control the speaker 120 to emit specific sound information. Furthermore, the interaction terminal of the processing unit 110 is connected to the interaction terminal of the central controller 200 for data interaction. Furthermore, the central controller 200 can also be connected to a serial port touch screen (not shown in the figure) for information display.
[0023] Here, before starting to control the collaborative robot 300, the device can be powered on and initialized. After the device is powered on, the voice interaction unit 100 and the central controller 200 complete the basic startup, execute core actions such as loading the control program, initializing the system registers, resetting the IO signals, and loading the voice dictionary, so that the voice interaction unit 100, the central controller 200 and the collaborative robot 300 are all restored to the initial stable state, and the whole device enters the ready state, preparing for the subsequent communication self-test, instruction reception and action execution process.
[0024] Furthermore, after the device is powered on and initialized, a self-test of the network communication status is performed on the entire device to check whether the internal communication links between the units in the device are normal. When a communication abnormality is detected, the corresponding alarm information is displayed on the serial port touch screen so that technicians can analyze and handle it. Through the above self-test, problems such as communication interruption and signal abnormality can be detected in advance, avoiding subsequent command failures or loss of control of the collaborative robot 300.
[0025] Furthermore, in the actual work process, such as Figure 1 As shown, the voice interaction unit 100 is used to receive voice command information issued by the staff, convert the voice command information into a preset operation function message, and send the operation function message to the central controller 200; wherein, the voice command information is command information in voice form; here, when it is necessary to issue a command, the relevant staff can operate the voice interaction unit 100 to make the microphone work and verbally say the voice command information, including "start moving the tobacco pack", "take a sample of tobacco at a fixed point", "send the sugar and flavoring container to the No. 1 feeder", etc.
[0026] Furthermore, the voice interaction unit 100 can collect voice commands spoken by the operator through its microphone and transmit the collected voice commands to its internal processing unit. The processing unit parses and recognizes the voice commands based on a preset Automatic Speech Recognition (ASR) model or Large Language Model (LLM) model, converting them into text commands and determining the corresponding job function message according to a preset mapping relationship. In this case, the processing unit within the voice interaction unit 100 can be pre-set with a function message lookup table, which records multiple preset text commands and the corresponding preset job message for each preset text command. The processing unit can identify the preset job message corresponding to the preset text command that is identical to the original text command as the job function message.
[0027] Furthermore, the voice interaction unit 100 sends the converted work function messages to the central controller 200 via RS485 communication or serial communication. Here, the work function messages are pre-compiled digital codes that correspond one-to-one with various robot work instructions. Different work function messages correspond to different process actions of the robot, such as material gripping, fixed-point handling, and equipment reset.
[0028] Furthermore, the voice interaction unit 100 can also convert the received voice command information into text-based command text and determine the semantic feature vector of the command text. Here, the automatic speech recognition module pre-set in the voice interaction unit 100 can convert the audio-based voice command information into text-based command text. Further, based on a pre-trained semantic encoding model such as the Word2Vec model, contextual feature extraction and vector mapping processing can be performed on the command text to output a fixed-dimensional semantic feature vector.
[0029] Furthermore, a preset message lookup table is obtained, wherein the message lookup table records multiple preset functional messages and a corresponding semantic feature vector for each preset functional message. Here, the standard instruction text corresponding to each task action that the collaborative robot 300 can execute, and the preset functional message that enables the collaborative robot 300 to execute the task action can be obtained. Further, feature extraction and vector conversion are performed on each standard instruction text using a semantic encoding model to generate a corresponding semantic feature vector for each task action. Then, the corresponding semantic feature vectors for the task actions and the corresponding preset functional messages are associated to finally construct a functional message lookup table. Here, the standard instruction text and the instruction text can be the action name of the corresponding task action.
[0030] Furthermore, the cosine similarity between the semantic feature vector and each of the corresponding semantic feature vectors is calculated to obtain the cosine similarity corresponding to each of the corresponding semantic feature vectors. Specifically, the voice interaction unit 100 can retrieve all the corresponding semantic feature vectors in the generated semantic feature vector and function message comparison table, calculate the cosine similarity between each corresponding semantic feature vector and the semantic feature vector, and associate the corresponding semantic feature vector with its corresponding cosine similarity.
[0031] Furthermore, among all the cosine similarities corresponding to the aforementioned semantic feature vectors, the target similarity with the largest value is determined, and the corresponding semantic feature vector is determined. The preset function message corresponding to the semantic feature vector is then determined as the operation function message.
[0032] Here, the voice interaction unit 100 can determine the target similarity value closest to 1 from the cosine similarities corresponding to all the reference semantic feature vectors, and compare this target similarity value with a preset similarity threshold. The similarity threshold is a critical value used to determine whether the currently recognized voice command effectively matches a preset command. This threshold value is in the range of 0 to 1 and is pre-stored in the processing unit. When the target similarity is greater than the similarity threshold, the match is deemed valid, and the corresponding preset function message is extracted as the work function message. If the target similarity is less than the similarity threshold, the match is deemed invalid, and no function message is sent to the lower-level central controller 200. Further, when the target similarity is greater than the similarity threshold, the reference semantic feature vector corresponding to the target similarity is determined, and the preset function message corresponding to this reference semantic feature vector is determined as the work function message.
[0033] The above method first converts voice commands into command text and generates their semantic feature vectors. Then, based on a pre-stored function message comparison table, cosine similarity comparison is performed to select the optimal matching item. Combined with a similarity threshold verification mechanism, valid commands are selected. This method is applicable to the colloquial expressions of operators. By determining the closest action command through semantic analysis, matching errors caused by environmental interference are reduced. The method accurately completes the automatic conversion of voice commands into operation function messages, effectively avoiding the situation where invalid commands are issued and causing robot malfunctions. This improves the voice recognition matching accuracy and operational safety of the device.
[0034] Furthermore, after receiving the operation function message, the central controller 200 receives the operation function message, obtains the preset action control instruction corresponding to the operation function message, and sends the preset action control instruction to the collaborative robot 300 so that the collaborative robot 300 performs the operation action corresponding to the preset action control instruction.
[0035] Specifically, the central controller 200 receives the task function message transmitted by the voice interaction unit 100, retrieves the pre-stored function message and motion control instruction mapping form in its local memory, searches for and matches the corresponding preset motion control instruction based on the received task function message, and then sends the preset motion control instruction to the collaborative robot 300, causing the collaborative robot 300 to perform the corresponding task action. The motion control instruction mapping form records the preset motion control instructions corresponding to each type of task function message. When a task function message received from the voice interaction unit 100 matches a task function message in the motion control instruction mapping form, the preset motion control instruction corresponding to that task function message is sent to the collaborative robot 300.
[0036] Here, before controlling the collaborative robot 300 to perform specific actions, relevant personnel can set and confirm the robot's operating speed via voice. The operator can issue voice commands such as "Set the robot speed to 30%" or "Confirm speed setting." After recognizing the command, the voice interaction unit 100 extracts the speed value described by the command and transmits it to the central controller 200. The central controller 200 then sends the speed parameter to the controller of the collaborative robot 300. Furthermore, after receiving the speed value, the collaborative robot 300, through the central controller 200, controls the voice interaction unit 100 to announce that the speed has been set. After the operator confirms via voice, the central controller 200 makes the speed setting of the collaborative robot 300 effective, avoiding safety accidents caused by incorrect speed settings.
[0037] The voice control device for the collaborative robot provided in this embodiment can cooperate with the central controller through a voice interaction unit to realize voice command control of the collaborative robot, eliminating the need for physical operations such as teach pendants, buttons, or panels. This significantly simplifies the operation process of the collaborative robot and effectively improves its control efficiency. Specifically, voice commands issued by relevant personnel can be directly converted into work function messages and drive the robot to execute corresponding actions, making the collaborative robot more responsive and easier to operate. This allows it to adapt to the high-paced, high-noise operating environment of tobacco production lines, better meeting the high-efficiency control requirements of intelligent manufacturing.
[0038] In an optional example, the voice interaction unit is further configured to receive state control voice commands, convert the state control voice commands into preset state function messages, and send the state function messages to the central controller; wherein, the operating state of the central controller includes automatic control state and robot speed limit state, etc.
[0039] Here, if the operator wishes to put the central controller into automatic control mode or robot speed limit mode, they can speak a specific status control voice command to the voice interaction unit. The voice interaction unit receives the status control voice command, generates the corresponding status control text, and compares the status control text with a preset status lookup table. The status lookup table includes multiple corresponding status control voice command texts and a preset status function message corresponding to each corresponding status control voice command text. Furthermore, when a status control text matches a corresponding status control voice command text, the status function message corresponding to that corresponding status control voice command text is sent to the central controller.
[0040] Furthermore, the central controller is also used to receive the status function message and cause the central controller to enter the target operating state specified by the status function message. Each status function message corresponds to a target operating state, and different status function messages correspond to different target operating states. Here, the association between status function messages and operating states can be preset. When the central controller receives a status function message, it causes the central controller to enter the operating state associated with that status function message.
[0041] Here, when the central controller enters the automatic control state, the central controller only controls the collaborative robot to perform the operation after receiving the operation function message. At this time, the manual operation of the central controller can be disabled, the central controller blocks the manual operation permission, and the staff can only control the robot by voice. By limiting the control method to distinguish the operating mode, the safety hazards caused by misoperation can be effectively prevented.
[0042] Furthermore, when the central controller enters the robot speed limiting state, the central controller limits the operating speed of the collaborative robot when performing work actions, so that the robot performs work actions at the set operating speed. Specifically, after the central controller switches to the robot speed limiting state, it receives speed parameter data transmitted by the voice interaction unit and completes parameter storage. When issuing action control commands to the collaborative robot, it simultaneously attaches the set speed limit parameters. After receiving the commands and speed parameters, the collaborative robot completes the corresponding work actions according to the limited operating speed.
[0043] The embodiments provided in this application rely on the voice-based status message switching controller to switch working modes. Automatic control and speed limiting modes can be conveniently switched by voice. In automatic mode, manual operation is locked and only voice commands are responded to. In speed limiting mode, the robot's operating speed is uniformly controlled, reducing human error and parameter missetting, and improving the safety and ease of use of the equipment.
[0044] In an optional embodiment, each job function message corresponds to a specified running state; here, job function messages can be pre-associated with specific running states to form a function message-running state mapping table, so that specific job function messages correspond to corresponding specified running states.
[0045] Furthermore, the central controller is also configured to receive the operation function message, determine the specified operating state corresponding to the operation function message, and determine whether the target operating state in which the central controller is located when it receives the operation function message is the same as the specified operating state. When the target operating state is the same as the specified operating state, the central controller acquires the preset action control instruction corresponding to the operation function message and sends the preset action control instruction to the collaborative robot.
[0046] Specifically, after receiving a task function message from the voice interaction unit, the central controller retrieves the locally stored function message-running status mapping table, queries the specified running status bound to the current task function message, and then reads its own real-time target running status for consistency verification. If the verification result shows that the target running status matches the specified running status bound to the task function message, the central controller determines the preset action control command corresponding to the task function message and sends it to the collaborative robot. For example, a task function message corresponding to "start moving cigarette packs" only passes verification when the controller is in automatic control mode. Only after successful verification can the cigarette pack moving action command be retrieved and sent to the collaborative robot. Conversely, if the target running status does not match the specified running status bound to the task function message, it indicates that the collaborative robot does not have the conditions to execute the action. The central controller does not send the preset action control command corresponding to the task function message to the collaborative robot and generates corresponding alarm information.
[0047] The embodiments provided in this application establish a binding relationship between function messages and operating status in advance. After receiving the operation function message, the central controller first performs a status matching verification and only issues an action command when the operating status meets the requirements. This can avoid accidental activation of the equipment under abnormal working conditions and improve the safety of robot operation.
[0048] In an optional embodiment, the voice control device of the collaborative robot further includes a position sensor and an angle sensor; specifically, the position sensor is disposed at the end effector such as the robotic arm of the collaborative robot, and the output end of the position sensor is connected to the receiving end of the central controller. The position sensor is used to collect the position information of the end effector in real time and send the position information to the central controller; here, a three-dimensional coordinate system can be established with the center of the base of the collaborative robot as the origin, the position sensor determines the coordinate point of the end effector in the three-dimensional coordinate system, and the coordinate point is determined as the position information.
[0049] Furthermore, the central controller is also used to receive the position information of the end effector and determine whether the position information is within a preset safe position range for the end effector. When the position information is not within the safe position range, the central controller controls the voice interaction unit to issue a position exceedance voice alarm. A safe position range can be set for the end effector where the position sensor is located, and the boundary of the safe position range can be defined in the aforementioned three-dimensional coordinate system. Specifically, the central controller can receive the spatial coordinate data of the end effector uploaded by the position sensor in real time, retrieve the internally pre-stored safe position boundary parameters of the end effector, and when the measured coordinates exceed the preset safe position range, the central controller generates an alarm trigger signal and sends it to the voice interaction unit, which then plays the corresponding position exceedance voice prompt based on the trigger signal.
[0050] Furthermore, the angle sensor is installed at the robot joint of the collaborative robot to collect the rotation angle of the robot joint and send the rotation angle to the central controller; Furthermore, the central controller is also used to receive the rotation angle of the robot joint and determine whether the rotation angle is within a preset angle range for the robot joint. When the rotation angle is not within the angle range, the controller controls the voice interaction unit to issue an angle over-limit voice alarm message.
[0051] Specifically, an angle range can be preset for the robot joint where the angle sensor is located. If the central controller finds that the angle of the robot joint exceeds the corresponding angle range, the central controller generates an alarm trigger signal and sends it to the voice interaction unit. The voice interaction unit then plays the corresponding angle over-limit voice alarm information according to the trigger signal.
[0052] Furthermore, the central controller can determine abnormal alarms based on feedback signals from the collaborative robot, identifying robot malfunctions such as servo errors and collision triggers. Simultaneously, it can identify peripheral sensor abnormalities such as material shortages, safety door openings, and low air pressure through corresponding supporting sensors, and identify whether there are communication failures such as interruptions in the link between the central controller and the collaborative robot. Furthermore, if no abnormalities are detected, the current work process ends normally. Once an abnormality is detected, it proceeds to the manual confirmation stage. After manual handling is completed, the voice interaction unit broadcasts an abnormality resolution prompt.
[0053] The embodiments provided in this application collect operational data in real time through end-effector position sensors and joint angle sensors, and the central controller compares the data with preset safety thresholds to automatically detect over-limits. When an abnormality is detected, the voice module is activated to sound an alarm, which can promptly detect problems such as the robot arm going out of bounds and the joint angle exceeding the range, and avoid mechanical collision failures in advance, thereby improving the operational safety of the equipment.
[0054] In an optional embodiment, when the collaborative robot performs a task, it can send the action identification information of the task to the central controller. Specifically, when the collaborative robot performs a certain task, it can send the action identification information corresponding to the task back to the central controller. For example, when the collaborative robot performs a point-to-point movement task, it can send the task function message corresponding to the point-to-point movement task back to the central controller as action identification information.
[0055] Furthermore, the central controller is also used to output a preset voice prompt command corresponding to the action identifier information to the voice interaction unit when it receives the action identifier information; here, each action identifier information corresponds to a preset voice prompt command, realizing a one-to-one correspondence between action identifier information and voice prompt command. When the central controller receives the action identifier information, it sends the corresponding voice prompt command to the voice interaction unit.
[0056] Furthermore, the voice interaction unit is also used to receive the preset voice prompt command and play pre-stored voice prompt information corresponding to the preset voice prompt command. Here, each voice prompt command corresponds to a pre-stored audio segment as voice prompt information. When the voice interaction unit receives the preset voice prompt command, it plays the corresponding voice prompt information through the speaker. As an example, when the collaborative robot is performing point-to-point movement actions, the voice interaction unit can play the audio information "Point-to-point movement action in progress".
[0057] Furthermore, upon completion of the task, the collaborative robot can send a completion instruction to the central controller. Upon receiving this instruction, the central controller sends an audio playback command corresponding to the instruction to the voice interaction unit. Upon receiving the audio playback command, the voice interaction unit plays the pre-stored audio information corresponding to that command. For example, if the collaborative robot completes the silk-picking task and sends a completion instruction to the central controller, the central controller outputs a completion playback command to the voice interaction unit. Upon receiving the completion playback command, the voice interaction unit plays the audio message "Silk picking task completed," informing relevant personnel of the robot's operational status.
[0058] Furthermore, when the collaborative robot malfunctions, it can send an error message to the central controller. Upon receiving this message, the central controller sends a corresponding audio playback command to the voice interaction unit. The voice interaction unit then retrieves and plays pre-stored audio. For example, when the collaborative robot fails to grasp materials, it sends an error message to the central controller. Upon receiving this message, the central controller sends an error playback command to the voice interaction unit. The voice interaction unit then plays the audio message "Material grasping error, please check material position," reminding staff to address the problem promptly.
[0059] Furthermore, when the collaborative robot reaches a designated task progress or key milestone, it can send corresponding feedback information to the central controller. Upon receiving the information, the central controller sends an audio playback command to the voice interaction unit, which then retrieves and plays a preset prompt audio message. For example, when the robot has completed half of its task, it uploads a progress signal, causing the central controller to control the voice interaction unit to announce that the task is 50% complete. After the end effector reaches a preset safety point, it sends a signal indicating it has reached the designated safe position, and then plays an audio message saying "Safe position reached," allowing staff to monitor the equipment's operational status in real time.
[0060] Furthermore, after detecting an operational anomaly and broadcasting the abnormal voice message, a voice-based handling process requiring manual confirmation can be initiated. The central controller first drives the voice interaction unit to broadcast the corresponding abnormal content, such as "Robot collision detected, please confirm and handle." The operator then issues a confirmation command via voice, such as "Confirmed, execute reset." The voice interaction unit recognizes the command and converts it into the corresponding control code, which is then transmitted to the central controller. The central controller then drives the robot to perform actions such as fault reset, return to zero, or emergency stop based on the command. After the anomaly is handled, the system returns to standby mode, waiting to receive new voice-based operational commands.
[0061] The embodiments provided in this application can monitor the operation status of collaborative robots in real time. During equipment operation, when the action is completed, or when an abnormality occurs, the corresponding prompt audio is output through the voice interaction unit. This allows operators to intuitively understand the equipment's operating status without having to look at the display screen or robot teach pendant, facilitating blind operation on-site and effectively improving the ease of use and control performance of the device.
[0062] In an optional embodiment, the voice interaction unit receives voice command information in a manner that further includes: First, the voice interaction unit receives voice audio information and inputs it into a pre-trained voice denoising model to obtain a voice mask output by the voice denoising model. The voice mask is a frequency domain weight matrix with values between 0 and 1 for each frequency point. Values close to 1 indicate that the corresponding frequency point is valid human voice and is preserved, while values close to 0 indicate that the corresponding frequency point is environmental noise and is suppressed. By multiplying the voice mask with the original noisy voice spectrum, noise can be filtered out to obtain the denoised clean voice spectrum.
[0063] Here, during the training process of the speech denoising model, standard command speech of on-site operators without environmental noise interference can be acquired, as well as on-site measured noise such as the operation of collaborative motors, pneumatic cylinders, equipment fans, and workshop environmental background noise. The standard command speech and the on-site measured noise are randomly superimposed according to various signal-to-noise ratios (such as 0dB, 5dB, 10dB, 15dB) to generate multiple noisy speech data.
[0064] Furthermore, a speech mask capable of restoring noisy speech data to standard command speech is used as the training label for that noisy speech data, thus obtaining a training label corresponding to each noisy speech data and establishing a mapping relationship between noisy speech data and training labels. Further, based on multiple noisy speech data and their matching training labels, iterative training of the neural network is completed, ultimately resulting in a speech denoising model that can output a corresponding speech mask when inputting noisy speech.
[0065] Furthermore, the speech mask is fused with the speech audio information to obtain denoised speech audio information, which is then identified as the speech command information. Specifically, after obtaining the speech mask output by the speech denoising model, the voice interaction unit performs a frequency-point weighted fusion operation on the speech mask and the acquired original noisy speech audio information. The speech mask suppresses noise frequencies in the original speech spectrum while retaining effective human voice frequencies, accurately filtering out interference noise such as workshop equipment operation noise and environmental background noise. This completes the denoising optimization processing of the original speech audio, resulting in denoised speech audio information with a higher signal-to-noise ratio and clearer human voice characteristics. Finally, the processed, clean, denoised speech audio information is identified as recognizable and parseable valid speech command information to ensure the accuracy of subsequent speech command recognition.
[0066] The embodiments provided in this application generate a speech mask through a pre-trained speech denoising model to denoise the collected original speech audio. This can effectively filter out interference noise such as equipment operation and environmental background noise in industrial workshops, optimize the quality of speech acquisition, improve the accuracy of speech command recognition in complex industrial scenarios, and ensure the stability and reliability of robot voice control command transmission and execution.
[0067] The voice control device for the collaborative robot provided in this application can be customized with exclusive voice control functions for various process stages such as tobacco processing and re-drying, adapting to specific production scenario requirements such as quality inspection sampling and tobacco pack handling. Simultaneously, the device is compatible with various robots with standard communication protocol interfaces, eliminating the need for complex programming and repetitive teaching; tasks can be quickly issued via natural language, effectively improving production flexibility. Compared to traditional teach pendants and programming control methods, this application does not require operators to have professional backgrounds, significantly reducing equipment learning and usage costs. Furthermore, this application constructs a two-way human-machine voice interaction system, capable of receiving human voice commands and broadcasting work progress and equipment malfunction status information via voice, eliminating the need for manual inspection of panel indicator lights and codes, facilitating rapid troubleshooting, and better aligning with human operating habits, significantly improving equipment operation convenience and production efficiency.
[0068] Furthermore, this application also provides a collaborative robot system, which includes a collaborative robot and a voice control device for the collaborative robot as described above; specifically, the central controller in the voice control device of the collaborative robot is connected to the control terminal of the collaborative robot and is used to control the collaborative robot to perform work actions.
[0069] The serial numbers in this application are for descriptive purposes only and do not represent the superiority or inferiority of any particular implementation scenario. The above disclosures are merely a few specific implementation scenarios of this application; however, this application is not limited thereto, and any variations conceived by those skilled in the art should fall within the protection scope of this application.
Claims
1. A voice control device for a collaborative robot, used to control the collaborative robot, characterized in that, The voice control device of the collaborative robot includes a voice interaction unit and a central controller. The output of the voice interaction unit is connected to the interaction terminal of the central controller, and the output of the central controller is connected to the control terminal of the collaborative robot. The voice interaction unit is used to receive voice command information, convert the voice command information into a preset operation function message, and send the operation function message to the central controller; The central controller is used to receive the operation function message, obtain the preset action control instruction corresponding to the operation function message, and send the preset action control instruction to the collaborative robot so that the collaborative robot performs the operation action corresponding to the preset action control instruction.
2. The voice control device for a collaborative robot according to claim 1, characterized in that, The voice interaction unit converts the voice command information into a preset operation function message, including: The voice interaction unit converts the voice command information into command text and determines the semantic feature vector of the command text. Obtain a preset message lookup table, wherein the message lookup table records multiple preset functional messages and a corresponding semantic feature vector for each preset functional message; Calculate the cosine similarity between the semantic feature vector and each of the corresponding semantic feature vectors to obtain the cosine similarity corresponding to each of the corresponding semantic feature vectors; Among all the cosine similarities corresponding to the semantic feature vectors of the comparison, the target similarity with the largest value is determined, and the preset function message corresponding to the target similarity is determined as the operation function message.
3. The voice control device for a collaborative robot according to claim 2, characterized in that, The voice interaction unit determines the target similarity with the largest value among all the cosine similarities corresponding to the compared semantic feature vectors, and determines the preset function message corresponding to the target similarity as the job function message, including: Among all the cosine similarities corresponding to the semantic feature vectors of the comparison, the target similarity with the largest value is determined, and the target similarity is compared with a preset similarity threshold; When the target similarity is greater than the similarity threshold, the preset function message corresponding to the target similarity is determined as the job function message.
4. The voice control device for a collaborative robot according to claim 1, characterized in that, The voice interaction unit is also used to receive status control voice commands, convert the status control voice commands into preset status function messages, and send the status function messages to the central controller. The central controller is also used to receive the status function message and cause the central controller to enter the target operating state specified by the status function message.
5. The voice control device for a collaborative robot according to claim 4, characterized in that, The target operating state includes the automatic control state; When the central controller enters the automatic control state, it only controls the collaborative robot to perform the operation after receiving the operation function message.
6. The voice control device for a collaborative robot according to claim 4, characterized in that, The operation function message corresponds to a specified running status; The central controller is also configured to receive the operation function message, determine the specified operating state corresponding to the operation function message, and determine whether the target operating state currently held by the central controller is the same as the specified operating state. When the target operating state is the same as the specified operating state, the central controller acquires the preset action control instruction corresponding to the operation function message and sends the preset action control instruction to the collaborative robot.
7. The voice control device for a collaborative robot according to any one of claims 1 to 6, characterized in that, The voice control device of the collaborative robot also includes a position sensor and an angle sensor; The position sensor is installed at the end effector of the collaborative robot. The output of the position sensor is connected to the receiving end of the central controller. The position sensor is used to collect the position information of the end effector in real time and send the position information to the central controller. The central controller is also used to receive the position information of the end effector and determine whether the position information is within a preset safe position range for the end effector. When the position information is not within the safe position range, the controller controls the voice interaction unit to issue a position over-limit voice alarm. The angle sensor is installed at the robot joint of the collaborative robot. The output end of the angle sensor is connected to the receiving end of the central controller. The angle sensor is used to collect the rotation angle of the robot joint and send the rotation angle to the central controller. The central controller is also used to receive the rotation angle of the robot joint and determine whether the rotation angle is within a preset angle range for the robot joint. When the rotation angle is not within the angle range, the controller controls the voice interaction unit to issue an angle over-limit voice alarm message.
8. The voice control device for a collaborative robot according to any one of claims 1 to 6, characterized in that, When the collaborative robot performs a task, it sends the action identification information of the task to the central controller. The central controller is also configured to, upon receiving the action identification information, determine a preset voice prompt command corresponding to the action identification information, and send the preset voice prompt command to the voice interaction unit; The voice interaction unit is also used to receive the preset voice prompt command and play the pre-stored voice prompt information corresponding to the preset voice prompt command.
9. The voice control device for a collaborative robot according to any one of claims 1 to 6, characterized in that, The voice interaction unit receives voice command information, including: The voice interaction unit receives voice audio information and inputs the voice audio information into a pre-trained voice denoising model to obtain the voice mask output by the voice denoising model. The voice mask is fused with the voice audio information to obtain the noise-reduced voice audio information, and the noise-reduced voice audio information is determined as the voice command information.
10. A collaborative robot system, characterized in that, The collaborative robot system includes a collaborative robot and a voice control device for the collaborative robot as described in any one of claims 1 to 9; The central controller in the voice control device of the collaborative robot is connected to the control terminal of the collaborative robot and is used to control the collaborative robot to perform work actions.