Object moving method and device, electronic equipment and robot

CN119188622BActive Publication Date: 2026-09-18BEIJING GALBOT AI CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411388180.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-09-30
Publication Date
2026-09-18
Estimated Expiration
2044-09-30

AI Technical Summary

Technical Problem

这种物体搬移方式,耗时耗力,不具备泛化性

Benefits of technology

[0070] In the technical application provided in this application, when performing object transfer, the starting and destination positions of the target object are determined by combining voice data and image data. That is, the starting and destination positions of the current transfer operation are determined, thereby controlling the robot to move the target object from the starting position to the destination position. In this object transfer method, even if the starting and/or destination positions of the object change, the device can actively analyze voice and image data to obtain the starting and destination positions of the current transfer operation. There is no need for manual planning and configuration of the robot's movement trajectory, reducing the time and manpower required for object transfer and improving the generalization of the object transfer method.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119188622B_ABST
    Figure CN119188622B_ABST
Patent Text Reader

Abstract

Embodiments of the present application provide an object moving method and device, electronic equipment and robot, and relate to the technical field of robots. The method comprises: obtaining target voice data of a target person collected by a voice receiving device, and obtaining target image data of the target person collected by an image collecting device; performing content recognition on the target voice data and the target image data to obtain a starting position and a destination position of a target object; and controlling a robot to move the target object at the starting position to the destination position. This scheme can reduce the time and manpower consumed by object moving when the starting position and / or the destination position of the object placement changes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of robotics, and in particular to a method, apparatus, electronic device, and robot for moving objects. Background Technology

[0002] In robot-assisted handling solutions, the starting and destination positions for stacking objects are fixed. The robot moves to the starting position according to a pre-planned trajectory and uses its robotic arm to transfer the object onto a transport vehicle. After the transport vehicle and the robot move to the destination position, the robot uses its robotic arm to unload the object from the transport vehicle and stack it at the destination position.

[0003] When the starting and / or destination positions of stacked objects change, the user needs to plan the robot's movement trajectory according to the new starting and destination positions and configure the trajectory into the robot to move the objects. This method of object moving is time-consuming, labor-intensive, and lacks generalizability. Summary of the Invention

[0004] The purpose of this application is to provide a method, apparatus, electronic device, and robot for moving objects, so as to reduce the time and manpower required for moving objects when the starting position and / or destination position of the object changes. The specific technical solution is as follows:

[0005] In a first aspect, embodiments of this application provide a method for moving an object, the method comprising:

[0006] Acquire target voice data of the target person collected by the voice receiving device, and acquire target image data of the target person collected by the image acquisition device;

[0007] Content recognition is performed on the target speech data and the target image data to obtain the starting position and destination position of the target object.

[0008] The robot is controlled to move the target object from the starting position to the destination position.

[0009] In some embodiments, the step of acquiring the target voice data of the target person collected by the voice receiving device includes:

[0010] Acquire the voice data collected by the voice receiving device;

[0011] Extract speech features from the acquired speech data;

[0012] If the voice feature matches the voice feature that triggers the relocation, then the voice data collected by the voice receiving device will continue to be acquired as the target voice data of the target person.

[0013] In some embodiments, the voice features include at least one of the following: a person's voice biometrics, a voice spectrogram that triggers key information, text data corresponding to the triggering key information, and voice data corresponding to the triggering key information.

[0014] In some embodiments, the step of acquiring the target voice data of the target person collected by the voice receiving device includes:

[0015] Obtain the wake-up command;

[0016] According to the wake-up command, the voice data collected by the voice receiving device is obtained as the target voice data of the target person.

[0017] In some embodiments, the step of obtaining a wake-up command includes:

[0018] If a click is detected on the wake-up button, a wake-up command is generated; or

[0019] Upon receiving a task instruction including key triggering information, a wake-up instruction is obtained based on the task instruction; or

[0020] Upon obtaining the image data of the target person, a wake-up command is generated.

[0021] In some embodiments, the step of acquiring target image data of the target person acquired by the image acquisition device includes:

[0022] Locate the current position of the target person;

[0023] The image acquisition device is instructed to track the target person at the current location and obtain target image data of the target person.

[0024] In some embodiments, the step of locating the current location of the target person includes:

[0025] Acquire the first image data captured by the image acquisition device;

[0026] Extract the biometric features of the personnel included in the first image data;

[0027] If the extracted image biometrics match the image biometrics of the target person, the location of the person included in the first image data is determined as the current location of the target person.

[0028] In some embodiments, the method further includes:

[0029] The image acquisition device is instructed to track the target person for a preset duration, or, upon obtaining the starting position and the destination position, to stop tracking the target person.

[0030] In some embodiments, the step of acquiring target image data of the target person by the image acquisition device is performed under any of the following conditions:

[0031] Received a task instruction that includes key trigger information;

[0032] The object stacking time was detected to be longer than the preset time.

[0033] Detect the presence of people within the designated area;

[0034] The voice data of the target person was obtained.

[0035] In some embodiments, the step of performing content recognition on the target speech data and the target image data to obtain the starting position and destination position of the target object includes:

[0036] Content recognition is performed on the target voice data and the target image data to obtain the intention and actions of the target person.

[0037] If the intention indicates moving an object, then the first starting position and the first destination position of the object indicated by the action are identified, and the starting position and destination position of the target object are obtained.

[0038] In some embodiments, the method further includes:

[0039] If the intention indicates moving an object, and the intention indicates a second starting position and a second destination position for placing the object, then the starting direction and destination direction for placing the object indicated by the action are identified; a first distance between the second starting position and the starting direction is determined, and a second distance between the second destination position and the destination direction is determined; if both the first distance and the second distance are less than a threshold distance, then the second starting position and the second destination position are respectively taken as the starting position and destination position for placing the target object.

[0040] In some embodiments, the step of identifying the first starting position and the first destination position of the object placed by the action includes:

[0041] Identify the starting and destination directions of the object placement indicated by the action;

[0042] Determine the object closest to the starting direction and the destination area closest to the destination direction; or, determine the object closest to the target person in the starting direction and the destination area closest to the target person in the destination direction.

[0043] The location of the determined object is taken as the first starting location, and the location of the determined destination area is taken as the first destination location.

[0044] In some embodiments, before acquiring the target speech data and the target image data, the method further includes:

[0045] Based on the robot's working environment, the types of obstacles are determined;

[0046] If the obstacle is the type of the object to be moved, then the position of the obstacle is located so that, if the obstacle is identified as the target object, the position of the obstacle is used as the starting position of the target object.

[0047] In some embodiments, the step of obtaining the type of obstacle based on the robot's working environment includes:

[0048] Acquire second image data of obstacles in the robot's working environment from the image acquisition device; identify the type of obstacle by analyzing the second image data; or

[0049] Read the tags carried by obstacles in the robot's working environment; determine the type of obstacle based on the tags.

[0050] In some embodiments, the step of obtaining the type of obstacle based on the robot's working environment includes:

[0051] The robot uses a detection device to detect obstacles in its working environment. The detection device includes at least one of the image acquisition device and the millimeter-wave antenna.

[0052] If the obstacle is detected, the type of the obstacle is identified.

[0053] In some embodiments, the method further includes:

[0054] Upon receiving a task instruction that includes triggering key information, the step of obtaining the type of obstacle based on the robot's working scenario is executed.

[0055] Secondly, embodiments of this application provide an object moving device, the device comprising:

[0056] The acquisition module is used to acquire target voice data of the target person collected by the voice receiving device and target image data of the target person collected by the image acquisition device.

[0057] The recognition module is used to perform content recognition on the target voice data and the target image data to obtain the starting position and destination position of the target object.

[0058] The control module is used to control the robot to move the target object from the starting position to the destination position.

[0059] Thirdly, embodiments of this application provide an electronic device, including a processor, a communication interface, a memory, and a communication bus, wherein the processor, the communication interface, and the memory communicate with each other through the communication bus;

[0060] The memory is used to store computer programs;

[0061] When the processor executes the program stored in the memory, it implements any of the object moving methods described above.

[0062] Fourthly, embodiments of this application provide a robot, including a voice receiving device, an image acquisition device, a processor, a communication interface, a memory, and a communication bus, wherein the voice receiving device, the image acquisition device, the processor, the communication interface, and the memory communicate with each other through the communication bus;

[0063] The voice receiving device is used to collect the target voice data of the target person;

[0064] The image acquisition device is used to acquire target image data of the target person;

[0065] The memory is used to store computer programs;

[0066] When the processor executes the program stored in the memory, it implements any of the object moving methods described above.

[0067] Fifthly, embodiments of this application provide a computer-readable storage medium storing a computer program, which, when executed by a processor, implements any of the object moving methods described above.

[0068] Sixthly, embodiments of this application provide a computer program product containing instructions that, when run on a computer, cause the computer to perform any of the object transfer methods described above.

[0069] Beneficial effects of the embodiments in this application:

[0070] In the technical application provided in this application, when performing object transfer, the starting and destination positions of the target object are determined by combining voice data and image data. That is, the starting and destination positions of the current transfer operation are determined, thereby controlling the robot to move the target object from the starting position to the destination position. In this object transfer method, even if the starting and / or destination positions of the object change, the device can actively analyze voice and image data to obtain the starting and destination positions of the current transfer operation. There is no need for manual planning and configuration of the robot's movement trajectory, reducing the time and manpower required for object transfer and improving the generalization of the object transfer method.

[0071] Of course, implementing any product or method of this application does not necessarily require achieving all of the advantages described above at the same time. Attached Figure Description

[0072] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other embodiments can be obtained based on these drawings.

[0073] Figure 1a This is a first schematic diagram of an object moving scenario in related technologies;

[0074] Figure 1b This is a second schematic diagram of an object moving scenario in related technologies;

[0075] Figure 1c This is a third schematic diagram of an object moving scenario in related technologies;

[0076] Figure 2 This is a schematic diagram of a first method for moving objects provided in an embodiment of this application;

[0077] Figure 3a This is a schematic diagram of a first deployment of the image acquisition device and the voice receiving device provided in the embodiments of this application;

[0078] Figure 3b This is a second deployment diagram of the image acquisition device and voice receiving device provided in the embodiments of this application;

[0079] Figure 4a A schematic diagram of a mobile device provided in an embodiment of this application;

[0080] Figure 4b A schematic diagram of a robot provided in an embodiment of this application;

[0081] Figure 5A schematic diagram of instruction input provided in an embodiment of this application;

[0082] Figure 6 A schematic diagram of a prompt information output scenario provided in an embodiment of this application;

[0083] Figure 7 This is a first schematic diagram of a robot working scenario provided in an embodiment of this application;

[0084] Figure 8 This is a schematic diagram of a second process for the object moving method provided in the embodiments of this application;

[0085] Figure 9a This is a second schematic diagram of a robot working scenario provided in an embodiment of this application;

[0086] Figure 9b This is a third schematic diagram of a robot working scenario provided in an embodiment of this application;

[0087] Figure 10 This is a schematic diagram of a third method for moving objects provided in an embodiment of this application;

[0088] Figure 11 This is a fourth schematic diagram of a robot working scenario provided in the embodiments of this application;

[0089] Figure 12 This is a schematic diagram of the fourth process of the object moving method provided in the embodiments of this application;

[0090] Figure 13 A signaling schematic diagram of an object moving system provided in an embodiment of this application;

[0091] Figure 14 A schematic diagram of the structure of the object moving device provided in the embodiments of this application;

[0092] Figure 15 A schematic diagram of the structure of an electronic device provided in an embodiment of this application;

[0093] Figure 16 This is a schematic diagram of the structure of a robot provided in an embodiment of this application. Detailed Implementation

[0094] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art based on this application are within the scope of protection of this application.

[0095] In robot-assisted handling solutions, the starting and destination positions of stacked objects are fixed. For example... Figure 1a and Figure 1b As shown, the robot moves to the starting position A according to a pre-planned trajectory and uses its robotic arm to transfer the object onto a transport vehicle. The transport vehicle can be a small, independent vehicle, such as... Figure 1a As shown; the transport vehicle can also be the robot itself, meaning the robot itself has storage space, such as... Figure 1b As shown. Figure 1b For example, after an object is transferred onto a transport vehicle, the transport vehicle and the robot move to the destination location B according to the planned trajectory. After the transport vehicle and the robot move to the destination location B, the robot uses its robotic arm to unload the object from the transport vehicle and stack the object in the destination location B.

[0096] In an object relocation scheme, if the starting position and / or destination position of the object stacking changes, such as Figure 1c As shown, the destination position changes to position C. The user needs to replan the robot's movement trajectory based on the initial position A and the destination position C of the stacked objects, and configure the movement trajectory into the robot to achieve object transfer. This method of object transfer is time-consuming and labor-intensive, and lacks generalization ability.

[0097] To address the aforementioned problems, embodiments of this application provide a method for moving objects, such as... Figure 2 As shown, the method includes the following steps:

[0098] Step S201: Obtain the target voice data of the target person collected by the voice receiving device, and obtain the target image data of the target person collected by the image acquisition device.

[0099] Step S202: Perform content recognition on the target speech data and target image data to obtain the starting position and destination position of the target object.

[0100] Step S203: Control the robot to move the target object from the starting position to the destination position.

[0101] In the technical application provided in this application, when performing object transfer, the starting and destination positions of the target object are determined by combining voice data and image data. That is, the starting and destination positions of the current transfer operation are determined, thereby controlling the robot to move the target object from the starting position to the destination position. In this object transfer method, even if the starting and / or destination positions of the object change, the device can actively analyze voice and image data to obtain the starting and destination positions of the current transfer operation. There is no need for manual planning and configuration of the robot's movement trajectory, reducing the time and manpower required for object transfer and improving the generalization of the object transfer method.

[0102] The object transfer method provided in this application can be applied to robots or electronic devices such as servers, clusters, and mobile terminals connected to robots. For ease of description, the following description uses electronic devices as the execution subject, but this is not intended to be limiting. The electronic device can be integrated into the robot or can be a standalone physical machine.

[0103] In step S201 above, the voice receiving device is used to collect voice data in the robot's working scene, and the image acquisition device is used to collect image data in the robot's working scene. There can be one or more voice receiving devices and image acquisition devices.

[0104] The voice receiving device and the image acquisition device can be independent physical machines, that is, independently deployed in the robot's working environment, such as... Figure 3a The camera and microphone shown; the voice receiving device and image acquisition device can also be integrated into robots or electronic devices, such as... Figure 3b As shown, the robot is equipped with four cameras (the camera on the back of the robot is not shown in the figure) to collect image data from various directions, and a microphone (i.e., a voice receiving device) is deployed on the top of the robot's head to collect voice data.

[0105] The target personnel are the people who direct the robot to move objects. This can be a specific person or any person in the robot's working environment.

[0106] During object movement, the voice receiving device acquires the voice data of the target personnel, i.e., target voice data; additionally, the image acquisition device acquires the image data of the target personnel, i.e., target image data. The electronic device acquires both the target voice data and the target image data. There can be one or more target image data sets. These target image data sets can come from the same image acquisition device or from different image acquisition devices.

[0107] In this embodiment, the voice receiving device can collect voice data emitted by all personnel and equipment in the robot's working environment. To improve the accuracy of object handling, the electronic device acquires the voice data collected by the voice receiving device, performs noise reduction and filtering on the acquired voice data, and obtains the target voice data of the target personnel.

[0108] Furthermore, the image acquisition device can collect image data from various directions within the robot's working environment and continuously acquires image data. To improve the accuracy of object handling, the electronic device acquires multiple image data sets from the image acquisition device and filters out target image data that includes the target person. For example, the electronic device performs image recognition on the acquired image data to obtain the biometric features of the person in each image data set; it then acquires image data whose biometric features match those of the target person and uses this as the target image data.

[0109] In step S202 above, the target object is the object that needs to be moved this time, the starting position is the current stacked position of the target object, and the destination position is the position where the target object needs to be moved.

[0110] After acquiring the target speech data, the electronic device performs content recognition on the speech data to obtain the speech content recognition results. This includes recognizing key information such as keywords and key phrases included in the target speech data, as well as the user intent indicated by the target speech data, the starting position of the target object stacking, and the destination position of the target object stacking. Key information may include, but is not limited to, preset vocabulary, words representing entities, words representing actions, such as words representing the object to be moved, words representing the robot's number or name, words representing geographical location, and words representing the moving operation. User intent indicates the purpose of the target person, such as the target person's purpose being to move an object, or moving an object from one location (i.e., the starting position) to another location (i.e., the destination position).

[0111] After acquiring the target image data, the electronic device performs content recognition on the target image data to obtain the content recognition results of the image, such as recognizing the action of the target person, as well as the starting direction and starting position of the action indication, the destination direction and destination position of the action indication, etc.

[0112] By combining the content recognition results of the target speech data and the target image data, the electronic device can determine the starting position and the destination position of the target object stacking.

[0113] For example, the target speech data is "Move the cardboard box from location A to location B". The electronic device performs content recognition on the target speech data, obtaining the following result: the starting position for stacking the cardboard boxes is location A, and the destination position for stacking the cardboard boxes is location B. Additionally, the electronic device performs content recognition on the target image data, obtaining the following result: the target person performed the specified object moving action. Combining the content recognition results of the target speech data and the target image data, the electronic device can determine that an object needs to be moved, the target object is a cardboard box, the starting position for placing the target object is location A, and the destination position for stacking the target object is location B.

[0114] In this embodiment of the application, the electronic device may also use other methods to determine the starting position and destination position of the target object, and there is no limitation on this.

[0115] In step S203 above, after determining the starting position and destination position of the target object, the electronic device controls the robot to move to the starting position, and uses the robotic arm to move the target object from the starting position to the destination position. Figure 1a The transport vehicle shown or such Figure 1b The robot's storage space is shown. Then, electronic devices control the robot to move to the target location. If the target object is moved to... Figure 1a On the transport vehicle shown, electronic equipment also controls the vehicle to move to its destination. The electronic equipment then controls a robot to unload the target object at the destination using its robotic arm. This achieves object transfer.

[0116] In this embodiment, when the electronic device is integrated onto the robot, the electronic device can directly control itself to move the target object from the starting position to the destination position. When the electronic device is a standalone physical machine, the electronic device sends control commands to the robot, which include the starting position and the destination position. The robot receives the control commands and, based on the starting position and destination position included in the control commands, controls itself to move the target object from the starting position to the destination position.

[0117] In this embodiment, the electronic device integrates voice data and image data to determine whether to move the object, as well as the starting and destination positions of the move. This effectively avoids accidental triggering and improves the accuracy of object movement.

[0118] In some embodiments, the electronic device may use the following two methods to implement the operation of acquiring target voice data in step S201 above.

[0119] Method a1, active acquisition method. Specifically, the electronic device acquires the voice data collected by the voice receiving device; extracts the voice features of the acquired voice data; if the voice features match the voice features that trigger the relocation, it continues to acquire the voice data collected by the voice receiving device as the target voice data of the target person.

[0120] The voice features can include at least one of the following: a person's voice biometrics, a voice spectrogram of triggering key information, text data corresponding to the triggering key information, and voice data corresponding to the triggering key information. The triggering key information is information that triggers the execution of an object-moving operation, such as words indicating the object to be moved, words indicating the robot's number or name, words indicating the geographical location, words indicating the moving operation, and statements indicating the intention to move. Voice biometrics can include, but are not limited to, timbre, pitch, and intensity. Using voice biometrics, electronic devices can distinguish the person speaking.

[0121] Speech feature matching means that the deviation between two speech features is less than a preset deviation threshold, or that the two speech features are the same.

[0122] In this embodiment, the electronic device pre-records the voice features that trigger the relocation. The voice receiving device collects voice data in real time. The electronic device acquires the voice data collected by the voice receiving device in real time and extracts the voice features from the voice data. The electronic device matches the extracted voice features with the voice features that trigger the relocation; if the two match, the subsequently acquired voice data is used as the target voice data of the target person, and the content recognition processing in step S202 above is performed on the target voice data; if the two do not match, no further processing, such as the content recognition processing in step S202 above, is performed on the voice data.

[0123] Taking voice features, including voice biometrics, as an example, the electronic device pre-records the voice biometrics of the target person. The electronic device extracts the voice biometrics from the voice data. If the extracted voice biometrics are the same as the pre-recorded voice biometrics of the target person, it means that the target person is currently speaking. The voice data subsequently collected by the voice receiving device is the voice data of the target person. The electronic device then continues to acquire the voice data collected by the voice receiving device as the target voice data of the target person.

[0124] Taking the robot's ID as the triggering key information and the voice data including the robot's ID as an example, the electronic device pre-records the voice data of the robot's ID. The electronic device extracts the voice data of the robot's ID. If the voice data of the robot's ID is extracted, it means that the robot may need to perform a moving operation. Then the electronic device continues to acquire the voice data collected by the voice receiving device, and uses the acquired voice data as the target voice data of the target person, so as to obtain complete target voice data and accurately determine the starting position and destination position of the target object.

[0125] For example, the robot is numbered 0001. The electronic device has pre-recorded the voice data of "0001". After acquiring the voice data of "0001" collected by the voice receiving device, the electronic device will use the voice data subsequently acquired by the voice receiving device as the target voice data of the target person.

[0126] In this embodiment, the voice receiving device collects voice data in real time, ensuring the integrity of the target voice data. The electronic device only performs subsequent transfer operations, such as steps S202 and S203, after acquiring voice data whose voice features match the voice features that trigger the transfer. This reduces the resource consumption of the electronic device.

[0127] In addition, in order to reduce the power consumption of electronic devices, the electronic devices are in standby mode; only when voice data matching the voice features that trigger the relocation is obtained will the electronic devices be woken up and then perform subsequent relocation operations.

[0128] Method a2, passive acquisition method. Specifically, the electronic device receives a wake-up command; based on the wake-up command, it acquires the voice data collected by the voice receiving device, which serves as the target voice data for the target person. The wake-up command is used to wake up the electronic device, enabling it to process voice and image data normally.

[0129] In this embodiment of the application, the electronic device may obtain the wake-up command in any of the following ways.

[0130] Method a21: The electronic device is equipped with a wake-up button. The electronic device is in standby mode. The user (e.g., the target person) clicks the wake-up button to wake up the electronic device. The electronic device detects the user clicking the wake-up button, i.e., after detecting the click operation, it generates a wake-up command. According to the wake-up command, the electronic device is awakened and then performs subsequent moving operations, such as the processing operations in steps S202 and S203 above.

[0131] For example, electronic devices are Figure 4a The mobile device shown is a smartphone connected to the robot, and the power button serves as the wake-up button. The user presses the power button. The mobile device is then activated and begins collecting voice data via its microphone (i.e., the voice receiver); the voice data collected by the microphone is then used as the target voice data for the target person.

[0132] For example, electronic devices are Figure 4bThe robot shown has a power button that serves as a wake-up button. When the user presses the power button, the robot is awakened, and its microphone (i.e., voice receiver) begins collecting voice data. The robot then uses the voice data collected by the microphone as the target person's voice data.

[0133] In method a22, the electronic device receives a task instruction and detects the information carried by the task instruction. If it is determined that the task instruction carries trigger key information, a wake-up instruction is obtained based on the task instruction. That is, when a task instruction including trigger key information is received, the electronic device obtains a wake-up instruction based on the task instruction. According to the wake-up instruction, the electronic device is woken up and then performs subsequent moving operations, such as the processing operations in steps S202 and S203 above.

[0134] In this embodiment, the task instruction can be in text or voice format. The source of the task instruction can be an electronic device sent from another device, or it can be directly input by the user into the electronic device. Upon receiving a task instruction that includes triggering key information, the electronic device can use the task instruction directly as a wake-up instruction, or it can use the task instruction to generate a specified wake-up instruction; there is no limitation on this.

[0135] For example, such as Figure 4b The robot shown has a display screen for human-computer interaction. Users can input task commands to perform movements on the screen, such as... Figure 5 Select "Object Move" in the command input box shown, and click the "OK" button. The robot will then receive the task command containing "Object Move". "Object Move" is the key trigger information, which allows the robot to determine that the task command is a wake-up command.

[0136] In method a23, upon acquiring image data of the target person, the electronic device generates a wake-up command. Based on the wake-up command, the electronic device is activated and subsequently performs the relocation operation, as described in steps S202 and S203 above.

[0137] In this embodiment, the image acquisition device acquires image data in real time. The electronic device obtains the image data acquired by the image acquisition device, and if it determines that the image data is of the target person, it generates a wake-up command.

[0138] In this embodiment, the voice receiving device can begin collecting voice data only after the electronic device is woken up, thereby reducing the power consumption of the voice receiving device. Alternatively, the voice receiving device can collect voice data in real time; it will not acquire voice data collected by the voice receiving device when the electronic device is in standby mode, but will acquire the voice data after the electronic device is woken up, thus reducing the delay in voice data acquisition.

[0139] In addition, after the electronic device is woken up, it acquires the voice data collected by the voice receiving device as the target voice data, which reduces the amount of voice data processed by the electronic device and saves the device's computing resources.

[0140] In some embodiments, the electronic device may use the following two methods to implement the operation of acquiring target image data in step S201 above.

[0141] Method b1, active acquisition method. Specifically, the electronic device locates the current position of the target person; it instructs the image acquisition device to track the target person at the current position and obtain the target image data of the target person. Then, subsequent relocation operations are performed, such as the processing operations in steps S202 and S203 above.

[0142] In this embodiment, the user can input the current location of the target person into the electronic device, and the electronic device can then obtain the current location of the target person input by the user. The electronic device can also use image data including the target person acquired by an image acquisition device to locate the current location of the target person.

[0143] In some embodiments, the electronic device can acquire first image data acquired by the image acquisition device; extract the image biometric features of the person included in the first image data; if the extracted image biometric features match the image biometric features of the target person, it indicates that the person included in the first image data is the target person, and the location of the person included in the first image data is determined as the current location of the target person. The image biometric features include facial features and body features, etc.

[0144] In this embodiment, the first image data can be any image data acquired by the image acquisition device. When the electronic device determines that the first image data includes the target person, it can use the calibration relationship between the image coordinate system and the world coordinate system, as well as the pixel coordinates of the target person included in the first image data, to determine the current position of the target person.

[0145] When the electronic device determines that the first image data includes the target person, it can also use the first image data to determine the direction of the target person relative to the electronic device, and use a millimeter-wave antenna to send a first signal to the target person in the direction of the electronic device, and receive a second signal reflected by the target person, the second signal being the echo signal of the first signal; the electronic device uses the first signal and the second signal to determine the current position of the target person.

[0146] In this embodiment of the application, the electronic device may also use other methods to locate the current location of the target person, and there is no limitation on this.

[0147] Once the current location of the target person is determined, the electronic device instructs the image acquisition device to track the target person at the current location; correspondingly, the image acquisition device tracks the target person, acquires target image data, and then the electronic device obtains the target image data.

[0148] In this embodiment of the application, after the electronic device instructs the image acquisition device to track the target person at the current location, the image acquisition device can continuously track the target person and collect target image data until the target person leaves the work scene.

[0149] In some embodiments, to save energy and avoid accidentally triggering object movement operations, the electronic device can instruct the image acquisition device to stop tracking the target person after a preset time has elapsed. This way, the image acquisition device will no longer track the target person but will instead collect image data of the work scene normally, saving energy and avoiding the problem of accidentally triggering object movement operations.

[0150] In some embodiments, in order to save energy, the electronic device may instruct the image acquisition device to stop tracking the target person once the starting position and destination position of the target object are obtained.

[0151] In this embodiment, the electronic device performs content recognition on each acquired target voice data and each target image data to determine the starting and destination positions of the target object. Once the starting and destination positions of the target object are determined, the electronic device instructs the image acquisition device to stop tracking the target personnel. Subsequently, the image acquisition device no longer tracks the target personnel but instead collects image data of the work scene normally, saving energy consumption and avoiding the problem of accidentally triggering object movement operations.

[0152] In this embodiment of the application, if the extracted image biometrics do not match the image biometrics of the target person, the electronic device may not perform further processing on the image data, that is, it will not perform subsequent transfer operations, such as the processing operations in steps S202 and S203 above.

[0153] In some embodiments, the image acquisition device acquires image data in real time. The electronic device acquires the image data acquired by the image acquisition device in real time and extracts the biometric features of the person included in each image data acquired by the image acquisition device. If the extracted biometric features match the biometric features of the target person, then the image data is used as the target image data of the target person. In this way, the electronic device can select one or more target image data of the target person from multiple image data acquired by the image acquisition device, and then perform subsequent transfer operations, such as the processing operations in steps S202 and S203 above.

[0154] Method b2, passive acquisition method. Specifically, the electronic device can use any of the following passive acquisition methods to acquire target image data.

[0155] In method b21, the electronic device receives a task instruction and detects the information carried by the task instruction; if it is determined that the task instruction carries triggering key information, it acquires the target image data of the target person acquired by the image acquisition device. That is, upon receiving a task instruction including triggering key information, the electronic device executes the step of acquiring the target image data of the target person acquired by the image acquisition device.

[0156] In this embodiment of the application, the acquisition of task instructions that trigger key information is included, as described above. Figure 5 A partial example. Upon receiving a task instruction that includes triggering key information, the electronic device can use the method described above, b1, to acquire target image data.

[0157] In some embodiments, the task instruction that triggers key information may also include the current location of the target person; based on the current location of the target person included in the task instruction, the electronic device uses the above-described method b1 to acquire target image data.

[0158] In method b22, the electronic device can detect the stacking time of objects in the work scene. If the stacking time of the objects is detected to be longer than a preset time, the electronic device performs the step of acquiring target image data of the target person collected by the image acquisition device. The electronic device can use method b1 described above to acquire the target image data.

[0159] In this embodiment of the application, the electronic device can pre-obtain the location of objects placed in the work scene. Based on the obtained object location, the electronic device can detect the corresponding object and then detect the stacking time of the object.

[0160] The location of the object can be manually entered by the user on the electronic device, such as the user being able to... Figure 5Enter the location of the object in the instruction input box shown. The location of the object can also be detected by electronic devices using image acquisition devices or millimeter-wave antennas.

[0161] For example, an image acquisition device collects image data in real time. An electronic device identifies the object in the image data collected by the acquisition device, and uses the calibration relationship between the image coordinate system and the world coordinate system, as well as the pixel coordinates of the object in the image data, to determine the object's position.

[0162] For example, the electronic device identifies the image data collected by the image acquisition device, obtains the placed object and the direction of the object relative to the electronic device, and uses a millimeter-wave antenna to send a first signal to the object in the direction of the electronic device, and receives a second signal reflected by the target person. The second signal is the echo signal of the first signal. The electronic device uses the first signal and the second signal to determine the position of the object.

[0163] After determining the location of the object, the electronic device can detect the duration of the object's stacking and decide whether to acquire the target image data.

[0164] In this embodiment of the application, in order to facilitate the acquisition of target image data, when the stacking time of the object is detected to be longer than a preset time, the electronic device can output prompt information indicating the relocation of the object image through a voice playback device (such as a speaker) or a display screen, such as... Figure 6 One possible scenario is shown: the robot's chest display shows the message "The object at position X needs to be moved" so that the target personnel can perform the moving operation in a timely manner, ensuring that the electronic equipment can quickly acquire the target image data and quickly realize the object moving.

[0165] In method b23, a designated area for moving is pre-planned in the work scenario. The electronic device will only perform target image data acquisition when the target personnel are located in the designated area and perform the moving operation. That is, when a person is detected in the designated area, the electronic device will take the person in the designated area as the target personnel and perform the step of acquiring the target image data of the target personnel collected by the image acquisition device.

[0166] In this embodiment, the image acquisition device can acquire image data in real time. The electronic device identifies the image data and obtains the identification result. If the identification result indicates that a person is located in a designated area, the device continuously acquires image data of the designated area, which is the target image data.

[0167] For example, robots are electronic devices, while people are located in... Figure 7 In the work scenario shown, area 1 is the designated area. Figure 7 In the middle, the personnel are located outside of area 1, such as Figure 7As shown in (1), the robot does not acquire target image data; when a person moves into area 1, as... Figure 7 As shown in (2), the robot continuously collects image data of region 1 and performs the above steps S202 and S203 based on the collected image data.

[0168] In method b24, the voice receiving device can collect voice data in real time. The electronic device acquires the voice data collected by the voice receiving device, and if it determines that the voice data belongs to the target person, it executes the step of acquiring the target image data of the target person collected by the image acquisition device. That is, upon acquiring the voice data of the target person, the electronic device executes the step of acquiring the target image data of the target person collected by the image acquisition device.

[0169] In this embodiment, the image acquisition device can begin acquiring image data only after the electronic device has acquired the voice data of the target person, thereby reducing the power consumption of the image acquisition device. Alternatively, the image acquisition device can acquire image data in real time; the electronic device will not acquire the image data acquired by the image acquisition device until after acquiring the voice data of the target person, thus reducing the delay in image data acquisition.

[0170] In some embodiments, such as Figure 8 As shown, a method for moving an object is also provided, which may include the following steps:

[0171] Step S801: Obtain the target voice data of the target person collected by the voice receiving device, and obtain the target image data of the target person collected by the image acquisition device; the same as step S201 above, and will not be repeated here.

[0172] Step S802: Perform content recognition on the target voice data and target image data to obtain the target person's intention and actions;

[0173] Step S803: If the intention indicates to move an object, then identify the first starting position and the first destination position of the object to be placed by the action, and obtain the starting position and destination position of the target object.

[0174] Step S804: Control the robot to move the target object from the starting position to the destination position. This is the same as step S203 above, and will not be described again here.

[0175] In the technical solution provided in this application embodiment, the electronic device combines the intent indicated by the target voice data and the position indicated by the target image data to determine the starting position and the destination position of the target object stacking, thereby reducing the misjudgment rate of determining the starting position and destination position by a single factor and improving the accuracy of object moving.

[0176] In step S802 above, the electronic device performs content recognition on the target voice data to obtain the target person's intention, and performs content recognition on the target image data to obtain the target person's actions.

[0177] After obtaining the target person's intentions and actions, the electronic device can determine whether to perform the relocation, as well as the starting and destination positions of the object relocation.

[0178] If the electronic device recognizes that the intention is to instruct the movement of an object, the electronic device executes step S803 to further identify the target person's actions in order to determine whether the target person's actions are actions to move an object.

[0179] In step S803 above, if the electronic device determines that the target person's intention is to move an object, it identifies the starting position (i.e., the first starting position) and the destination position (i.e., the first destination position) of the object placement indicated by the target person's action.

[0180] If the first starting position and the first destination position are identified, it means that the target personnel have performed the object moving operation, and the first starting position and the first destination position are then used as the starting position and destination position for stacking the target object.

[0181] For example Figure 9a In the work scenario shown, the target person said, "Move the thing from there to there," and at the same time, the target person waved their arm in two directions, that is, from one direction to another, as shown. Figure 9a The robot receives voice data stating "Move the object from there to there" and image data showing the person waving their arms in both directions. Voice data of "Move the object from there to there" and image data of the person waving their arms in both directions indicate the intention to move the object. Image data of the waving arms indicate the action of moving the object, with the starting position at position 1 and the destination at position 2. Combining the intention and action, the robot determines that the object needs to be moved, with the starting position at position 1 and the destination at position 2.

[0182] In some embodiments, the step of identifying, by the electronic device, the first start position and the first destination position where an object is placed as indicated by an action may comprise: identifying a start direction and a destination direction where an object is placed as indicated by the action; determining an object closest to the start direction, and determining a destination area closest to the destination direction; taking the position of the determined object as the first start position, and taking the position of the determined destination area as the first destination position. Herein, the destination area may be identified by a wireframe drawn on the ground, or identified by an identification plate, or identified by an obstacle such as a shelf, which is not limited herein.

[0183] For example Figure 9b In the working scenario shown in , the target person says "move the thing from there to there", and at the same time, the target person swings his arm to point to two directions, as shown in Figure 9b direction a and direction b in . There are 3 positions where objects are placed on direction a, as shown in Figure 9b positions 1 to 3 in ; there are 3 destination areas on direction b, as shown in Figure 9b positions 4 to 6 in . The robot acquires image data of the target person swinging an arm to point to two directions. The robot performs content recognition on the image data of the target person swinging an arm to point to two directions, and can obtain that the action of the target person is an object transferring action, the start direction for placing the object is direction a, and the destination direction for placing the object is direction b; wherein, the distances between positions 1 to 3 and direction a are d1 to d3 respectively, d1<d2<d3, the distances between positions 4 to 6 and direction b are d4 to d6 respectively, d4<d5<d6. Further, the robot determines position 1 as the first start position, that is, the start position where the target object is placed, and determines position 4 as the first destination position, that is, the destination position where the target object is placed.

[0184] In some other embodiments, the step of identifying, by the electronic device, the first start position and the first destination position where an object is placed as indicated by an action may comprise: identifying a start direction and a destination direction where an object is placed as indicated by the action; determining an object closest to the target person in the start direction, and determining a destination area closest to the target person in the destination direction; taking the position of the determined object as the first start position, and taking the position of the determined destination area as the first destination position.

[0185] For example Figure 9b In the working scenario shown in , the target person says "move the thing from there to there", and at the same time, the target person swings his arm to point to two directions, as shown in Figure 9b direction a and direction b in . There are 3 positions where objects are placed on direction a, as shown in Figure 9b positions 1 to 3 in ; there are 3 destination areas on direction b, as shown in Figure 9bpositions 4 to 6. The robot acquires image data of a target person waving an arm to point in two directions. The robot performs content recognition on the image data of the target person waving an arm to point in two directions, and can obtain that the action of the target person is an object moving action, the starting direction for placing the object is direction a, and the destination direction for placing the object is direction b; wherein distances from positions 1 to 3 on direction a to the target person are d11 to d13 respectively, d11<d12<d13, and distances from positions 4 to 6 on direction b to the target person are d14 to d16 respectively, d14<d15<d16. Further, the robot determines position 1 as the first starting position, that is, the starting position for placing the target object, and determines position 4 as the first destination position, that is, the destination position for placing the target object.

[0186] In the embodiments of the present application, the electronic device can uniquely identify a starting position and a destination position by using the intention and action of the target person, without manually configuring the starting position and the destination position, which reduces the labor consumption of object moving.

[0187] In some embodiments, as Figure 10 shown, an object moving method is also provided, which may include the following steps:

[0188] step S1001: acquiring target voice data of a target person collected by a voice receiving device, and acquiring target image data of the target person collected by an image collecting device; which is the same as the foregoing step S201, and will not be repeated herein.

[0189] step S1002: performing content recognition on the target voice data and the target image data to obtain the intention of the target person and the action of the target person; which is the same as the foregoing step S802, and will not be repeated herein.

[0190] step S1003: if the intention indicates moving an object, recognizing a first starting position and a first destination position for placing the object indicated by the action to obtain the starting position and the destination position for placing the target object; which is the same as the foregoing step S803, and will not be repeated herein.

[0191] step S1004: if the intention indicates moving an object and the intention indicates a second starting position and a second destination position for placing the object, recognizing a starting direction and a destination direction for placing the object indicated by the action;

[0192] step S1005: determining a first distance between the second starting position and the starting direction, and determining a second distance between the second destination position and the destination direction;

[0193] step S1006: if both the first distance and the second distance are smaller than a threshold distance, taking the second starting position and the second destination position as the starting position and the destination position for placing the target object respectively.

[0194] Step S1007: Control the robot to move the target object from the starting position to the destination position. This is the same as step S203 above, and will not be described again here.

[0195] In the technical solution provided in this application embodiment, the electronic device combines the intention and location indicated by the target voice data and the direction indicated by the target image data to determine the starting position and destination position of the target object, thereby reducing the misjudgment rate of determining the starting position and destination position by a single factor and improving the accuracy of object movement.

[0196] In step S1002 above, the electronic device performs content recognition on the target voice data to obtain the target person's intention, and performs content recognition on the target image data to obtain the target person's actions.

[0197] After obtaining the target person's intentions and actions, the electronic device can determine whether to perform the relocation, as well as the starting and destination positions of the object relocation.

[0198] If the electronic device recognizes that the intention is to instruct the movement of an object, and the intention indicates the starting position (i.e., the second starting position) and the destination position (i.e., the second destination position) of the object, the electronic device executes step S1004 to further identify the action of the target person in order to determine whether the action of the target person is the action of moving an object.

[0199] In step S1004 above, when the electronic device determines that the target person's intention is to move an object and that the intention is to indicate the second starting position and the second destination position of the object placement, it identifies the starting direction of the object placement and the destination direction of the object placement indicated by the target person's actions.

[0200] If the starting direction and the destination direction are identified, it means that the target person has performed the object moving action. The electronic device calculates the distance between the second starting position and the starting direction (i.e., the first distance) and the distance between the second destination position and the destination direction (i.e., the second distance).

[0201] If the starting and destination directions are not identified, it means that the target personnel have not performed the object moving operation, and the electronic device can end the object moving process.

[0202] In the foregoing step S1005, the threshold distance can be set according to actual requirements. The electronic device compares the first distance and the second distance with the threshold distance respectively. If the first distance is greater than or equal to the threshold distance, or the second distance is greater than or equal to the threshold distance, or both the first distance and the second distance are greater than or equal to the threshold distance, it indicates that a recognition error occurs, and the electronic device can end the current object moving process. If both the first distance and the second distance are smaller than the threshold distance, it indicates that the recognition is correct, and then the electronic device uses the second start position and the second destination position as the start position and the destination position where the target object is placed respectively.

[0203] For example Figure 11 in the working scenario shown, a target person says "move the thing from position 1 to position 2", and at the same time, the target person waves an arm to point to two directions, such as Figure 11 direction a and direction b therein. The robot obtains voice data of "move the thing from position 1 to position 2", and obtains image data of the target person waving an arm to point to two directions. The robot performs content recognition on the voice data of "move the thing from position 1 to position 2", and can obtain that the intention of the target person is to move an object, the second start position for placing the object is position 1, and the second destination position for placing the object is position 2. The robot performs content recognition on the image data of the target person waving an arm to point to two directions, and can obtain that the action of the target person is an object moving action, the start direction for placing the object is direction a, and the destination direction for placing the object is direction b; wherein, the distance between position 1 and direction a is d21, the distance between position 2 and direction b is d22, and the threshold distance is d0. If d21<d0 and d22<d0, the robot determines that position 1 is the second start position, that is, the start position where the target object is placed, and determines that position 2 is the second destination position, that is, the destination position where the target object is placed.

[0204] In the embodiments of the present application, the electronic device can uniquely recognize one start position and one destination position by using the intention and action of the target person, without manually configuring the start position and the destination position, thereby reducing the manpower consumption for object moving.

[0205] In some embodiments, such as Figure 12 shown, an object moving method is further provided, which may include the following steps:

[0206] Step S1201, acquiring types of obstacles based on a working scenario of a robot;

[0207] When moving objects, the electronic equipment pre-identifies the types of obstacles in the work environment to determine which obstacles (such as the target object mentioned above) need to be moved. The obstacle types can be categorized as stacked objects, single objects, or object names. The obstacle type can be determined based on the robot's work environment. For example, in a work environment requiring the movement of stacked objects, the obstacle types can be categorized as stacked placement types and non-stacked placement types; in a work environment requiring the movement of a specific object, the obstacle type can be categorized as the specific item name.

[0208] In the implementation of this application, the electronic device may acquire the type of obstacle in any of the following ways.

[0209] In method c1, the user inputs a task command into the electronic device, which includes the types of obstacles in the work environment. The electronic device then uses the task command to determine the types of obstacles in the work environment.

[0210] In method c2, the image acquisition device acquires image data of obstacles in the robot's working scene, i.e., the second image data; the electronic device acquires the second image data acquired by the image acquisition device; and the second image data is identified to obtain the type of obstacle.

[0211] In this embodiment, the user can input a task command into the electronic device, which includes the location of obstacles in the work scene. The electronic device instructs an image acquisition device to acquire image data at that location, and then the image acquisition device acquires image data of the obstacles at that location to obtain second image data; the electronic device identifies the type of obstacle from the second image data.

[0212] In this embodiment, the electronic device can also utilize a detection device to detect obstacles in the robot's working environment. The detection device may include at least one of an image acquisition device and a millimeter-wave antenna. Upon detecting an obstacle, the type of obstacle is identified. For example, as described above, the image acquisition device acquires second image data at the obstacle's location, and the second image data is then identified to determine the type of obstacle. The method of detecting obstacles is the same as the method of detecting target objects described above; details can be found in the relevant descriptions above and will not be repeated here.

[0213] Method c3 involves obstacles carrying tags, such as radio frequency identification (RFID) tags. Electronic devices read the tags carried by obstacles in the robot's working environment and determine the type of obstacle based on the read tags.

[0214] In this embodiment, the user can input a task command into the electronic device, which includes the location of obstacles in the work scene. The electronic device reads the RFID tag carried by the obstacle at that location to determine the type of obstacle.

[0215] In this embodiment, the electronic device can also utilize image acquisition devices and / or millimeter-wave antennas to detect obstacles in the robot's working environment. Upon detecting an obstacle, the type of obstacle is identified, such as by reading the RFID tag at the obstacle's location, and then determining the type of obstacle based on the read tag. The method of detecting obstacles is the same as the method of detecting target objects, as detailed in the above description, and will not be repeated here.

[0216] In this embodiment of the application, the electronic device may also use other methods to detect the position of the obstacle, such as using sensors installed on the robot chassis to detect the position of the obstacle, and there is no limitation on this.

[0217] In this embodiment, the electronic device can acquire the location or image data of obstacles in advance, but does not identify the type of obstacle. Upon receiving a task instruction including triggering key information, the type of obstacle is then identified; that is, the step of acquiring the obstacle type is only executed upon receiving a task instruction including triggering key information. If no task instruction including triggering key information is received, the electronic device may not identify the type of obstacle to reduce unnecessary energy consumption.

[0218] For example, an electronic device might be a robot, with an image acquisition device mounted on it. The robot moves continuously within a work environment; the image acquisition device moves with the robot, collecting image data of the work environment. The robot identifies the image data to obtain image data of obstacles. The robot can record the obstacle image data, obtaining an image dataset. When the robot receives a task instruction including triggering key information, it identifies the type of obstacle in each image in the image dataset.

[0219] Step S1202: If the type of obstacle is the type of object to be moved, then locate the position of the obstacle so that, when the obstacle is identified as the target object, the position of the obstacle is used as the starting position of the target object.

[0220] In this embodiment, the type of object to be moved can be set according to actual needs. For example, if the type of object to be moved is stacked, and multiple objects are stacked together, then the type of these multiple objects is stacked. Another example is that the type of object to be moved is a specified item type, such as cardboard boxes, food, or daily necessities.

[0221] Electronic devices can pre-record the type of object to be moved. After identifying the type of obstacle, the electronic device can determine whether the obstacle's type matches the type of the object to be moved. If so, it indicates that the obstacle is the object to be moved, such as the target object mentioned above, and then locates the obstacle's position. Subsequently, based on the obstacle's position, the electronic device can determine whether the obstacle is the target object, and if it is, uses the obstacle's position as the starting position of the target object. For example, the electronic device can locate the positions of one or more obstacles, and use these obstacle positions to... Figure 8 and Figure 10 The process shown determines the position of an obstacle as either the first starting position or the second starting position.

[0222] The location of the obstacle can be input by the user into the electronic device, or it can be obtained by the electronic device using an image acquisition device and / or a millimeter-wave antenna or other detection device. There is no limitation on this.

[0223] Step S1203: Obtain the target voice data of the target person collected by the voice receiving device, and obtain the target image data of the target person collected by the image acquisition device; the same as step S201 above, and will not be repeated here.

[0224] Step S1204 involves performing content recognition on the target speech data and target image data to obtain the starting position and destination position of the target object; this is the same as step S202 above and will not be repeated here.

[0225] Step S1205: Control the robot to move the target object from the starting position to the destination position. This is the same as step S203 above, and will not be described again here.

[0226] In this embodiment of the application, to facilitate the location of the starting and destination positions of the target object, the electronic device can be pre-configured with an electronic map. The electronic device marks the position of the object to be moved (such as the position of the obstacle mentioned above) on the electronic map, and determines the starting and destination positions of this moving operation by combining the electronic map.

[0227] In the technical solution provided in this application embodiment, the electronic device pre-locates the position of obstacles in the work scene, which facilitates the subsequent determination of the starting position and destination position of the target object and the accurate movement of the object.

[0228] The following is combined Figure 13 The system signaling diagram shown illustrates in detail the object handling method provided in this application embodiment. The electronic equipment, image acquisition device, and voice receiving device can be configured on the robot or are independent physical machines. Figure 13This example illustrates the use of electronic devices, image acquisition devices, and voice receiving devices mounted on a robot. Here, the electronic devices can be understood as the robot's processor.

[0229] In step S1301, the image acquisition device acquires image data of the working scene in real time and sends the acquired image data to the electronic device.

[0230] In this embodiment, the robot moves within the work environment. As the robot moves, the image acquisition device can collect image data of the entire work environment, forming an image dataset.

[0231] In step S1302, the electronic device identifies the image data acquired by the image acquisition device to obtain the image data of the obstacle.

[0232] In this embodiment of the application, the electronic device filters out image data including obstacles from the image dataset collected by the image acquisition device to form an obstacle image dataset.

[0233] In step S1303, the electronic device further identifies the image data of the obstacle, obtains image data of the obstacle as the object to be moved, and locates the position of the object to be moved.

[0234] In step S1304, the voice receiving device collects voice data in real time and sends the collected voice data to the electronic device.

[0235] In step S1305, the electronic device recognizes the voice data collected by the voice receiving device; if the voice data includes trigger key information, then for the subsequent image data, step S1306 is executed, and for the subsequent image data and voice data, step S1307 is executed; if the voice data does not include trigger key information, then the current moving operation is terminated, that is, subsequent steps S1306 and S1307 are not executed.

[0236] In step S1306, the electronic device identifies the image data to obtain the image data of the target person.

[0237] Once the voice data is determined to include key trigger information, the electronic device identifies the target person's image data from subsequently acquired image data.

[0238] In step S1307, the electronic device performs content recognition on the voice and image data of the target personnel to obtain the starting and destination positions of this relocation operation. The voice data of the target personnel includes voice data that triggers key information.

[0239] In step S1308, the electronic device controls the robot to move the target object from the starting position to the destination position.

[0240] In this embodiment, after obtaining the starting and destination positions of the current transfer operation, the electronic device can output response information using the robot's display screen or voice playback device, such as the display screen showing "Transfer instruction received, object transfer completed," etc. After completing the transfer operation, the electronic device can also output response information using the robot's display screen or voice playback device, such as the display screen showing "Object transfer completed," etc.

[0241] By outputting response information, users can promptly determine the status of various operations and adjust the robot's status accordingly. For example, after the display shows "Object moved complete," the user can promptly assign the next moving task.

[0242] Corresponding to the above-described object moving method, this application also provides an object moving device, such as... Figure 14 As shown, the device includes:

[0243] The acquisition module 1401 is used to acquire the target voice data of the target person collected by the voice receiving device and the target image data of the target person collected by the image acquisition device.

[0244] The recognition module 1402 is used to perform content recognition on the target voice data and target image data to obtain the starting position and destination position of the target object.

[0245] The control module 1403 is used to control the robot to move the target object from the starting position to the destination position.

[0246] In some embodiments, the acquisition module 1401 may be specifically used for:

[0247] Acquire voice data collected by the voice receiving device;

[0248] Extract speech features from the acquired speech data;

[0249] If the voice features match the voice features that triggered the relocation, then the voice data collected by the voice receiving device will continue to be acquired as the target voice data of the target person.

[0250] In some embodiments, the voice features include at least one of the following: a person's voice biometrics, a voice spectrogram that triggers key information, text data corresponding to the triggering key information, and voice data corresponding to the triggering key information.

[0251] In some embodiments, the acquisition module 1401 may be specifically used for:

[0252] Obtain a wake-up command; based on the wake-up command, obtain the voice data collected by the voice receiving device as the target voice data for the target person.

[0253] In some embodiments, the acquisition module 1401 may be specifically used for:

[0254] If a click is detected on the wake-up button, a wake-up command is generated; or

[0255] Upon receiving a task instruction that includes key trigger information, a wake-up instruction is obtained based on the task instruction; or

[0256] Once image data of the target person is obtained, a wake-up command is generated.

[0257] In some embodiments, the acquisition module 1401 may be specifically used for:

[0258] Locate the current location of the target personnel;

[0259] The image acquisition device is instructed to track the target person at the current location and obtain the target image data of the target person.

[0260] In some embodiments, the acquisition module 1401 may be specifically used for:

[0261] Acquire the first image data captured by the image acquisition device;

[0262] Extract the biometric features of the personnel included in the first image data;

[0263] If the extracted image biometrics match the image biometrics of the target person, the location of the person included in the first image data is determined as the current location of the target person.

[0264] In some embodiments, the acquisition module 1401 may be specifically used for:

[0265] When the image acquisition device is instructed to track the target person for a preset time, or when the starting and destination positions are obtained, the image acquisition device is instructed to stop tracking the target person.

[0266] In some embodiments, the acquisition module 1401 can also be used for:

[0267] The step of acquiring target image data of the target person captured by the image acquisition device shall be performed if any of the following conditions are met:

[0268] Received a task instruction that includes key trigger information;

[0269] The object stacking time was detected to be longer than the preset time.

[0270] Detect the presence of people within the designated area;

[0271] The voice data of the target personnel was obtained.

[0272] In some embodiments, the identification module 1402 may be specifically used for:

[0273] Content recognition is performed on target speech data and target image data to obtain the target person's intent and actions;

[0274] If the intention is to move an object, the first starting position and the first destination position of the object to be placed are identified, and the starting position and destination position of the target object are obtained.

[0275] In some embodiments, the identification module 1402 may be specifically used for:

[0276] If the intention is to move an object, and the intention is to place the object at a second starting position and a second destination position, then the starting direction and destination direction of the object placement indicated by the action are identified.

[0277] Determine the first distance between the second starting position and the starting direction, and determine the second distance between the second target position and the target direction;

[0278] If both the first distance and the second distance are less than the threshold distance, then the second starting position and the second target position will be used as the starting position and target position for placing the target object, respectively.

[0279] In some embodiments, the identification module 1402 may be specifically used for:

[0280] Identify the starting and destination directions of the object placed according to the action instruction;

[0281] Identify the object closest to the starting direction and the destination area closest to the destination direction; or, identify the object closest to the target person in the starting direction and the destination area closest to the target person in the destination direction.

[0282] The location of the determined object is taken as the first starting location, and the location of the determined destination area is taken as the first destination location.

[0283] In some embodiments, the acquisition module 1401 can also be used for:

[0284] Before acquiring the target voice data and target image data, the type of obstacle is determined based on the robot's working scenario;

[0285] If the obstacle is the same type as the object to be moved, then the location of the obstacle is determined so that, if the obstacle is identified as the target object, the location of the obstacle is used as the starting position of the target object.

[0286] In some embodiments, the acquisition module 1401 may be specifically used for:

[0287] Acquire second image data of obstacles in the robot's working environment from the image acquisition device; identify the types of obstacles by analyzing the second image data; or

[0288] Read the tags carried by obstacles in the robot's working environment; determine the type of obstacle based on the read tags.

[0289] In some embodiments, the acquisition module 1401 may be specifically used for:

[0290] Using detection equipment, obstacles in the robot's working environment are detected. The detection equipment includes at least one of an image acquisition device and a millimeter-wave antenna.

[0291] When an obstacle is detected, identify the type of obstacle.

[0292] In some embodiments, the acquisition module 1401 can also be used for:

[0293] Upon receiving a task instruction that includes key trigger information, the robot executes steps based on the robot's work scenario to determine the type of obstacle.

[0294] In the technical application provided in this application, when performing object transfer, the starting and destination positions of the target object are determined by combining voice data and image data. That is, the starting and destination positions of the current transfer operation are determined, thereby controlling the robot to move the target object from the starting position to the destination position. In this object transfer method, even if the starting and / or destination positions of the object change, the device can actively analyze voice and image data to obtain the starting and destination positions of the current transfer operation. There is no need for manual planning and configuration of the robot's movement trajectory, reducing the time and manpower required for object transfer and improving the generalization of the object transfer method.

[0295] Corresponding to the above-described object moving method, this application also provides an electronic device, such as... Figure 15 As shown, it includes a processor 1501, a communication interface 1502, a memory 1503, and a communication bus 1504, wherein the processor 1501, the communication interface 1502, and the memory 1503 communicate with each other through the communication bus 1504.

[0296] Memory 1503 is used to store computer programs;

[0297] The processor 1501 is used to execute the program stored in the memory 1503 to implement any of the above-mentioned object moving methods.

[0298] Corresponding to the above-described object handling method, this application also provides a robot, such as... Figure 16 As shown, it includes a voice receiving device 1605, an image acquisition device 1606, a processor 1601, a communication interface 1602, a memory 1603, and a communication bus 1604, wherein the processor 1601, the communication interface 1602, and the memory 1603 communicate with each other through the communication bus 1604.

[0299] The voice receiving device 1605 is used to collect the target voice data of the target personnel.

[0300] Image acquisition device 1606 is used to acquire target image data of the target person;

[0301] Memory 1603 is used to store computer programs;

[0302] The processor 1601 is used to execute the program stored in the memory 1603 to implement any of the above-mentioned object moving methods.

[0303] The aforementioned communication bus can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. This communication bus can be divided into address bus, data bus, control bus, etc. For ease of illustration, only one thick line is used to represent it in the diagram, but this does not indicate that there is only one bus or one type of bus.

[0304] The communication interface is used for communication between the aforementioned electronic devices or robots and other devices.

[0305] The memory may include random access memory (RAM) or non-volatile memory (NVM), such as at least one disk storage device. Optionally, the memory may also be at least one storage device located remotely from the aforementioned processor.

[0306] The processors mentioned above can be general-purpose processors, including central processing units (CPUs), network processors (NPs), etc.; they can also be digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components.

[0307] In another embodiment provided in this application, a computer-readable storage medium is also provided, which stores a computer program that, when executed by a processor, implements any of the above-described object moving methods.

[0308] In another embodiment provided in this application, a computer program product containing instructions is also provided, which, when run on a computer, causes the computer to perform any of the above-described object moving methods.

[0309] In the above embodiments, implementation can be achieved entirely or partially through software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented entirely or partially as a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid state disk (SSD)).

[0310] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0311] The various embodiments in this specification are described in a related manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the embodiments of devices, electronic devices, robots, storage media, and program products are basically similar to the method embodiments, so the descriptions are relatively simple; relevant parts can be referred to the descriptions of the method embodiments.

[0312] The above description is merely a preferred embodiment of this application and is not intended to limit the scope of protection of this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application are included within the scope of protection of this application.

Claims

1. A method for moving an object, characterized in that, The method includes: Acquire target voice data of the target person collected by the voice receiving device, and acquire target image data of the target person collected by the image acquisition device; Content recognition is performed on the target voice data and the target image data to obtain the intention and actions of the target person. In the case where the intention is to move the object, Identify the starting and destination directions of the object placement indicated by the action; determine the object closest to the starting direction and the destination area closest to the destination direction; or, determine the object closest to the target person in the starting direction and the destination area closest to the target person in the destination direction; use the determined object's position as the first starting position and the determined destination area's position as the first destination position to obtain the starting and destination positions of the target object placement; or, If the intention indicates a second starting position and a second destination position for object placement, then the starting direction and destination direction of object placement indicated by the action are identified; a first distance between the second starting position and the starting direction is determined, and a second distance between the second destination position and the destination direction is determined; if both the first distance and the second distance are less than a threshold distance, then the second starting position and the second destination position are respectively taken as the starting position and destination position for the target object placement; if both the first distance and the second distance are less than the threshold distance, it indicates that the identification is correct; if neither the first distance nor the second distance is less than the threshold distance, it indicates that the identification is incorrect. The robot is controlled to move the target object from the starting position to the destination position.

2. The method according to claim 1, characterized in that, The step of acquiring the target voice data of the target person collected by the voice receiving device includes: Acquire the voice data collected by the voice receiving device; Extract speech features from the acquired speech data; If the voice feature matches the voice feature that triggers the relocation, then the voice data collected by the voice receiving device will continue to be acquired as the target voice data of the target person.

3. The method according to claim 2, characterized in that, The voice features include at least one of the following: the person's voice biometrics, the voice spectrogram that triggers key information, the text data corresponding to the triggering key information, and the voice data corresponding to the triggering key information.

4. The method according to claim 2, characterized in that, The step of acquiring the target image data of the target person by the image acquisition device shall be performed if any of the following conditions are met: Received a task instruction that includes key trigger information; The object stacking time was detected to be longer than the preset time. Detect the presence of people within the designated area; The voice data of the target person was obtained.

5. The method according to claim 1, characterized in that, The step of acquiring the target voice data of the target person collected by the voice receiving device includes: acquiring a wake-up command; and acquiring the voice data collected by the voice receiving device according to the wake-up command, as the target voice data of the target person.

6. The method according to claim 5, characterized in that, The step of obtaining the wake-up command includes: generating a wake-up command when a click operation is detected on the wake-up button; or, obtaining a wake-up command based on the task command when a task command including trigger key information is received; or, generating a wake-up command when image data of the target person is obtained.

7. The method according to any one of claims 1 to 6, characterized in that, The step of acquiring target image data of the target person acquired by the image acquisition device includes: locating the current position of the target person; instructing the image acquisition device to perform target tracking on the target person at the current position, thereby obtaining the target image data of the target person.

8. The method according to claim 7, characterized in that, The step of locating the current location of the target person includes: acquiring first image data acquired by an image acquisition device; extracting the image biometric features of the person included in the first image data; if the extracted image biometric features match the image biometric features of the target person, then locating the location of the person included in the first image data as the current location of the target person.

9. The method according to claim 7, characterized in that, The method further includes: instructing the image acquisition device to track the target person for a preset time, or, upon obtaining the starting position and the destination position, instructing the image acquisition device to stop tracking the target person.

10. The method according to any one of claims 1 to 6, characterized in that, Before acquiring the target voice data and the target image data, the method further includes: acquiring the type of obstacle based on the robot's working scene; if the type of obstacle is the type of object to be moved, locating the position of the obstacle so that, if the obstacle is identified as the target object, the position of the obstacle is used as the starting position of the target object.

11. The method according to claim 10, characterized in that, The method further includes: upon receiving a task instruction including triggering key information, executing the step of obtaining the type of obstacle based on the robot's working scenario.

12. The method according to claim 10, characterized in that, The step of obtaining the type of obstacle based on the robot's working scenario includes at least one of the following: Acquire second image data of obstacles in the robot's working scene from the image acquisition device; identify the type of obstacle by analyzing the second image data; or, read the tags carried by obstacles in the robot's working scene; determine the type of obstacle based on the tags. Using a detection device, obstacles in the robot's working environment are detected; the detection device includes at least one of the image acquisition device and the millimeter-wave antenna; and when the obstacle is detected, the type of the obstacle is identified.

13. An object transfer device, characterized in that, The apparatus is used to implement the steps of the method according to any one of claims 1 to 12, including: The acquisition module is used to acquire target voice data of the target person collected by the voice receiving device and target image data of the target person collected by the image acquisition device. The recognition module is used to perform content recognition on the target voice data and the target image data to obtain the intention and action of the target person; when the intention indicates moving an object, it identifies the starting direction and destination direction of the object placement indicated by the action; it determines the object closest to the starting direction and the destination area closest to the destination direction; or, it determines the object closest to the target person in the starting direction and the destination area closest to the target person in the destination direction; it uses the position of the determined object as the first starting position and the position of the determined destination area as the first destination position to obtain the target object placement... The starting position and the destination position of the object are determined; or, if the intention indicates a second starting position and a second destination position for object placement, the starting direction and the destination direction for object placement indicated by the action are identified; a first distance between the second starting position and the starting direction are determined, and a second distance between the second destination position and the destination direction is determined; if both the first distance and the second distance are less than a threshold distance, the second starting position and the second destination position are respectively taken as the starting position and the destination position for the target object placement; if both the first distance and the second distance are less than the threshold distance, the identification is correct; if neither the first distance nor the second distance is less than the threshold distance, the identification is incorrect. The control module is used to control the robot to move the target object from the starting position to the destination position.

14. An electronic device, characterized in that, It includes a processor, a communication interface, a memory, and a communication bus, wherein the processor, the communication interface, and the memory communicate with each other through the communication bus; Memory, used to store computer programs; A processor, when executing a program stored in memory, implements the steps of the method described in any one of claims 1 to 12.

15. A robot, characterized in that, It includes a voice receiving device, an image acquisition device, a processor, a communication interface, a memory, and a communication bus, wherein the voice receiving device, the image acquisition device, the processor, the communication interface, and the memory communicate with each other through the communication bus; The voice receiving device is used to collect the target voice data of the target person; The image acquisition device is used to acquire target image data of the target person; The memory is used to store computer programs; When the processor executes the program stored in the memory, it implements the steps of the method described in any one of claims 1 to 12.

Citation Information

Patent Citations

  • Self-transfer system of mechanical arm based on combination of single-lens camera and double-lens camera

    CN111267083A

  • Picking system, picking apparatus, control apparatus, storage medium, and method

    CN116020768A