A robot control method and system based on computer vision technology

Through the robot control method based on computer vision technology, autonomous mobile robots can identify task events and make decisions independently in an unset environment, solving the problem that robots in the prior art are unable to adapt to unset environments, and achieving stronger autonomous decision-making and environmental adaptability.

CN119717884BActive Publication Date: 2025-05-16SHENZHEN EPS TECH CO LTD +2
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510223175.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-27
Publication Date
2025-05-16
Estimated Expiration
2045-02-27

AI Technical Summary

Technical Problem

Existing autonomous mobile robots cannot make and adapt independently in unset environments and scenario events, and rely on preset rules to perform tasks.

Method used

Using a robot control method based on computer vision technology, by configuring work task descriptions and work areas, performing patrols and identifying matching task events, obtaining basic information and status information of associated objects, and performing work tasks independently.

Benefits of technology

The autonomous mobile robot has achieved independent decision-making and environmental adaptability in an unset environment, and enhanced its flexibility and effectiveness in independently performing tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119717884B_ABST
    Figure CN119717884B_ABST
Patent Text Reader

Abstract

The present invention is based on the above-mentioned problems and proposes a robot control method and system based on computer vision technology. By configuring the work task description and work area of ​​the autonomous mobile robot, periodic inspections are performed within the scope of the work area, or temporary inspections are performed according to the user's control instructions. During the inspection process, task events matching the work task description are identified. When a task event matching the work task description is identified, basic information and status information of the associated object of the task event are obtained. The status information is a characteristic parameter of the associated object associated with the task event. The work task corresponding to the work task description is performed based on the basic information and status information of the associated object, thereby achieving more powerful autonomous decision-making ability and environmental adaptability by relying on computer vision technology.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of autonomous mobile robot control, and in particular to a robot control method and system based on computer vision technology. Background Art

[0002] Autonomous mobile robots are devices with environmental perception, autonomous mobility, autonomous execution and dynamic decision-making capabilities. Since their introduction, they have been widely used in industrial production, transportation and logistics, catering, hotels and other service fields, as well as household fields. Compared with other intelligent devices, autonomous mobile robots have expanded many functions that can be realized autonomously due to their autonomous mobility, such as multi-point transportation of items in complex space environments, floor cleaning, equipment inspection and home security, etc. An important basis for realizing these functions is the environmental perception capability of autonomous mobile robots, which is provided by various sensors integrated on autonomous mobile robots, such as visual sensors, lidars, ultrasonic sensors, accelerometers, gyroscopes and contact sensors.

[0003] Although autonomous mobile robots have certain autonomous decision-making capabilities, especially with the support of artificial intelligence technology, they can make some relatively simple decisions autonomously in the process of completing tasks in fixed environmental scenarios, such as autonomous obstacle avoidance and path planning. However, in most cases, they still follow pre-set rules to execute, judge and make decisions on various tasks, and are unable to deal with unset scene events and matters outside the established rules. Summary of the invention

[0004] Based on the above problems, the present invention proposes a robot control method and system based on computer vision technology, which realizes more powerful autonomous decision-making ability and environmental adaptability by relying on computer vision technology.

[0005] In view of this, a first aspect of the present invention proposes a robot control method based on computer vision technology, comprising:

[0006] Configuring a work task description and a work area of ​​the autonomous mobile robot, wherein the work task description is a combination of keywords representing the features of the work task or a sentence containing keywords representing the features of the work task;

[0007] Perform periodic inspections within the scope of the working area, or perform temporary inspections according to user control instructions;

[0008] Identify task events that match the work task description during the inspection process;

[0009] When a task event matching the work task description is identified, basic information and state information of an object associated with the task event are obtained, wherein the state information is a characteristic parameter of the associated object associated with the task event;

[0010] The work task corresponding to the work task description is executed based on the basic information and status information of the associated object.

[0011] Furthermore, the step of identifying a task event matching the work task description during the inspection process specifically includes:

[0012] Acquire a preconfigured first frame rate, a second frame rate, a first resolution, and a second resolution, wherein the second frame rate is greater than the first frame rate, and the second resolution is greater than the first resolution;

[0013] During the inspection process, a low-resolution environmental video of the working area is recorded at the first frame rate and the first resolution;

[0014] Performing object contour recognition on the video image in the low-resolution environment video to identify the suspected task object;

[0015] When a suspected task object is identified in the video image in the low-resolution environment video, recording a high-resolution environment video of the area where the suspected task object is located at the second frame rate and the second resolution;

[0016] The suspected task object is identified based on the video image of the high-resolution environment video.

[0017] Determining whether the suspected task object is a task object that matches the work task description;

[0018] When the suspected task object is a task object that matches the work task description, obtaining basic information and status information of the task object;

[0019] It is determined whether there is a task event matching the work task description according to the basic information and status information of the task object.

[0020] Furthermore, the step of performing object contour recognition on the video image in the low-resolution environment video to identify the suspected task object specifically includes:

[0021] Performing edge detection on each frame of video image in the low-resolution environment video to generate an edge image for each frame of video image;

[0022] extracting a contour image from the edge image to generate a first contour image list;

[0023] Calculate the object size corresponding to each contour image in the first contour image list;

[0024] Performing time-sequential matching on the first contour image list according to a pre-configured size range of the task object to obtain a second contour image list;

[0025] The contour images in the second contour image list are matched with a pre-configured contour template of the task object to identify the suspected task object.

[0026] Furthermore, the step of performing edge detection on each frame of video image in the low-resolution environment video to generate an edge image of each frame of video image specifically includes:

[0027] Performing grayscale processing on the video image to obtain a corresponding grayscale image;

[0028] Generate a gradient image of the grayscale image, wherein each pixel value in the gradient image is the modulus of the horizontal gradient and the vertical gradient of the corresponding pixel position on the grayscale image;

[0029] Non-maximum suppression processing and edge connection processing are performed on the gradient image to obtain an edge image of the video image.

[0030] Furthermore, the step of performing time-series matching on the first contour image list according to the pre-configured size range of the task object to obtain the second contour image list specifically includes:

[0031] Determining the video image being processed as a target video image, and determining the shooting time of the target video image as a reference time;

[0032] Traversing each contour image in the first contour image list of the target video image;

[0033] The traversed contour image is used to determine the target contour image, and the following processing is performed on the target contour image:

[0034] Determine the continuous existence interval of the target contour image based on the reference time, wherein the continuous existence interval is a time interval covered by continuous video image frames in which the object corresponding to the contour image continuously exists on the video image;

[0035] Calculate the object size of the target contour image on each frame of the video image in the persistent interval;

[0036] When the object size of the target contour image on each frame of video image in the persistent interval falls within the size range of the task object, the target contour image is put into the second contour image list.

[0037] Furthermore, the step of recording a high-resolution environmental video of the area where the suspected task object is located at the second frame rate and the second resolution specifically includes:

[0038] Positioning the suspected task object in the first several frames of video images of the high-resolution environment video;

[0039] Extracting state information of the suspected task object from the plurality of frames of video images, the state information including motion state information and occlusion status of the suspected task object;

[0040] The recording duration of the high-resolution environment video is determined according to the status information of the suspected task object.

[0041] Further, the step of identifying the suspected task object based on the video image of the high-resolution environment video to determine whether the suspected task object is the task object specifically includes:

[0042] Determining an object boundary of the suspected task object on an edge image, wherein the edge image is an edge image generated by performing edge detection on a low-resolution video image, and the low-resolution video image is a last frame of a video image in the low-resolution environment video;

[0043] Locating the suspected task object in the high-resolution video image based on the position of the object boundary in the low-resolution video image;

[0044] Extracting and generating a local image containing only the suspected task object from the high-resolution video image;

[0045] Inputting the local image into a pre-trained deep learning model to identify the type of the suspected task object;

[0046] Whether the suspected task object is a task object is determined according to the type of the suspected task object.

[0047] Furthermore, the step of determining whether there is a task event matching the work task description according to the basic information and status information of the task object specifically includes:

[0048] Generate a task description instance based on the work task description, wherein the task description instance includes a task object, a first state feature sequence, an action sequence, and a second state feature sequence;

[0049] Matching the keywords in the first state feature sequence with the basic information and state information of the task object to calculate a feature matching degree;

[0050] When the matching degree is greater than a preset matching degree threshold, it is determined that there is a task event matching the work task description.

[0051] Furthermore, the step of matching the keywords in the first state feature sequence with the basic information and state information of the task object to calculate the feature matching degree specifically includes:

[0052] The basic information and status information of the task object are combined to generate an object feature word list;

[0053] Traverse each keyword in the first state feature sequence to perform the following steps:

[0054] Determine the currently traversed keyword as the target keyword;

[0055] Sequentially calculate the semantic similarity between the target keyword and each feature word in the object feature word list ,in 1 to A positive integer between 1 to A positive integer between

[0056] is the number of keywords in the first state feature sequence, is the number of feature words in the object feature word list;

[0057] The maximum value of the semantic similarities is determined as the keyword matching degree of the target keyword:

[0058] ;

[0059] After the keyword traversal in the first state feature sequence is completed, the feature matching degree is calculated:

[0060] .

[0061] The second aspect of the present invention proposes a robot control system based on computer vision technology, comprising a first image sensor for recording low-resolution environmental video, a second image sensor for recording high-resolution environmental video, and a control device connected to the first image sensor and the second image sensor, wherein the control device is configured to implement the robot control method based on computer vision technology as described in any one of the first aspect of the present invention.

[0062] The present invention is based on the above-mentioned problems and proposes a robot control method and system based on computer vision technology. By configuring the work task description and work area of ​​the autonomous mobile robot, periodic inspections are performed within the scope of the work area, or temporary inspections are performed according to the user's control instructions. During the inspection process, task events matching the work task description are identified. When a task event matching the work task description is identified, basic information and status information of the associated object of the task event are obtained. The status information is a characteristic parameter of the associated object associated with the task event. The work task corresponding to the work task description is performed based on the basic information and status information of the associated object, thereby achieving more powerful autonomous decision-making ability and environmental adaptability by relying on computer vision technology. BRIEF DESCRIPTION OF THE DRAWINGS

[0063] Figure 1 The present invention provides a flowchart of a robot control method based on computer vision technology according to an embodiment of the present invention. DETAILED DESCRIPTION

[0064] In order to more clearly understand the above-mentioned purpose, features and advantages of the present invention, the present invention is further described in detail below in conjunction with the accompanying drawings and specific embodiments. It should be noted that the embodiments of the present application and the features in the embodiments can be combined with each other without conflict.

[0065] In the following description, many specific details are set forth to facilitate a full understanding of the present invention. However, the present invention may also be implemented in other ways different from those described herein. Therefore, the protection scope of the present invention is not limited to the specific embodiments disclosed below.

[0066] In the description of the present invention, the term "multiple" refers to two or more. Unless otherwise clearly defined, the orientation or positional relationship indicated by the terms "upper", "lower", etc. is based on the orientation or positional relationship shown in the drawings. It is only for the convenience of describing the present invention and simplifying the description, and does not indicate or imply that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation. Therefore, it cannot be understood as a limitation of the present invention. The terms "connection", "installation", "fixation", etc. should be understood in a broad sense. For example, "connection" can be a fixed connection, a detachable connection, or an integral connection; it can be directly connected or indirectly connected through an intermediate medium. For ordinary technicians in this field, the specific meanings of the above terms in the present invention can be understood according to the specific circumstances. In addition, the terms "first", "second", etc. are only used for descriptive purposes, and cannot be understood as indicating or implying relative importance or implicitly indicating the number of technical features indicated. Therefore, the features defined as "first", "second", etc. can explicitly or implicitly include one or more of the features. In the description of the present invention, unless otherwise specified, the meaning of "multiple" is two or more.

[0067] In the description of this specification, the description of the terms "one embodiment", "some implementations", "specific embodiments", etc. means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representation of the above terms does not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in any one or more embodiments or examples in a suitable manner.

[0068] A robot control method and system based on computer vision technology provided according to some embodiments of the present invention will be described below with reference to the accompanying drawings.

[0069] like Figure 1 As shown, the first aspect of the present invention proposes a robot control method based on computer vision technology, comprising:

[0070] Configuring a work task description and a work area of ​​the autonomous mobile robot, wherein the work task description is a combination of keywords representing the features of the work task or a sentence containing keywords representing the features of the work task;

[0071] Perform periodic inspections within the scope of the working area, or perform temporary inspections according to user control instructions;

[0072] Identify task events that match the work task description during the inspection process;

[0073] When a task event matching the work task description is identified, basic information and state information of an object associated with the task event are obtained, wherein the state information is a characteristic parameter of the associated object associated with the task event;

[0074] The work task corresponding to the work task description is executed based on the basic information and status information of the associated object.

[0075] Specifically, in the step of configuring the work task description of the autonomous mobile robot, the user can configure the work task description to the autonomous mobile robot by means of voice commands, or can configure the work task description to the autonomous mobile robot by sending control commands through a remote control program running in a remote control device such as a computer or a smart phone. The work task characteristics include but are not limited to the actions performed by the autonomous mobile robot, the basic information and status information of the associated object, etc. The basic information of the associated object includes one or more keywords such as the name, type, shape, size, color, etc. of the associated object for describing the main features of the associated object, and the status information of the associated object includes one or more keywords such as the position, posture, direction, etc. of the associated object for describing the state features of the associated object. Of course, the basic information of the associated object may also include an item number as a unique identity identifier of the associated object, etc. In the technical solutions of some embodiments of the present invention, the work task description may be expressed as a combination of several of the keywords such as "action", "name of the associated object", "type of the associated object", "current location of the associated object", "current status of the associated object", "target status of the associated object" or "target location of the associated object".

[0076] The working area is the geographical range in which the autonomous mobile robot can move autonomously when completing the work task corresponding to the work task description. Specifically, a specific communication device in a specified location can be used as a reference to determine the geographical area within a certain range around the communication device as the working area. The communication device is usually placed in a relatively fixed position and is not often moved, such as a routing device or other gateway device. Similarly, in the step of configuring the working area of ​​the autonomous mobile robot, the user can configure the activity range of the autonomous mobile robot through a remote control program running in a remote control device such as a computer or a smart phone.

[0077] Furthermore, the step of performing periodic inspection within the scope of the working area specifically includes moving within the working area according to a pre-configured inspection period and inspecting the working area through a visual sensor. Similarly, the step of performing temporary inspection within the scope of the working area according to a user's control instruction specifically includes moving within the working area and inspecting the working area through a visual sensor when receiving an inspection control instruction input by the user through voice, gesture or other means.

[0078] Furthermore, the step of identifying a task event matching the work task description during the inspection process specifically includes:

[0079] Acquire a preconfigured first frame rate, a second frame rate, a first resolution, and a second resolution, wherein the second frame rate is greater than the first frame rate, and the second resolution is greater than the first resolution;

[0080] During the inspection process, a low-resolution environmental video of the working area is recorded at the first frame rate and the first resolution;

[0081] Performing object contour recognition on the video image in the low-resolution environment video to identify the suspected task object;

[0082] When a suspected task object is identified in the video image in the low-resolution environment video, recording a high-resolution environment video of the area where the suspected task object is located at the second frame rate and the second resolution;

[0083] Identifying the suspected task object based on the video image of the high-resolution environment video to determine whether the suspected task object is a task object that matches the work task description;

[0084] When the suspected task object is a task object that matches the work task description, obtaining basic information and status information of the task object;

[0085] It is determined whether there is a task event matching the work task description according to the basic information and status information of the task object.

[0086] In the technical solution of the above embodiment, the autonomous mobile robot continuously records low-resolution environmental video of the working area at a lower frame rate during the inspection process to identify suspected task objects during movement. It should be known that the inspection area of ​​the autonomous mobile robot can be a designated part of the working area.

[0087] In the technical solutions of some embodiments of the present invention, when the suspected task object is a static object and the area where the suspected task object is located is a fixed area, then one or a small number of high-resolution images of the area where the suspected task object is located can be taken at the second resolution. When the suspected task object is a dynamic object, a step of recording a high-resolution environmental video of the area where the suspected task object is located at the second frame rate and the second resolution is performed.

[0088] Furthermore, the step of performing object contour recognition on the video image in the low-resolution environment video to identify the suspected task object specifically includes:

[0089] Performing edge detection on each frame of video image in the low-resolution environment video to generate an edge image for each frame of video image;

[0090] extracting a contour image from the edge image to generate a first contour image list;

[0091] Calculate the object size corresponding to each contour image in the first contour image list;

[0092] Performing time-sequential matching on the first contour image list according to a pre-configured size range of the task object to obtain a second contour image list;

[0093] The contour images in the second contour image list are matched with a pre-configured contour template of the task object to identify the suspected task object.

[0094] The edge image is a simplified image that only retains the contour information of each object in the image content after the color information and content details of the video image are discarded. In the technical solution of the above-mentioned implementation, the first contour image list contains the contour images of all objects on the video image, and the so-called objects are the people or objects on the video image. The object size refers to the actual size of the corresponding person or object in the real space, rather than its size on the image, and specifically may include the height, width, etc. of the corresponding person or object.

[0095] In the technical solution of the above-mentioned embodiment, a contour template of the task object at different shooting angles is pre-configured, which can be a plurality of task object pictures obtained by shooting the task object at different angles, and a plurality of contour images corresponding to different shooting angles are generated by performing edge detection on the pictures.

[0096] Furthermore, the step of performing edge detection on each frame of video image in the low-resolution environment video to generate an edge image of each frame of video image specifically includes:

[0097] Performing grayscale processing on the video image to obtain a corresponding grayscale image;

[0098] Generate a gradient image of the grayscale image, wherein each pixel value in the gradient image is the modulus of the horizontal gradient and the vertical gradient of the corresponding pixel position on the grayscale image;

[0099] Non-maximum suppression processing and edge connection processing are performed on the gradient image to obtain an edge image of the video image.

[0100] The grayscale image is the image after the color information of the video image is lost. More specifically, a color image is usually processed into a corresponding grayscale image using a mean method or a weighted average method, that is, the three RGB color values ​​are averaged or weighted averaged to obtain a grayscale value of the corresponding pixel in the range of 0 to 255.

[0101] Furthermore, before the step of calculating the gradient magnitude and gradient direction of each pixel in the grayscale image, the method further includes using a Gaussian filter to smooth the grayscale image to filter out noise information in the grayscale image.

[0102] The horizontal gradient of each pixel position on the grayscale image is the horizontal first-order derivative of the grayscale value of the pixel position. Similarly, the vertical gradient of each pixel position on the grayscale image is the vertical first-order derivative of the grayscale value of the pixel position. The modulus of the horizontal gradient and the vertical gradient of the corresponding pixel position on the grayscale image refers to the positive square root of the sum of the squares of the horizontal first-order derivative and the vertical first-order derivative of the corresponding pixel position.

[0103] Furthermore, the step of performing time-series matching on the first contour image list according to the pre-configured size range of the task object to obtain the second contour image list specifically includes:

[0104] Determining the video image being processed as a target video image, and determining the shooting time of the target video image as a reference time;

[0105] Traversing each contour image in the first contour image list of the target video image;

[0106] The traversed contour image is used to determine the target contour image, and the following processing is performed on the target contour image:

[0107] Determine the continuous existence interval of the target contour image based on the reference time, wherein the continuous existence interval is a time interval covered by continuous video image frames in which the object corresponding to the contour image continuously exists on the video image;

[0108] Calculate the object size of the target contour image on each frame of the video image in the persistent interval;

[0109] When the object size of the target contour image on each frame of video image in the persistent interval falls within the size range of the task object, the target contour image is put into the second contour image list.

[0110] In the step of performing object contour recognition on the video image in the low-resolution environment video to identify the suspected task object, the autonomous mobile robot processes each frame of the video image in the recorded low-resolution environment video of the working area frame by frame in time sequence, or performs sampling processing, so generally only one frame of the video image is processed at the same time (without multi-threading).

[0111] Those skilled in the art know that even if multi-threaded processing is adopted, it will not affect the implementation of the technical solution of this embodiment).

[0112] The contour images in the second contour image list are contour images of objects whose object sizes match the size of the preconfigured task object, that is, the contour images in the second contour image list are contour images of objects whose object sizes fall within the size range of the preconfigured task object.

[0113] Furthermore, the step of recording a high-resolution environmental video of the area where the suspected task object is located at the second frame rate and the second resolution specifically includes:

[0114] Positioning the suspected task object in the first several frames of video images of the high-resolution environment video;

[0115] Extracting state information of the suspected task object from the plurality of frames of video images, the state information including motion state information and occlusion status of the suspected task object;

[0116] The recording duration of the high-resolution environment video is determined according to the status information of the suspected task object.

[0117] The first several frames of video images of the high-resolution environment video refer to the earliest several frames of video images recorded in the high-resolution environment video generated by the step of recording the high-resolution environment video of the area where the suspected task object is located at the second frame rate and the second resolution, wherein the several frames can be a certain pre-configured number of frames, such as 3 frames or 5 frames, etc., for extracting the motion state information of the suspected task object at the early stage of the high-resolution environment video. The motion state information includes but is not limited to the motion direction and motion speed of the suspected task object.

[0118] In the step of determining the recording duration of the high-resolution environment video according to the status information of the suspected task object, the moving direction, moving speed and occlusion of the suspected task object determine the recording duration of the suspected task object. When the suspected task object is not occluded, a shorter recording duration of the high-resolution environment video can be configured. When the suspected task object is occluded, the change of the occlusion of the suspected task object is determined according to the moving direction of the suspected task object, that is, the situation in which the occluded area becomes larger or smaller, so as to calculate the recording duration of the high-resolution environment video according to the moving speed of the suspected task object.

[0119] Further, the step of identifying the suspected task object based on the video image of the high-resolution environment video to determine whether the suspected task object is the task object specifically includes:

[0120] Determining an object boundary of the suspected task object on an edge image, wherein the edge image is an edge image generated by performing edge detection on a low-resolution video image, and the low-resolution video image is a last frame of a video image in the low-resolution environment video;

[0121] Locating the suspected task object in the high-resolution video image based on the position of the object boundary in the low-resolution video image;

[0122] Extracting and generating a local image containing only the suspected task object from the high-resolution video image;

[0123] Inputting the local image into a pre-trained deep learning model to identify the type of the suspected task object;

[0124] Whether the suspected task object is a task object is determined according to the type of the suspected task object.

[0125] The last frame of video image in the low-resolution environmental video refers to the last frame of video image in the low-resolution environmental video recorded before the reference time, taking the time when the step of recording the high-resolution environmental video of the area where the suspected task object is located at the second frame rate and the second resolution is performed.

[0126] Furthermore, after the step of performing object contour recognition on the video image in the low-resolution environment video to identify the suspected task object, when the suspected task object is detected, the suspected task object is continuously tracked in the low-resolution environment video and the suspected task object is continuously tracked within the reference time.

[0127] The position of the object boundary of the suspected task object is determined on each subsequent frame of the low-resolution video image.

[0128] Preferably, the deep learning model is a convolutional neural network model obtained by deep training using a large amount of sample data of task objects prepared in advance.

[0129] Furthermore, the step of determining whether there is a task event matching the work task description according to the basic information and status information of the task object specifically includes:

[0130] Generate a task description instance based on the work task description, wherein the task description instance includes a task object, a first state feature sequence, an action sequence, and a second state feature sequence;

[0131] Matching the keywords in the first state feature sequence with the basic information and state information of the task object to calculate a feature matching degree;

[0132] When the matching degree is greater than a preset matching degree threshold, it is determined that there is a task event matching the work task description.

[0133] In the technical solution of the above-mentioned embodiment, the action sequence includes a sequence of one or more actions that the autonomous mobile robot needs to perform when completing the work task corresponding to the work task description. The first state feature sequence is used to describe the initial state of the task object, which includes one or more keywords that describe the state characteristics of the task object before the autonomous mobile robot executes the action sequence. The second state feature sequence is used to describe the target state of the task object, which includes one or more keywords that describe the state characteristics of the task object after the autonomous mobile robot executes the action sequence.

[0134] Furthermore, the step of executing the work task corresponding to the work task description based on the basic information and status information of the associated object specifically includes controlling the autonomous mobile robot to execute the action sequence so that the task object changes from the initial state described by the first state feature sequence to the target state described by the second state feature sequence.

[0135] Furthermore, the step of matching the keywords in the first state feature sequence with the basic information and state information of the task object to calculate the feature matching degree specifically includes:

[0136] The basic information and status information of the task object are combined to generate an object feature word list;

[0137] Traverse each keyword in the first state feature sequence to perform the following steps:

[0138] Determine the currently traversed keyword as the target keyword;

[0139] Sequentially calculate the semantic similarity between the target keyword and each feature word in the object feature word list ,in 1 to A positive integer between 1 to A positive integer between is the number of keywords in the first state feature sequence,

[0140] is the number of feature words in the object feature word list;

[0141] The maximum value of the semantic similarities is determined as the keyword matching degree of the target keyword:

[0142] ;

[0143] After the keyword traversal in the first state feature sequence is completed, the feature matching degree is calculated:

[0144] .

[0145] Specifically, the target keyword and the feature words in the object feature word list may be vectorized and then the cosine similarity between the two may be calculated as the semantic similarity between the two. The vectorization may be implemented based on a pre-trained word vector file such as GloVe.

[0146] The second aspect of the present invention proposes a robot control system based on computer vision technology, comprising a first image sensor for recording low-resolution environmental video, a second image sensor for recording high-resolution environmental video, and a control device connected to the first image sensor and the second image sensor, wherein the control device is configured to implement the robot control method based on computer vision technology as described in any one of the first aspect of the present invention.

[0147] In the technical solutions of some embodiments of the present invention, the first image sensor is an image sensor integrated on the autonomous mobile robot, and the second image sensor is an image sensor installed in the working area. The control device can be a processor installed in the autonomous mobile robot, or an external control device that is communicatively connected to the autonomous mobile robot.

[0148] By adopting the technical solution of the above-mentioned implementation mode, the user only needs to configure the work task description of the autonomous mobile robot, and the autonomous mobile robot can autonomously complete related tasks in the configured working area without the need for cumbersome configuration and complex control operations, so that the autonomous mobile robot has more powerful autonomous decision-making ability and environmental adaptability.

[0149] It should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the statement "comprise a ..." do not exclude the presence of other identical elements in the process, method, article or device including the elements.

[0150] According to the embodiments of the present invention as described above, these embodiments do not describe all the details in detail, nor do they limit the invention to the specific embodiments described. Obviously, many modifications and changes can be made based on the above description. This specification selects and specifically describes these embodiments in order to better explain the principles and practical applications of the present invention, so that those skilled in the art can make good use of the present invention and the modified use based on the present invention. The present invention is limited only by the claims and their full scope and equivalents.

Claims

1. A robot control method based on computer vision technology, characterized in that: include: Configuring a work task description and a work area of ​​the autonomous mobile robot, wherein the work task description is a combination of keywords representing the features of the work task or a sentence containing keywords representing the features of the work task; Perform periodic inspections within the scope of the working area, or perform temporary inspections according to user control instructions; Identify task events that match the work task description during the inspection process; When a task event matching the work task description is identified, basic information and status information of an object associated with the task event are obtained, wherein the status information is a characteristic parameter of the associated object associated with the task event; Execute the work task corresponding to the work task description based on the basic information and status information of the associated object; The steps of identifying task events matching the work task description during the inspection process specifically include: Acquire a preconfigured first frame rate, a second frame rate, a first resolution, and a second resolution, wherein the second frame rate is greater than the first frame rate, and the second resolution is greater than the first resolution; During the inspection process, a low-resolution environmental video of the working area is recorded at the first frame rate and the first resolution; Performing object contour recognition on the video image in the low-resolution environment video to identify the suspected task object; When a suspected task object is identified in the video image in the low-resolution environment video, recording a high-resolution environment video of the area where the suspected task object is located at the second frame rate and the second resolution; Identifying the suspected task object based on the video image of the high-resolution environment video to determine whether the suspected task object is a task object that matches the work task description; When the suspected task object is a task object that matches the work task description, obtaining basic information and status information of the task object; Determine whether there is a task event matching the work task description according to the basic information and status information of the task object; The step of judging whether there is a task event matching the work task description according to the basic information and status information of the task object specifically includes: Generate a task description instance based on the work task description, wherein the task description instance includes a task object, a first state feature sequence, an action sequence, and a second state feature sequence; Matching the keywords in the first state feature sequence with the basic information and state information of the task object to calculate a feature matching degree; When the matching degree is greater than a preset matching degree threshold, it is determined that there is a task event matching the work task description.

2. The robot control method based on computer vision technology according to claim 1, characterized in that: The step of performing object contour recognition on the video image in the low-resolution environment video to identify the suspected task object specifically includes: Performing edge detection on each frame of video image in the low-resolution environment video to generate an edge image for each frame of video image; extracting a contour image from the edge image to generate a first contour image list; Calculate the object size corresponding to each contour image in the first contour image list; Performing time-sequential matching on the first contour image list according to a pre-configured size range of the task object to obtain a second contour image list; The contour images in the second contour image list are matched with a pre-configured contour template of the task object to identify the suspected task object.

3. The robot control method based on computer vision technology according to claim 2, characterized in that: The step of performing edge detection on each frame of video image in the low-resolution environment video to generate an edge image of each frame of video image specifically includes: Performing grayscale processing on the video image to obtain a corresponding grayscale image; Generate a gradient image of the grayscale image, wherein each pixel value in the gradient image is the modulus of the horizontal gradient and the vertical gradient of the corresponding pixel position on the grayscale image; Non-maximum suppression processing and edge connection processing are performed on the gradient image to obtain an edge image of the video image.

4. The robot control method based on computer vision technology according to claim 2, characterized in that: The step of performing temporal matching on the first contour image list according to the size range of the pre-configured task object to obtain the second contour image list specifically includes: Determining the video image being processed as a target video image, and determining the shooting time of the target video image as a reference time; Traversing each contour image in the first contour image list of the target video image; The traversed contour image is used to determine the target contour image, and the following processing is performed on the target contour image: Determine the continuous existence interval of the target contour image based on the reference time, wherein the continuous existence interval is a time interval covered by continuous video image frames in which the object corresponding to the contour image continuously exists on the video image; Calculate the object size of the target contour image on each frame of the video image in the persistent interval; When the object size of the target contour image on each frame of video image in the persistent interval falls within the size range of the task object, the target contour image is put into the second contour image list.

5. The robot control method based on computer vision technology according to claim 1, characterized in that: The step of recording a high-resolution environmental video of the area where the suspected task object is located at the second frame rate and the second resolution specifically includes: Positioning the suspected task object in the first several frames of video images of the high-resolution environment video; Extracting state information of the suspected task object from the plurality of frames of video images, the state information including motion state information and occlusion status of the suspected task object; The recording duration of the high-resolution environment video is determined according to the status information of the suspected task object.

6. The robot control method based on computer vision technology according to claim 1, characterized in that: The step of identifying the suspected task object based on the video image of the high-resolution environment video to determine whether the suspected task object is the task object specifically includes: Determining an object boundary of the suspected task object on an edge image, wherein the edge image is an edge image generated by performing edge detection on a low-resolution video image, and the low-resolution video image is a last frame of a video image in the low-resolution environment video; Locating the suspected task object in the high-resolution video image based on the position of the object boundary in the low-resolution video image; Extracting and generating a local image containing only the suspected task object from the high-resolution video image; Inputting the local image into a pre-trained deep learning model to identify the type of the suspected task object; Whether the suspected task object is a task object is determined according to the type of the suspected task object.

7. The robot control method based on computer vision technology according to claim 1, characterized in that: The step of matching the keywords in the first state feature sequence with the basic information and state information of the task object to calculate the feature matching degree specifically includes: The basic information and status information of the task object are combined to generate an object feature word list; Traverse each keyword in the first state feature sequence to perform the following steps: Determine the currently traversed keyword as the target keyword; Sequentially calculate the semantic similarity between the target keyword and each feature word in the object feature word list ,in 1 to A positive integer between 1 to A positive integer between is the number of keywords in the first state feature sequence, is the number of feature words in the object feature word list; The maximum value of the semantic similarities is determined as the keyword matching degree of the target keyword: ; After the keyword traversal in the first state feature sequence is completed, the feature matching degree is calculated: 。 8. A robot control system based on computer vision technology, characterized in that: The invention comprises a first image sensor for recording low-resolution environment video, a second image sensor for recording high-resolution environment video, and a control device connected to the first image sensor and the second image sensor, wherein the control device is configured to implement the robot control method based on computer vision technology as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Security automation in a mobile robot

    US20200053324A1

  • Intelligent visual perception system

    WO2021253961A1