Computer-implemented method for generating a set of predefined text descriptions for a machine learning model trained for object recognition with an open vocabulary.
An automated method for generating predefined text descriptions using a text encoder improves the accuracy and efficiency of object recognition in machine learning models with an open vocabulary, addressing the challenges of manual selection and language barriers.
Patent Information
- Application Number
- DE102024207609
- Authority / Receiving Office
- DE · DE
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-08-09
- Publication Date
- 2026-02-12
AI Technical Summary
The accuracy of machine learning models with an open vocabulary for object recognition is dependent on terminology, leading to inefficiencies in manual selection and increased costs due to prompt engineering, and risks from spelling errors and language barriers.
An automated method for generating a set of predefined text descriptions using a text encoder to determine the most similar dictionary text descriptions, reducing the need for manual selection and improving recognition accuracy.
Automated selection of text descriptions enhances recognition accuracy, reduces costs, and minimizes errors, enabling efficient object recognition in dynamic environments.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
State of the art
[0001] A trained machine learning model with an open vocabulary can be able to recognize objects in images that it was not trained on. For this purpose, the machine learning model can use a dictionary with a large number of dictionary text descriptions (as an open vocabulary). Specifically, the dictionary can contain dictionary text descriptions that refer to objects not present in the training images.
[0002] However, the accuracy with which the machine learning model recognizes these objects corresponding to dictionary text descriptions can depend significantly on the terminology, so that supposed synonyms can lead to considerably different recognition rates. For example, the recognition rate of the machine learning model in an image might be higher for the term "office chair" than for the term "chair." The manual selection of suitable terminology (with a high recognition rate) is also known as prompt engineering. Disclosure of the invention
[0003] This disclosure relates to methods for generating a set of predefined text descriptions for a machine learning model trained for object recognition with an open vocabulary, enabling automated selection of the appropriate terminology. This eliminates the need for manual selection during prompt engineering, significantly reducing both costs (e.g., personnel costs for prompt engineering) and time. It also reduces the likelihood of spelling errors and the risk of performance loss due to a limited vocabulary resulting from operator language barriers.
[0004] Several aspects concern a computer-implemented method for generating a set of predefined text descriptions for a machine learning model (pre-)trained for object recognition with an open vocabulary, such that for each predefined text description, a corresponding area in an image input into the machine learning model is displayed if the area shows an object represented by the predefined text description. The method comprises: providing a multitude of images and a multitude of initial text descriptions, each of which is assigned to an area (e.g., as a mask, bounding box, etc.) in a corresponding image of the multitude of images and indicates what is shown in the area; and determining a multitude of coded dictionary text descriptions.by generating a coded dictionary text description for each dictionary text description from a multitude of dictionary text descriptions using a text encoder of the machine learning model; for each initial text description from the multitude of initial text descriptions: determining a coded initial text description using the text encoder; selecting (a predefined number of) one or more coded dictionary text descriptions from the multitude of coded dictionary text descriptions which are most similar to the coded initial text description according to a first (semantic) similarity measure; for each text description of the initial text description and each dictionary text description that belongs to one of the (selected) one or more coded dictionary text descriptions, as a predefined text description.Input at least the image corresponding to the initial text description into the machine learning model and determine a similarity between an output of the machine learning model and the area assigned to the initial text description according to a second similarity measure, and add the text description with the greatest (determined) similarity to the set of predefined text descriptions.
[0005] The following are various examples of implementation.
[0006] Example 1 is the procedure for generating a sentence from given text descriptions as described above.
[0007] Example 2 is set up according to Example 1, wherein the area determined in the input image for the specified text description using the machine learning model and the area assigned to the initial text description are represented by a mask (e.g. segmentation mask) and / or a bounding box.
[0008] Example 3 is set up according to Example 1 or 2, wherein the first similarity measure has a cosine similarity; and / or wherein the second similarity measure has an intersection set over union set.
[0009] Example 4 is set up according to one of Examples 1 to 3, wherein the one or more coded dictionary text descriptions are selected according to a predefined number.
[0010] Example 5 is set up according to one of Examples 1 to 4, wherein the input of at least the image associated with the initial text description into the machine learning model features that each image of the multitude of images is entered into the machine learning model and the similarity across all images is determined.
[0011] This ensures that the text description is added to the set of predefined text descriptions for which the greatest similarity (e.g., a summed similarity) is determined in the totality of all images of the multitude of images.
[0012] Example 6 is a data processing unit that is set up to execute the procedure according to any one of Examples 1 to 5.
[0013] Example 7 is a method for controlling a (navigable) robotic device, comprising: while the robotic device navigates in its environment, generating a map of the environment using simultaneous localization and mapping (SLAM) and capturing images representing the environment; performing semantic object recognition for each captured image using the machine learning model with the set of predefined text descriptions generated according to one of Examples 1 to 5; generating a semantic map of the environment by integrating a result of the semantic object recognition into the map of the environment; and controlling the robotic device using the semantic map of the environment.
[0014] Example 8 is a control device which has one or more processors configured to perform the procedure according to Example 7.
[0015] Example 9 is a robot device comprising: the control device according to Example 8; and at least one imaging sensor configured to capture images of the environment of the robot device.
[0016] Example 10 is a computer program with instructions which, when executed by a processor, cause the processor to perform the procedure according to one of Examples 1 to 5 or 7.
[0017] Example 11 is a computer-readable medium that stores instructions which, when executed by a processor, cause the processor to perform the procedure according to one of Examples 1 to 5 or 7.
[0018] In the drawings, similar reference numerals generally refer to the same parts in all the different views. The drawings are not necessarily to scale, with the emphasis generally placed on illustrating the principles of the invention. Various aspects are described in the following description with reference to the drawings. Fig. Figure 1 shows a robot device arrangement according to various aspects; Figure 2 shows a flowchart of a procedure for generating a set of predefined text descriptions for a machine learning model trained for object recognition with an open vocabulary according to various aspects; Fig. Figure 3 shows a schematic flowchart of the procedure for generating the set of given text descriptions according to various aspects; and Fig. Figure 4 shows the use of the method to generate a semantic map of the robot device's environment according to various aspects.
[0019] The following detailed description refers to the accompanying drawings, which illustrate specific details and aspects of this disclosure in which the invention can be implemented. Other aspects may be used, and structural, logical, and electrical modifications may be made without deviating from the scope of the invention. The various aspects of this disclosure are not necessarily mutually exclusive, as some aspects of this disclosure may be combined with one or more other aspects of this disclosure to form new aspects.
[0020] Several examples are described in more detail below.
[0021] Fig. Figure 1 shows a robot device arrangement 100 according to various aspects. The robot device arrangement 100 can include a robot device 102 (hereinafter referred to as "robot"). The in Fig. The robot device 102 shown and described below as an example represents an exemplary robot device for illustrative purposes and can, for example, be a transport robot for transporting objects (e.g., goods) 104 within its (dynamic) environment, such as a factory or a warehouse. The robot device 102 can, for example, be an Active Shuttle from Bosch Rexroth. The robot device 102 can be a robot that moves on the ground. For this purpose, the robot device 102 can, for example, have wheels or any other suitable component (e.g., caterpillar tracks, support legs, etc.). It should be noted that this robot device serves for illustrative purposes and can generally be any type of computer-controlled device that is capable of navigating (e.g., autonomously or at least semi-autonomously) in its environment, such as a household robot (e.g., a vacuum cleaner).a cleaning robot), a vehicle that is at least partially automated, etc.
[0022] To illustrate, shows Fig. 1 a warehouse as the (dynamic) environment of the robot device 102, which may contain static objects, such as the objects 104, one or more shelves 110 etc., and / or dynamic objects, such as one or more other robot devices 112, one or more forklifts 108, one or more people (e.g. workers) 106 etc.
[0023] To control the robot device 102, the robot device arrangement 100 can include a (robot) control device 114, which is configured to implement the interaction with the environment according to a control program. In some aspects, the robot device 102 can include the control device 114. In other aspects, the robot device arrangement 100 can include a (central) control device 114, which can be configured to control the robot device 102 and optionally one or more further robot devices (e.g., including the further robot device 112). In still other aspects, the robot device 102 can include a control device that implements part of the control device 114, and another (central) control device can implement another part of the control device 114 described herein.
[0024] The term "control device" (also referred to as "control unit") can be understood as any type of logical implementation unit, which may include, for example, a circuit and / or a processor capable of executing software, firmware, or a combination thereof stored on a memory medium and capable of issuing instructions, e.g., to an actuator in this example. The control device can, for example, be configured by program code (e.g., software) to control the operation of a system, in this example, a robot.
[0025] In the present example, the control device 114 can comprise a computer 116 and a memory 118, which stores code and data on the basis of which the computer 116 controls the robot device 102. According to various embodiments, the control device 114 can control the robot device 102 on the basis of a robot control model 120 stored in the memory 118.
[0026] To enable the robot device 102 to navigate its (dynamic) environment, the control device 114 can use sensor data representing the environment of the robot device 102. For example, the sensor data can include images of the environment of the robot device 102, provided by one or more imaging sensors 122. At least one of the one or more imaging sensors 122 can be attached to the robot device 102 and / or at least one of the one or more imaging sensors 122 can be separate from the robot device 102 (e.g., to allow a view for observing more than one robot device 102).
[0027] An imaging sensor, as used herein, can be, for example, a camera (e.g., a standard camera, a digital camera, an infrared camera, a stereo camera, etc.), a radar sensor, a lidar sensor, an ultrasonic sensor, etc. Therefore, an image can be an RGB image, an RGB-D image, or a depth image (also called a D-image). A depth image described herein can be any type of image that contains depth information. Intuitively, a depth image can contain three-dimensional information about one or more objects. A depth image described herein can, for example, contain a point cloud provided by a lidar sensor and / or a radar sensor. A depth image can, for example, be an image with depth information provided by a lidar sensor.
[0028] It is understood that the one or more imaging sensors 122 serve as examples and that the robot device arrangement 100 can include any other type of one or more perception sensors.
[0029] The control device 114 can be configured to control the robot device 102 based on an output of the control model 120 in response to the input of at least one image into the control model 120.
[0030] To detect objects in the vicinity of the robot device 102 (e.g., to detect the objects 104, the one or more people 106, the one or more forklifts 108, the one or more other robot devices 112, etc.), the control model 120 can include a machine learning model trained for object detection with an open vocabulary.
[0031] A machine learning model trained for object recognition with an open vocabulary may also be able to recognize objects for which it was not explicitly trained. For this purpose, the machine learning model can use a set of predefined text descriptions. As explained herein, the accuracy with which the machine learning model recognizes objects can depend on the terminology of the text descriptions in the set of predefined text descriptions.
[0032] Fig. Figure 2 shows a flowchart of a (computer-implemented) procedure 200 for generating a set of predefined text descriptions for a machine learning model trained for object recognition with an open vocabulary according to various aspects.
[0033] The machine learning model can be trained to output a specific area in an image input to the machine learning model for each given text description, provided that area represents an object defined by the given text description. For example, a given text description could be the term "chair," and the machine learning model could recognize an area in an image input if that area depicts a "chair."
[0034] The procedure 200 enables automated determination of the set of given text descriptions.
[0035] Method 200 (in 202) can involve providing a multitude of images and a multitude of initial text descriptions. Each initial text description can be assigned to an area (e.g., as a mask, bounding box, etc.) within a corresponding image and indicate what is shown in that area. For example, an image might show a chair, which is represented by a corresponding area (e.g., as a mask or bounding box) within the image, and the corresponding initial text description might be "chair". The object recognition described herein can also include image segmentation.
[0036] The procedure 200 can (in 204) feature a determination of a large number of coded dictionary text descriptions by generating a coded dictionary text description for each dictionary text description of a large number of dictionary text descriptions using a text encoder of the machine learning model.
[0037] Method 200 can (in 206) exhibit for each initial text description of the plurality of initial text descriptions: - Determining a coded initial text description using the text encoder (in 206A); - Selecting (a predefined number of) one or more coded dictionary text descriptions from the multitude of coded dictionary text descriptions which are most similar to the coded initial text description according to a first (semantic) similarity measure (in 206B); - for each text description of the initial text description and each dictionary text description belonging to one of the (selected) one or more coded dictionary text descriptions, as a given text description, input at least the image belonging to the initial text description into the machine learning model and determine a similarity between an output of the machine learning model and the area assigned to the initial text description according to a second similarity measure (in 206C); and - Adding the text description with the greatest (determined) similarity to the set of given text descriptions.
[0038] The following section describes various aspects of Procedure 200 in more detail. Fig. Figure 3 shows a schematic flowchart 300 with various aspects of the procedure 200.
[0039] An open-vocabulary machine learning model 302 can generally include a text encoder 312 and an image encoder 314. The image encoder 314 can be configured to encode an input image (i.e., to generate a coded image; also referred to as a latent representation of the image). The text encoder 312 can be configured to encode an input text description (i.e., to generate a coded text description; also referred to as a latent representation of the text description). The text encoder 312 can be trained to map two texts with similar semantic meaning (e.g., synonyms) to two latent representations that have a high degree of similarity according to a predefined similarity measure.
[0040] The multitude of images 304(n=1...N) provided in 202 can contain a number N of images (where N is any integer greater than or equal to one). An initial text description 308(n) can be assigned to an area (e.g., as a mask, bounding box, etc.) 306(n) within an image 304(n).
[0041] Furthermore, the multitude of dictionary text descriptions 310(m=1...M) can be provided or be. Here, "M" can be any integer greater than or equal to two. According to various aspects, M can be greater than or equal to 100, e.g., greater than or equal to 1000, e.g., greater than or equal to 10000.
[0042] In 204, for each dictionary text description 310(m) of the multitude of dictionary text descriptions 310(m=1...M), a corresponding coded dictionary text description 316(m) can be generated by entering the dictionary text description 310(m) into the text encoder 312. In this way, the multitude of coded dictionary text descriptions 316(m=1...M) can be generated intuitively.
[0043] In 206A, the initial text description 308(n) can be entered into the text encoder 312 to determine the coded initial text description 318(n).
[0044] In 206B, for each coded dictionary text description 316(m) of the plurality of coded dictionary text descriptions 316(m=1...M), a similarity (e.g., represented by a similarity value) between the coded dictionary text description 316(m) and the coded initial text description 318(n) can be determined according to a first (semantic) similarity measure. The first (semantic) similarity measure can, for example, be a cosine similarity. It is understood that the cosine similarity is exemplary and any other similarity measure (e.g., a similarity metric, e.g., a distance metric) can be used, provided it corresponds to the distance measure used by the text encoder. From the M coded dictionary text descriptions 316(m=1...M) a predefined number K of one or more coded dictionary text descriptions 320(k=1...K) can then be selected which are most similar to the coded initial text description 318(n).“K” can be an integer greater than or equal to one. The one or more dictionary text descriptions 322(k=1...K) can correspond to the dictionary text descriptions from the multitude of dictionary text descriptions 310(m=1...M) from which the one or more coded dictionary text descriptions 320(k=1...K) were determined.
[0045] Selecting the dictionary text descriptions that most closely resemble K can reduce the computational effort required for subsequent object recognition.
[0046] In 206C, a recognition rate for at least the object shown in the area 306(n) can then be determined for the initial text description 308(n) and for each dictionary text description 322(k) of the one or more dictionary text descriptions 322(k=1...K). In some aspects, a respective recognition rate for each image 304(n) of the plurality of images 304(n=1...N) can be determined for the initial text description 318(n) and for each dictionary text description 322(k) of the one or more dictionary text descriptions 322(k=1...K).
[0047] For the sake of illustration and easier understanding, various explanations will only refer to object recognition for image 304(n) belonging to the initial text description 308(n).
[0048] The initial text description 318(n) and each dictionary text description 322(k) of the one or more dictionary text descriptions 322(k=1...K) can be considered as predefined text descriptions based on which object recognition is performed using the machine learning model 302. The output 324(l=1...K+1) can specify a respective area in the image 304(n) for each text description of the initial text description 318(n) and each dictionary text description 322(k) if the object represented by the text description in the image 304(n) has been recognized by the machine learning model 302.
[0049] To determine the recognition rate, the respective output 324(l) for each of the K+1 text descriptions can be compared with the range 306(n). For this purpose, a similarity between these can be determined according to a second similarity measure. For example, the range 306(n) could be a (segmentation) mask or a bounding box, and the range output by the machine learning model 302 could also be a mask or a bounding box. In this case, the second similarity measure could, for example, be an intersection set over a union of the ranges to be compared. It is understood that if the object is not recognized for one of the text descriptions, no range is output, and therefore no intersection set exists.
[0050] In 206D, the text description 326 with the greatest (determined) similarity to the set of given text descriptions can then be added. This added text description 326 can therefore be either the initial text description 308(n) or one of the dictionary text descriptions from the multitude of coded dictionary text descriptions 316(m=1...M).
[0051] In this way, a text description 326 can be determined automatically, for which the trained machine learning model 302 has a high recognition rate of the associated object.
[0052] By performing steps 206A to 206D on all initial text descriptions, the set of predefined text descriptions is generated. Determining this set of predefined text descriptions after training the machine learning model can be considered hyperparameter optimization.
[0053] Method 200 can be used to determine the set of predefined text descriptions for a variety of object recognition applications. For example, the robot device 102 can be configured for object recognition. Optionally, Method 200 can be used during the operation of the robot device 102. For example, the machine learning model 302 can receive feedback (e.g., from an operator of the robot device 102) for an object shown in an image, because it was previously recognized incorrectly. The feedback can include the initial text description of the object, and the area of the image showing the object can be marked (e.g., by means of a bounding box, a mask, etc.). Method 200 can then be performed on this image and the associated initial text description to add an (optimized) text description to the set of predefined text descriptions.Method 200 clearly enables online hyperparameter optimization.
[0054] A method for controlling a robot (e.g., the robot device 102) may involve capturing an image (e.g., using one or more imaging sensors described herein) showing one or more objects (in the environment of the robot device 102). The method for controlling the robot may involve inputting the image into the machine learning model 302 with a set of predefined text descriptions as an open vocabulary to recognize the one or more objects, and may then involve controlling the robot taking into account the recognized one or more objects (e.g., navigating the robot device 102 in its (dynamic) environment). The object recognition may, for example, be semantic and / or panoptic.
[0055] With reference to Fig.4. The robot device 102 can use the machine learning model 302 with the set of predefined text descriptions as an open vocabulary to perform object recognition in the image 402 captured at a specific time, in order to generate a segmentation image 404 (e.g., semantic or panoptic) based on the object recognition. The machine learning model 302 can be used for semantic object recognition. This can be done for multiple images that are captured while the robot device 102 navigates in their environment. Furthermore, the robot device 102 (while navigating its environment) can generate a map 406 of the environment using simultaneous localization and mapping (SLAM). The map 406 can enable the robot device 102 to navigate in its environment.
[0056] Depending on various aspects, a semantic map 408 of the environment can then be generated from one or more segmentation images 404 and the map 406 of the environment (by integrating a result of the semantic object recognition into the map 406 of the environment). The robot device 102 can then be controlled using this semantic map 408 of the environment. In this context, the use of an open-vocabulary machine learning model generally allows for the addition of new object classes (during operation), and the method 200 enables the optimization of the text description for such a new object class. This eliminates the need to replace the semantic object recognition model; instead, it can be adapted during operation.
Claims
[1] Computer-implemented method (200) for generating a set of predefined text descriptions for an open-vocabulary machine learning model (302) trained for object recognition such that for each predefined text description, an area in an image input into the machine learning model (302) is output if the area shows an object represented by the predefined text description, comprising the method (200): • Providing (202) a plurality of images (304) and a plurality of initial text descriptions (308), each initial text description being assigned to an area (306) in an associated image of the plurality of images (304) and indicating what is shown in the area (306); • Determine (204) a multitude of coded dictionary text descriptions (316) by generating a coded dictionary text description for each dictionary text description of a multitude of dictionary text descriptions using a text encoder (312) of the machine learning model (302); • for each initial text description of the multitude of initial text descriptions: ◯ Determining (206A) a coded initial text description (318) using the text encoder (312), ◯ Select (206B) one or more coded dictionary text descriptions (320) from the multitude of coded dictionary text descriptions which are most similar to the coded initial text description (318) according to a first similarity measure, ◯ for each text description of the initial text description and each dictionary text description belonging to one or more coded dictionary text descriptions, as a given text description, input (206C) at least the image belonging to the initial text description into the machine learning model (302) and determine a similarity between an output (324) of the machine learning model (302) and the area (306) associated with the initial text description according to a second similarity measure, and o Adding (206D) the text description (326) with the greatest similarity according to the second similarity measure to the set of given text descriptions. [2] Method (200) according to claim 1, wherein the area determined in the input image for the specified text description by means of the machine learning model (302) and the area assigned to the initial text description are represented by means of a mask and / or a bounding frame. [3] Method (200) according to claim 1 or 2, wherein the first similarity measure has a cosine similarity; and / or wherein the second similarity measure has an overlap set over union set. [4] Method (200) according to one of claims 1 to 3, wherein the input of at least the image belonging to the initial text description into the machine learning model (302) comprises that each image of the plurality of images is entered into the machine learning model (302) and the similarity across all images is determined. [5] Method for controlling a robot device (102), comprising the method: • while the robot device (102) navigates in its environment, generating a map (406) of the environment by simultaneous localization and mapping and capturing images (402) representing the environment; • Performing semantic object recognition (404) for each captured image using the machine learning model (302) with the set of predefined text descriptions generated according to any one of claims 1 to 4; • Generating a semantic map (408) of the environment by integrating a result of semantic object recognition (404) into the map (406) of the environment; and • Controlling the robot device (102) using the semantic map (408) of the environment. [6] Data processing unit configured to perform the method (200) according to any one of claims 1 to 4. [7] Control device (114) comprising one or more processors configured to perform the method (200) according to claim 5. [8] Robot device comprising: • the control device (114) according to claim 7; and • at least one imaging sensor (122) configured to capture images of the environment of the robot device (102). [9] Computer program with instructions which, when executed by a processor, cause the processor to perform the method (200) according to any one of claims 1 to 5. [10] Computer-readable medium storing instructions which, when executed by a processor, cause the processor to perform the method (200) according to any one of claims 1 to 5.