Method, Apparatus, and System for Automating Machine Operation - Patent application
Patent Information
- Application Number
- JP2024537962
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2021-12-21
- Filing Date
- 2022-12-21
- Publication Date
- 2026-01-06
AI Technical Summary
The automation of machine operation is difficult and complicated due to the need for manual interaction with human-designed interfaces, which requires complex automation systems.
A robot system is used to automate machine operation by employing a robot arm with actuators, a visual source, and a processor that extracts information from images to identify and interact with machine interfaces, performing physical tasks through a two-way medium.
Enables efficient and automated interaction with machine interfaces, allowing robots to perform tasks typically requiring human input, enhancing operational efficiency and reducing manual intervention.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
[Technical field]
[0001] (CROSS REFERENCE TO RELATED APPLICATIONS) This application claims the benefit of and priority to U.S. Provisional Application No. 63 / 292,316, filed December 21, 2021, the entire teachings of which are incorporated herein by reference.
[0002] The present application relates generally to automating machine operations, and more specifically, to methods, systems, and apparatus for automating machine operations using robotic systems. [Background technology]
[0003] Automating the operation of a machine, which is typically operated manually, can be a difficult and complex process because it requires the use of an automated system to interact with an interface designed for use by a human operator. Summary of the Invention [Means for solving the problem]
[0004] A robotic system for operating a machine having an interface with an interactive medium for use in performing a physical task using the machine can include an arm, an actuator coupled to the arm for indicating changes in the interactive medium when the actuator interfaces with the interactive medium, a visual source for providing an image of the machine within a field of view, a processor coupled to the arm and the visual source, and a memory coupled to the processor and configured to store a computer-executable program. The computer-executable program can include instructions, upon execution by the processor, to extract information from the image indicative of a metric for identifying an interface within the field of view, indicative of changes in the interactive medium, extract information indicative of a characteristic of the interactive medium based on a signal from the actuator, operate the arm, actuate the interactive medium, and perform the physical task using the machine.
[0005] The arm can be a robotic arm having a proximal end and a distal end.
[0006] An actuator can be coupled to the arm. For example, the actuator can be coupled to a distal end of the arm. The actuator can include a pressure sensor coupled to its distal end that indicates a change by sensing a change in resistance when the distal end of the actuator interfaces with a bidirectional medium. Although described with respect to a pressure sensor, any suitable sensor available in the art can be used in conjunction with the embodiments disclosed herein (e.g., an optical sensor). The pressure sensor and / or the optical sensor can be a digital sensor. Additionally, the pressure sensor can include a piezoresistive sensor.
[0007] The processor can extract information indicative of a characteristic of the interactive medium based on a signal from the pressure sensor indicative of a change in resistance exhibited by the pressure sensor.
[0008] The visual source can include a camera, such as an infrared camera. For example, the camera can be coupled to an arm. The camera can be configured to rotate around the machine and capture images of the machine. Additionally, the visual source can include a source that stores visual information regarding the field of view. Generally, any visual source available in the art can be used. For example, the visual source can include at least one of an ultrasonic camera, an ultrasonic sensor, and a LIDAR.
[0009] Further, the visual source can be coupled to the robotic system via a communication network. For example, the visual source can be configured to provide a video stream of the field of view. Alternatively, or in addition, the visual source can include a pre-captured image.
[0010] Further, the arm can be configured to move in two dimensions within a field of view proximate the machine. Alternatively, or in addition, the arm can be configured to move in three dimensions within a field of view proximate the machine. The field of view can be a three-dimensional field of view.
[0011] The actuator can be configured to move in a dimension that is additional to the dimension in which the arm can be configured to move. The arm can include one or more sections connected at one or more joints, where each section can be configured to translate relative to an adjacent section relative to its respective joint. The arm can further include one or more sections connected at one or more joints, where each section can be configured to rotate relative to an adjacent section relative to its respective joint. Additionally, the arm can be configured to rotate within a near field of view of the machine.
[0012] The robotic system may include one or more additional cameras configured to acquire two or more images of the machine from two or more fields of view. The two or more images may be configured to form a stream of images. The stream of images may provide a three-dimensional image of the machine.
[0013] The memory can be configured to store information for determining a metric for identifying an interface. The information can include predefined data regarding dimensions of a machine within the field of view. The information can include predefined data regarding dimensions of an interface within the field of view. The metric can include dimensions of the machine within the field of view. For example, the metric can include a location of the machine within the field of view. Further, the metric can include an orientation of the machine within the field of view. Alternatively or in addition, the metric can include a location of the interface within the field of view. Further, the metric can include a boundary of the interface within the field of view.
[0014] The processor may further identify dimensions of the interactive media. The characteristics may include dimensions of the interactive media. For example, the characteristics may include a color of the interactive media. Alternatively or in addition, the characteristics may include a texture of the interactive media. Further, the characteristics may include visual characteristics of the interactive media. Further, the characteristics may include text appearing on the interactive media.
[0015] The system may store the metric in memory. For example, the system may store the characteristic in memory.
[0016] Further, the processor may be coupled to a communications network and configured to receive instructions from a remote entity via the communications network.
[0017] The system can adjust the camera and acquire additional images in the additional field of view. Further, the system can extract information including metrics for identifying an interface in the field of view, analyze the extracted information, determine whether additional information is needed to identify the interface, and if the additional information is needed, adjust the camera and acquire additional images in the additional field of view. Adjusting the camera can include adjusting an angle of the camera.
[0018] The robotic system may further include instructions that, depending on execution, may adjust an intensity of light emitted by the light emitter prior to acquiring the additional images. The system may extract information including a metric for identifying an interface in the field of view, analyze the extracted information, determine whether additional information is needed to identify the interface, and acquire the additional images at different times if additional information is needed. Additionally, the system may analyze the acquired images and score each image based on the amount of information available in the image for identifying the interface or interactive media in the field of view. Additionally, the system may score the images based on the probability of correct detection of a characteristic of the interactive media, respectively. The system may also exclude images having a score lower than a predetermined score from being used to identify the interface or interactive media in the field of view. Additionally, the system may employ images having a score higher than a predetermined score to identify the interface or interactive media in the field of view.
[0019] The memory can be configured to store previously verified data for at least one of the metrics for identifying the interface or interactive medium. The previously verified data can include data provided by a human operator of the machine. Additionally or alternatively, the previously verified data can include data provided from a remote entity via a communication network. The previously verified data can include data obtained from at least one of an original instruction manual for the machine and an online resource. The system can analyze at least one of the original instruction manual or the online resource and extract information for identifying the interface or interactive medium.
[0020] The previously verified data may include a glyph dictionary. The glyph dictionary may include a collection of images. The system may be configured to update the glyph dictionary using metrics obtained from the images and pressure sensor to identify characteristics of the interface and interactive media. The instructions may be configured to update the glyph dictionary using images that have a score higher than a predetermined score. The instructions may score each image by comparing the image to previously verified data.
[0021] The memory can be configured to store a list of physical tasks to perform using the interactive medium. The memory can be configured to store a ranking for each physical task from the list of physical tasks. The arm can be instructed to perform the physical tasks based on a ranking assigned to each task by the processor. Physical tasks having a higher ranking can be performed before physical tasks with a lower ranking. The physical task with the highest ranking can be a preferred task to perform using the machine. The robotic system further includes an actuation arm connected to a distal end of the arm and capable of performing the physical tasks via the interactive medium.
[0022] The memory can be configured to store a library of physical tasks for execution using the machine and through interaction with the interactive media. The processor can select a physical task for execution from the library of physical tasks. The instructions executed by the processor can further include defining the list of physical tasks based on data obtained from a human operator who previously operated the machine. The instructions executed by the processor can further include defining the list of physical tasks based on data obtained from at least one of original instructions for the machine and an online resource.
[0023] The processor can be configured to execute instructions in response to a voice command. The physical tasks can include at least one of lifting a heavy object, opening a door, lowering a heavy object, pushing a heavy object, and sliding a heavy object. The actuator can be configured to record physical tasks performed by the machine in response to an interaction with the interactive medium. The actuator can be configured to perform all actions available through interaction with the interactive medium and record physical tasks performed in response to the actions. The physical tasks can be recorded in a database. The database can be a glyph dictionary.
[0024] The robotic system can include a user interface connected to the processor for use in controlling the robotic system. The user interface can include a graphical interface for initiating execution of a computer-executable program. The processor can be configured to use an image processing scheme to extract information indicative of a metric for identifying an interface within a field of view.
[0025] The processor may be configured to receive previously verified images of interfaces within the field of view, use an image processing scheme to extract information indicative of a metric for identifying interfaces within the field of view, and employ a deep learning framework to identify the interfaces. The deep learning framework may include supervised deep learning. Alternatively or additionally, the deep learning framework may include unsupervised deep learning. The deep learning framework may include reinforcement deep learning. The robotic system may further include scoring images from a visual source based on an amount of information available within the image for identification of a machine or interactive interface, and images having a score higher than a predetermined value within the deep learning framework may be used.
[0026] The robotic system may further include one or more sensors configured to measure at least one property of the machine, the interactive element, and the field of view. The robotic system may further include an optical sensor configured to measure light intensity within the field of view, and the light intensity obtained from the optical sensor may be used to extract a metric for identifying the interface. The robotic system may further include an optical sensor configured to measure light intensity within the field of view, and the system may adjust the image based on the light intensity.
[0027] The robotic system may further include a sensor configured to measure a characteristic of an area surrounding the machine and adjust the extracted information based on the measured characteristic. The measured characteristic may be at least one of glare, audible noise, humidity, and tactility.
[0028] The instructions, upon execution, may compare the light intensity to a predetermined threshold and cause the visual source to capture additional images of the field of view if the light intensity is less than the predetermined threshold. The additional images may be captured at different times. The robotic system may further include a light emitter coupled to the light sensor, the light sensor responsive to light emitted from the light emitter to measure light intensity reflected within the field of view. The light emitter may be configured to emit light in response to the light intensity being less than the predetermined threshold.
[0029] A robotic system for operating a machine having an interactive medium for use in performing a physical task using the machine and having an interface with the interactive medium for operating the machine can include an arm configured to move in a first dimension, an actuator coupled to the arm and configured to move in a second dimension, a vision source configured to acquire an image of the machine, a processor coupled to the arm and the actuator, and a memory coupled to the processor and configured to store a computer executable program. The executable program can include instructions that, upon execution by the processor, extract information from the image indicative of a metric for identifying an interface in a field of view based on the metric, move the arm in the first dimension proximate the interface and move the actuator in the second dimension proximate the interface, record observed changes by the actuator in the second dimension, extract information indicative of a characteristic of the interactive medium based on the observed changes, operate the interactive element, and actuate the interactive medium to perform the physical task using the machine. [Brief description of the drawings]
[0030] [Figure 1A] FIG. 1A is a block diagram of a robotic system according to some embodiments disclosed herein. [Figure 1B]FIG. 1B is a schematic diagram of a system for testing an embedded system of a device according to some embodiments disclosed herein. [Figure 1C] FIG. 1C is a block diagram of an example of a viewable user interface of a device under test according to certain embodiments disclosed herein. [Figure 1D] FIG. 1D is a block diagram of an example of a viewable user interface of a device under test according to certain embodiments disclosed herein. [Figure 1E] FIG. 1E is a block diagram of an example state of a viewable user interface of a device under test, according to certain embodiments disclosed herein. [Figure 1F] FIG. 1F is a block diagram of an example state of a viewable user interface of a device under test, according to certain embodiments disclosed herein. [Figure 1G] FIG. 1G is a flow diagram of a method for constructing a descriptor according to some embodiments disclosed herein. [Figure 1H] FIG. 1H illustrates exemplary steps of a method for processing an image according to some embodiments disclosed herein. [Figure 1I] FIG. 1I is an example of a procedure for processing an image according to some embodiments disclosed herein. [Figure 1J] FIG. 1J is an example of a procedure for processing an image according to some embodiments disclosed herein. [Figure 1K] FIG. 1K is an exemplary graph of a comparison of time of evaluation, according to certain embodiments disclosed herein. [Diagram 2] FIG. 2 is a block diagram of electronic circuitry according to some embodiments disclosed herein. [Diagram 3] FIG. 3 is a flow diagram of a procedure for operating a robotic system according to some embodiments disclosed herein. [Figure 4]FIG. 4 is a flow diagram of a procedure for calibrating a robotic system according to some embodiments disclosed herein. [Diagram 5] FIG. 5 is a flow diagram for preparing a dictionary of interactive media elements, their associated tasks, and / or associated physical outcomes according to some embodiments disclosed herein. [Figure 6] FIG. 6 is a flow diagram of a procedure for identifying preferred physical tasks to perform using a machine, according to some embodiments disclosed herein. [Figure 7] FIG. 7 is a flow diagram of a procedure for carrying out physical tasks that are preferred for implementation using a machine in a cyclical manner, according to some embodiments disclosed herein. [Figure 8] FIG. 8 is a block diagram of an example of a robotic system according to some embodiments disclosed herein. [Figure 9] FIG. 9 is a block diagram of an example of a robotic system according to one embodiment disclosed herein. [Figure 10] FIG. 10 is a block diagram of an example of a robotic system according to one embodiment disclosed herein. [Figure 11] FIG. 11 is a block diagram of an example of a robotic system according to one embodiment disclosed herein. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0031] Detailed Description The present disclosure relates to methods, systems, and apparatus for operating machines using robotic systems. The disclosed methods, systems, and apparatus can be used to perform and automate various functions and tasks using machines. For example, the disclosed methods, systems, and apparatus can be used to configure an arm (e.g., a robotic arm) to perform tasks and functions that are typically performed by a human operator through direct physical interaction with the machine.
[0032] 1A is a high-level block diagram of a robotic system 100 according to aspects disclosed herein. The robotic system 100 can be used to perform a physical task using a machine 150. The physical task can generally be any task that can be performed by a human operator via interaction (e.g., physical interaction) with the machine 150 (e.g., pressing a button). The machine 150 can also generally be any suitable machine available in the art. For example, the machine 150 can be a vending machine, a copy machine, a keyboard, a tablet, a phone, a kiosk, a manufacturing machine, a health care device, a 3D printer, etc.
[0033] The machine 150 may include an interactive medium 152 through which the machine 150 is operated. For example, the machine 150 may include a keypad or an LCD screen through which operation of the machine 150 is performed.
[0034] The interactive medium 152 can generally be any suitable interactive medium known / available in the art. For example, the interactive medium 152 can be a keypad, a light switch, a button, a liquid crystal display (LCD) with a manually pressed button, etc. The interactive medium 152 can be a smart and / or touch-sensitive interactive medium, such as an interactive medium with touch-sensitive elements (e.g., a touch screen interface with digital buttons). Additionally or alternatively, the interactive medium 152 can include a manual interface with buttons that are manually actuated (e.g., a button with physical feedback, a wheel, and a hold switch).
[0035] The robotic system 100 may further include an arm 170 for performing the operations of the system 100. The arm 170 may generally be any suitable mechanical element capable of performing the functions disclosed herein. For example, the arm 170 may be a robotic arm and / or a bracket. Furthermore, the arm 170 may be secured, fixed, and / or coupled to the system 100 using any suitable technique known in the art. For example, the arm 170 may be secured / fixed to a frame 172 and coupled to the system via the frame 172.
[0036] In some implementations, arm 170 can include a proximal end 173 attached to frame 172 and a distal end 175 configured to physically interface with machine 150 (e.g., via interactive medium 152). Arm 170 can also be configured to move in more than one dimension and physically interface with machine 150 and interactive medium 152.
[0037] The arm 170 can be movable and / or rotatable along one or more directions / dimensions. Specifically, the arm 170 can be configured to move / translate along one or more directions and rotate along one or more dimensions. Additionally, the arm can include one or more sections 170A, 170B configured to move (e.g., about joint 170J) in one or more directions and / or rotate (about joint 170J) along one or more dimensions. The movement and / or rotation of the various sections 170A, 170B of the arm 170 can provide the arm 170 with the ability to translate and / or rotate across various dimensions (3, 4, 5, 6, etc.).
[0038] The robotic system 100 may further include an actuator 176 configured to interface with the interactive medium 152. The actuator 176 may generally be any element capable of interacting with the interactive medium 152. For example, the actuator 176 may comprise a protrusion, a sensor (e.g., a pressure sensor), etc. In some implementations, the actuator 176 may be coupled to a distal end 175 of an arm (e.g., a robotic arm) and configured to perform a function of the robotic system (e.g., operate a machine). Alternatively or in addition, the actuator 176 may be a protrusion extending from a portion of the arm 170 and may interact with the interactive medium 150.
[0039] The actuator 176 can be configured to translate / move and / or rotate about at least one additional direction and / or dimension than the directions and dimensions through which the arm 170 can translate / rotate. This additional direction / dimension of movement can provide the arm 170 and the actuator 176 with movement / rotation along the added dimension / direction. For example, the arm 170 can be configured to move in five dimensions (e.g., X, Y, Z, 45 degrees left along the Z-direction, and 25 degrees right along the X-direction), and the actuator 176 can be configured to rotate about its attachment to the arm 170, thereby providing at least one additional direction of movement / rotation to the system 100.
[0040] The actuator 176 can be configured to detect changes in the interactive medium 152 in response to interfacing with the interactive medium 152. For example, the actuator 176 can physically contact the interactive medium and / or interact with the interactive medium through any suitable means available in the art (e.g., by directing a light beam, such as a laser, at the interactive medium). The actuator 176 can be configured to detect changes in the location of a button 199 in the interactive medium 152. For example, in implementations in which the interactive medium 152 comprises a physical push button 199, the actuator 176 can be configured to detect changes in the interactive medium 152 by observing that the interactive medium 152 is displaced (e.g., along the Y-direction) when the actuator 176 contacts the push button 199. This allows the actuator 176 to detect the position of the push button 199 by contacting various spots on the surface 152′ of the interactive medium 152 and detecting the position of the push button 199 in response to observing changes in the interactive medium 152 (e.g., in displacement along the Y-direction).
[0041] In addition to detecting physical displacements within the interactive medium 152, the actuator 176 may further comprise a sensor (e.g., a pressure sensor) configured to detect the position of various elements of the interactive medium 152. For example, the sensor 174 may detect a change in a characteristic of the screen 152'' (e.g., an LCD screen) of the interactive medium 152 in response to contacting the touch-sensitive button 199. For example, the sensor 174 may be a pressure sensor that detects a change in resistance of the LCD screen 152' of the interactive medium 152 in response to contacting the touch-sensitive button 199. Generally, any suitable sensor known in the art may be used in conjunction with the embodiments disclosed herein. For example, the sensor 174 may be a piezoelectric sensor, a pressure sensor, a piezoelectric pressure sensor, a piezoelectric mechanical sensor, a piezoelectric mechanical pressure sensor, an optical sensor, a digital sensor, an audio sensor, or the like. Further, in addition to or instead of measuring changes in resistance, the sensor 174 may detect changes in other characteristics of the interactive medium 152, such as changes in dimension, color, texture, the presence of digital text, and may detect the button 199 by distinguishing the touch-sensitive display over other surfaces of the interactive medium 152 or machine 150.
[0042] As shown in FIG. 1A , the arm 170, the actuator 176, and / or the sensor 174 can be coupled to a distal end 175 of the arm 170 and configured to move in the vicinity of the arm 170. For example, the arm 170 can be movable along at least one dimension about the bidirectional medium 152, and the actuator 176 can be movable along at least one other dimension. Additionally or alternatively, the actuator 176 can also be configured to move in at least one dimension other than the dimension in which the arm 170 moves. For example, the arm 170 can also be configured to move along X and Y (horizontal and vertical) in the vicinity of the machine 150, while the actuator 176 moves along the Z dimension across the bidirectional medium 152. Additionally or alternatively, the arm 170 can also be configured to move in three dimensions (x, y, z) in the vicinity of the machine 150, while the actuator 176 rotates across the bidirectional medium 152.
[0043] 1A , the robotic system 100 can further include a vision source 160 configured to provide an image of the machine 150. The vision source 160 can be any suitable source capable of providing an image of the machine 150, such as a memory that stores previously acquired images of the machine 150, and / or a camera that captures a live image of the machine 150 while the machine 150 is within the camera's field of view 101. For example, the vision source 160 can comprise a suitable static camera that captures an image of the machine 150 within its field of view, a 360° camera that rotates around the field of view and captures an image of the field of view, a three-dimensional camera that captures a three-dimensional image of the field of view, a dynamic video camera configured to provide a video stream of the field of view. Alternatively or in addition, the vision source 160 can comprise an infrared camera configured to capture an image of the field of view under low ambient lighting conditions using infrared radiation and thermography, and / or a dynamic video camera with a lens for capturing an image or video of the field of view. In some implementations, visual source 160 may include two or more cameras providing two or more images from two or more fields of view (e.g., as a stream of images).
[0044] The images can be digital and / or analog images that are converted to a digital format and provided to the system 100 by the visual source 160. As described in more detail below, the images from the visual source 160 are automatically transferred to the electronics network 110, which has a processor 120 that processes the images and extracts relevant information from the images about the machine 150 and the interactive interface 152. A variety of image processing schemes can be used.
[0045] Visual source 160 may be connected to system 100 via any suitable connection. For example, visual source 160 may be coupled to system 100 via communications network 148 and configured to transmit one or more images via network 148 to processor 120 for processing. For example, visual source 160 may transmit two or more images of machine 150 within a field of view to processor 120 for use in preparing a three-dimensional image of machine 150.
[0046] The system 100 may further include any suitable additional elements that may facilitate the capture of images from the machine 150, the recognition of various elements of the interactive medium 152 by the actuator 176, and / or the mitigation of the effects of environmental factors during the calibration and / or operation of the system 100. For example, the system 100 may include a light emitter 162 configured to adjust the light intensity in the vicinity of the machine 150 prior to and / or during the capture of an image by the visual source 160. The light emitter 162 may vary the intensity of the delivered light and / or control the angle at which the light is delivered before / while the image is captured by the visual source 160 to mitigate the adverse effects of environmental factors caused by the light intensity. Similarly, the system 100 may include an audio sensor that may detect audible ambient noise and inform the processor 120 in the digital circuitry 110 of the system of the presence of the ambient noise. In response, the processor 120 may initiate one or more actions that reduce and mitigate the audible ambient noise. For example, the processor 120 may adjust the volume of the audible instructions provided to the system 100 to mitigate environmental noise.
[0047] In general, any suitable noise sensor and / or noise reduction / mitigation scheme can be used in conjunction with the embodiments disclosed herein. For example, noise reduction techniques can include, but are not limited to, adjusting the intensity of the light at which the image is captured, adjusting the angle of the light delivered for image capture, in the case of an audible command, adjusting the volume, which may affect the physical feedback of the button, adjusting humidity.
[0048] As described, the system 100 may further include a light sensor 163 that measures the light intensity of the area surrounding the machine 150 and automatically transfers the measured value to the processor 120 and electronic circuitry 110. The processor 120 compares the measured light intensity to a predefined threshold (e.g., stored in the database 140) and determines whether the measured light intensity is sufficient / suitable for capturing an image. If the light intensity value is lower / higher than the predefined value, the processor 120 may instruct the light emitter 162 to emit additional / less light and / or instruct the visual source 160 to capture additional images with adjusted lighting conditions and / or at a different time when the lighting conditions are expected to improve. Additionally or alternatively, the processor 120 may cause the visual source 160 to capture one or more additional images of the machine 150 and / or the interactive medium 152. For example, the visual source 160 may obtain additional images of the interactive medium 152 from different perspectives (e.g., by rotating a camera) or by adjusting the intensity of the light (e.g., using a light emitter).
[0049] FIG. 1B is a block diagram of a system for testing an embedded system 6 of a device 1 according to some aspects disclosed herein. The system for testing an embedded system 6 of a device is shown on FIG. 1B, comprising a device under test 1, a test robot 2, and a central control unit 3. The device under test 1 comprises at least an observable user interface 61 (OUI) (FIGS. 1C and 1D) adapted to indicate a state of the device under test 1. As shown in FIGS. 1B-1F, the OUI 61 comprises at least one button 5 and / or visual indicator 63. The button 5 is a physical device that is used for user-device interaction and thus may be subjected to pressing, rotating, moving, etc., which is used to send commands or requests to the embedded system 6 of the device under test 1. The visual indicator 63 may be an LED or another light source and is used to identify the state of the device under test 1. The state of the device under test 1 means whether it is turned off or on, the task being realized, the option being selected, etc. Settings and tasks performed by the device under test 1 are enabled and activated by buttons 5 and with the aid of visual indicators 63. The device under test 1 may further comprise a display 4, which is used to provide visual feedback to the user, indicate the time or other relevant information, etc.
[0050] The device under test 1 may comprise a display 4, preferably embodied as a touch screen, at least one button 5, and an embedded system 6. The display 4 is adapted to show screens of a graphical user interface 62 (GUI) of the embedded system 6. The screens of the GUI 62 may comprise an initial screen, a loading screen, a shutdown screen, and other screens showing various functionalities such as possible settings of the device under test 1, a list of selectable actions to be performed by the device under test 1, an error screen, a menu screen, etc. At least one screen has at least one action element 7, which is preferably embodied as a button on the touch screen of the device, i.e. it may be, for example, a menu screen, which has several action elements 7, each of which may represent a selectable option to form a menu screen. There may also be screens without action elements 7, such as a loading screen, which generally does not require any user action. The action element 7 itself may not have a physical form, but rather may be implemented as an icon on the touch screen of the display 4. Both the buttons 5 and the action elements 7 are used to interact with the embedded system 6 and are assigned actions, which may be different for the buttons 5 and the action elements 7 and may vary from screen to screen. The actions are performed in response to interacting with the buttons 5 or the action elements 7, the interaction being performed by pressing the button 5 or touching (pressing) the action element 7. The action may then result in a change of the screen shown, or selecting a task from a list of tasks to be performed, cancelling an action, changing the settings of the device under test 1, logging in or out of the device, etc. At least one screen has a connection to at least one other screen. Screen connection means that by performing an action on one screen, the display then shows another screen.Such connections may, for example, be presented by a screen with a list of possible tasks to be performed. Each task may be assigned an action element 7, and upon pressing the action element 7 the user is transferred to another screen showing further details about the selected task.
[0051] The system may include a screen showing a list of possible configurations of the device under test 1, each configuration option being assigned an action element 7, and upon pressing the action element 7 or button 5, the user is transferred to another screen associated with the selected configuration. Some screens may not have a connection to another screen, such as a loading screen, which does not require any user input and serves only as a transition screen. It is necessary for the device under test 1 to contain at least a processor 8 and a computational memory 9 in order to smoothly start the embedded system software. For example, the device under test 1 is communicatively coupled to a central control unit 3. The communication coupling can be realized either physically by a cable, such as an Ethernet cable or a USB cable, or as a contactless coupling, via the Internet, Bluetooth, or by connecting it to a network. However, the communication coupling is not necessary and may not be implemented. In that case, the device under test 1 may act independently of the central control unit 3.
[0052] The test robot 2 (FIG. 1B) may comprise at least an effector 10, a camera 11, a processor 12 with an image processing unit 13, and a memory unit 14. The effector 10 may be adapted to move along at least two, preferably three orthogonal axes. For example, the effector 10 is attached to a first robot arm 15, which guides the movement of the effector 10. The effector 10 is used to simulate a physical contact of a user with the display 4, the buttons 5, and the action elements 7. If the display 4 is embodied as a touch screen, the effector 10 may be adapted to interact with the display 4 (e.g., by being made from a material that can interact with a touch screen). Furthermore, the camera 11 is installed on the test robot 2 in such a way that it faces towards the display 4. Thus, the display 4 and the buttons 5 are within the field of view of the camera 11, and the current screen shown on the display 4 is clearly captured by the camera 11. In some implementations, the camera 11 is not mounted on the test robot 2, but is mounted in such a way that the display 4 and the buttons 5 are within the field of view of the camera 11. For example, the camera 11 can be mounted, for example, on a tripod. The processor 12 of the test robot 2 is adapted to control the actions of the test robot 2 via software installed thereon. The actions of the test robot 2 that the processor 12 controls include controlling the movement of the first robot arm 15 and the effector 10, communicating with the central control unit 3, and operating the camera 11. The test robot 2 is communicatively coupled to the central control unit 3. The communication coupling can be realized either physically by a cable, such as an Ethernet cable or a USB cable, or as a contactless coupling, via the Internet, Bluetooth, or by connecting it to a network.
[0053] As shown in FIG. 1B, the test robot 2 further comprises a second robot arm 16 equipped with a communication tag 17, and the device under test 1 further comprises a communication receiver 18. The second robot arm 16 is adapted to place the communication tag 17 in the vicinity of the communication receiver 18 when commanded by the processor 12. The communication tag 17 is preferably embodied as an RFID tag in a chip or plastic card. In other embodiments, any communication tag 17 based on RFID or NFC communication protocols may be used. Other forms of wireless communication or connection standards such as BlueTooth, infrared, Wifi, LPWAN, 3G, 4G, or 5G are also possible. The communication tag 17 includes an RFID chip, which is activated when placed in proximity to the communication receiver 18, which continuously emits a signal in the radio frequency spectrum. In response to activating the communication tag 17, an action in the device under test 1 can be performed. A task to be performed by the device under test 1, such as for example printing a document, is waiting for confirmation, which occurs by user verification. Each user may have a communication tag 17, assigned to them in the form of a card. Verification is then performed by placing the communication tag 17 in the vicinity of a communication receiver 18.
[0054] The test robot 2 may further comprise a sensor 19 adapted to detect the action performed. The sensor 19 may be a camera, a photocell, a motion sensor, a noise sensor, etc. The sensor 19 may be used to detect the action performed if its result cannot be captured on the display 4. For example, when the device under test 1 is a printer, the action may be to print a paper. A prompt for the user to print a document may be shown on the display 4, which may be confirmed either by pressing the action element 7 or by placing the communication tag 17 near the communication receiver 18. After confirmation, the document is printed. However, this action cannot be captured by the camera 11 since it is mainly focused on the display 4 of the device under test 1 (e.g., the printer). The printing itself may be detected by the sensor 19, for example, if the sensor 19 is a camera, the printed paper may be captured by the camera, or the movement may be detected by a movement sensor, such as a photocell, within a certain area where the document emerges from the printer.
[0055] The central control unit 3 may be in a personal computer with at least a monitor and a PC mouse, and therefore comprises a processor 20 and a memory unit 21. The central control unit 3 is used to control the overall process and ensure communication between itself and the test robot 2, as will be further described herein. The device under test 1 may be separate and may run independently on the central control unit 3. However, the device under test 1 may be communicatively coupled to the central control unit 3 if the device under test 1 allows such a connection.
[0056] Furthermore, the device under test 1 can be any suitable device (e.g., a printer). Furthermore, the device can be any device with an embedded system 6 implemented, either with a GUI 62 or with an OUI 61 in other forms. The OUI 61 may not have a touch screen or even a display 4, but may have another visual indicator 63, such as an LED or other form of light indication. The following list of devices does not limit the scope of protection given by the claims. Devices with a touch screen may be printers, tablets, smartphones, control panels for industrial applications, cash registers, information kiosks, ATMs, terminals, smart home appliances, etc. Devices with only a display 4, not embodied as a touch screen, may be microwave ovens, radios, calculators, any of the devices named in the previous list, etc.
[0057] Descriptor Construction For the characterization of an image, numerical values are used. These numerical values, stored in a vector format, are called descriptors and can provide information about the color, texture, shape, motion, and / or location of the image. To describe the image in more detail, more specific descriptors need to be used. Descriptors can be of various shapes and lengths and provide different numerical or verbal descriptions of the image being processed. For example, a descriptor providing information about the color distribution in an image can have different numerical values, lengths, and shapes than a descriptor providing information about the shape of the image, so that the color distribution can be described by one alphanumeric value with information about the average color of the image, such as #FF0000, #008000, #C0C0C0, etc., and the shape can be described by a vector with two numerical values for the length and width of the image. Another descriptor can be used to describe the average color and shape of the image. This descriptor can be represented by a vector with three values, one for the color and two for the size. This method of constructing the descriptor 70 is illustrated in FIG. 1G and described in more detail below.
[0058] To build the descriptor 70, a photo 72 is acquired with an empty descriptor 71. An image 50 of the device under test 1 comprises a region of interest, such as a control panel 60. To extract the region of interest 73, the image 50 is cropped so that only the region of interest containing relevant information is shown in the cropped image 51, which is preferably the OUI 61. Cropping is used to remove any part of the image, which does not contain useful information and is therefore not relevant for further processing, making the image processing part of the method faster.
[0059] The descriptors can be used to describe images of the OUI 61 of various devices with displays. As the optical sensor 11 captures the images, a part of the device under test 1 on which the display 4 is placed can also be captured. The image of the device under test 1 itself can be removed from the picture without loss of important information, since it does not carry relevant information. To make the process of finding the area of interest easier, various marks 54 can be placed on the device under test 1 to highlight the location of the area of interest. An example of highlighting the area of interest in an image is shown in Figs. 1C-1F. In these figures, the front panel of the device under test 1 with the display 4 is shown. Typically, the display 4 has a rectangular shape. To highlight the location of the OUI 61 on the device under test 1, a set of marks 54 can be arranged, which are placed at the corners of the rectangle forming the OUI 61. By using these marks 54, it is easier for the image processing unit 13 to determine the location of the area of interest faster and with better accuracy.
[0060] Once the region of interest has been cropped from the original image 50, the image 50 is converted into a binary edge image 75. In this format, the pixels forming the image can achieve only two values, 0 and 1, marking either black or white pixels, respectively. In general, the image 50 can be converted into a grayscale image 74. Depending on the average intensity of the pixels, a threshold value of the pixel intensity is chosen. Pixels with values lower than the threshold are then converted into black pixels, while pixels with values higher than the threshold are converted into white pixels. The threshold value is not constant for all processed images, but rather varies with the overall quality of the image and the surrounding conditions. The binary edge image can be divided into rows and columns, the maximum value of the rows and columns being given by the size of the image in pixels, for example an image of format 640x480 can be divided into 640 rows and 480 columns, etc.
[0061] In some implementations, the image can be divided into a different number of rows and columns. However, with a decreasing number of divided rows and columns, the information value of the descriptor also decreases rapidly. Furthermore, two histograms can be generated to count the non-zero valued pixels in each row and column 76. The first histogram counts the number of non-zero valued pixels in each row, and the second histogram counts the number of non-zero valued pixels in each column. Both of these histograms are normalized 78, and after this step, two sets of numbers are available. The first set of numbers contains numbers that represent the normalized histogram of the non-zero valued pixels in each row, and the second set of numbers contains numbers that represent the normalized histogram of the non-zero valued pixels in each column. The normalized values are added to the empty descriptor 78. These two sets of numbers are the first two values of the constructed descriptor. Further, the cropped image 51 can be divided into equal sectors 79 using horizontal and vertical dividing lines, the non-zero valued pixels in each row and column in each sector are counted 80, the counted values are normalized 81 and added to a descriptor 82 until a desired iteration is reached 83, and the procedure is completed 84. The procedure can be repeated as detailed below as 85. The cropped image 51 can be split into two halves, left and right, using a vertical dividing line 52. The halves do not need to be of equal size. The left and right halves can be split into columns, with the total number of columns being the same as the number of columns in the unsplit image. A histogram can be constructed and normalized that represents the number of non-zero valued pixels in the columns of the first half. Additionally, a histogram can be constructed and normalized that represents the number of non-zero valued pixels in the columns in the second half. These procedures can be performed in any order.
[0062] The cropped image 51 can be split into two halves, top and bottom, using a horizontal split line 52. The halves do not need to be of the same size. The top and bottom halves can be split into rows, the total number of rows being the same as the number of rows in the unsplit image. Furthermore, a histogram representing the number of non-zero valued pixels in the rows of the first half can be constructed and normalized. A histogram representing the number of non-zero valued pixels in the rows in the second half is also constructed and normalized. These two steps are interchangeable, meaning it does not matter which histogram is generated first. Furthermore, it does not matter whether the image is first split using a vertical or horizontal line.
[0063] Depending on completing the procedure, four more sets of numbers become available. The first set of numbers contains numbers that represent the normalized histogram of non-zero valued pixels in the left half columns of the image, the second set of numbers contains numbers that represent the normalized histogram of non-zero valued pixels in the right half columns of the image, the third set of numbers contains numbers that represent the normalized histogram of non-zero valued pixels in the top half rows, and the fourth set of numbers contains numbers that represent the normalized histogram of non-zero valued pixels in the bottom half rows. These four sets of numbers are added to the descriptor, which now has a total of six values, i.e. sets of numbers. The order of the sets of numbers in the descriptor is not relevant. However, all the descriptors must have the same form, and therefore the order of the sets of values must be the same for all the constructed descriptors.
[0064] An additional set of numbers can be constructed according to the process described above. The cropped image 51 can be divided into three thirds, namely left, center, and right, using two vertical dividing lines 52. The thirds do not need to be of the same size. The left third, center third, and right third can be divided into columns, the total number of columns being the same as the number of columns in the undivided image and the image divided into two halves. Furthermore, a histogram representing the number of non-zero valued pixels in the columns of the left third is constructed and normalized. A histogram representing the number of non-zero valued pixels in the columns of the center third is also constructed and normalized. Furthermore, a histogram representing the number of non-zero valued pixels in the columns of the right third is also constructed and normalized. These three steps are interchangeable, meaning it does not matter which histogram is generated first.
[0065] The cropped image 51 is divided into three thirds, top, middle and bottom, using two horizontal dividing lines 52. The thirds do not have to be of equal size. The top, middle and bottom thirds are divided into rows, the total number of rows being the same as the number of rows in the undivided image and the image divided into two halves. A histogram is constructed and normalized that represents the number of non-zero valued pixels in the rows of the top third. Additionally, a histogram is constructed and normalized that represents the number of non-zero valued pixels in the rows of the middle third. A histogram is constructed and normalized that represents the number of non-zero valued pixels in the rows of the bottom third. These three steps are interchangeable, meaning it does not matter which histogram is generated first.
[0066] In response to completing the procedure, six more sets of numbers become available: a first set of numbers contains numbers that represent a normalized histogram of non-zero valued pixels in columns of the left third of the image, a second set of numbers contains numbers that represent a normalized histogram of non-zero valued pixels in columns of the middle third of the image, a third set of numbers contains numbers that represent a normalized histogram of non-zero valued pixels in columns of the right third of the image, a fourth set of numbers contains numbers that represent a normalized histogram of non-zero valued pixels in rows of the top third of the image, a fifth set of numbers contains numbers that represent a normalized histogram of non-zero valued pixels in rows of the middle third of the image, and a sixth set of numbers contains numbers that represent a normalized histogram of non-zero valued pixels in rows of the bottom third of the image. These six sets of numbers are added to a descriptor, which has a total of 12 values, i.e., sets of numbers. The order of the sets of numbers in the descriptor is not relevant. However, all of the descriptors may have the same form, and therefore the order of the set of values may be the same for all of the constructed descriptors.
[0067] The image can be further divided, analogously, into quarters, fifths, etc., up to the maximum applicable division number, which is the lower number of the image's size in pixels. Assuming that the division ordinal is N, it is clear that the division generates 2N new histograms to be added to the descriptor. A preferred embodiment of the method divides the image into a maximum of three parts, however, the number of divisions is not a limitation of the subject of the present invention.
[0068] The descriptor constructed according to the method described in the paragraph above can further be used to identify an image. For example, the descriptor can be implemented in either the test robot 2 or the central control unit 3. The implementation can be performed in software. For example, the test robot 2 or the central control unit 3 can have software installed thereon, the purpose of the software being to generate a descriptor of a given image. As input, the software can receive an image, and the output of the software can be a descriptor of the given image. For example, the image processing unit 13 can receive an input in the form of an image 50 captured by the camera 11, a part of the image 50 forming an area of interest, such as the display 4 or the OUI 61 of the device under test 1. The image 50 can further be cropped, such that only the area of interest is in the cropped picture. The descriptor is constructed as detailed above and stored in the memory unit 14. The output of the image processing unit 13 can thus be a descriptor of the cropped image 51 that is input.
[0069] Since each set of numbers represents the number of non-zero valued pixels in a row or column of a given sector, the inverse method of constructing a descriptor of an image is also applicable: the inverse method may work with numbers representing zero valued pixels as opposed to using the number of non-zero valued pixels.
[0070] Furthermore, counting only non-zero valued pixels also represents zero valued pixels since the total number of pixels in each row or column is given by the sum of these two numbers, and the interrelationship of these two numbers can be given by the simple equations N0=N-N1 and N1=N-N0, where N0 is the number of zero valued pixels, N1 is the number of non-zero valued pixels, and N is the total number of pixels in a given row or column.
[0071] Furthermore, the cropped image 51 can be divided into so-called sectors 53 using horizontal or vertical dividing lines 52. The sectors can be halves, thirds, quarters, etc. of the cropped image 51 and should be of the same size. Each sector 53 comprises lines and rows of pixels forming the cropped image 51. The rows and columns may further be divided into groups, where a group consists of at least one row or column. The histogram is counted for each group of each sector.
[0072] Image Classification The embedded system 6 of the device under test 1 can use the display 4 to show the user information about task progress, device status, options, settings, etc. in a graphical user interface (GUI) 62. The user can interact with the user interface, and the feedback of the embedded system 4 is usually shown in the form of a screen on the GUI 62. The user interface can comprise input elements such as buttons 5, touchpad, trackpoint, joystick, keyboard, etc. to receive input from the user. If the display 4 of the device under test 1 is a touch screen, it can also comprise an action element 7 implemented in the GUI 62. Touching the input element or action element 7 can result in feedback of the device under test 1, an action can be performed, the screen on the display 4 can be changed, the device under test 1 can be turned on or shut down, etc. For further use, all possible screens, states of the visual indicators 63, current settings, etc. constitute the state of the device under test 1. A change in the screen shown in the GUI 62, illuminating the visual indicators 63, performing an action, etc. means that the state of the device under test 1 has been changed.
[0073] The screens shown on the display 4 of the user interface can be sorted into classes that describe their purpose or meaning, for example a title screen showing a company logo with a load bar, a login screen asking the user to enter credentials, an error screen informing the user of a task failure, a screen with a list of tasks, actions available to the user, a screen showing the settings of the test device, language options, a list of information about the state of the device, etc. Some of these screens can be shown in various forms with only slight differences between them, such as a login screen with an empty slot for the username and password, with only the username partially filled in, or completely filled in. The usernames and passwords for various users can have different lengths, so the screens on which the credentials are filled in can be subtly different. For the purposes of testing the device under test 1, the differences between these pictures can be noted. Furthermore, the test system can also note that these pictures are similar and all belong to the same class of login screens.
[0074] Furthermore, a list of tasks to be performed can be shown on the display of the test device. If the device is a printer, for example, the task list can show a queue of documents to be printed, their order, size, number of pages, etc. The task list can change its form depending on the number of tasks to be performed, and some of the tasks may be selected to be removed from the queue, so that they can be highlighted or a selected symbol can appear in their vicinity. The test system should note that these pictures are similar and all belong to the same class task list screen.
[0075] Every possible screen of the GUI 62 displayed on the display 4 that needs to be recognized should be manually assigned to a class by a person. In this way, it is ensured that all images of the screen are assigned to the correct class. Furthermore, the authorized person can review as many images of the GUI 62 screens as possible to cover all classes, which will be used to sort the screens of the GUI 62. During the manual sorting, the authorized person can be prompted to mark and highlight the action elements 7 on the screens that they are assigning to a class. Furthermore, the authorized person can manually mark the position and size of the action element 7 and assign an action to this action element 7. An action may refer to an instruction that the action element 7 passes to the embedded system 6. The action element 7 can thus be, for example, an OK button that confirms an action performed by the user. Another example of an action element 7 can be a "login" button that validates the user's credentials and logs them into the system. Further examples of action elements 7 can be a cancel button, a button to close a window, a sign-up button, an action confirmation button, each task can act as an action element 7, an arrow button for navigating around the GUI, or any text field, etc. When assigning a screen to a class, it is therefore necessary to mark all the action elements 7 in the screen and to assign to each of them an action to be performed. The action can lead to closing a window, logging up, changing the screen, selecting a task to be performed, etc. It is therefore possible to change the screen by pressing an action element.
[0076] Furthermore, the display 4 of the device under test 1 may be mounted on a control panel 60 of the device under test 1. The device under test 1 may comprise an embedded system 6 adapted to receive user input in the form of instructions via action elements 7 implemented in a touch screen or via interactive elements such as buttons 5. A user may use the action elements 7 and the interactive elements to interact with the device under test 1 and perform tasks.
[0077] Furthermore, at least one screen may have a connection to at least one other screen, meaning that upon interaction with an action element 7 or an interaction element, the original screen is transferred to the subsequent screen to which it is connected. There may be some screens to which no action element 7 is assigned, which therefore do not have a connection to any other screen. This screen may be, for example, a loading screen, an initial screen after the device under test 1 is turned on, or a shutdown screen after the device under test 1 is shut down. Each class of screens may therefore contain even more screens in their diversity or with slight variations, as discussed above.
[0078] In some implementations, the device under test 1 may not include a touch screen or GUI 62 requiring user-device interaction. Feedback to the user is rather provided by visual indicators 63, such as LEDs, LCD panels, etc. In this case, the device under test 1 may be installed in multiple states. The states of the device under test 1 include the actions being performed by the device under test 1, its information feedback, its current settings, etc. For example, if the device 1 is an oven (FIGS. 1E and 1F), the control panel 63 includes two knobs 64, a first knob 64 adapted to change the mode of the oven (turn on grill, increase temperature, decrease temperature, hot air mode, etc.) and a second knob 64 adapted to change the temperature setting, and two visual indicators 63, the first one indicating whether the oven is turned on and the second one indicating whether the heating is on or off, in other words whether the desired temperature has been reached. The oven may further include a display 4 with a clock and two buttons 7 used to set the oven timer, the time of the clock, etc. The selected mode, the current setting, the set temperature, and the state of the visual indicator 63 constitute all possible states of the oven. The oven is in a first state when it is shut down. By turning the knob 64, the user selects the operating mode of the oven and thus changes the state of the oven, which usually involves switching on the first visual indicator 63. By turning the second knob 64, the user sets the desired temperature and thus changes the state of the oven, which usually involves switching on the first visual indicator 63 (see FIG. 1F). It is therefore necessary to take pictures of all possible states of the oven, locate the buttons 7 or knobs 64 and assign them a function that will change the state of the oven to another state. The pictures are assigned a descriptor and classified.
[0079] Decision-making module The central control unit 3 comprises a decision-making module 22, preferably embodied as a neural network. The decision-making module 22 is implemented in the central control unit 3 in software form. The decision-making module 22 comprises a trained classifier. The classifier is trained on a data set comprising either a set of descriptors of classified images of states of the device 1 under test, or classified images of states of the device 1 under test, or a set of descriptors of classified images of states of the device 1 under test together with classified images of states of the device 1 under test.
[0080] As described above, all possible states of the device under test 1 are manually identified and sorted into classes. Each state is assigned a descriptor according to the method described above. To generate an even larger dataset for training the classifier of the decision-making module 22, images of the screen can be captured in various conditions such as illumination, size, angle of the camera 11 relative to the device under test 1, brightness, contrast of the display 4, image sharpness, etc. Images of the same state of the device under test 1 taken under various conditions form a batch. For all classified images of the states of the device under test 1, identification descriptors are generated and assigned. It is clear that the descriptors for the images of a given batch will be similar but not identical since the images were taken under different conditions.
[0081] The classifier is trained on a data set that comprises either the descriptors of sorted images of states of the device 1 under test, images of states of the device 1 under test, or a combination of both. The larger the training data set, the more accurate the trained classifier will be. After training, the classifier is prepared to sort the images or the descriptors into a given class. In the sorting process, the classifier takes as input the images of states of the device 1 under test or the descriptors of a given image, and is able to assign the image or the descriptor to the correct class. With the understanding that the classifier is not completely accurate all the time, in order to improve the accuracy of the classifier, it is recommended to expand the training data set by capturing one image of a given state of the device 1 under test under various conditions such as brightness of the display, color contrast, lighting conditions, camera angle and distance from the display, etc. These images depict the same state of the device 1 under test. When cropped, their information value will be the same even if the images themselves are slightly different. For that reason, the descriptors of these images will also be slightly different. In this way, after training of the classifier is performed, it will classify the images of the screen more accurately.
[0082] Method for testing an embedded system - Patents.com The method for testing the embedded system 6 of the device under test 1 comprises a series of steps. In particular, a set of images can be generated, depicting at least two states of the device under test 1. This set of images is stored in the memory unit 21 of the central control unit 3. The set of images can comprise all possible possible states of the device under test 1. For at least one state of the device under test 1, at least one action element 7 and a connection to at least one other state of the device under test 1 can be assigned. The connection is applied when the action element 7 or button 5 of the current state of the device under test 1 is pressed. After pressing, the current state of the device under test 1 changes to another state of the device under test 1 to which the initial one is connected. An identification descriptor is then assigned to each of the images of the state of the device under test 1. The descriptor is a set of numerical values, preferably in vector form, that describe the current state of the device under test 1. The descriptor of the state can vary slightly depending on the current conditions under which the image was acquired. In general, the descriptors for the different states of the device 1 under test are different and should not be alternated.
[0083] In the next step (FIGS. 1H-1J), an image of the device under test 1 is captured using the camera 11 of the test robot 2. The photo taken by the robot 2 is cropped so that only a photo of the OUI 62 itself is shown. The image is then saved and stored in the memory unit 14 of the test robot 2 or in the memory unit 21 of the central control unit 3. A current descriptor is assigned to the image of the stored state and compared with the identification descriptor stored in the memory unit 21 of the central control unit 3. The comparison process is performed by a decision-making module. The current descriptor of the current image of the device under test 1 is used as input for a neural network, and as output, the current state of the device under test 1 is determined with a certain degree of accuracy. Since the current descriptor of the acquired image, which describes the current state of the device under test 1, differs slightly from the identification descriptor stored in the central control unit 3 and associated with the state of the device under test 1, a direct comparison is not as effective and accurate as one using a neural network. After the current state of the device under test 1 is determined, the position of an action element 7 on the display 4 or a button 5 on the control panel 60 is determined and the effector 10 is moved so that the action element 7 or button 5 is pressed.
[0084] The method for testing the embedded system 6 of the device under test 1 can further be used to measure the response and reaction times of the embedded system 6. When the action element 7 or button 5 of the current state of the device under test 1 is pressed, the time measurement is started. Once the screen on the display 4 or the state of the device under test 1 in general has been changed, the time measurement is ended and the value of the elapsed time is stored in the memory unit 21 of the central control unit. The state of the device under test 1 can then be determined. In this way, the time taken by the embedded system 6 to perform various operations can be measured, such as the time required for the change of various screens on the GUI 62 or the change of the state of the device under test 1 (as shown in FIG. 1K).
[0085] The test robot 2 may further comprise a sensor 19, such as a motion sensor, a heat sensor, a noise sensor or an additional camera. This sensor 19 is adapted to detect whether an action has been performed. In a preferred embodiment, the device under test 1 is a printer. The action may therefore be the printing of a document. The sensor 19 is then used to detect whether the document has been printed. For example, a camera or a motion sensor is placed to monitor the area for which the printed document is placed and, upon detecting the movement, it becomes evident that the document has been successfully printed. This can be used to measure the time to print. In an embodiment of the method, the user is prompted to confirm the print action via the GUI 62. In other words, a screen with a print command is shown on the display 4. This command can be manually confirmed by pressing an action element 7, which is associated with a confirmation action, such as an OK button.
[0086] For example, the device under test 1 may be an oven and the sensor 19 may be a thermal sensor installed inside the oven. The test robot 2 gives the oven a task to heat up, for example, to 200° C., which changes the oven's state by illuminating a visual indicator. Once the oven reaches the required temperature, the sensor 19 sends information to the test robot 2 about the action to be performed. It is clear that the above examples are merely illustrative and do not limit the scope of the invention to only printers and ovens.
[0087] The robot 2 may comprise a second robot arm 16, which comprises a communication tag 17, such as a card or chip, based on the RFID or NFC communication protocol. The printer or other device under test 1 then further comprises a communication receiver 18, which is adapted to communicate with the communication tag 17. Verification of the action may then be performed by placing the communication tag 17 in the vicinity of the communication receiver 18. Once the action has been verified by any of the methods described in this paragraph, a time measurement is initiated. Since the action is, in an embodiment, to print a document, the document is printed and detected by the sensor 19. Upon detection, the time measurement is terminated and the measured time value is stored either in the memory unit 14 of the robot 2 or in the memory unit 21 of the central control unit 3.
[0088] 2 is a high-level block diagram of electronic circuitry 110 that may be used with embodiments disclosed herein. As described above, electronic circuitry 110 may include a processor 120 configured to monitor operation of robotic system 100 and to send and / or receive signals related to operation of robotic system 100 and machine 150.
[0089] Processor 120 can be configured to collect or receive information and data regarding the operation of robotic system 100 and / or machine 150, and / or store or automatically transfer the information and data to another entity (e.g., another part of the robotic system). Processor 120 can further be configured to control, monitor, and / or perform various functions required for analysis, interpretation, tracking, and reporting of information and data used or collected by arm 170, vision source 160, and other components of system 100 (e.g., actuators, sensors, light emitters, etc.).
[0090] Generally, these functions can be performed and implemented by any suitable computer system and / or in digital circuitry or computer hardware, and the processor 120 can implement and / or control various functions and methods described herein. The processor 120 can comprise a central processing unit (CPU) 122, which is connected to a main memory 130 and includes processing circuitry configured to operate on instructions received from the main memory 130 and execute various instructions. The CPU 122 can be any suitable processing unit known in the art. For example, the CPU 122 can be a general-purpose and / or special-purpose microprocessor, such as an application-specific instruction set processor, a graphics processing unit, a physics processing unit, a digital signal processor, an image processor, a co-processor, a floating-point processor, a network processor, and / or any other suitable processor that can be used in digital computing circuitry. Alternatively or in addition, the processor 120 can comprise at least one of a multi-core processor and a front-end processor.
[0091] Generally, the processor 120 and CPU 122 can be configured to receive instructions and data from a main memory 130 (e.g., a read-only memory or a random access memory or both) and execute the instructions. The instructions and other data can be stored in the main memory 130. The processor 120 and the main memory 130 can be contained in or supplemented by special purpose logic circuitry. The main memory 130 can be any suitable form of volatile memory, non-volatile memory, semi-volatile memory, or virtual memory contained in a machine-readable storage device suitable for embodying data and computer program instructions. For example, the main memory 130 can comprise one or more of a magnetic disk (e.g., internal or removable disk), a magneto-optical disk, a semiconductor memory device (e.g., an EPROM or EEPROM), a flash memory, a CD-ROM, and / or a DVD-ROM disk.
[0092] The main memory 130 may include an operating system 132 configured to implement various operating system functions. For example, the operating system 132 may be responsible for controlling access to various devices, memory management, and / or implementing various functions of the robotic system 100. In general, the operating system 132 may be any suitable system software capable of managing computer hardware and software resources and providing general services for computer programs.
[0093] The main memory 130 may also hold application software 134. For example, the main memory 130 and the application software 134 may include various computer-executable instructions, application software, and data structures, such as computer-executable instructions and data structures, that implement various aspects of the embodiments described herein. For example, the main memory 130 and the application software 134 may include computer-executable instructions, application software, and data structures that may be employed to operate the machine 150 with the robotic system 100, and image processing software used to process images acquired by a visual source to extract information about the machine and / or visual interfaces.
[0094] Generally, the functions performed by the robotic system 100 can be implemented in digital electronic circuitry, or in computer hardware executing software, firmware, or a combination thereof. The implementation can be as a computer program product (e.g., a computer program tangibly embodied in a non-transitory machine-readable storage device) for execution by or to control the operation of a data processing apparatus (e.g., a computer, a programmable processor, or multiple computers).
[0095] The main memory 130 may also be connected to a cache unit (not shown) configured to store copies of the most frequently used data stored in the main memory 130. Program code that may be used with the embodiments disclosed herein may be implemented and written in any form of programming language, including compiled or interpreted languages, and may be deployed in any form, including as a stand-alone program or as a component module, subroutine, or other unit suitable for use in a computing environment. A computer program may be configured to be executed on a computer at one facility, or on multiple computers, or distributed across multiple facilities and interconnected by a communications network such as the Internet.
[0096] The processor 120 may further be coupled to a database or data storage device 140. The data storage device 140 may be configured to store information and data related to various functions and operations of the robotic system 100 and / or the machine 150. For example, the data storage device 140 may store data collected by the robotic system 100, detected changes in ambient environmental conditions during image capture of the machine 150 (e.g., changes in lighting detected by the visual source 160), etc.
[0097] The processor 120 may further be coupled to an interface 142. The interface 142 may be configured to receive information and instructions from the processor 120. The interface 142 may generally be any suitable display available in the art, such as a liquid crystal display (LCD) or a light emitting diode (LED) display. For example, the interface 142 may be a smart and / or touch sensitive display 145 that may receive instructions from a user and / or provide information to a user to operate the machine 150 using the robotic system 100.
[0098] The processor 120 can further be connected to various interfaces. Connection to the various interfaces can be established through a system or input / output (I / O) interface 144 (e.g., Bluetooth, USB connector, audio interface, Firewire, interface for connecting to peripheral devices, etc.). The I / O interface 144 can be directly or indirectly connected to the robotic system 100.
[0099] The processor 120 can further be coupled to a communication interface 146, such as a network interface. The communication interface 146 can be a communication interface included within the robotic system 100 and / or a remote communication interface 146 configured to communicate with the robotic system 100. For example, the communication interface 146 can be a communication interface configured to provide the robotic system 100 with a connection to a suitable communication network 148, such as the Internet. Transmission and reception of data, information, and instructions can occur via the communication network 148. Furthermore, the communication interface 148 can be an interface configured to enable communication between the electronic circuitry 110 (e.g., a remote computer) and the robotic system 100 (e.g., via any suitable communication means, such as wired or wireless communication protocols, including Wifi and Bluetooth® communication schemes).
[0100] 3 is a high-level flow diagram of a procedure for operating the robotic system 100, according to some embodiments disclosed herein. To define the dimensions of the interactive element (e.g., the screen to be operated) 201, the procedure can calibrate the vision source (e.g., camera) 202 and the arm (e.g., robot) 204 to ensure accurate interaction between the arm 170 and the machine 150.
[0101] During the calibration phase, images acquired by visual source 160 can be assessed for quantity and quality of information. For example, images from visual source 160 can be reviewed to determine whether they provide sufficient data to locate machine 150 and / or interactive medium 152 within field of view 101, and additional images can be acquired, if necessary.
[0102] Additionally, the processor 120 can evaluate and score images based on their individual quality and information provided by a particular image, and the processor 120 can discard images having a score below a predetermined level. Scores can generally be assigned to images based on any suitable metric for evaluating information in an image. For example, scores can be assigned based on the probability of the presence of sufficient data in the image to extract information about the machine 150 and / or the interactive medium 152. In particular, the processor 120 can evaluate the images and assign a score to each image that indicates the probability that the image may provide sufficient data to detect information about the machine 150 and / or the interactive medium 152.
[0103] Alternatively, or in addition, the images can be scored based on the amount of information available in each image for identifying the machine 150, interface, or interactive medium 152 in the field of view 101 and / or based on the probability of correct detection of characteristics of the interactive medium 152. As described, the processor 120 can eliminate images having a score below a predetermined score from being used to identify the interface or interactive medium 152 in the field of view and / or employ images having a score above a predetermined score for identifying the interface or interactive medium 152 in the field of view 101.
[0104] As described, in response to determining that additional data is needed for accurate detection of the machine 150 or interface, the processor 120 can adjust the camera to capture additional images in the additional field of view 101. Specifically, the processor 120 can be coupled to a frame on which the camera is movably mounted. If the processor 120 determines that additional images from other perspectives are needed, it can instruct the camera frame to change the position and / or orientation of the camera to capture additional images in the other field of view 101. The adjustment of the camera can comprise any suitable adjustment, for example, adjusting the angle of the camera.
[0105] Additionally, as described above, the processor 120 may be connected to one or more light emitters and / or light sensors to adjust the intensity of light emitted by the light emitters prior to acquiring the additional image. Other sensors may also be used to collect additional visual data. For example, light sensors, infrared cameras, optical sensors, light emitters, and other actuators may be used to collect data regarding light intensity, lighting conditions, device display settings, and distortion. In response to the collected data, the processor 120 may make further adjustments to the visual source 160.
[0106] 1A , arm 170 can further be activated to actuate actuator 176, which moves in proximity to machine 150 and obtains further data regarding machine 150 and interactive medium 152. Specifically, during arm calibration 204, arm 170 can interface with interactive medium 152 and machine 150 (under direction of processor 120) to locate the boundaries of machine 150 and the location / position of interactive medium 152 and its various elements.
[0107] The pressure sensor 174 can be coupled to a distal end 175 of the arm 170 and configured to indicate a change in resistance when the distal end 175 of the arm 170 interfaces with the interactive medium 152 of the machine 150. Additionally or alternatively, the actuator 176 can detect a change in the location of the button 199 within the interactive medium 152. For example, the system 100 can be configured such that the arm 170 moves in a first dimension while the actuator 176 coupled to the arm 170 moves in a second dimension, different from the first dimension. To calibrate the machine 150, the processor 120 can analyze images provided by the visual source 160 and extract certain information about the interface within the field of view 101 from the images (e.g., the location or orientation of the interface). Using this information, the arm 170 and actuator 176 can move about the identified area of the machine 150 (e.g., in the vicinity of the machine and interface) and identify the interface (e.g., by physically contacting the machine and interface and recording its characteristics in response to the physical contact). The actuator 176 can further move in the vicinity of the identified interface and observe possible changes in the interface in response to contacting the interface. For example, the actuator 176 can sense changes in response to contacting a physical key (e.g., key movement) or a touch screen key (e.g., a pressure sensor senses a change in the amount of pressure applied). The changes observed by the actuator 176 can be recorded and used to determine characteristics of the interactive medium 152. Additionally, information obtained from the visual source 160 can be used to determine characteristics of the interactive medium 152, such as shape, color, etc.
[0108] The arm calibration 204 can be a semi-automated process 206 with at least some user input for interfacing with the interactive medium 152 of the machine 150. The user input for the arm calibration 204 can be provided by a user through the communications network 148 and the communications interface 146.
[0109] Additionally or alternatively, arm calibration 204 can be a fully automated process 208 without requiring user input. For example, as discussed above, the vision source 160 of the robotic system 100 can include two or more cameras configured to provide two or more images from two or more fields of view to form a stream of images and capture a three-dimensional image of the machine 150. The three-dimensional image of the machine 150 can be used to determine metrics for the interactive medium 152 of the machine 150 in three dimensions, which can be processed by the processor 120 and application software 134 to automate the arm calibration 204.
[0110] In general, the metrics obtained using the arm 170, actuator 176, and vision source 160 may include any information that may assist the system 100 in determining characteristics of the machine 150 and interactive medium 152 within the field of view (e.g., the location of the machine 150 within the field of view, the orientation of the machine 150 within the field of view, the dimensions of the interface with the interactive medium 152 within the field of view, the location of the interface with the interactive medium 152 within the field of view, and the boundaries of the interface with the interactive medium 152 within the field of view). The extracted metrics may be stored in the main memory 130 of the robotic system 100 for use in performing tasks performed by the system.
[0111] Metrics (e.g., dimensions, position, location in space, orientation) of elements of the interactive medium 152 of the machine 150 and the UI interface of the interactive medium 152 can be determined through a fully automated process executed by the application software 134 and the processor 120. For example, the screen and user-interface (UI) capture 210 can be performed by the processor 120 by comparing and matching images captured by the visual source 160 with previously captured images of the screen and UI 212 that have been extracted from data provided to the robotic system 100. The data can be provided to the system 100 by a user or retrieved from a database that stores such data. For example, the data can be extracted from documents (e.g., the original manual of the machine 150 or online resources) provided to the robotic system 214 by a user. The robotic system can be equipped with text recognition capabilities that provide the robotic system with the ability to extract information about the machine 150, the screen, and / or the user-interface and apply the information in identifying these elements.
[0112] Additionally, the processor 120 of the robotic system 100 may perform an image scoring process 216 on the pre-captured images of the screen and UI 212. The image scoring process 216 may include scoring the images based on the amount of information contained in the image and / or the quality of the image (e.g., based on the probability of correct detection of the characteristics of the interactive media 152 in the image). The image scoring process 216 may further include filtering out images with scores 218 below a pre-defined score used to identify the interactive media 152 in the field of view and reserving images with scores above a pre-defined score used to identify the interactive media 152 in the field of view for further processing by the processor 120. The retained images may be stored in a memory for future use by the robotic system 220. The memory may be a local memory of the robotic system or a remote memory or database (e.g., a database stored in the cloud) connected to the system via a communication network.
[0113] The robotic system can further define and determine tasks 222 that can be performed using the machine 150 via the interface. These actionable tasks can be tasks previously defined (e.g., by a user) and / or tasks determined by the arm 170 and actuators 176 via interaction with the machine 150. Specifically, the robotic system 100 can determine, based on data provided to the system (e.g., by a user), via the interface, that performing a certain task (e.g., action A) will result in a certain outcome (e.g., outcome A'). Additionally or alternatively, the robotic system can determine, based on interaction with the interface, that performing a certain task (e.g., action B) can result in a certain outcome (e.g., outcome B') 222.
[0114] Further, the robotic system can determine and develop such a list of actionable tasks by repeatedly interacting with the machine 150 via the interactive medium 152, performing tasks, and recording outcomes associated with each task. This provides the machine 150 with the ability to learn from repeated interactions with the interface and develop a library of actions and corresponding actionable tasks. Specifically, the system can interact with the interactive medium 152 and determine that performing certain actions (e.g., actions C, D, and E) results in certain corresponding outcomes (e.g., outcomes C', D', and E'). The system can develop an action dictionary / database that stores certain information about the actions, such as the visual appearance or location of the actions and / or the keys involved in performing the actions. The system can further develop an outcome dictionary / database that stores outcomes associated with each action. The dictionary (database) developed by the system can store any suitable information for identifying actions and outcomes (actionable items). For example, the dictionary may comprise a glyph dictionary that stores information about interactive element items used to perform actions 224 (e.g., characteristics of buttons on interactive medium 152). At any point during the learning process, the user may provide input to the system to correct or modify the action. For example, the user may provide information about any features or elements of the interactive interface that were not correctly identified by machine 150 during learning process 226. Alternatively, or in addition, the user may provide information about actions that have not been identified by the system.
[0115] As described, when defining actionable items, reserved images that score above a predetermined score can be used to identify actions and their corresponding actionable items that can be performed via the interface.
[0116] Additionally, UI elements contained within the reserved images can be compared to a dictionary 224 (e.g., a glyph dictionary stored in memory 130). The dictionary 224 can include a database comprising a collection of images that include as images at least one of common UI elements, actions corresponding to characteristics of the interface shown in the image, actionable items shown in the image, and other capabilities of interactive media. As described, user input regarding the meaning of the unrecognized UI element 226 can be requested for UI elements within the reserved images that cannot be matched in the dictionary.
[0117] The system can further assign scores to tasks and their corresponding outcomes based on evaluating the tasks and outcomes for their optimality. Specifically, the system can determine that two or more tasks can result in the same outcome, review the two or more tasks, and rank / score the tasks based on factors such as time, cost, efficiency, number of steps involved, etc. In some implementations, tasks with scores below a predetermined score can be discarded and tasks with scores above a predetermined threshold can be adopted to achieve the desired outcome. Alternatively or in addition, the system can use the task with the highest score to achieve the desired outcome.
[0118] The system can further determine a flow generation indicating the order of procedures to follow to perform the task, and store this order for use in performing the task. The tasks can further be ranked and scored similarly as described above, with higher scored tasks being prioritized over lower scored tasks, as detailed above.
[0119] 4 is a flow diagram of a procedure for calibrating a robotic system according to some embodiments disclosed herein. For example, procedure 300 can be used to obtain additional information from visual source 160. As explained above, the system can obtain data received from various sensors and information received from visual source 160, analyze the data and information, and determine various characteristics of the interface and machine (302-304). If the system determines that additional data may be required to identify the interface and / or machine 150, the system 100 can prompt the visual source 160 (e.g., a camera) to obtain additional images and / or prompt the sensor to obtain additional data 306. The system 100 can further employ procedures to improve the capture of the additional data (e.g., improving lighting conditions, changing device display settings, improving image distortion, changing the angle of the camera 308, etc.) to improve the capture of the additional data 308.
[0120] As described above, any suitable sensors can be employed to acquire the data. For example, light sensors, infrared cameras, optical sensors, light emitters, and / or other actuators can be used to collect data including light intensity, lighting conditions, device display settings, and distortion. Additionally, in some implementations, other sensors 178 (e.g., digital sensors) can be used to identify, determine, and / or adjust display settings of the interactive medium 152 of the machine 150.
[0121] As described above, data received from other sensors 178 may be processed by processor 120 to assess the conditions under which the images were captured and to improve the quality and quantity of information obtained from the additional images. For example, the system may employ a light sensor to measure the light intensity of the environment surrounding machine 150 and compare the light intensity to a stored predetermined threshold of light intensity. If the light intensity is determined to be below the predetermined threshold, the system may trigger, by visual source 160, a recapture of an image of the field of view of machine 150 under changed lighting conditions (e.g., by employing a light emitter to emit additional light, capturing an image at a different time, and / or capturing an image after changing the brightness of interactive medium 152).
[0122] Further, the robotic system 100 can be adjusted to meet desired conditions for image capture. For example, the visual source 160 of the robotic system can include a light emitter 162 that can be triggered to provide a light source for image capture by the visual source 160 in response to the light sensor measuring an ambient light intensity below a predetermined threshold of light intensity. The light emitter 162 of the visual source 160 can also be adjusted to provide light with a range of different intensities and / or provide light from different angles. The visual source 160 of the robotic system 100 can also be adjusted to capture images of the machine 150 at times when the light sensor measures an ambient light intensity higher than a predetermined threshold light intensity. Additionally or alternatively, the visual source 160 can include an infrared camera configured to capture images under low light intensity when the ambient light intensity is determined to be lower than the predetermined threshold.
[0123] 5 illustrates a flow diagram 400 for preparing a dictionary 224 (e.g., a glyph dictionary) of elements of interactive media 152, their associated tasks, and / or associated physical outcomes (actionable items). As shown in FIG. 5, the dictionary 224 can be prepared using supervised learning 402, observational learning 408, or unsupervised learning 412.
[0124] The supervised learning dictionary 224 can be prepared by providing the system with input received from a user, including information related to known interactive media 152 elements (e.g., their images), their associated meanings 404, their associated tasks, and physical outcomes. The user input can optionally include glyph images and associated meanings present in a document (e.g., the machine 150 manual or an online resource) provided by the user to the robotic system 100. In some implementations, the processor 120 may present the dictionary elements (e.g., glyph images) and their associated meanings, tasks, and outcomes to the user for confirmation 406. Further, the dictionary 224 can be updated following user confirmation 406. In addition, the updated dictionary 224 can be used for identification of actionable items and their associated capabilities contained within the reserved images, as discussed in the above section.
[0125] Additionally or alternatively, the dictionary 224 can be formed by observing the user's actions 410 in response to interactions with UI elements in one or more images provided by the visual source 160. In particular, the visual source 160 can store images showing a user interacting with elements of an interactive element to perform tasks and achieve outcomes. The processor 120 of the robotic system 100 can process such images and identify relevant UI elements and their associated tasks and outcomes. This information can be used to update the dictionary 224 and identify the UI elements, their associated tasks, and outcomes.
[0126] Additionally or alternatively, the dictionary 224 can be generated and updated through direct interaction of the arms 170 and actuators 176 with the machine 150 and interactive elements, without requiring user input. The unsupervised learning approach 408 allows the system to interact with the machine 150 and discover elements of the interactive medium 152 and their associated tasks and outcomes. In addition to learning from direct interaction with the machine 150, the system can also update the dictionary by examining images and / or associated tasks and outcomes obtained from previously acquired images 418 of the machine 150 and interactive medium 152, expanding its knowledge of the machine 150 and interactive medium 152. As described above, the system can use pre-recorded documentation (e.g., user instructions), online resources, and / or other available images and information 418.
[0127] In some embodiments, a glyph dictionary that stores images of elements of interactive media 152, associated tasks, and / or their outcomes can be formed by the system and used to perform functions 420, 422, and 424 disclosed herein. For example, glyphs can be used to describe actionable items and desired outcomes that correspond to tasks performed using elements on interactive media 152.
[0128] 6 illustrates a flow diagram 500 of a procedure for identifying preferred physical tasks to perform using machine 150. As shown in FIG. 6, the procedure for identifying preferred physical tasks to perform using machine 150, which includes identifying physical tasks to perform using interactive medium 152 for operation of machine 150, includes capturing a state of machine 502, collecting data from sensors 504, interpreting data from sensors 506, determining a list of physical tasks 508, ranking the list of physical tasks 510, generating a task plan using the ranked list of physical tasks 512, identifying a highest ranked physical task 514, and executing the highest ranked physical task 516.
[0129] As discussed above, the robotic system 100 can be used to identify a current status of the machine 150 related to performing a preferred physical task by collecting data from the sensor 504 and interpreting data from the sensor 506. The sensors for collecting data from the sensor 504 and interpreting data from the sensor 506 can include pressure sensors, light sensors, infrared cameras, optical sensors, light emitters, and other actuators for collecting data including light intensity, lighting conditions, device display settings, and distortion. For example, data from the pressure sensor 174 coupled to the distal end 175 of the arm 170 can be configured to indicate a change in resistance when the distal end 175 of the arm 170 interfaces with the interactive medium 152. Additionally, data from the light sensor can be used to measure ambient light intensity. In interpreting the data from the sensor 506, the data collected while collecting data from the sensor 504 is processed by the processor 120 of the robotic system 100 to present a status of the machine 150. Identifying the current status of the machine 150 can be used to determine a list of physical tasks 508 that are capable of being performed by the machine 150 .
[0130] The list of physical tasks 508 that can be performed by the machine 150 can be used to rank a list of physical tasks 510 from the physical task with the lowest ranking to the physical task with the highest ranking to generate a task plan with a ranked list of physical tasks 512. For example, the ranking of the list of physical tasks 510 is performed by the processor 120 of the robotic system 100 by comparing the list of physical tasks with the machine 150 operation goal. Furthermore, the machine operation goal can be user-defined, and the list of physical tasks can be compared to the glyph dictionary 224 stored in the main memory 130 of the robotic system 100 to rank the list of physical tasks with respect to achieving the machine 150 operation goal. Furthermore, the ranking of the list of physical tasks 510 can also be changed by user input.
[0131] A task plan with a ranked list of physical tasks 512 can be prepared by storing a list of physical tasks ordered from the physical task with the lowest ranking to the physical task with the highest ranking to achieve an operation goal of the machine 150. Furthermore, the robotic system 100 can be configured to perform physical tasks with a higher ranking before physical tasks with a lower ranking.
[0132] The robotic system 100 can be used to identify a highest ranked physical task 514 from a list of physical tasks, arranged from the physical task with the lowest ranking to the physical task with the highest ranking, and execute the highest ranked physical task 516 to achieve the operation goal of the machine 150. For example, the physical task having the highest ranking can be identified as a preferred task to perform using the machine 150 to achieve the operation goal of the machine 150. Furthermore, the arm 170 can be instructed by the processor 120 of the robotic system 100 to perform the preferred task to achieve the operation goal of the machine 150.
[0133] 7 illustrates a flow diagram 600 of a procedure for executing preferred physical tasks to perform using machine 150 in a cyclical manner. As shown in FIG. 7, the procedure for executing preferred physical tasks to perform using machine 150 in a cyclical manner includes establishing an operational goal for machine 602, executing a preferred physical task for achieving the operational goal for machine 604, requesting input in case of error detection during execution of preferred physical task 606, resuming execution of the preferred physical task for achieving the operational goal for the machine after resolving error 608, and evaluating achievement of the operational goal for machine 610.
[0134] The robotic system 100 can be configured to request user input for the operation goal of the machine 602. Specifically, high frequency desired goals for the operation of the machine 150 can be stored in the main memory 130, from which preferred physical tasks for achieving the operation goal of the machine 150 can be identified. The robotic system 100 can be configured to execute preferred physical tasks for achieving the operation goal of the machine 604 in response to the user input for the operation goal of the machine 602. For example, the operation goal of the machine 150 can be to maintain the temperature of a room at a fixed temperature, and the robotic system 100 can be configured to execute the preferred physical task of increasing or decreasing the temperature of a thermostat to maintain the temperature of the room at the fixed temperature.
[0135] The robotic system 100 can be configured to request input in case of error detection during the execution of the preferred physical task 606. For example, the robotic system 100 can be configured to detect an error in accomplishing a goal for operating the machine 150 by receiving data from the sensors described in the above section. Furthermore, the robotic system 100 can be configured to request input in case of an error to resolve the error by comparing the preferred physical task with the dictionary 224 (glyph dictionary) stored in the main memory 130 of the robotic system 100, matching the actionable items in the preferred physical task, and updating the actionable items of the preferred physical task. The robotic system 100 can be configured to update the glyph dictionary 224 and the list of physical tasks (including the preferred physical task) stored in the main memory 130 of the robotic system 100 after resolving the error. Furthermore, the robotic system 100 can also be configured to request user input in response to detecting an error during the execution of the preferred physical task. The robotic system 100 can then be configured to resume performance of the preferred physical task to achieve the machine 150 operation goal after resolving the error 608.
[0136] The robotic system 100 can be used to evaluate the achievement of the machine 150's operation goal by performing the preferred physical task. The evaluation of the achievement of the preferred physical task to achieve the machine 610's operation goal can be performed by evaluating data from a sensor. For example, in the case of maintaining the temperature at a fixed temperature, the robotic system 100 can be configured to detect the temperature of the room using data from a thermal sensor. Furthermore, the evaluation of the achievement of the preferred physical task to achieve the machine 610's operation goal can further include requesting a user input for confirmation of the achievement of the machine 610's operation goal. The evaluation of the achievement of the preferred physical task to achieve the machine 610's operation goal can also further include an idle function to stop the robotic system 100 from performing the preferred physical task to achieve the machine 604's operation goal once the machine 150's operation goal has already been achieved.
[0137] The robotic system 100 can further be used to operate one or more machines in a serial or parallel sequence. FIG. 8 depicts a block diagram of an example of the use of the robotic system to operate one or more machines 150, 150'. As shown, one or more machines 150 / 150' can be arranged on a shelving unit 880 or similar structure available in the art that can accommodate one or more machines. The shelving unit includes one or more shelves 881, 881' or similar structure that can receive one or more machines 150, 150'. The machines can be arranged on the shelving unit / shelf in any suitable manner. For example, the machines can be arranged on shelves 881, 881' such that their individual interactive media are readily available for access by the arm 170 of the robotic system. As shown in FIG. 8, the robotic system 800 can include one or more arms 170, 170′, each configured to operate as detailed above. The robotic arm 170 can be connected to the system using a rail-shaped frame 871 that allows the robotic arm to move along the rail-shaped frame 871 (e.g., along a direction D1). The rail-shaped frame 871 can extend along the length of the shelf 881 and be arranged to allow the robotic arm 170 to move along the rail 871 and across the shelf 881 and interact with the interactive medium of the machine. For example, the robotic arm 170 can be configured to move along the horizontal (X) and / or vertical (Y) directions about the frame 871. Alternatively or in addition, the robotic arm 170 can be configured to move in two or more dimensions about the frame 871. For example, the robotic arm 170 can be configured to rotate in multiple directions while also moving horizontally and / or vertically along the rail 871.As described above, this arrangement of the robotic arm 170 can enable the robotic arm 170 to move in three, four, five, or more dimensions about the frame 871.
[0138] 8, the robotic arm is configured to move along rails 871 in multiple dimensions (e.g., X, Y, and Z dimensions) to interface with each of one or more machines 150 / 150'. As shown, the rails can be rails that extend along the shelving unit 880 from one side 881A to another 881B of the shelving unit 880 and provide a medium for the robotic arm to move from the first side 881A to the second side 881B. The vision source 160 can further be arranged along another rail 861 (or similar structure) and configured to move similarly to the robotic arm 170 and capture images of the field of view.
[0139] As detailed above, the actuator 176 (e.g., a plotter or stylus) of the robot arm 170 can be configured to interface with the interactive medium 152 of each of the one or more machines 150 / 150′. As described above, the actuator can move along the rail 871 and act on the machines in a serial and / or parallel sequence. For example, the actuator can move along the rail 871 from one side 881A to another 881B and initiate each machine with a task. The actuator can then return to the first side 881A and continue acting on the machines, subsequently performing another task and / or continuing a task previously initiated on each machine. For example, assuming a robotic arm is acting on several tablets arranged on a shelf, the robotic arm can start at a first side 881A and begin turning on each tablet, work on the other tablets by moving along the rail to a second side 881B for each tablet as the previous tablet starts up, return to the first side 881A, perform a subsequent task on each tablet (e.g., initiate an operating system update on a device and continue updating the other devices while moving to the second side 881B as the previous device updates), return to the first side 881A, and continue and / or perform other tasks.
[0140] As described above, the vision source 160 can be a static or video camera that provides a live video stream of the interactive media 152 of each of the one or more machines 150 / 150' within the field of view 801. The vision source 160 can further be configured to move in multiple dimensions (e.g., X, Y, and Z dimensions) to capture images or live video of each of the one or more machines 150 / 150'. For example, the vision source 160 can be configured to translate in one or more directions along the camera rail 861 and rotate in one or more dimensions along the camera rail 861. The robot arm 170 can translate / rotate independently of the vision source 860 along the rail 871. Alternatively and / or in addition, the robot arm 170 and the vision source 160 can be arranged relative to one another (e.g., along the same or similar path).
[0141] 9 depicts a block diagram of an example of a robotic system 900 according to some aspects disclosed herein. As shown, the robotic arm 170 can be configured to move in two or more directions (e.g., along direction P1) about a rail frame 872. The frame further includes an extension 873 configured to support the visual source 160. This configuration allows the visual source 160 to move in conjunction with the robotic arm 170 along at least one direction.
[0142] As described above, the visual source 160 can capture at least one image or live video of the one or more machines 150 / 150′ within the field of view 801. The at least one image or live video can be provided to the processor 120 ( FIG. 1A ) for processing the at least one image or live video of the interactive medium 152 (e.g., a smartphone or tablet icon) of each of the one or more machines 150 / 150′ and extracting relevant information about each of the one or more machines 150 / 150′.
[0143] Additionally, as detailed above, the visual source 160 and the robotic arm 170 can be coupled to the processor 120 via the communications network 148 of FIG. 1A. In response to processing at least one image or live video, the processor 120 can be configured to control the robotic arm 170 and interface with the interactive medium 152 of each machine 150. For example, the processor 120 can be configured to translate the robotic arm 170 along rails 871 to interface with different locations of the interactive medium 152 of each of the one or more machines 150 / 150′.
[0144] FIG. 10 depicts another example of a robotic system 1000 according to some aspects disclosed herein. As shown, the system 1000 can include multiple vision sources 160 / 160' disposed along a camera rail 861 to capture at least one image of each of the one or more machines 150 / 150' within a field of view 801. For example, each of the one or more machines 150 / 150' can have at least one of the multiple vision sources 160 / 160' allocated to capture at least one image of that machine 150. The multiple vision sources 160 can be static or video cameras that provide a live video stream of the interactive medium 152 of each of the one or more machines 150 / 150'. Additionally, the multiple vision sources 160 / 160' can be configured to rotate about their position on the camera rail 861 to capture additional images of the machines within the field of view.
[0145] Additionally, each of the multiple visual sources 160 / 160′ can capture at least one image or live video of each of the one or more machines 150 / 150′ within the field of view 801. As detailed above, the live video can be provided to the processor 120 of FIG. 1A for processing the at least one image or live video of the interactive medium 152 (e.g., a smartphone or tablet icon) of each of the one or more machines 150 / 150′ and extracting relevant information about each of the one or more machines 150 / 150′.
[0146] FIG. 11 depicts another example of a robotic system 1100 that may be used to operate one or more machines 150 / 150′. As shown in FIG. 11, one or more machines 150 / 150′ may be arranged on a movable belt 1080 (e.g., a conveyor belt) and move along a direction D1. A frame 1072 of the robotic arm may be disposed on the movable belt 1080 such that the robotic arm 170 may interface with the interactive medium 152 of each machine as the machine moves along the movable belt. For example, the robotic arm 170 may be disposed on a fixed location on the belt 1080 such that the robotic arm 170 may interface with the interactive medium 152 of each machine as the machine moves on the belt and comes close to the robotic arm 170. The frame 1072 further includes an extension 1061 configured to support a visual source 160.
[0147] As described above, the actuator 176 can be configured to translate / move and / or rotate about at least one additional direction and / or dimension relative to the robot arm 170 disposed on the rail 1071. The actuator 176 can be configured to rotate about its attachment to the distal end 175 of the robot arm 170, thereby providing at least one additional direction of translation / rotation for the robot arm 170. The protrusion 174 of the actuator 176 can also be configured to translate / rotate about its attachment to the actuator 176, thereby providing a further additional direction of translation / rotation for the robot arm 170.
[0148] Those skilled in the art will appreciate that various modifications can be made to the above-described embodiments without departing from the scope of the present invention. Although the present specification discloses advantages in the context of certain illustrative non-limiting embodiments, various changes, substitutions, permutations, and alterations may be made without departing from the scope of the present specification as defined by the appended claims. Furthermore, any feature described in relation to any one embodiment may also be applicable to any other embodiment.
Claims
1. A robotic system for operating a machine, the machine having an interface having a bidirectional medium for use in performing a physical task using the machine, the robotic system comprising: a robotic arm having a proximal end and a distal end; an actuator coupled to the robotic arm at the distal end, the actuator configured to physically interact with the interactive medium; a visual source configured to provide an image of the machine within a field of view; a processor coupled to the robotic arm and the visual source; a memory coupled to the processor, the memory configured to store a computer-executable program comprising instructions; and Equipped with The instructions, when executed by the processor, extracting information from the image indicative of a metric for identifying the interface within the field of view, the metric comprising at least one of a location, an orientation, a dimension, and a boundary of the interface within the field of view; using the information extracted from the image to identify the interface within the field of view; operating the robotic arm to interact with the identified interface via the actuator and receiving a signal from the robotic arm indicative of detection of a change in the interactive medium by the actuator; extracting information indicative of a characteristic of the interactive medium based on the signal from the actuator indicative of the detection of the change in the interactive medium, the characteristic of the interactive medium including a position of at least one element of the interactive medium; operating the robotic arm to perform the physical task using the machine by actuating the at least one element of the interactive medium; A robot system that performs the following:
2. The robot system of claim 1, wherein when the instructions are executed, the instructions determine whether additional information is required to identify the interface by analyzing the information indicating the metrics for identifying the interface within the field of view, and if the additional information is required, obtain at least one additional image from the visual source.
3. The robot system of claim 2, wherein the visual source comprises a camera, the camera being configured to acquire the image of the machine within its field of view by rotating around the machine within its field of view.
4. The robotic system of claim 3, wherein the system is configured to acquire the at least one additional image by adjusting the visual source.
5. The robot system of claim 1, wherein the actuator includes a pressure sensor positioned at a distal tip of the actuator, and the pressure sensor, upon sensing at least one of a resistance change or displacement of the bidirectional medium, generates the signal indicating detection of the change in the bidirectional medium by the actuator.
6. The robotic system of claim 1, wherein the visual source is configured to provide a plurality of images, and the instructions, when executed, analyze each image provided by the visual source and score each image based on at least one of the amount of information indicative of the metric for identifying the interface in the field of view or the probability of correct identification of the interface in the field of view, and exclude images having a score lower than a predetermined score from being used to identify the interface in the field of view.
7. The robotic system of claim 6, wherein the instructions, when executed, score the plurality of images based on previously verified data.
8. The robotic system of claim 7, wherein the previously verified data includes at least one of data obtained from a glyph dictionary, data provided by a human operator of the machine, data about the machine obtained from a remote entity via a communications network, data obtained from an original description of the machine, and data obtained from an online resource for the machine.
9. The robotic system of claim 8, wherein the glyph dictionary comprises a collection of images of common interactive media elements.
10. The robot system of claim 8, wherein the instruction, when executed, updates the glyph dictionary to include the metrics for identifying the interface within the field of view obtained from the information extracted from at least one image.
11. The robotic system of claim 1, wherein the memory is configured to store a list of physical tasks to be performed using the machine's interactive medium.
12. The robotic system of claim 11, wherein the robotic system is configured to obtain the list of physical tasks from at least one of a human operator, an original description of the machine, or an online resource for the machine.
13. The robot system of claim 12, wherein the memory is configured to store a ranking for each physical task in the list of physical tasks, and the instructions, when executed, cause the robot arm to operate to perform physical tasks having a higher ranking before physical tasks having a lower ranking.
14. The robotic system of claim 1, wherein the instructions, when executed, extract from the image the information indicative of the metrics for identifying the interface within the field of view using an image processing scheme that employs a deep learning framework to identify the interface.
15. The robotic system of claim 14, wherein the deep learning framework comprises at least one of supervised deep learning, unsupervised deep learning, or reinforcement deep learning.
16. The robotic system of claim 1, further comprising an optical sensor configured to measure light intensity within the field of view.
17. The robotic system of claim 16, wherein the instructions, when executed, extract and use the light intensity measured by the optical sensor to extract the metric for identifying the interface within the field of view.
18. The robotic system of claim 16, wherein the visual source comprises a camera configured to acquire the image of the machine within a field of view, and the system is configured to adjust the light intensity prior to acquiring the image.
19. The robotic system of claim 1, wherein the visual source is coupled to the robotic system via a communication network.
20. The robotic system of claim 1, wherein the processor is configured to execute the instructions in response to a voice command.
21. The robotic system of claim 1, further comprising a user interface coupled to the processor, the user interface configured to be used to control the robotic system.