Machine, System
The integration of sensor technology and AI into agricultural machines addresses the lack of automation in crop recognition and environmental assessment, enhancing the efficiency and quality of agricultural operations.
Patent Information
- Application Number
- JP2025062830
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2025-04-06
- Publication Date
- 2025-06-12
- Estimated Expiration
- 2045-04-06
AI Technical Summary
Existing agricultural systems lack efficient automation capabilities, limiting their ability to optimize crop recognition, environmental assessment, and decision-making for agricultural operations.
A machine equipped with sensors and AI capabilities that recognizes crops and environmental conditions, determining whether agricultural operations are feasible and generating adaptive operations through reinforcement learning.
Enables efficient automation of agricultural tasks by accurately recognizing crops and environmental conditions, improving the quality and efficiency of agricultural operations.
Abstract
Description
Technical Field
[0001] The present invention relates to a machine or a system.
Background Art
[0002] The statements in this section only provide background information related to this disclosure and do not necessarily constitute prior art.
[0003] Patent Document 1 discloses a temperature control system for an agricultural greenhouse.
Prior Art Documents
Patent Documents
[0004]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0005] However, the inventor recognized that at least in the above-described embodiment, there is a disadvantage that there is no configuration capable of efficiently automating agriculture.
Means for Solving the Problems
[0006] At least one disclosure of the present invention is a machine related to agriculture, recognizing crops based on information obtained from one or more sensors, recognizing the environment around the machine, judging whether the operation of agricultural work is possible or not by AI based on the crop information and the environmental information, a machine is provided.
Effects of the Invention
[0007] In this configuration, at least, there is usefulness in efficiently automating agriculture.
[0008] These and other aspects, features, and advantages of the present disclosure will become apparent from the following detailed written description of the preferred embodiments and aspects taken in conjunction with the following drawings, but modifications and variations can be made without departing from the spirit and scope of the novel concepts of the present disclosure. Aspects in one embodiment of the present disclosure can be combined with, or replaced by, one or more of the aspects in another embodiment of the present disclosure, as long as there is no contradiction.
DETAILED DESCRIPTION OF THE INVENTION
[0009] In the following disclosure, many different embodiments and examples are provided for implementing different features of the presented subject matter. To simplify the present disclosure, specific examples of components and arrangements are disclosed below. Of course, these are merely examples and are not intended to be limiting. For example, a structure in which a first feature is covered by, or in contact with, a second feature disclosed subsequently may include embodiments in which the first feature and the second feature are formed so as to be in direct contact, as well as embodiments in which additional features are formed between the first feature and the second feature so that the first feature and the second feature are not in direct contact. Further, in the present disclosure, reference numerals and / or letters may be repeated in various examples. Such repetition is for the sake of brevity and clarity and does not in itself require a relationship between the various embodiments and / or the configurations being described. Further, when a first element is described as being "connected" or "coupled" to a second element, such description includes embodiments in which the first element and the second element are directly connected or coupled to each other, as well as embodiments in which the first element and the second element are indirectly connected or coupled to each other with one or more other elements intervening therebetween.
[0010] As used herein, the recitation "at least one of" encompasses all exemplified variations. For example, the recitation "comprises at least one of A, B, or C" is synonymous with "consisting of A, B, C and combinations thereof". And it encompasses all possible variations of A, B, C, A + B, A + C, B + C, A + B + C. In the present disclosure, the disclosure of an embodiment combining two or more components can be implemented as an embodiment in which any one or more components are separated, as long as there is no contradiction or unless otherwise stated in this specification. For example, the recitation "performing A, B, and C" is synonymous with "comprising of A, B or C and combinations thereof". And it encompasses all possible variations of A, B, C, A + B, A + C, B + C, A + B + C.
[0011] In the present disclosure, the disclosure of using a machine, an electronic operator, or a computer can include embodiments of a method, a recording medium, an apparatus, or a program. The description "A is B" used herein can be replaced with "A includes B", as long as there is no contradiction or unless otherwise stated in this specification.
[0012] The terms in the present disclosure, including the terms described in the claims, can be interpreted in consideration of the descriptions and drawings described in the specification, and furthermore, as long as there is no contradiction with the suggestions in the present disclosure, based on what one or more members of the public have so named, indicated, understood, or implemented, or what is possible in the past, present, or future. Regarding the operating method used in at least one embodiment, the following embodiments can be adopted. The description of JP6456303, which well explains at least one embodiment, is cited for explanation (hereinafter, citation starts).
[0013] As used herein, the term "computer", as is known in the art, generally includes a processor, a memory such as a hard drive, disk drive, flash drive, or memory stick, or other non-transitory computer-readable medium or non-transitory storage device, at least one information storage / search device, such as a keyboard, mouse, pointing and touch device, touch screen, or microphone, at least one input device, and a display structure such as a well-known computer screen. Additionally, a computer may include one or more network connections, such as a wired or wireless connection. As is known in the art, such a computer or computer system may include more or less of the items listed above, and is not limited to, for example, tablet computers or smart devices, but encompasses other electronic media and electronic devices.
[0014] As used herein, the term "cloud" or "cloud computing" refers to a centralized and virtualized computing facility where all computing resources are shared. For application systems and subsystems, since they are all "in the cloud", it is no longer possible to refer to a specific machine.
[0015] As used herein, the term "Distributed Internet Service System" refers to a distributed Internet service platform that transforms Internet applications for execution in various computing environments. The DIS system delivers Internet applications, including content, data, and logic, to any number and type of device, to whatever extent appropriate, and along the network, via a Component Distribution Server / Asset Distribution Server. Through DIS, Internet applications can be hosted and centrally managed as services based on each user's needs, locally cached and executed at the user's device or nearby location while maintaining their integrity. Web-enabled computing devices can be upgraded with DIS software to become DIS-compliant to enjoy and execute distributed Internet services. The Distributed Internet Service System is fully described in any one of the patent families of U.S. Patent Nos. 7,136,857; 7,150,015; 7,181,731; 7,209,921; 7,430,610; 7,685,183; 7,685,577; 7,752,214; 8,326,883; 8,386,525; 8,443,035; 8,458,142; 8,458,222; 8,473,468; 8,527,545; and 8,650,226, and U.S. Patent Publications Nos. 2012 / 0005205 and 2013 / 0091252, all of which are jointly owned by OPIE40, Holdings, Inc. and are hereby incorporated by reference. (End of citation)
[0016] Regarding the operation method used in at least one or more embodiments, the following embodiments can be adopted for the conventional Internet method that does not use a distributed Internet. The description of JP7113047, which well explains at least one or more embodiments, will be cited for explanation (hereinafter, citation starts).
[0017] Embodiments including the matters specifically disclosed in this specification can provide an automatic response system realized in a form that actually converses with humans based on artificial intelligence, thereby realizing a more natural conversation with users while quickly and conveniently processing inquiries, reservations, delivery orders, etc.
[0018] The plurality of electronic devices 110, 120, 130, 140 may be fixed terminals or mobile terminals realized by a computer system. Examples of the plurality of electronic devices 110, 120, 130, 140 include AI speakers, smartphones, mobile phones, navigation devices, PCs (personal computers), notebook PCs, digital broadcast terminals, PDAs (Personal Digital Assistants), PMPs (Portable Multimedia Players), tablets, game consoles, wearable devices, IoT (internet of things) devices, VR (virtual reality) devices, AR (augmented reality) devices, etc. As an example, in FIG. 1, an AI speaker is shown as the electronic device 110. However, in the embodiments of the present invention, the electronic device 110 may mean one of various physical computer systems that can communicate with other electronic devices 120, 130, 140 and / or servers 150, 160 via the network 170 using substantially wireless or wired communication methods.
[0019] The communication method is not limited, and it may include not only a communication method using a communication network that the network 170 can include (for example, a mobile communication network, a wired Internet, a wireless Internet, a broadcast network, a satellite network, etc.), but also short-range wireless communication between devices. For example, the network 170 may include any one or more of networks such as a PAN (personal area network), a LAN (local area network), a CAN (campus area network), a MAN (metropolitan area network), a WAN (wide area network), a BBN (broadband network), and the Internet. Further, the network 170 may include any one or more of network topologies including a bus network, a star network, a ring network, a mesh network, a star-bus network, a tree or hierarchical network, etc., but is not limited thereto.
[0020] Servers 150 and 160 may each be implemented by one or more computer devices that communicate with a plurality of electronic devices 110, 120, 130, 140 via a network 170 to provide instructions, code, files, content, services, etc. For example, server 150 may be a system that provides a first service to a plurality of electronic devices 110, 120, 130, 140 connected via network 170, and server 160 may also be a system that provides a second service to a plurality of electronic devices 110, 120, 130, 140 connected via network 170. As a more specific example, server 150 may provide, as the first service, a service (such as an automatic response service, for example) targeted by the corresponding application to a plurality of electronic devices 110, 120, 130, 140 through an application that is a computer program installed and executed in the plurality of electronic devices 110, 120, 130, 140. As another example, server 160 may provide, as the second service, a service that distributes files for installation and execution of the above-described application to a plurality of electronic devices 110, 120, 130, 140.
[0021] FIG. 2 is a block diagram for explaining the internal configurations of an electronic device and a server in an embodiment of the present invention. In FIG. 2, the internal configuration of electronic device 110 and the internal configuration of server 150 are described as examples for the electronic device. Also, the other electronic devices 120, 130, 140 and server 160 may have the same or similar internal configurations as the above-described electronic device 110 or server 150.
[0022] The electronic device 110 and the server 150 may include memories 211 and 221, processors 212 and 222, communication modules 213 and 223, and input / output interfaces 214 and 224. The memories 211 and 221 may be non-transitory computer-readable recording media, and may include non-transitory mass storage devices such as RAM (random access memory), ROM (read only memory), disk drives, SSDs (solid state drives), flash memories, etc. Here, non-transitory mass storage devices such as ROM, SSD, flash memory, and disk drives may be included in the electronic device 110 or the server 150 as separate non-transitory recording devices distinct from the memories 211 and 221. Also, the memories 211 and 221 may record an operating system and at least one program code (for example, code for a browser installed and executed in the electronic device 110, an application installed in the electronic device 110 for providing a specific service, etc.). Such software components may be loaded from a computer-readable recording medium different from the memories 211 and 221. Such another computer-readable recording medium may include computer-readable recording media such as floppy (registered trademark) drives, disks, tapes, DVD / CD-ROM drives, memory cards, etc. In other embodiments, the software components may be loaded into the memories 211 and 221 through the communication modules 213 and 223 which are not computer-readable recording media. For example, at least one program may be loaded into the memories 211 and 221 based on a computer program (for example, the above-described application) installed by a file distributed by a file distribution system (for example, the above-described server 160) that distributes developer or application installation files via the network 170.
[0023] The processors 212 and 222 may be configured to process instructions of a computer program by performing basic arithmetic, logic, and input / output operations. The instructions may be provided to the processors 212 and 222 by the memories 211 and 221 or the communication modules 213 and 223. For example, the processors 212 and 222 may be configured to execute instructions received according to program code recorded in a recording device such as the memories 211 and 221.
[0024] The communication modules 213 and 223 may provide functions for the electronic device 110 and the server 150 to communicate with each other via the network 170, or may provide functions for the electronic device 110 and / or the server 150 to communicate with other electronic devices (e.g., the electronic device 120) or other servers (e.g., the server 160). As an example, a request generated by the processor 212 of the electronic device 110 according to program code recorded in a recording device such as the memory 211 may be transmitted to the server 150 via the network 170 under the control of the communication module 213. Conversely, control signals, instructions, contents, files, etc. provided under the control of the processor 222 of the server 150 may be received by the electronic device 110 through the communication module 213 of the electronic device 110 via the communication module 223 and the network 170. For example, control signals, instructions, contents, files, etc. received through the communication module 213 from the server 150 may be transmitted to the processor 212 and the memory 211, and the contents and files, etc. may be recorded in a recording medium (the above-mentioned non-transitory recording device) that the electronic device 110 may further include.
[0025] The input / output interface 214 may be means for interfacing with an input / output device 215. For example, the input device may include devices such as a keyboard, a mouse, a microphone, a camera, etc., and the output device may include devices such as a display, a speaker, a tactile feedback device, etc. As another example, the input / output interface 214 may be means for interfacing with a device in which functions for input and output are integrated into one, such as a touch screen. The input / output device 215 may be composed of the electronic device 110 and one device. Also, the input / output interface 224 of the server 150 may be means for interfacing with a device (not shown) for input or output that can be connected to or included in the server 150. As a more specific example, when the processor 212 of the electronic device 110 processes the instructions of a computer program loaded in the memory 211, a service screen or content configured using the data provided by the server 150 or the electronic device 120 may be displayed on the display through the input / output interface 214.
[0026] Also, in other embodiments, the electronic device 110 and the server 150 may include more components than the components shown in FIG. 2. However, it is not necessary to clearly show most of the conventional components in the figure. For example, the electronic device 110 may be realized to include at least a part of the input / output device 215 described above, or may further include other components such as a transceiver, a camera, various sensors, a database, etc. As a more specific example, when the electronic device 110 is an AI speaker, various components such as various sensors, a camera module, physical buttons, buttons using a touch panel, an input / output port, a vibrator for vibration, etc., which are generally included in the AI speaker, may be realized to be further included in the electronic device 110. (End of citation)
[0027] A machine is disclosed. According to at least one embodiment, the user terminal consists of a control unit, a RAM, a storage unit, a graphics processing unit, a communication interface, and an interface unit, which are respectively connected by an internal bus. In at least one embodiment, the user terminal includes terminals owned by the user. On the other hand, it includes not only the terminals owned by the user, but also terminals whose owners are other than the user (including sellers and traders of goods and services, governments, and local governments). As an example, in providing or advertising goods or services (hereinafter referred to as "such provision, etc." in this paragraph), terminals provided for use by the recipients of such provision, etc. (including things to be transferred or lent), terminals provided for such provision, etc., and terminals related to the provision of such goods or services by the recipients of such provision, etc. That is, it includes terminals owned by others to whom the user has only temporarily been permitted to use, and terminals lent to the user.
[0028] According to at least one embodiment, the control unit is composed of a CPU and a ROM. The control unit executes the program stored in the storage unit to control the user terminal. The RAM is the work area of the control unit. The storage unit is a storage area for storing programs and data. The control unit reads the program and data from the RAM for processing. The control unit outputs a drawing command to the graphics processing unit by processing the program and data loaded into the RAM.
[0029] According to at least one embodiment, the graphics processing unit is connected to the display unit. The display unit has a display screen. When the control unit outputs a drawing command to the graphics processing unit, the graphics processing unit outputs a video signal for displaying an image on the display screen. Here, the display unit may be a touch panel equipped with a touch sensor. The touch panel of this display unit functions as an input unit.
[0030] According to at least one embodiment, the communication interface can be connected to a communication network wirelessly or by wire, and can transmit and receive data to and from a server device via the communication network. The data received via the communication interface is loaded into the RAM and processed by the control unit. An external memory (e.g., an SD card, etc.) is connected to the interface unit.
[0031] According to at least one embodiment, the user terminal is not particularly limited as long as it is a computer device having a display screen and an input unit. Examples of the user terminal include a conventional mobile phone, a tablet terminal, a smartphone, a desktop or notebook personal computer, etc. It may also be composed of a VR goggle, i.e., a screen (or two display panels, one for each eye) attached to a frame (or headset) fixed or attached to the head with a strap. The user terminal has an audio output unit.
[0032] According to at least one embodiment, the user terminal can be communicatively connected to a server device via a communication network. It can communicate via the communication network to transmit or receive information.
[0033] According to at least one embodiment, the server device includes at least a control unit, a RAM, a storage unit, and a communication interface, which are connected by an internal bus respectively.
[0034] According to at least one embodiment, the control unit is composed of a CPU and a ROM, executes a program stored in the storage unit, and controls the server device. The control unit also includes an internal timer for measuring time. The RAM is a work area of the control unit. The storage unit is a storage area for storing programs and data. The control unit reads the program and data from the RAM and performs program execution processing based on information received from the user terminal, etc.
[0035] Disclose about AI. According to at least one embodiment, artificial intelligence includes machine learning, deep learning, generative AI, large language models, LLM, foundation models, generative AI. Generative AI uses transformers and employs a number of mechanisms called attention. It uses self-supervised learning, Extract Prediction. In this case, the AI can predict the next word. Given a sentence, it predicts the next word from the text up to that point. It creates a large number of supervised learning problems. As a result, an AI that can predict the next word can be created. Generative AI can predict grammatical structures, topic connections, and the likelihood that a person with such a style will write such a sentence. Furthermore, generative AI can learn the structure, causal relationships, and knowledge behind just predicting the next sentence. Generative AI has a high speed of scaling, and the larger the number of parameters, the higher the accuracy. Ordinary statistics and machine learning will overfit if the model parameters are made too large compared to the data sample size. LLM will have higher accuracy as the number of parameters is increased. One generative AI has 175 billion parameters. Generative AI is trained with supervised learning to have smooth conversations. It is taught not to say strange things. It writes reviews or acts as a call center operator.
[0036] According to at least one embodiment, a large language model (LLM) is, non-exhaustively, a natural language processing model of machine learning constructed using a large amount of dataset and deep learning technology. Generally, a method called "fine-tuning" for training on a specific task is used to adapt it to various natural language processing (NLP) tasks such as text classification / generation, sentiment analysis, text summarization, and question answering. According to at least one embodiment, self-supervised learning is close to human essential intelligence. When humans act, they always predict the next event that will occur and predict the next input. In that process, they can learn the structure of the external world. Predicting the next word is an essential intelligence and is considered to be close to what the cerebral cortex does. According to at least one embodiment, a large language model memorizes the input information but generalizes to the extent necessary to predict the next word. It does not generalize all the information from the beginning. A large language model requires capacity to memorize information. Also, parameters are required for that purpose. According to at least one embodiment, a large language model is equipped with 8 models with 175 billion parameters or 220 billion parameters.
[0037] According to at least one embodiment, videos and images are represented as a set of visual patches, which are small data units similar to the text tokens of an LLM. Patches can effectively represent the model of visual data and are used as a very scalable and effective representation for training generative models with various types of videos and images. First, the video is compressed into a low-dimensional latent space, and then the representation is decomposed into spatio-temporal patches to convert the video into patches.
[0038] According to at least one embodiment, a Video compression network is a network that reduces the dimension of visual data. It receives raw videos as input and outputs a temporally and spatially compressed latent representation. AI is trained in this compressed latent space and then generates videos within this compressed latent space.
[0039] According to at least one embodiment, given a compressed input video, Spacetime Latent Patches extracts a series of spatio-temporal patches that function as transformer tokens. With the patch-based representation, Sora can be trained on videos and images of various resolutions, lengths, and aspect ratios, and at inference time, randomly initialized patches are placed in a grid of appropriate size to control the size of the generated video.
[0040] According to at least one embodiment, the AI is a diffusion model and is trained to predict the original "clean" patches when noisy patches (and conditional information such as text prompts) are input. The AI is a diffusion transformer, which exhibits remarkable scaling properties in various areas such as language modeling, computer vision, and image generation. The diffusion transformer is also effective as a video generation model. As the computational cost of training increases, the quality of the samples improves significantly.
[0041] According to at least one embodiment, the AI applies caption regeneration technology to train a highly explanatory caption model and then uses it to generate text captions for all videos in the training set. Training highly explanatory captions improves not only the overall quality of the generated videos but also the faithfulness of the text. Utilize GPT to convert short user prompts into long detailed captions and send them to the model. This enables the AI to generate high-quality videos that exactly follow the user's prompts.
[0042] According to at least one embodiment, in natural language processing, vectorization can be performed by AI along the following process. First, cleaning processing of the text given as preprocessing is performed. In the cleaning processing, unnecessary words such as JavaScript code and HTML tags included in the text are deleted. Since these codes are codes used for display on the Internet, they are generally information not used in natural language processing. Subsequently, the text is split into word levels by morphological analysis. Morphological analysis is to classify into the smallest language units with meaning in a sentence of natural language written in characters. As morphological analysis tools, "MeCab", "JUMAN", and "JANOME" can be used. In normalization, words with the same meaning such as writing variations are unified into one word. Stop words are words that are not subject to processing for reasons such as being unusable in natural language processing. Examples of stop words include those that do not have meaning alone, such as auxiliary words and auxiliary verbs among words. When calculating vectors, these may be removed and only meaningful words may be targeted. Vectorization may also be performed without removing these stop words. Vectorization is a process of converting a word, which is a character string, into a vector. By vectorization, word data is converted into numerical data. When converting a word into a vector, it is carried out by a method called Bag of Words or distributed representation. Bag of Words is a method of vectorizing a text using the number of occurrences of words that appear in the given text. Since it focuses on how many words appear in the text, the order of words and texts is not considered. Distributed representation is a method of vectorizing by focusing on the meaning of a word. By vectorizing the meaning of a word, it is possible to give a vector close to words with similar meanings and usage, and the relationship between words can also be expressed by a vector. By the expression in a vector, addition and subtraction of the meanings of words are possible. The application process can utilize the natural language converted into numerical data as the input of machine learning. Specifically, the vectorized natural language is input into a classifier to perform text classification.Tools that can be used here include "TensorFlow", "scikit-learn", "PyTorch", etc.
[0043] Disclose a machine. In at least one embodiment, the machine may exist as a combination of one or more embodiments or functions in the present disclosure.
[0044] The machine includes a machine that can move itself or a machine equipped with a power to move itself. The machine may be equipped with a video acquisition device. The machine includes a robot. The machine has communication means and can communicate with other computers, computing facilities, the cloud, the Internet, applications, and information. As an example, the machine includes a machine that moves by at least one or more wheels, including a bicycle. A bicycle is a vehicle or machine that moves autonomously without a driver. Furthermore, the machine includes a machine that walks or runs with at least one or more movable parts. As an example, it is capable of bipedal walking. As an example, it includes a humanoid robot. A humanoid robot has at least two movable parts, and these movable parts correspond to feet. The machine includes a robot or a humanoid robot that combines artificial intelligence (AI) and robotics. These machines can utilize autonomous driving technology and machine learning algorithms to act autonomously while recognizing the surrounding situation. The machine is equipped with AI technology and sensors. Thereby, it can recognize the surrounding environment in real time and select appropriate actions. For example, it can move while avoiding obstacles or communicate with people around it. The machine includes a machine that flies using power. As an example, it includes a drone. An aircraft includes an airplane, a rotary-wing aircraft, a glider, a balloon, and other devices specified by government ordinance that can be used for aviation. The aircraft may be equipped with one or more propellers. A drone is equipped with a video acquisition device. A drone can take pictures. The video includes thermography.
[0045] A video acquisition device or means is disclosed. In at least one embodiment, the video acquisition device can be implemented by any of the machines and devices described in this disclosure. As a non-exhaustive example, the video acquisition device includes a device that captures video and converts it into digital data. Digital cameras (which capture still images and videos and store them on a recording medium), video cameras (which mainly capture videos), webcams (which are connected to a personal computer and used for video conferencing, live streaming, etc.), surveillance cameras (which are installed for security purposes and record video constantly), and smartphone cameras. "Video" in this disclosure is a term with a very broad meaning. Generally, it refers to an image formed by light, that is, visual information that can be captured by seeing with the eyes. That is, "video" includes thermography (a technology that visualizes invisible heat). Thermography uses an infrared camera to detect the infrared rays emitted from an object and captures the temperature distribution as an image by displaying its intensity in different colors. Naturally, the "video acquisition device" includes a device that acquires thermography.
[0046] A method for acquiring human motion data using motion capture is disclosed. In at least one embodiment, motion capture includes technologies for recording the movements of humans or objects as digital data. The optical method digitizes the three-dimensional positions and postures of markers attached to the subject using multiple infrared cameras. The inertial method measures the displacement and posture of inertial sensors attached to the subject. Markerless (video-based) reads the silhouette of the subject with a video camera and estimates the position of the bones. Humans can choose according to the purpose of the service provided. For example, a person performing a task. The task can be any physical movement regardless of its mode, including light work. For example, transporting luggage to a predetermined location, loading and unloading, cooking, farming, welding, performing a predetermined task on a factory line, and cleaning. This includes cases where two or more people cooperate to perform these tasks. In this case, it includes performing a certain task and delivering the resulting product to the next predetermined item.
[0047] Disclosed is a method for training motion data by AI. In at least one embodiment, actual human motion capture clips are collected. Next, reinforcement learning is used to train a control policy that mimics human motion. The policy is trained in a physical simulation to track the pose of the reference motion at each time step. Next, by using various reference motions in the reward function, the simulated robot can be trained to mimic various skills.
[0048] Disclosed is a method for generating adaptable motions according to environmental changes around a machine. In at least one embodiment, since the simulator generally provides only a rough approximation of the real world, the policy trained in the simulation may have degraded performance when deployed to an actual robot. Therefore, a high sample efficiency latent space adaptation technique is used to transfer the policy trained in the simulation to the real world. First, to encourage the policy to learn robust motions against changes in dynamics, physical quantities such as the mass and friction of the robot are changed to randomize the dynamics of the simulation. Since the values of these parameters can be accessed during training in the simulation, they can also be mapped to a low-dimensional representation using the learned encoder. This encoding is passed as an additional input to the policy during training. Since the physical parameters of the actual robot are not known in advance, when deploying the policy to the actual robot, the encoder is removed and a series of parameters that allow the robot to normally execute the target skill in the real world are directly searched in the latent space. This technique enables the adaptation of the policy to the real world using real-world data. In the above approach, the policy can be trained in simulation and adapted to the real world.
[0049] However, when the task involves complex and diverse physical phenomena, it is also necessary to learn directly from real-world experience. Develop an automatic learning system with software and hardware components using multi-task learning procedures, learners with safety constraints, and several carefully designed hardware and software components. Multi-task learning generates a learning schedule that directs the robot towards the center of the workspace, preventing the robot from leaving the training area. Also, design safety constraints to reduce the number of falls. This safety constraint is solved by the double gradient descent method. In each rollout, the scheduler selects a task where the desired walking direction is towards the center. For example, if there are two tasks, forward and backward, when the robot is at the back of the workspace, the forward task is selected, and vice versa for the backward task. During the episode, the learner executes the double gradient descent procedure to iteratively optimize rather than treating both the task objective and the safety constraint as a single goal. If the robot falls, the automatic stand-up controller is called to proceed to the next episode.
[0050] Obtain information on environmental changes around the machine and perform training using that variable. By this method, an operation adaptable to environmental changes can be generated. Obtain information regarding human work, and based on that information, an operation adaptable to environmental changes around the machine can be generated. For example, a factory line refers to a production method of flow work for processing and assembling products and parts flowing on a belt conveyor or the like. In an automobile factory line, the motion data of the worker is obtained using motion capture. The machine learns the motion data by AI. The machine performing the work includes any of the machines described in this specification. There may be a humanoid robot (hereinafter referred to as a "working machine"). The working machine responds to environmental changes around the machine. The working machine is given physical parameters regarding the environment. For example, it includes information on the shape of an article that is the work target flowing on the factory line, the coefficient of friction, parts that should not be touched, and combinations of one or more of these. The working machine uses the physical parameters to perform actions adapted to the real world. The physical parameters may be given to the working machine, or it may search for them by itself using sensors or the like. In this way, the working machine can autonomously take actions adapted to the actual environment while imitating the human behavior as a reference operation. According to this embodiment, there are the convenience and industrial applicability that at least automation applicable based on human actions can be realized.
[0051] When generating an operation, a method for recognizing the environment around a machine based on information obtained from one or more sensors and changing the operation based on the obtained information is disclosed. In at least one embodiment, the sensors include those that can acquire, in data form, variables or physical parameters regarding the environment around the machine, regardless of their type. As an example, it includes LiDAR, cameras, force sensors, and combinations of one or more of these. LiDAR irradiates laser light and measures the distance to an object, the shape of the object, etc. based on the information of the reflected light. LiDAR (lidar) can grasp the distance, shape, and positional relationship of a preceding vehicle, pedestrians, buildings, etc. in three dimensions. The machine searches for or acquires physical parameters by itself. For example, by using a camera or the like, understand the speed of the work target article flowing on the factory line, calculate how many seconds the operation must be performed within that speed, and increase the driving speed of the machine so that the work can be completed within that time. As a result of the inventor's earnest consideration, it has been found that by changing the operation based on the obtained information, the quality of the work of the working machine is improved. In particular, in the case of LiDAR, extremely accurate physical parameters regarding the shape and distance of the object can be obtained. Therefore, it is difficult for noise to occur in learning and calculations based on the physical parameters, and as a result, the actions based on the results of the learning are of high quality. This effect was not known at all in the prior art and has novelty and remarkable effects. According to this embodiment, there are the convenience and industrial applicability that at least automation applicable based on human actions can be realized.
[0052] Disclosed is a method for acquiring motion data of human motion using an image, learning the motion data by reinforcement learning, and generating an adaptive motion for a motion similar to or different from the motion. In at least one embodiment, motion data of human motion is acquired using an image. For example, there is an image acquisition device above a person's head, and the image acquisition device captures the person's motion. The person's motion may be any physical motion regardless of its form, including speech and gestures. There is an image acquisition device on the side of the person, including acquiring an image of the way of walking while carrying luggage. There is an image acquisition device on the ceiling of a large warehouse, including acquiring an image of the path a person walks. The machine performs region estimation. Based on the image, when a feature point of the person's body enters a specified region, the person's motion is detected. For example, it is detected that the wrist has entered the region for picking up a part. The machine performs pose estimation. The feature points of the person's body are learned to detect the person's pose. For example, it is detected how the person takes out and assembles a tool. The machine performs background estimation. The feature image in the background is learned to classify the background. For example, when working, it is detected whether a driver is being used or not. The machine performs object detection. An image that matches the learned image is detected from a set region in the image. For example, it is detected whether a predetermined tool is being used. The machine learns the motion data by reinforcement learning.
[0053] Reinforcement learning maximizes the score by feeding back the score for the output. In reinforcement learning, pre-prepared training data such as supervised learning and unsupervised learning is not used. Two functions, an agent and an environment, are utilized. The agent is the AI model to be developed. The agent receives the "state" from the environment as an input. In many cases, a simulator is used for the environment, and the agent receiving the input from the environment is sometimes called "observation". The agent that has observed the environment returns an output corresponding to it or a random output (takes an action). What kind of output to produce and what kind of operation or action to take vary depending on whether the development target is robot control. Basically, they are all numerical values or codes output by a function, but the values and meanings vary depending on whether the AI to be developed is a block-breaking game, a Go or Shogi program. The environment receives the output of the agent and returns a new state as its reaction. At this time, a "reward" is also given to the agent depending on the state. The reward (Reward) is also called a score and is numerical information. The agent observes the environment, takes some action, and selects the action that maximizes the reward. The relationship between the evaluation of the action and the reward (and sometimes a penalty) is considered by humans and given to the AI as parameters. This becomes an important tuning point for determining functions and systems in reinforcement learning. In reinforcement learning, through this repetition, the agent is taught what actions to take to effectively achieve the task. Instead of humans programming the control of the servo motors for a robot to walk without falling or the operation of the paddle to eliminate many blocks without dropping the ball in a block-breaking game, the machine (agent) is allowed to try and error. For example, for walking, if moving the motor that extends one foot forward causes loss of balance and falling, no reward is obtained or it becomes negative. Next, move another motor. If that control that shifts the center of gravity of the pivot foot does not cause a fall, the reward increases, so that operation is incorporated into the control as an effective one. The basis of the processing in reinforcement learning is an algorithm or method for observing the environment and selecting / determining the next action. The amount of information given by the environment is diverse. In the case of the Othello game, the amount of information is small immediately after the start.As the board progresses, the number, position, and arrangement of stones become more diverse. At this time, mathematical methods such as regression analysis and probability theory, or evaluation by a neural network, are used to grasp and evaluate the board state. Reinforcement learning can also be regarded as a "function". An agent is a function that receives values (states) from the environment as inputs and outputs the next action. The environment can also be regarded as a function that receives the agent's action as an input and returns a new state. Different from AI that learns using static data (supervised learning and unsupervised learning), reinforcement learning can also be said to be a dynamic learning that adjusts the next process according to the output result (reward) of the function. The reward can be said to be "an evaluation of the environment's reaction to the agent's random action". Maximizing the reward means comparing the state change (this is called "value") when the agent does not take a random action with the state change (value) when taking a random action, and adopting the one with a higher value as a "policy (a series of successful actions)". Calculating the value is done by two functions, the state value function and the action value function. The state value function calculates the value when not taking a random action. The action value function calculates the value when taking a random action. As a result, when the action value function is higher, the random action is incorporated into the policy as the correct action. As a result, the AI can learn the moves that lead to victory for each board state and the control method for the motors for the robot to walk.
[0054] The machine learns operation data through reinforcement learning. The machine learns the optimal actions through its own trial and error. For example, in a logistics warehouse, it learns about the task of transporting a predetermined item to a predetermined location. It performs reinforcement learning on how to hold the item, how to move to minimize the required time, etc. As a result, it generates actions adaptable to actions similar to or different from human actions. According to this embodiment, there are the convenience and industrial applicability of being able to achieve at least more efficient automation than when humans work.
[0055] The machine further discloses a method of recognizing a feedback action of a human gesture, voice, gaze, or a combination of one or more of these when performing the generated action, evaluating the action generated based on the feedback action, and changing the action based on the evaluation. In at least one embodiment, the machine acquires an image of a human around the machine by an image acquisition device. The machine acquires sound by a sound acquisition device. The sound to be acquired may be around the machine or may be the sound of the location around the device by a sound acquisition device physically separated from the machine. The machine recognizes the feature amount of a human face from the image and recognizes the human face. Further, the feature amount of the expression of the human face is recognized, and the gaze of the human is estimated or recognized from the image of the human eyes. Further, the gesture of the person is estimated from the image of the person. A gesture includes body movements and hand gestures, but is not limited to the hands, and may be any physical movement regardless of its form. For the respective actions, a value serving as a reward is set. For example, a high reward is set for the action of making a circle with the hand, and a low reward is set for the action of making a cross with the hand. In the case of sound, a high reward is set for the sound "OK", and a low reward is set for the sound "no". In the case of gaze, a high reward is set for the gaze directed towards the machine, and a low reward is set for the gaze not directed towards the machine. In such a manner, the machine can recognize the feedback action. The machine evaluates the action generated based on the feedback action, and by performing reinforcement learning with reference to the reward, changes the action based on the evaluation. For example, the machine performs a predetermined operation on the factory line. A human looks at the action and gives an "OK" or makes a cross gesture. Then, the machine recognizes the feedback action, evaluates the action generated based on the feedback action, and changes the action based on the evaluation. As a result of the inventor's earnest consideration, the results of self-learning or trial and error of only the machine are not necessarily the best results considered by humans. The results of learning different from the human intention may be beneficial to humans in some cases, but may also produce conflicting results. Therefore, by a human giving feedback on the machine's behavior and incorporating the information as a reward into reinforcement learning, it is possible to produce the best learning result in line with the human intention.This effect was not known at all in the prior art and has novelty and remarkable effects.
[0056] The machine further discloses a method of acquiring an execution video of the generated operation, learning by AI based on the execution video, and changing the operation based on the acquired information. In at least one embodiment, the video acquisition device includes a case where it is provided in the machine and a case where it exists independently of the machine. The video acquisition device acquires an execution video of the operation executed by the machine. This video includes still images and moving images. In the case of moving images, it includes a series of moving images from before the operation is generated to after it is generated. The machine learns by AI based on the video. For example, when the video is about a humanoid robot and the work at the robot's hand, it includes acquiring the video with the video acquisition device of the machine. When the video is about the robot moving inside a large warehouse or an operation using the whole or part of the robot, depending on the operation, the video is acquired with a video acquisition device different from the robot. This is because it may be possible to acquire data with a large amount of information about the overall movement of the robot. The learning method includes, firstly, reinforcement learning. In the case of reinforcement learning, the machine autonomously evaluates the operation appearing in the video, sets a reward value, and based on that reward value, learns whether to imitate the operation, not to imitate it, and in which part of the overall operation to imitate / not to imitate. For example, if it has fallen, the reward value of that operation is low. Secondly, there is supervised learning or unsupervised learning. Supervised learning processes learning data with correct answers and outputs the correct answers. Unsupervised learning extracts patterns and features of input data by computational processing. The user of the machine gives information on whether the video is correct or not. For example, the user gives a correct answer to the video judged to be the best among the videos of a number of machines. In addition, for prohibited actions that should not be performed in the video, an incorrect answer is given to the video in which such an action was performed. The machine learns the actions related to the correct video based on this information. In the case of unsupervised learning, clustering analysis (clustering) and dimensionality reduction of videos and the like are performed. Then, the user gives information on whether the cluster is correct or not. The machine learns based on this information.
[0057] Disclosed is learning without a teacher. The "cluster" in cluster analysis means "cluster" or "lump" in English, and refers to a state where things with similar characteristics are gathered. That is, cluster analysis is an analysis method for grouping similar things (things with similar feature amounts) from a large amount of data into several groups. Creating a group of data with similar characteristics by this method is called "clustering". Clustering is a classification in a state where there is no correct answer data. Classification can be performed, but the meaning of each group may not be clear. The result may need to be interpreted by humans. The types of cluster analysis include "hierarchical clustering" used when the number of classification targets is small and "non-hierarchical clustering" applied when there are a large number of classification targets. Hierarchical clustering hierarchically groups data with similar features one by one in the order of clusters, and repeats until finally becoming one large cluster. Since the process is visualized in a diagram like a tournament table (tree diagram), it is an analysis method that makes it easy to grasp the characteristics of the data. Also, non-hierarchical clustering does not have a hierarchical structure. It is only necessary to set in advance how many clusters to divide into, and the data is divided according to the number of those clusters. There is also a method in which the machine automatically divides without determining the number.
[0058] Learning is not limited to only the machine that performed the operation. The machine acquires an execution video of the generated operation, transmits the video to a machine other than the machine that performed the operation, and a machine other than the machine that performed the operation learns by AI based on the execution video and changes the operation based on the acquired information. As a result of the inventor's sincere consideration, in the conventional method, since each robot learns machine learning independently, data obtained by the machine's trial and error behavior cannot be learned by other robots. Therefore, the collective intelligence of the robots could not be improved at the shortest speed. According to this method, when a plurality of machines perform work, other machines can also learn the learning materials in a certain machine, so the intelligence of the entire machine is improved synergistically. This effect was not known at all in the conventional technology and has novelty and remarkable effects.
[0059] It can also be implemented as a program for operating the machine described in any of this specification.
[0060] Disclosed is a method for changing the operation of a machine based on a received request or learning data, comprising communication means for transmitting and receiving information with a server or other machine. In at least one embodiment, the machine comprises communication means for transmitting and receiving information with a server or other machine. As described above, the machine can acquire images or the like from another machine or a server and learn based on the images or the like. As a result of the learning, the operation of the machine can be changed. The machine can change its operation based on a request from another machine. For example, this is the case when two or more machines cooperate to perform a task. In a logistics facility, Machine A transports an article to a predetermined position and delivers the article to Machine B. In this case, Machine A requests Machine B to receive the article. The method of request includes a method of transmitting a request or a signal by an electromagnetic method. Machine B receives the request and receives the article. As a result of the inventor's earnest consideration, it has been found that, by the method of changing the operation of a machine based on a received request or learning data, the operation of cooperation between two or more machines can also be learned and the behavior can be improved. This effect was not known at all in the prior art and has novelty and remarkable effects.
[0061] A method for avoiding collisions with humans, obstacles, or machines other than itself is disclosed. In at least one embodiment, the machine uses the video acquired by a video acquisition device to detect an obstacle or a machine other than itself (hereinafter referred to as "obstacle, etc.") in the traveling direction of the machine. When there is an obstacle, etc., the machine changes its route or temporarily stops moving until the obstacle, etc. disappears. In another embodiment, the machine has a device for transmitting its own position information. This device includes GPS. The machine transmits its own position information to other machines or a server. The machine receives the position information of other machines or the position information of obstacles, etc. from other machines or the server. The position information of obstacles, etc. includes information transmission from the machine that detected its existence or a method by which the user registers the fact. Based on the received position information of other machines, the machine changes its route to avoid collisions with obstacles, etc. or temporarily stops moving until the obstacles, etc. disappear. According to this embodiment, there are the convenience and industrial applicability that the work in which at least two or more machines cooperate can be automated.
[0062] A method with control means for cooperating with machines or humans other than itself is disclosed. In at least one embodiment, the method described anywhere in this specification is adopted. The machine acquires the video of the humans around the machine by a video acquisition device. When there are humans around, the machine changes its route so as not to collide with the humans or temporarily stops moving until the obstacles, etc. disappear. The video is analyzed, and based on the video, the work being done by the humans is taken over. For example, if a human is carrying an article and the machine takes over the article. According to this embodiment, there are the convenience and industrial applicability that at least the machine and the human can cooperate in work. There are the convenience and industrial applicability that at least the work in which the machine and the human cooperate can be automated.
[0063] Disclosed is a method for equipping with energy efficiency means for improving its own energy consumption, planning its own operation based on calculations by the efficiency means, or performing control or charging to minimize the motor output or operating speed according to the work content. In at least one embodiment, the machine sets a reward regarding its own energy consumption amount during reinforcement learning. The machine performs an operation in a certain work and obtains the energy consumption amount regarding the operation. The power measurement means can measure the consumed power regardless of the means. It includes a power meter, a watt checker, an energy monitor. The power measurement means includes a method mounted on the machine itself and a method not mounted on the machine itself but mounted on a supply device that supplies power to the machine. The machine identifies the amount of power consumed during its own operation. The calculation can be obtained by calculating the difference in power consumption before and after the machine performs a certain operation. The supply device obtains the power consumption amount related to the operation for each machine and each operation. The machine performs reinforcement learning with a reward based on the power consumption amount. The lower the power consumption amount, the higher the reward. The machine will take actions with less power consumption through reinforcement learning. As a result, the same work can be carried out with less power consumption. According to this embodiment, there are convenience and industrial applicability in reducing at least the power consumption related to the work of the machine or realizing the efficiency improvement of the power consumption amount.
[0064] Disclosed is a method for transmitting its own learning data to a machine other than itself. In at least one embodiment, the method described in any of this specification is employed. The machine transmits its own learning data to a machine or server other than itself. The machine receives learning data related to a machine other than itself from a machine or server other than itself. The machine changes its operation based on the received learning data or learns using the learning data. By this method, the machine can inherit the learning data of other machines, so there are convenience and industrial applicability in improving the learning efficiency.
[0065] The content disclosed in this specification can also be configured as a system.
[0066] Disclosed is agriculture. In at least one embodiment, agriculture, regardless of its form, is an organic production industry that uses land to cultivate crops, raises livestock to produce materials necessary for food, clothing, and shelter, utilizes the power of the land to cultivate useful plants, raises useful animals, engages in agricultural processing, forestry, or any combination of these.
[0067] Disclosed is a method for recognizing crops based on information obtained from one or more sensors. In at least one embodiment, a sensor, regardless of its form, includes a machine that acquires information around the device as some variable. The sensor includes an image acquisition device. The machine includes a humanoid robot. The working machine may be equipped with a sensor itself, or the sensor may be separate from the working machine. For example, separately from the working machine, the sensor may be provided on a farm, another working machine may be equipped with the sensor, the information obtained from the sensor may be transmitted, and a machine without a sensor may receive it. The machine analyzes images. The machine can recognize crops by background estimation or image analysis. For example, the machine refers to an image of a plant without tomatoes and detects the difference from the acquired image. The machine detects the red video as the difference. The machine stores or searches for data and searches for what the red video is identical or similar to. If it is similar to tomatoes, it is recognized as tomatoes, and if it is similar to strawberries, it is recognized as strawberries.
[0068] A method for recognizing the environment around a machine is disclosed. In at least one embodiment, the methods described anywhere in this specification are employed. The sensors include those that can acquire data on variables or physical parameters regarding the environment around the machine, regardless of their type. As an example, it includes LiDAR, cameras, force sensors, and combinations of one or more of these. The sensors include hygrometers and soil moisture meters. A soil moisture meter is a device that measures the amount of moisture contained in the soil. As a result of the inventors' earnest consideration, in a farm, the environment around the machine varies daily depending on the season and weather. Therefore, when the machine moves within the farm under the same conditions, it cannot be applied to the different environments. For example, in the case of a humanoid robot, it will fall on muddy ground. In a farm, the strength of the footing can be predicted by machines that measure moisture, such as hygrometers and soil moisture meters. The machine recognizes the possibility that the ground under its feet is muddy based on the information from the soil moisture meter. This effect was not known at all in the prior art and has novelty and remarkable effects.
[0069] Disclosed is a method for an AI to determine whether an agricultural operation can be performed based on crop information and environmental information. In at least one embodiment, the machine provides crop information to the AI. The crop information includes images of the crops. Further, the machine may be equipped with a refractometer and a hardness meter. The machine recognizes the crops around it, uses these machines on the crops, and obtains information on their sugar content, hardness, and combinations thereof. The refractometer may be any type as long as it can measure the sugar content. It includes a refractometer and non-destructive sugar content measurement. The refractometer calculates the sugar content by measuring the refractive index of light traveling straight through water or air. The non-destructive sugar meter irradiates near-infrared light and measures it with a sensor. It utilizes the property that sugar is easily absorbed by light of a specific wavelength and measures without damaging the crop. The hardness meter may be any type as long as it can measure the hardness. A durometer presses against a sample and reads the indicated value. There is also a method of pressing a needle against the object to be measured. As a result of the inventor's earnest consideration, if any of the refractometer, hardness meter, or combinations thereof are used, the machine can sufficiently determine whether the crop is ripe. In the prior art, even when it was known that the color of the crop had changed, the machine could not determine its texture, taste, or softness. A machine equipped with a refractometer and a hardness meter can automatically determine the ripeness of the crop. This effect was not known at all in the prior art and has novelty and remarkable effects. The environmental information includes any information related to meteorology, temperature, humidity, environmental images, and combinations of one or more of these. For example, the temperature has reached a certain level or above or below, the humidity has reached a certain level or above or below, a typhoon is approaching, it is raining, images of plants (e.g., the plant has turned brown and is withering), and information on combinations of these. Such information is related to the maturity of the crop or takes precedence over the information on the maturity of the crop and can be information for determining whether the crop should be harvested. Whether the agricultural operation can be performed may be any agricultural operation regardless of its form. For example, it includes harvesting the crop, not harvesting it, applying fertilizer, watering, weeding, thinning, and combinations of one or more of these. The AI includes AI learned by supervised learning, reinforcement learning, unsupervised learning, and combinations of one or more of these.For example, learn whether to harvest for the combinations of information disclosed in this specification. For example, the condition for harvesting can be that the sugar content exceeds a predetermined value and the strength is below a predetermined value. Also, if the sugar content is 10% below the predetermined value and the strength is 10% above the predetermined value, but according to the image of the plant, the plant is withering, this can be a condition for harvesting. The user of the machine can set the conditions for harvesting and perform supervised learning. Even without being given the correct answer, the machine learns the conditions for harvesting based on the data sent from the machine daily, with the reward being to harvest under the condition of the highest sugar content. As a result of the inventor's earnest consideration, by combining a machine that acquires data in this way with an AI that learns, agricultural work can be correctly and efficiently automated. This effect was not known at all in the prior art and has novelty and remarkable effects.
[0070] When it is determined that the AI is operable, disclose a method for generating the operation. In at least one embodiment, the machine generates an operation based on the determination of the AI. For example, when the AI determines that a crop should be harvested, the machine harvests that crop.
[0071] The machine includes a humanoid robot. As a result of the inventor's earnest consideration, in agriculture, a humanoid robot may be particularly useful. As a result of the inventor's earnest consideration, it was found that when the machine is a humanoid robot, the convenience is greatly improved. Since the farm is designed to suit a two-legged walking human, a bicycle or a drone may not be able to act satisfactorily. For example, it is difficult for a bicycle to reach a terraced field, and it is difficult for a drone to fly into a place where branches are intricate or near the ground. It was found that a two-legged walking robot is optimal as a configuration of a machine that enters such places and all farms to harvest crops. That is, since a humanoid robot has a movable range and size similar to or the same as that of a human, it can generally surely reach plants that a human can harvest, and it is the most excellent in versatility. Therefore, one of the most preferable forms for providing services is a humanoid robot. This effect was not known at all in the prior art and has novelty and remarkable effects.
[0072] Disclosed is a method for learning operation data related to past generated operations through reinforcement learning and generating adaptive operations for crops targeted by the past operations or operations or environments similar to or different from the environments that generated the past operations. The method described anywhere in this specification is incorporated. The machine learns operation data related to past generated operations through reinforcement learning. For example, when the ground is muddy, the variables of the ground strength and friction are different. Suppose the machine loses balance and falls in such a case of different variables. No reward is obtained or it becomes negative. Next, if another motor is operated and it does not fall in the control that shifts the center of gravity of the leg, the reward increases, so that operation is incorporated into the control as an effective one. In this way, adaptive operations are generated for operations or environments similar to or different from the environments that generated the past operations. Suppose the machine tries to harvest oranges and accidentally crushes them. In this case, no reward is obtained or it becomes negative. Next, if the strength of the motor is changed and the oranges can be harvested without being crushed, the reward increases, so that operation is incorporated into the control as an effective one. Next year, when harvesting oranges again, if the oranges are crushed in the past way, further reinforcement learning is performed to harvest the soft oranges of that year without crushing them. In this way, adaptive operations are generated. According to this embodiment, there are at least the convenience and industrial applicability that at least agricultural operations can be automated with high quality.
[0073] Disclosed is a method in which two or more of the above-described machines are present in the same field, and the machines cooperate in their operations by communicating the operation or position of one machine in the field. In at least one embodiment, the method described anywhere in this specification is employed. The machine communicates with other machines. When another machine is harvesting a row of crops, the machine harvests a different row of crops. "Two or more machines cooperate in their operations" can also be replaced with "two or more machines are present in the same field, and the operation of the machine is changed based on the information obtained by communicating the operation or position of one machine in the field." That is, if the machine changes its actions based on the information obtained from other machines, it can naturally take cooperative actions. For example, when harvesting a row of crops, when a machine that harvests crops from the front of the same row approaches, the machine changes its actions to move to another row so as not to collide with that machine. According to this embodiment, there is the convenience and industrial applicability that farming work can be efficiently performed using at least two or more machines.
[0074] Disclosed is a method of determining a movement route within a field based on information about the environment around the machine and moving while avoiding obstacles. In at least one embodiment, the method described anywhere in this specification is employed. The obstacles include other machines. The information about the environment around the machine may include fallen trees and puddles. This information is generated from an image acquisition device of the machine, an image acquisition device provided on the farm, or an image acquisition device of another machine. The image acquisition device of another machine includes, for example, a mode in which a drone acquires an image of the farm. The machine recognizes obstacles through background estimation and image analysis. The machine determines a movement route within the field based on the information about the obstacles. The machine identifies a route that does not pass through the obstacles in order to go around all the places to walk within the field. As a result of the inventor's earnest consideration, a system combining a humanoid robot that performs work and a drone that acquires images is particularly effective. That is, farming work cannot be completed solely by a drone, and a humanoid robot cannot sufficiently obtain information about the farm environment. By combining the two, it becomes possible to automatically generate actions adapted to the environment. This effect was not known at all in the prior art and has novelty and remarkable effects.
[0075] Disclose a method for a machine to be a walking robot. Incorporate the method described anywhere in this specification.
[0076] The embodiments described in this specification can also be implemented as a system.
[0077] The embodiments described in this specification can also be implemented as a program.
[0078] Disclose a method for determining the growth state of crops or the presence or absence of diseases by image recognition, and using AI to determine whether to generate operations such as harvesting, fertilizing, pest control, or a combination of one or more of these according to the determination. In at least one embodiment, the machine determines the growth state of the crops. For example, a video acquisition device acquires variables of the pigments of the crops. If the crops are withering, this pigment changes. If the crops are healthy, the pigment changes. Diseases can be determined by comparison with an image of a normal plant. If Japanese beetles are reflected, it can be known that there are pests. The machine causes these information to be determined by AI. The AI performs supervised learning by being labeled as correct or incorrect for the actions taken in the past in such a situation. Or perform reinforcement learning. The AI determines whether to perform operations such as harvesting, fertilizing, pest control, or a combination of one or more of these. The AI includes an AI that has been supervised or reinforcement learned for the results of combining two or more of these. As a result of the inventor's earnest consideration, as a special circumstance in agriculture, the combination of harvesting, fertilizing, and pest control changes the yield and quality of the crops. Therefore, learning based on two or more variables is effective. This effect was not known at all in the prior art and has novelty and remarkable effects.
[0079] Disclosed is a method for obtaining meteorological data and using AI to identify the operations to be generated based on meteorological information, environmental information around the machine, and crop information. In at least one embodiment, meteorological data includes data representing the state of the Earth regardless of its form. It includes data representing phenomena observed in the atmosphere and ocean, such as rain, clouds, wind, waves, etc. The current state of the Earth is being observed using various sensors and machines from land, sea, and space around the world. And starting from the current observation data, the future state of the Earth can be predicted by calculating the time change of the Earth using a supercomputer according to physical laws. This includes these information and the information processed based on this information. The meteorological data includes meteorological data covering two or more days. The meteorological data includes meteorological prediction data for the day after tomorrow. The machine may obtain meteorological data by itself, or may receive meteorological data transmitted from the outside. As a result of the inventor's sincere consideration, as agricultural-specific circumstances, meteorological information, environmental information around the machine, and crop information interact with each other. Even if it can be determined based on the environmental information around the machine and crop information that fertilization should be carried out, there may be a case where it should be determined based on meteorological information that fertilization should not be carried out. Conversely, even if it can be determined based on crop information and meteorological information that harvesting should be carried out, there may be a case where it should be determined based on the environmental information around the machine that harvesting should not be carried out. In other words, even if machine learning or reinforcement learning is performed using only one or two of these, the behavior of the machine may fail due to the presence of other variables, resulting in poor learning efficiency. Thus, by comprehensively learning these three variables, the machine can automatically perform appropriate agricultural actions. This effect was not known at all in the prior art and has novelty and remarkable effects.
[0080] Disclosed is a method for changing the operation of a machine based on the information of a farm acquired by a drone. In at least one embodiment, the method described anywhere in this specification is adopted. For example, the drone can acquire weather information, environmental information around the machine, crop information, or one or more combinations thereof. The drone is equipped with a hygrometer to measure the humidity of the atmosphere, recognizes the presence of fallen trees in the farm using an image acquisition device, and approaches the crop to measure the sugar content using a refractometer. Based on these variables, the machine changes its behavior. According to this embodiment, there are the convenience and industrial applicability of optimally automating agriculture by emphasizing at least machines with multiple different attributes.
[0081] Disclosed is a method for a machine to communicate and connect with a terminal other than the machine and change the operation of the machine based on a request from the terminal and a judgment by AI. In at least one embodiment, the method described anywhere in this specification is adopted.
[0082] Disclosed is a method for a machine equipped with lifting means to recognize the environment around the machine and approach a crop based on crop information and environmental information. In at least one embodiment, the lifting means includes a mechanical mechanism that raises its own coordinates to a higher position regardless of its form. This includes a stair lift and a platform lift. For example, it includes an elevator. The machine recognizes that the crop to be harvested is above the machine based on the information obtained by the image acquisition device. The machine recognizes that the crop is not located within its movable range. In that case, the machine uses the lifting means to harvest the crop. For example, if the ground is not muddy based on the environmental information, a humanoid robot uses its own lifting means to harvest. If the ground is muddy, it communicates with a drone, which is another machine, to harvest the crop. The robot is equipped with an arm and a cutter, and harvests the crop by grasping the crop with the arm and cutting the stem with the cutter. Since the robot is equipped with image acquisition means, it can distinguish and recognize the crop and the stem. According to this embodiment, there are the convenience and industrial applicability of efficiently harvesting at least crops at high positions.
[0083] Disclosed is a method of issuing an audio alert that warns that a machine is in operation when the approach of a person is detected by a sensor. In at least one embodiment, the machine detects the approach of a person by a sensor or an image acquisition device. When a person approaches to a predetermined distance, the machine issues an audio alert that warns that the machine is in operation by an audio output unit. The audio alert may be any sound regardless of its form. Any sound is sufficient because a person can determine that it is a warning. It may be a stored voice of "Working!", or simply a buzzer. As a result of the inventor's sincere consideration, due to the specific circumstances of the field, the view is blocked by plants, so there is a risk of collision when a person and a machine work intensively. Furthermore, since the machine sprays blades during harvesting and weeding, and poisonous chemicals during pest control, the risk is particularly high. In this case, it may be preferable that the machine itself performing the work has a function of issuing an alert. In other embodiments, even if it is not the audio output unit provided in the machine itself performing the work, an alert is issued from the audio output unit of another machine. For example, the earphone worn by a person may be the audio output unit, or the audio output unit provided in the field may be used. This effect was not known at all in the prior art and has novelty and remarkable effects.
[0084] The following discloses an outline of the embodiments described above.
[0085] A machine related to agriculture, recognizes crops based on information obtained from one or more sensors, recognizes the environment around the machine, judges whether an agricultural operation can be performed by AI based on crop information and environmental information, machine
[0086] The above machine, when AI determines that it is operable, generates the operation, machine
[0087] The above machine, Learn the operation data related to the past generated operations through reinforcement learning, and generate an adaptive operation for a crop that was the target of the past operation or an operation or environment that is similar to or different from the environment in which the past operation was generated. Machine.
[0088] An agricultural system, where two or more of the above machines exist in the same field, and two or more machines cooperate to generate operations by communicating the operations or positions of one machine in the field. System
[0089] The above machine, determines a movement path within the field based on information about the environment around the machine, and moves while avoiding obstacles. Machine
[0090] The above machine, where the machine is a walking robot. Machine
[0091] An agricultural system, where based on the information obtained by the above machine, the AI determines whether an operation is possible, and a working machine that performs each operation receives communication from the machine and executes the operations determined to be possible. System.
[0092] A program for moving the above machine.
[0093] An agricultural machine, which determines the growth state of a crop or the presence or absence of a disease by image recognition, and uses AI to determine whether it is necessary to generate operations such as harvesting, fertilizing, pest control, or a combination of one or more of these in response to the determination. Machine.
[0094] An agricultural machine, which acquires weather data, Based on meteorological information, environmental information around the machine, and crop information, the AI identifies the operation to be generated. Machine.
[0095] An agricultural system, Based on the information of the field acquired by the drone, Changes the operation of the above-mentioned machine. System.
[0096] The above-mentioned system, The machine communicates and connects with terminals other than the machine itself, Based on the request from the terminal and the judgment by the AI, changes the operation of the machine. Machine.
[0097] An agricultural machine, Equipped with lifting means, Recognizes the environment around the machine, Approaches the crop based on the information of the crop and the information of the environment. Machine.
[0098] The above-mentioned machine, When a human approach is detected by the sensor, The voice output unit issues a voice alert warning that it is in operation. Machine.
[0099] The invention according to the present disclosure only needs to be able to achieve at least one of the above-described effects.
Claims
1. A machine relating to agriculture, Recognizing crops based on information from one or more sensors; Recognize the environment around the machine, Based on crop and environmental information, AI will determine whether or not agricultural work can be performed. The machine is a humanoid robot, Acquire human agricultural behavior data; AI learns motion data, Generate actions based on learned data, crop information, and environmental information. machine
2. An agricultural machine, Obtain weather information, Based on weather information, environmental information around the machine, and crop information, AI determines the action to be taken. The machine is a humanoid robot, Acquire human agricultural behavior data; AI learns motion data, Generate actions based on the learned data, meteorological information, environmental information, and crop information. machine
3. An agricultural machine, The machine is a humanoid robot, Obtain crop information, Obtain environmental information, Obtaining weather data, We require the AI to learn based on the above three pieces of information, Generate actions based on the learned data and the above three pieces of information. machine
4. An agricultural machine, Recognize the environment around the plant or machine based on information obtained from one or more sensors Based on crop information or environmental information, AI will determine whether or not agricultural work can be performed. The machine is a humanoid robot whose movements are generated by AI. machine
5. A machine as claimed in any one of claims 1 to 4, further comprising: When the generated action is performed, the system recognizes human feedback actions and Modifying generated actions based on feedback behavior; machine
6. A machine as claimed in any one of claims 1 to 4, further comprising: Acquire an execution video of the generated action, The AI learns from the execution video and changes the behavior based on the learned information. machine
7. A machine as claimed in any one of claims 1 to 4, further comprising: Acquire execution video of another machine that executed the generated operation, The AI learns from the execution video and changes the behavior based on the learned information. machine
8. A machine as claimed in any one of claims 1 to 4, further comprising: Plan actions to make your energy consumption more efficient, Generate actions based on the plan, machine
9. A machine as claimed in any one of claims 1 to 4, further comprising: Sending its own learning data to machines other than itself; machine
10. A machine according to any one of claims 1 to 4, The environmental information includes information about soil moisture; Generating walking movements according to soil moisture information; machine
11. A machine according to any one of claims 1 to 4, The crop information includes information regarding the sugar content of the crop, Generate actions depending on sugar content information; machine
12. A machine according to any one of claims 1 to 4, The crop information includes information regarding crop intensity; Generate actions according to the intensity information; machine
13. A machine as claimed in any one of claims 1 to 4, The crop information includes information regarding the sugar content or strength of the crop; The environmental information includes information on the climate, Generate an action according to the sugar content or strength of the crop and environmental information; machine
14. A machine as claimed in any one of claims 1 to 4, The crop information includes information regarding the sugar content or strength of the crop; The environmental information includes information on the climate, Prioritize environmental information over information about the sugar content or strength of crops, and generate appropriate behavior. machine
15. A machine according to any one of claims 1 to 4, The crop information includes information regarding crop intensity; Varying the strength of the harvesting action according to the strength information; machine
16. A machine as claimed in any one of claims 1 to 4, The environmental information includes information about soil moisture; When soil moisture information exceeds a certain value, it instructs a flying machine to harvest. machine
Citation Information
Patent Citations
Self-propelled type device, control method for self-propelled type device, and control program for self-propelled type device
JP2015225507A
Information processing system and program
JP2019082765A
Harvesting work system
JP2020018255A
Route management system and its management method
JP2022519465A
Method and device for multi-machine collaborative farming
JP2022548503A
Cited By
Spatial information integrated platform system, information processing method, and program
JP7849936B1