Support apparatus and support method
The system addresses the limitation of text-only input by using a generation AI, neural network, and knowledge database to refine image-based outputs, achieving enhanced support through iterative learning and human feedback.
Patent Information
- Application Number
- JP2024084743
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-05-24
- Publication Date
- 2025-12-05
AI Technical Summary
Existing systems, such as those described in Non-Patent Document 1, are limited to text input and struggle to process image inputs, making it difficult to extract knowledge from images in real-world scenarios like construction sites.
A system comprising a generation AI, neural network with reinforcement learning, and a knowledge and logical structure database that processes image inputs, iteratively refines outputs to approach a normal distribution, and selects the best result using human feedback.
Enables the generation of optimized outputs by combining inductive learning from observed events with deductive reasoning based on knowledge and logical structures, providing the best possible support.
Smart Images

Figure 2025177692000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to an assistance device and an assistance method. [Background technology]
[0002] For example, Non-Patent Document 1 RLHF (Reinforcement Learning with Human Feedback) describes that after reinforcement learning of a sentence, human knowledge is fed back. [Prior art documents] [Non-patent literature]
[0003] [Non-Patent Document 1] A rough guide to RLHF (reinforcement learning with human feedback) https: / / www.brainpad.co.jp / doors / contents / 01_tech_2023-05-31-160719 / Summary of the Invention [Problem to be solved by the invention]
[0004] However, the system of Non-Patent Document 1 allows text input but does not allow image input, so it is difficult to take images at a construction site, for example, and easily obtain knowledge on the spot. [Means for solving the problem]
[0005] The system is characterized by comprising a generation AI that inputs an image, a neural network that is a reward-shaped reinforcement learning that inputs the output results of the generation AI, and a knowledge and logic structure DB to which the output results of the neural network are increased to approach a normal distribution and the results are referenced, the reference results of the knowledge and logic structure DB are fed back to the generation AI and output from the generation AI, the results are input again to the neural network, the output results from the neural network are increased to approach a normal distribution and the best result is selected from among them. [Effects of the Invention]
[0006] The present invention provides the best possible support. [Brief explanation of the drawings]
[0007] [Figure 1] 1 is a schematic diagram showing the configuration of a support device 10 according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0008] (1) Configuration of the support device
[0009] FIG. 1 is a schematic diagram showing the configuration of a support device 10 according to this embodiment.
[0010] Generative AI11 is a technology that outputs new data and information based on the data that a computer has learned. It allows AI to perform the "thinking" and "planning" tasks that have traditionally been performed by humans, generating ideas and content. For example, gpt-4.
[0011] Neural network 12 uses reinforcement learning with reward shaping, where when the AI makes a choice, it is given an evaluation (reward) for that action, and it learns how to behave in a way that increases the evaluation. Unlike the correct answer label in supervised learning, the reward does not have to be given immediately for a single action, but can be given as a result of performing several actions depending on the situation.
[0012] The knowledge and logical structure DB 13 is a DB of knowledge structures and logical structures, such as hypergraphs, Boolean algebras, rules, regulations, and heuristics.
[0013] Reinforcement learning14 is reinforcement learning from the feedback of human knowledge based on cause and effect.
[0014] The user interface 15 mainly represents devices such as personal computers and smartphones, and is the part that connects the device with the user, such as applications and websites.
[0015] (2) Operation of the support device The operation of the support device 10 according to this embodiment will be described with reference to Fig. 1. In this embodiment, for example, a photograph showing a construction site is used.
[0016] First, input an image of a construction site (Figure 2) into the generation AI 11 (for example, gpt-4). The results are as follows: "This image was taken at a construction site. Two workers are wearing construction helmets and safety vests and are crouching on the floor working. They are wearing workman's uniforms and one appears to be wearing knee pads, which are intended to reduce strain on the knees over long periods of time. In the background, scaffolding and unfinished concrete structures typical of construction sites can be seen. They appear to be measuring or marking something. The surroundings are bright and natural light streams in, giving the construction site an open, indoor feel. The overall image suggests that carefully planned and professional work is being carried out."
[0017] Next, this result is input to the neural network 12, which is a reinforcement learning for reward shaping. Then, in A of Figure 1, the result is used to increase the output as much as possible by the central limit theorem, so that the output approaches a normal distribution. The output at this time is Wearing a construction helmet Wearing a safety vest Wearing craftsman's clothing Wearing knee pads - Working on measuring or marking something is.
[0018] Next, the results of A are organized by referring to the knowledge and logical structure DB13. Construction helmets, safety vests, measuring or marking are rules / regulations at construction sites, so those are excluded. (1) Wearing workman's clothing (2) Wearing knee pads is fed back to the generation AI 11.
[0019] The results of (1) and (2) above are input again into the generation AI11. If you input the above (1) into generation AI11, the result will be as follows: "To describe a construction worker wearing a workman's uniform, one could use the following: The construction workers were dressed in sturdy, work-appropriate workwear. They covered themselves from head to toe and were characterized by a simple yet robust design. Hats or helmets protected their heads, and their arms and legs were covered with sturdy fabric. They had work gloves on their hands and appeared ready to handle tools and materials. Their clothing emphasized safety and functionality, and it was clear that it was designed to withstand the harsh environment of a construction site.
[0020] Also, if you input the above (2) into generation AI11, the result will be as follows. "When describing a construction worker wearing knee pads, you might say: Construction sites often require workers to work on the ground or kneel for long periods of time, and workers wear knee pads to prepare for such situations. The tightly fitted, cushioned pads protect the knees from pressure and friction that occurs during work. The knee pads, incorporated as part of the workers' uniforms, not only improve comfort but also work efficiency. It appears that workers are taking care of their lower legs safely and effectively.
[0021] Then, the data is input again into the neural network 12, and the result is made to approach the normal distribution shown in Figure 1A. (1) Whether or not to use the term "workman's uniform" at a construction site (2) Are knee pads required under the rules? is selected, and the output is obtained from the user interface 15. This is a deductive approach based on knowledge and logical structures that cannot be explained by the results obtained from the inductive approach of ordinary generative AI from the events observed. Hybridizing this with generative AI and reinforcement learning leads to the optimization of output.
[0022] Alternatively, the output may be obtained after repeating the above loop of the support device.
[0023] Furthermore, reinforcement learning 14 based on human feedback may be performed and fed back to the neural network 12 or the knowledge and logic structure DB 13 .
[0024] (3) Effects According to this embodiment, the best possible output can be expected by combining the results obtained inductively from the events observed by the generative AI with a deductive approach based on knowledge and logical structures.
[0025] <Other embodiments> The present disclosure is not limited to the above-described embodiments as they are. The present disclosure can be embodied by modifying the components within the scope of the gist of the disclosure in the implementation stage. Furthermore, the present disclosure can be formed into various disclosures by appropriately combining multiple components disclosed in the above-described embodiments. For example, some components may be deleted from all the components shown in the embodiments. Furthermore, the components may be appropriately combined in different embodiments. [Explanation of symbols]
[0026] 10 Support equipment 11 Generation AI 12 Neural Networks 13 Knowledge / logical structure DB 14 Reinforcement Learning 15 User Interface
Claims
1. A generation AI that inputs an image; A neural network that is a reinforcement learning reward-shaping network that inputs the output result of the generation AI; a knowledge / logical structure DB in which the output results of the neural network are increased to approximate a normal distribution and the results are referenced; Equipped with The support device feeds back the reference results of the knowledge / logical structure DB to the generation AI, outputs them from the generation AI, inputs the results back into the neural network, increases the output results from the neural network, brings them closer to a normal distribution, and selects the best result from among them.
2. 2. The support device according to claim 1, wherein the support device loop is repeated.
3. 2. The support device according to claim 1, wherein the neural network and the knowledge / logical structure DB are input with the results of reinforcement learning based on feedback of human knowledge.
4. An image is input to a generating AI; A step of inputting the output result of the generation AI into a neural network that is a reinforcement learning of reward shaping; a step of increasing the output result of the neural network to approach a normal distribution, and referencing the result to a knowledge / logical structure DB; Equipped with A support method in which the reference results of the knowledge / logical structure DB are fed back to the generation AI and output from the generation AI, and the results are input again into the neural network, increasing the number of output results from the neural network and bringing them closer to a normal distribution, and selecting the best result from among them.
5. 5. The support method according to claim 4, further comprising repeating a loop of said support method.
6. 5. The support method according to claim 4, wherein the neural network and the knowledge / logical structure DB are input with the results of reinforcement learning based on feedback of human knowledge.