Information processing method, information processing system, and program
The information processing method addresses the limitation of existing systems by using a language generation model to output danger prediction messages based on construction site images, enhancing safety management and accident prevention.
Patent Information
- Application Number
- JP2023206445
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-12-06
- Publication Date
- 2025-06-18
AI Technical Summary
Existing danger prediction systems at construction sites cannot output danger prediction messages using language generation models.
An information processing method that acquires an image of a construction site, generates an input prompt with relevant information, and inputs it into a language generation model to output a danger prediction message.
Enables the output of danger prediction messages at construction sites using language generation models, improving safety management and preventing accidents.
Smart Images

Figure 2025091268000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to an information processing method, an information processing system, and a program.
Background Art
[0002] In recent years, the development of technologies related to danger prediction at construction sites has been actively promoted. For example, Patent Document 1 discloses a danger prediction activity support system that supports the cultivation of a safety culture through danger prediction activities in on-site work.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] However, the invention according to Patent Document 1 has a problem that it cannot output a danger prediction message related to danger prediction by a language generation model.
[0005] In one aspect, it is to provide an information processing method and the like capable of outputting a danger prediction message related to danger prediction by a language generation model.
Means for Solving the Problems
[0006] An information processing method according to one aspect acquires an image of a construction site, inputs an input prompt including information related to the acquired image of the construction site into a language generation model that uses danger prediction information related to danger prediction at the construction site, and executes a process of outputting a danger prediction message related to danger prediction at the construction site.
Effects of the Invention
[0007] On one hand, it becomes possible to output a danger prediction message related to danger prediction by a language generation model.
Brief Description of Drawings
[0008]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Figure 13
Figure 14
Figure 15
Figure 16
Figure 17
Figure 18
Figure 19
Figure 20
Figure 21
Figure 22
Figure 23
Figure 24
Mode for Carrying Out the Invention
[0009] Hereinafter, the present invention will be described in detail based on the drawings showing its embodiments.
[0010] (Embodiment 1) Embodiment 1 relates to a mode of outputting a danger prediction message regarding danger prediction at a construction site by a language generation model. FIG. 1 is an explanatory diagram showing an overview of a safety management system at a construction site. The system of the present embodiment includes an information processing device 1 and a unit 2, and each device transmits and receives information via a network N such as the Internet.
[0011] The information processing device 1 is an information processing device that performs processing, storage, and transmission / reception of various information. The information processing device 1 is, for example, a server device, a personal computer, or a general-purpose tablet PC (personal computer), etc. In the present embodiment, the information processing device 1 is assumed to be a server device, and hereinafter it will be read as server 1 for simplicity.
[0012] The unit 2 is a device (unit) that transmits an image of the construction site, receives and outputs a danger prediction message regarding danger prediction at the construction site, etc. The unit 2 has a camera, a microphone, a speaker, etc., and is attached to a helmet worn by a worker at the construction site. Note that the unit 2 may be an information processing device such as a smartphone, a mobile phone, or a wearable device such as an Apple Watch (registered trademark). Note that the unit 2 is preferably of a helmet-mounted type, but may also be of a chest pocket type or the like.
[0013] The server 1 according to the present embodiment acquires an image of the construction site from the unit 2. The server 1 generates an input prompt including information regarding the acquired image of the construction site. The server 1 inputs the generated input prompt into a language generation model (message output model 151 described later) that uses danger prediction information regarding danger prediction at the construction site, and outputs a danger prediction message regarding danger prediction at the construction site. The server 1 transmits the output danger prediction message to the unit 2. Note that the input prompt will be described later.
[0014] FIG. 2 is a block diagram showing a configuration example of the server 1. The server 1 includes a control unit 11, a storage unit 12, a communication unit 13, a reading unit 14, and a mass storage unit 15. Each configuration is connected by a bus B.
[0015] The control unit 11 includes an arithmetic processing device such as a CPU (Central Processing Unit), MPU (Micro-Processing Unit), GPU (Graphics Processing Unit), FPGA (Field Programmable Gate Array), DSP (Digital Signal Processor), or quantum processor. By reading and executing the control program 1P (program product) stored in the storage unit 12, the control unit 11 performs various information processing, control processing, etc. related to the server 1.
[0016] Note that the control program 1P can be deployed to be executed on a single computer, or placed at one site, or distributed across multiple sites and executed on multiple computers interconnected by a communication network. In FIG. 2, the control unit 11 is described as a single processor, but it may also be a multi-processor.
[0017] The storage unit 12 includes memory elements such as RAM (Random Access Memory) and ROM (Read Only Memory), and stores the control program 1P or data etc. necessary for the control unit 11 to execute processing. Also, the storage unit 12 temporarily stores data etc. necessary for the control unit 11 to execute arithmetic processing. The communication unit 13 is a communication module for performing communication-related processing, and transmits and receives information to and from the unit 2 etc. via the network N.
[0018] The reading unit 14 reads a portable storage medium 1a including a CD (Compact Disc)-ROM or DVD (Digital Versatile Disc)-ROM. The control unit 11 may read the control program 1P from the portable storage medium 1a via the reading unit 14 and store it in the mass storage unit 15. Also, the control unit 11 may download the control program 1P from another computer via the network N etc. and store it in the mass storage unit 15. Furthermore, the control unit 11 may read the control program 1P from the semiconductor memory 1b.
[0019] The large-capacity storage unit 15 includes a recording medium such as an HDD (Hard Disk Drive) or an SSD (Solid State Drive). The large-capacity storage unit 15 includes a message output model 151, an operator DB (database) 152, a danger prediction result DB 153, case data DB 154, and reference data DB 155.
[0020] The message output model (language generation model) 151 is an output device that outputs a danger prediction message related to danger prediction at the construction site based on information related to the image of the construction site, and is a learned model generated by machine learning. The operator DB 152 stores information about the operators at the construction site.
[0021] The danger prediction result DB 153 stores danger prediction results such as danger prediction messages related to danger prediction at the construction site output from the message output model 151. The case data DB 154 stores a plurality of case data (disaster quick report data or recurrence prevention countermeasure reports, etc.) including dangers at the construction site, countermeasures (e.g., prior safety measures or accident prevention cards), and similar cases. The reference data DB 155 stores a plurality of reference data related to laws, regulations, rules, or guidelines at the construction site.
[0022] In this embodiment, the storage unit 12 and the large-capacity storage unit 15 may be configured as an integrated storage device. Also, the large-capacity storage unit 15 may be composed of a plurality of storage devices. Furthermore, the large-capacity storage unit 15 may be an external storage device connected to the server 1.
[0023] The server 1 may execute various information processing and control processing, etc. by a single computer, or may execute them distributedly by a plurality of computers. Also, the server 1 may be realized by a plurality of virtual machines provided in one server, or may be realized using a cloud server.
[0024] FIG. 3 is an explanatory diagram showing an example of the record layout of the worker DB 152 and the danger prediction result DB 153. The worker DB 152 includes a worker ID column, a unit ID column, a name column, an affiliation column, an occupation type column, and a responsibility column. The worker ID column stores the worker ID of the worker uniquely specified to identify each worker. The unit ID column stores the unit ID of the unit uniquely specified to identify the unit attached to each helmet.
[0025] The name column stores the name of the worker. The affiliation column stores the affiliation (company, department, etc.) of the worker. The occupation type column stores the occupation type of the worker (for example, technical occupation or skilled occupation). The responsibility column stores the responsibility of the worker. For example, responsibilities include, for technical occupations, on-site agent, supervisor, or person in charge, etc., and for skilled occupations, foreman, electrician, or assistant, etc.
[0026] The danger prediction result DB 153 includes a construction date column, a worker ID column, an image data column, a message column, and an output time column. The construction date column stores the date on which construction was carried out at the construction site. The worker ID column stores the worker ID for identifying the worker. The image data column stores the image data of the construction site. The message column stores the danger prediction message at the construction site output from the message output model 151. The output time column stores the time when the danger prediction message was output.
[0027] FIG. 4 is an explanatory diagram showing an example of the record layout of the case data DB 154. The case data DB 154 includes a case data ID column, a danger column, a countermeasure column, a case data column, a file path column, a similar case data column, and a similar case link column. The case data ID column stores the ID of the case data uniquely specified to identify each case data. The danger column stores the type of danger at the construction site. The countermeasure column stores the countermeasures against the danger.
[0028] The case data series stores detailed information of case data (disaster quick report data). The detailed information of case data includes date and time information of the disaster occurrence, disaster classification, construction type, type of disaster, victim information, construction name, disaster occurrence status, or special notes (e.g., implementation status of power inspection confirmation), etc. The file path series stores the file paths of documents (e.g., disaster quick reports) containing case data.
[0029] The similar case data series stores detailed information of case data of similar cases. The similar case link series stores link information for accessing similar cases. The link information has forms such as, for example, a URL (Uniform Resource Locator) for accessing similar cases, a QR code (registered trademark), highlighted text, or an icon.
[0030] Note that not limited to the above-mentioned columns, appropriate columns may be provided as long as they are information related to case data.
[0031] Figure 5 is an explanatory diagram showing an example of the record layout of the reference data DB155. The reference data DB155 includes a reference data ID column, a type column, a name column, a reference data column, and a file path column. The reference data ID column stores the ID of the reference data that is uniquely specified to identify each reference data. The type column stores the type of reference data. The type of reference data includes laws, regulations, rules, or guidelines, etc. Note that the type of reference data may include construction specifications or design drawings, etc.
[0032] The name column stores names such as laws (e.g., "Electrical Equipment Technical Standards"), regulations (e.g., JIS standards), rules (e.g., handling of the input panel of UPS), or guidelines (e.g., Electrical Equipment Handling Guidelines).
[0033] The reference data sequence stores detailed data of reference data such as laws, regulations, rules, or guidelines (contents described for each article, chapter, section, or item, etc.). For example, when the law is the "Electrical Equipment Technical Standards", the contents of each article (Article 1, Article 2, etc.) of the "Electrical Equipment Technical Standards" are stored in the reference data sequence. Or, when the rule is the "Handling of the Input Panel of UPS", the contents of each chapter (Chapter 1, Chapter 2, etc.) are stored in the reference data sequence.
[0034] The file path sequence stores the file paths of documents describing various reference data.
[0035] Note that each of the case data DB154 and the reference data DB155 may store a case data vector obtained by vectorizing case data and a reference data vector obtained by vectorizing reference data.
[0036] Regarding the vectorization process, for example, by using a pre-trained BERT language model, it can be converted into a feature vector represented by a numerical value of 768 dimensions for each string data. Note that as the vectorization process, known technologies such as Word2Vec, GloVe, fastText, BERT, XLNet, or ALBERT can be used.
[0037] Note that the storage form of each of the above-mentioned DBs is an example, and other storage forms may be used as long as the relationship between the data is maintained.
[0038] Figure 6 is a block diagram showing a configuration example of unit 2. Unit 2 includes a control unit 21, a storage unit 22, a communication unit 23, a camera 24, a camera control unit 25, an operation unit 26, a speaker 27, a microphone 28, and a display unit 29. Each component is connected by a bus B.
[0039] The control unit 21 includes an arithmetic processing device such as a CPU or an MPU, and performs various information processing and control processing related to the unit 2 by reading and executing a control program 2P (program product) stored in the storage unit 22. In FIG. 6, the control unit 21 is described as a single processor, but it may be a multi-processor.
[0040] The storage unit 22 includes a memory element such as a RAM or a ROM, and stores a control program 2P or data necessary for the control unit 21 to execute processing. Further, the storage unit 22 temporarily stores data and the like necessary for the control unit 21 to execute arithmetic processing.
[0041] The communication unit 23 is a communication module for performing communication-related processing, and transmits and receives information to and from the server 1 or the like via the network N. The camera 24 is an imaging device such as a CCD (Charge Coupled Device) camera or a CMOS (Complementary Metal Oxide Semiconductor) camera. The camera control unit 25 performs camera control processing such as imaging of an image and switching between the camera on mode / off mode.
[0042] The operation unit 26 receives a user operation such as imaging of an image and switching between the camera on mode / off mode, and notifies the received operation to the camera control unit 25. For example, the operation unit 26 receives a user operation by an input device such as a mechanical button. When the unit 2 has a display unit, the operation unit 26 may receive a user operation by an input device such as a touch panel or a pen tablet provided on the surface of the display unit.
[0043] The speaker 27 is a device that converts an electrical signal into sound. The microphone 28 is a device that converts sound into an electrical signal. Note that the speaker 27 and the microphone 28 may be a headset connected to the unit 2 by a short-range wireless communication method such as BLUETOOTH (registered trademark).
[0044] The display unit 29 is a liquid crystal display, an organic EL (electroluminescence) display, or the like, and displays various information according to the instructions of the control unit 21. Note that the display unit 29 may be wearable glass (smart glass). Note that the display unit 29 is not essential.
[0045] FIG. 7 is an explanatory diagram for explaining the process of outputting a danger prediction message. The unit 2 attached to the helmet worn by the worker at the construction site acquires an image of the construction site by the camera 24.
[0046] Note that the unit 2 may always perform shooting using the camera 24, or may shoot an image of the construction site every predetermined period (for example, 1 minute). Alternatively, the unit 2 may shoot an image of the construction site by the camera 24 when it receives a shooting operation of the worker by the operation unit 26.
[0047] Furthermore, the unit 2 may receive a shooting instruction (for example, "Please shoot") based on the voice data of the worker via the microphone 28. In this case, the unit 2 uses a known voice recognition technology to recognize a shooting instruction from the received voice data, and shoots an image of the construction site according to the recognized shooting instruction.
[0048] The unit 2 transmits the acquired image of the construction site to the server 1. The server 1 receives the image of the construction site transmitted from the unit 2. As shown in the figure, the server 1 receives an image of the construction site indicating that the worker is going down the stairs.
[0049] The server 1 generates an input prompt including information about the image of the construction site. Specifically, the server 1 acquires information about the image from the received image of the construction site. The information about the image includes the feature amount of the image, the object information included in the image (for example, stairs or a transformer, etc.), or the natural language description of the image (for example, "The worker is going down the stairs" or "The worker is approaching the high-voltage charging unit").
[0050] Regarding the feature amounts of the images, the server 1 may extract the feature amounts of the images of the construction site taken using a local feature amount extraction method such as, for example, A-KAZE (Accelerated KAZE), SIFT (Scale Invariant Feature Transform), SURF (Speeded-Up Robust Features), ORB (Oriented FAST and Rotated BRIEF), or HOG (Histograms of Oriented Gradients).
[0051] Alternatively, the server 1 may extract the feature amounts from the images of the construction site using a learned model obtained by performing machine learning such as, for example, deep learning. The learned model is configured using an algorithm such as, for example, CNN (Convolutional Neural Network). Note that, in addition to CNN, the learned model may be configured using any object detection algorithm such as R-CNN (Region Based Convolutional Neural Network), Fast R-CNN, Faster R-CNN, Mask R-CNN, SSD (Single Shot Multibook Detector), or YOLO (You Only Look Once), or may be configured by combining some of these models.
[0052] Regarding the objects included in the images, the objects included in the images can be recognized using object detection technology. The object detection technology is, for example, in addition to pattern matching, A-KAZE, SIFT, etc., and the objects included in the images may be detected by extracting the feature amounts using a local feature amount extraction method.
[0053] Alternatively, the server 1 may detect the objects included in the images using an object detection model generated using a neural network structure such as, for example, CNN, R-CNN, or YOLO.
[0054] Regarding the natural language description of an image, for example, a pre-trained natural language description generation model for an image may be used. The natural language description generation model is, for example, a model constructed by combining an encoder convolutional neural network and a decoder LSTM neural network. The server 1 inputs the acquired image of the construction site into the natural language description generation model and outputs a natural language description (a sequence of multiple words) for the image.
[0055] For example, the server 1 may input an image of the construction site indicating that a worker is going down the stairs into the natural language description generation model and output a natural language description of "a worker is going down the stairs".
[0056] Note that the image of the construction site may be a video. For example, when a video of the construction site is captured by the camera 24 of the unit 2, the server 1 extracts a plurality of frame images from the captured video of the construction site at a predetermined interval (for example, 10-frame images). The server 1 may use the above-described natural language description generation model or Transformer, etc., to obtain information regarding each of the extracted frame images.
[0057] Note that in this embodiment, the acquisition process of information regarding an image will be described as an example of a method for natural language description of an image.
[0058] When the server 1 receives an image of the construction site, it uses a pre-trained natural language description generation model for an image to obtain information regarding the image (for example, the natural language description of the image). The server 1 obtains an image information vector of the information regarding the image by vectorizing the obtained information regarding the image.
[0059] Based on the obtained image information vector, the server 1 searches (extracts) data corresponding to the query from the case data DB154 and the reference data DB155. The query is a preset stereotyped sentence. For example, the query may be "Please tell me if there is any problem".
[0060] Specifically, the server 1 calculates the similarity between the acquired image information vector and the vectors of the case data (such as danger, countermeasures, or similar cases) stored in the case data DB154 respectively. The server 1 compares each calculated similarity with a predetermined similarity threshold. The server 1 extracts from the case data DB154 the case data corresponding to the similarity that is equal to or higher than the predetermined similarity threshold, or the file path of the document describing the case data, etc.
[0061] The server 1 calculates the similarity between the acquired image information vector and the vectors of the reference data (such as laws, regulations, rules, or guidelines) stored in the reference data DB155 respectively. The server 1 compares each calculated similarity with a predetermined similarity threshold. The server 1 extracts from the reference data DB155 the reference data corresponding to the similarity that is equal to or higher than the predetermined similarity threshold, or the file path of the document describing the reference data, etc.
[0062] In addition, in the above-described data extraction process, an example of comparing the similarity with a predetermined similarity threshold has been described, but it is not limited to this. For example, the server 1 may extract the data corresponding to the highest similarity among the calculated similarities.
[0063] For example, based on the acquired image information vector, the server 1 extracts the file path of "Construction Site Equipment Guidelines, Chapter 2, Working Stairs" from the reference data DB155, and extracts the file path of the disaster report with the case data ID of "D-00001" from the case data DB154.
[0064] The server 1 generates an input prompt including the query, the information about the acquired image, the various extracted data, and the file paths of the documents describing the various data.
[0065] FIG. 8 is an explanatory diagram showing an example of the input prompt and the danger prediction message. FIG. 8A is an explanatory diagram showing an example of the input prompt. The input prompt includes the query, the information about the image, and the various data extracted from the DB.
[0066] As shown in the illustration, the query is "Tell me if there are any dangerous problems and also tell me the countermeasures." The information about the image is "The worker is going down the working stairs." The various data extracted from the DB includes the file path of Disaster Bulletin 00001 with case data ID D-00001, the file path of "Construction Site Equipment Guidelines, Chapter 2, Working Stairs" with reference data ID S-00005, etc.
[0067] Subsequently, returning to FIG. 7, the server 1 inputs the generated input prompt into the message output model 151 and outputs a danger prediction message related to danger prediction at the construction site. The danger prediction message output from the message output model 151 includes the predicted dangers and countermeasures at the construction site.
[0068] The message output model 151 is a language generation model that uses danger prediction information related to danger prediction at the construction site and is utilized as a program module that is part of artificial intelligence software. The message output model 151 is a constructed language generation model that takes an input prompt including an image of the construction site and outputs a danger prediction message related to danger prediction at the construction site.
[0069] The message output model 151 is a language generation model constructed by performing pre-training with a large amount of text data (dataset). As the message output model 151, for example, large language models (LLMs) such as Transformer, ALBERT (A Lite BERT), GPT (Generative Pre-trained Transformer)-2, GPT-3, GPT-4, LLaVA (Large Language and Vision Assistant), MiniGPT-4, or BERT (Bidirectional Encoder Representations from Transformers) can be used.
[0070] Note that instead of storing in the mass storage unit 15, the message output model 151 may be configured such that the server 1 accesses and reads from an external language processing server or language processing platform, etc.
[0071] The input prompt is created in a format that can be understood by the message output model 151 and is the prompt given as the input to the message output model 151. The message output model 151 interprets the input input prompt and outputs an appropriate response (for example, a danger prediction message).
[0072] As an example, the message output model 151 splits the input prompt into tokens because it is converted into a format that the message output model 151 can process. The message output model 151 performs context understanding processing by calculating the relationship between each token in the input prompt and other tokens.
[0073] The message output model 151 performs response generation processing on the input prompt based on the language knowledge obtained through pre-training and fine-tuning, etc. For example, the message output model 151 uses generation methods such as greedy decoding, beam search, or sampling to select the optimal tokens. The message output model 151 performs decoding processing on the selected tokens to convert them back to text format and generates a danger prediction message as output data.
[0074] The server 1 outputs the predicted dangers and countermeasures at the construction site as a danger prediction message from the message output model 151.
[0075] FIG. 8B is an explanatory diagram showing an example of a danger prediction message. For the working staircase, a danger prediction message, similar cases associated with the danger prediction message, and link information (for example, URL) for accessing the similar cases are output from the message output model 151.
[0076] As shown in the illustration, the danger prediction message is: "Danger: When going down the stairs, disasters such as slipping and falling occur repeatedly. Pay sufficient attention to your feet. Countermeasure: If there is a handrail on the stairs, use the handrail." The similar cases are: "August 10, 2022, ankle sprain and fall URL:*** April 1, 2019, right ankle fracture URL:***".
[0077] Note that in this embodiment, an example of outputting both similar cases and link information has been described, but it is not limited to this. For example, either similar cases or links may be output from the message output model 151.
[0078] Server 1 stores the acquired danger prediction message in the danger prediction result DB153 in association with the operator ID. Server 1 transmits the acquired danger prediction message, similar cases, and link information to unit 2 based on the unit ID of unit 2 attached to the helmet worn by the operator.
[0079] Unit 2 receives the danger prediction message, similar cases, and link information transmitted from server 1. Unit 2 outputs the received danger prediction message, similar cases, and link information through speaker 27. Note that unit 2 may also display the received similar cases and link information on display unit 29.
[0080] Note that it is not limited to the use of the learned message output model 151 described above. For example, server 1 may use image information vectors obtained by vectorizing information related to images of a large number of construction sites, case data (such as dangers, countermeasures, and similar cases), or reference data (such as laws, regulations, rules, or guidelines) as a data set (learning data) to generate (construct) the message output model 151 through pre-training.
[0081] FIG. 9 is a flowchart showing a processing procedure when transmitting a danger prediction message to unit 2. The control unit 21 of unit 2 executes a subroutine of a process of acquiring an image of the construction site via the camera control unit 25 (step S201). Note that the subroutine of the process of acquiring an image of the construction site will be described later. The control unit 21 transmits the acquired image of the construction site to the server 1 by the communication unit 23 (step S202).
[0082] The control unit 11 of the server 1 receives the image of the construction site transmitted from unit 2 by the communication unit 13 (step S101). The control unit 11 executes a subroutine of a process of generating an input prompt (step S102). Note that the subroutine of the process of generating an input prompt will be described later. The control unit 11 inputs the generated input prompt to the message output model 151 (step S103), and outputs a danger prediction message at the construction site from the message output model 151 (step S104).
[0083] The control unit 11 stores the danger prediction message output from the message output model 151 in the danger prediction result DB 153 of the mass storage unit 15 in association with the worker ID (step S105). Specifically, the control unit 11 stores the construction date, the image data of the construction site, the danger prediction message, and the output time as one record in the danger prediction result DB 153 in association with the worker ID.
[0084] The control unit 11 transmits the danger prediction message to unit 2 by the communication unit 13 based on the unit ID (step S106). The control unit 21 of unit 2 receives the danger prediction message transmitted from the server 1 by the communication unit 23 (step S203). The control unit 21 outputs the received danger prediction message by the speaker 27 (step S204). The control unit 21 ends the process.
[0085] Figure 10 is a flowchart showing the processing procedure of a subroutine for acquiring an image of a construction site. The camera control unit 25 of unit 2 determines whether it is in the camera-off mode (step S01). If the camera control unit 25 determines that it is not in the camera-off mode (NO in step S01), it proceeds to the process of step S04 described later.
[0086] If the camera control unit 25 determines that it is in the camera-off mode (YES in step S01), it receives, via the operation unit 26, a setting operation by the operator to switch to the camera-on mode (step S02). The camera control unit 25 sets the mode from the camera-off mode to the camera-on mode according to the received setting operation (step S03).
[0087] The camera control unit 25 receives, via the operation unit 26, a shooting operation to shoot an image of the construction site (step S04). The camera control unit 25 shoots an image of the construction site via the camera 24 (step S05). The camera control unit 25 acquires the shot image of the construction site (step S06). The camera control unit 25 ends the subroutine for the acquisition process of the image of the construction site and returns.
[0088] Figure 11 is a flowchart showing the processing procedure of a subroutine for generating an input prompt. The control unit 11 of the server 1 acquires information about the image (for example, a natural language description of the image) from the image of the construction site (step S11). The control unit 11 performs a vectorization process on the acquired information about the image (step S12). The control unit 11 calculates a similarity based on the image information vector obtained by the vectorization process (step S13).
[0089] Specifically, the control unit 11 calculates the similarity between the acquired image information vector and the vectors of the case data (such as danger, countermeasures, or similar cases) stored in the case data DB154 respectively. The control unit 11 calculates the similarity between the acquired image information vector and the vectors of the reference data (such as laws, regulations, rules, or guidelines) stored in the reference data DB155 respectively.
[0090] Based on the similarity between the calculated image information vector and the vector of the case data, the control unit 11 extracts from the case data DB 154 of the mass storage unit 15 a similarity that is equal to or higher than a predetermined similarity threshold, or case data corresponding to the highest similarity, or a file path of a document describing the case data, etc. (step S14).
[0091] Based on the similarity between the calculated image information vector and the vector of the reference data, the control unit 11 extracts from the reference data DB 155 of the mass storage unit 15 a similarity that is equal to or higher than a predetermined similarity threshold, or reference data corresponding to the highest similarity, or a file path of a document describing the reference data, etc. (step S15).
[0092] The control unit 11 extracts the text included in the extracted case data or reference data (step S16). When the control unit 11 extracts the file path of a document, it reads the data stored in each document. The control unit 11 may extract the text included in the read data.
[0093] The control unit 11 generates an input prompt including a preset query (for example, "Tell me if there is any problem"), the extracted text, and information regarding the image of the construction site (step S17). The control unit 11 ends the subroutine of the input prompt generation process and returns.
[0094] FIG. 12 is a flowchart showing a processing procedure when outputting similar cases associated with a danger prediction message. Regarding the content overlapping with FIG. 9, the same reference numerals are given and the description is omitted. After executing the process of step S103, the control unit 11 of the server 1 outputs from the message output model 151 a danger prediction message, similar cases associated with the danger prediction message, and link information for accessing the similar cases (step S111).
[0095] The control unit 11 executes the process of step S105. The control unit 11 transmits the acquired danger prediction message, similar cases, and link information to the unit 2 via the communication unit 13 based on the unit ID of the unit 2 (step S112). The control unit 21 of the unit 2 receives the danger prediction message, similar cases, and link information transmitted from the server 1 via the communication unit 23 (step S211). The control unit 21 outputs the received danger prediction message, similar cases, and link information via the speaker 27 (step S212). The control unit 21 ends the process.
[0096] According to this embodiment, based on the image of the construction site, it is possible to output a danger prediction message regarding danger prediction at the construction site by using the message output model 151.
[0097] According to this embodiment, by utilizing the message output model 151, it is possible to realize the improvement of safety management at the construction site without depending on the individual danger prediction ability of the workers.
[0098] According to this embodiment, by outputting similar cases associated with the danger prediction message, it is possible to prompt reconfirmation such as attracting attention and prevent the occurrence of dangerous accidents. Therefore, it is possible to realize safe and secure decent work (work that is fulfilling and makes people feel human).
[0099] According to this embodiment, based on the image of the construction site and the case data related to the image, it is possible to output a danger prediction message regarding danger prediction at the construction site by using the message output model 151.
[0100] <Modification Example 1> A process of sharing a danger prediction message for a worker with a plurality of other workers will be described. At the construction site, a plurality of workers including the first worker are grouped. The danger prediction message for the first worker can be output to the unit 2 mounted on the helmets of the other grouped workers.
[0101] FIG. 13 is an explanatory diagram showing an example of the record layout of the worker DB 152 in Modification 1. Note that the same reference numerals are given to the contents overlapping with FIG. 3, and the description thereof is omitted. The worker DB 152 includes a group ID column. The group ID column stores a group ID for identifying a group in which a plurality of workers are grouped at the construction site.
[0102] The server 1 acquires an image of the construction site from the unit 2 attached to the helmet worn by the first worker. The server 1 acquires a danger prediction message related to danger prediction at the construction site by using the message output model 151 based on the acquired image of the construction site. The server 1 transmits the acquired danger prediction message to the unit 2 attached to the helmet of the first worker.
[0103] The server 1 identifies the group to which the first worker belongs based on the worker ID of the first worker. The server 1 acquires the worker IDs of other multiple workers included in the identified group. The server 1 transmits the danger prediction message for the first worker to the unit 2 attached to the helmet of each worker based on the acquired worker ID of each worker. Each unit 2 receives the danger prediction message transmitted from the server 1. Each unit 2 outputs the received danger prediction message by the speaker 27.
[0104] FIG. 14 is a flowchart showing a processing procedure when sharing a danger prediction message for a first worker with other workers. Note that the same reference numerals are given to the contents overlapping with FIG. 9, and the description thereof is omitted.
[0105] After executing the process of step S106, the control unit 11 of the server 1 acquires the group ID of the group to which the first worker belongs from the worker DB 152 of the mass storage unit 15 based on the worker ID of the first worker (step S121). The control unit 11 acquires the worker IDs of other multiple workers included in the group from the worker DB 152 based on the acquired group ID (step S122).
[0106] Based on the acquired operator IDs, the control unit 11 transmits, via the communication unit 13, a danger prediction message for the first operator to the unit 2 attached to the helmet of each operator (step S123). The control unit 11 ends the process.
[0107] Similar to the processing of the unit 2 attached to the helmet of the first operator, the unit 2 attached to the helmet of other operators receives the danger prediction message transmitted from the server 1. Each unit 2 outputs the received danger prediction message by means of the speaker 27.
[0108] Note that in this first modification example, an example in which a danger prediction message for the first operator is shared with a plurality of other operators within the group to which the first operator belongs has been described, but the present invention is not limited thereto. For example, by using BLUETOOTH (registered trademark) or GPS (Global Positioning System) or the like, it may be shared with other operators in the vicinity.
[0109] For example, when the unit 2 attached to the helmet of the first operator receives a danger prediction message transmitted from the server 1, it uses BLUETOOTH to detect the unit 2 attached to the helmets of other operators at the construction site. Based on the unit ID of the detected other unit 2, the unit 2 attached to the helmet of the first operator transmits a danger prediction message for the first operator to the corresponding unit 2.
[0110] According to this first modification example, by sharing a danger prediction message for the first operator with a plurality of other operators, for example, when the first operator approaches a danger area, other operators can call the first operator's attention to the danger. <Modification Example 2> The images of the construction site may include not only still images but also videos. In Modification 2, a process of using a video of the construction site will be described. A video is composed of a plurality of frame images that are temporally continuous. When a video of the construction site is captured, the feature amount of the image can be extracted for each frame image or for every several frame images.
[0111] Specifically, the unit 2 mounted on the helmet worn by the worker at the construction site captures a video of the construction site. The unit 2 transmits the captured video of the construction site to the server 1. The server 1 receives the video of the construction site transmitted from the unit 2. The server 1 extracts a plurality of frame images from the received video of the construction site at a predetermined interval (for example, 10 frame images).
[0112] The server 1 acquires information regarding each of the extracted frame images. For example, when the information regarding the frame image is a feature amount, the server 1 may input each of the acquired frame images into a learning model generated using a neural network structure such as a CNN, and extract the feature amount of each frame image. Alternatively, the server 1 may extract the feature amount of each frame image using A-KAZE or the like.
[0113] Note that, similar to the process of Embodiment 1, the server 1 may also acquire information regarding the frame image including object information (for example, a staircase) included in the frame image, or a natural language description of the image (for example, the worker is going down the staircase) in addition to the feature amount of the frame image.
[0114] Thereafter, similar to the process of Embodiment 1, the server 1 extracts case data and reference data for generating an input prompt from the DB based on the information regarding the acquired image. The server 1 extracts the sentences included in the extracted case data and reference data.
[0115] Server 1 generates an input prompt including a preset query, the extracted text, and information about the image of the construction site. Server 1 inputs the generated input prompt into the message output model 151 and outputs a danger prediction message for the construction site.
[0116] According to this Modification 2, it is possible to output a danger prediction message using the message output model 151 based on the video of the construction site.
[0117] (Embodiment 2) Embodiment 2 relates to a form in which a danger prediction message is output by the message output model 151 based on the attributes of the worker and the image of the construction site. Note that descriptions of content overlapping with Embodiment 1 are omitted. The attributes of the worker include age, affiliation, job type, number of years of practical experience, or responsibilities, etc.
[0118] FIG. 15 is an explanatory diagram showing an example of the record layout of the worker DB 152 in Embodiment 3. Note that the same reference numerals are given to the content overlapping with FIG. 3 and the description thereof is omitted. The worker DB 152 includes an age column and a column for the number of years of practical experience. The age column stores the age of the worker. The column for the number of years of practical experience stores the number of years of practical experience of the worker.
[0119] FIG. 16 is an explanatory diagram for outputting a danger prediction message according to the attributes of the worker. FIG. 16A is an explanatory diagram showing an example of an image of the construction site. Server 1 acquires an image of the construction site from the unit 2 attached to the helmet worn by the worker at the construction site. As shown in the figure, Server 1 acquires an image of the construction site including the high-voltage input panel of the UPS (Uninterruptible Power Supply).
[0120] Server 1 acquires the attributes of the worker including the job type, the number of years of practical experience, or responsibilities, etc. of the worker from the worker DB 152 based on the worker ID. For example, the acquired attributes are "job type: skilled worker, number of years of practical experience: 1, responsibilities: on-site construction, electrical work".
[0121] Server 1 identifies the text corresponding to the obtained attributes of the worker. For example, Server 1 may identify the text "I am an electrician responsible for on-site construction. My job type is a skilled job type. I am a beginner with less than 2 years of practical experience." according to the attributes of a worker with "Job type: Skilled job, Years of practical experience: 1, Responsibilities: On-site construction, Electrician".
[0122] Alternatively, Server 1 may identify the text "I am a site agent. I am responsible for on-site construction management. My job type is a technical job type. I am a senior with 8 years of practical experience." according to the attributes of a worker with "Job type: Technical job, Years of practical experience: 8, Responsibilities: Construction management, Site agent".
[0123] Similar to Embodiment 1, Server 1 performs acquisition processing of information regarding the image of the construction site, and extraction (search) processing of case data and reference data based on the information regarding the image. Server 1 generates an input prompt including the identified text, a preset query, the sentences included in the extracted case data and reference data, and the information regarding the image of the construction site.
[0124] FIG. 16B is an explanatory diagram showing an example of an input prompt. The input prompt includes a query, information regarding the image, text corresponding to the attributes of the worker, and various data extracted from the DB.
[0125] As shown in the figure, the query is "Today's work is to check the status of the high-voltage input panel of the UPS." The information regarding the image is "The door of the high-voltage input panel of the UPS is open. The charging part is also visible in the front." The text corresponding to the attributes of the worker is "I am an electrician responsible for on-site construction. My job type is a technical job type. I am a beginner with less than 2 years of practical experience." The various data extracted from the DB include the file path of "Chapter 1 of UPS Input Panel Handling" and the file path of "Case Data D-00002", etc.
[0126] Server 1 inputs the generated input prompt into the message output model 151, and outputs a danger prediction message according to the attributes of the operator and an image associated with the danger prediction message.
[0127] Figure 16C is an explanatory diagram showing an example of a danger prediction message. For today's work of checking the status of the high-voltage input panel of the UPS, a danger prediction message and an image associated with the danger prediction message are output from the message output model 151.
[0128] As shown in the figure, the danger prediction message is "There is a high risk of electric shock in the photo part, so do not touch it. When checking, do not touch the power supply of the UPS." Also, an image (photo) associated with the danger prediction message is displayed together with the display of the danger prediction message.
[0129] The image associated with the danger prediction message may be extracted from, for example, the case data DB154. Based on the danger prediction message output from the message output model 151, Server 1 performs a search (extraction) process for images included in the similar case data from the case data DB154. When the search process for the image associated with the danger prediction message is performed, a part of the danger prediction message or keywords (such as electric shock, inspection, and power supply) included in the danger prediction message may be used.
[0130] Specifically, Server 1 obtains a danger prediction message vector by performing vectorization processing on the danger prediction message output from the message output model 151. Server 1 calculates the similarity between the obtained danger prediction message vector and the image information vector included in the similar case data stored in the case data DB154 respectively. Based on the similarity between the calculated danger prediction message vector and the image information vector, Server 1 extracts an image corresponding to a similarity equal to or higher than a predetermined similarity threshold or the highest similarity from the similar case data column of the case data DB154.
[0131] Note that the link information (e.g., URL) described in the similar case link column of the case data DB154 can be used. For example, the server 1 acquires the URL of a similar case from the case data DB154. The server 1 acquires the image posted on the site specified by the acquired URL. The server 1 obtains an image vector by performing vectorization processing on the acquired image. Similar to the above-described processing, the server 1 may extract an image associated with the danger prediction message based on the similarity between the danger prediction message vector and the obtained image information vector.
[0132] The server 1 transmits the danger prediction message output from the message output model 151 and the image associated with the danger prediction message to the unit 2 mounted on the helmet worn by the operator. The unit 2 receives the danger prediction message and the image transmitted from the server 1. The unit 2 outputs the received danger prediction message through the speaker 27 and displays the received image on the display unit 29.
[0133] FIG. 17 is a flowchart showing the processing procedure of a subroutine for generating an input prompt in the second embodiment. Note that the same reference numerals are given to the contents overlapping with those in FIG. 11, and the description thereof is omitted.
[0134] After executing the processing of step S16, the control unit 11 of the server 1 acquires the attributes (job type, number of years of practical experience, or responsibilities, etc.) of the operator from the operator DB152 in the mass storage unit 15 based on the operator ID (step S21). The control unit 11 specifies the text corresponding to the acquired attributes of the operator (step S22). For example, the control unit 11 may generate text describing the job type, number of years of practical experience, or responsibilities, etc. based on the acquired job type, number of years of practical experience, or responsibilities, etc. of the operator.
[0135] The control unit 11 generates an input prompt (step S23). Specifically, the control unit 11 generates an input prompt including the identified text, the preset query, the extracted sentence, and information regarding the image of the construction site. The control unit 11 ends the subroutine of the input prompt generation process and returns.
[0136] According to the present embodiment, based on the attributes of the workers at the construction site and the image of the construction site, it is possible to output a danger prediction message according to the attributes of the workers by using the message output model 151.
[0137] (Embodiment 3) Embodiment 3 relates to a form in which a danger prediction message is output by the message output model 151 based on the construction work content of the day at the construction site and the image of the construction site. Note that descriptions of the contents overlapping with Embodiments 1 to 2 are omitted.
[0138] FIG. 18 is a block diagram showing a configuration example of the server 1 in Embodiment 3. Note that the same reference numerals are given to the contents overlapping with FIG. 2 and the descriptions are omitted. The mass storage unit 15 includes a construction work content DB 156. The construction work content DB 156 stores the construction work content of the workers.
[0139] FIG. 19 is an explanatory diagram showing an example of the record layout of the construction work content DB 156. The construction work content DB 156 includes a construction date column, a worker ID column, a work name column, a work time zone column, a work location column, a work procedure column, and a danger caution column.
[0140] The construction date column stores the date on which construction is carried out at the construction site. The worker ID column stores the worker ID for identifying the worker. The work name column stores the name of the work. The work time zone column stores the time zone of the work. The work location column stores the location of the work. The work procedure column stores the procedure of the work. The danger caution column stores the danger caution items in the work.
[0141] Server 1 acquires an image of the construction site from Unit 2 attached to the helmet worn by the worker at the construction site. Server 1 acquires the construction work details of the worker for the current day from the construction work details DB156 in the mass storage unit 15 based on the date of the current day and the worker ID. The construction work details include the work name, work time period, work location, work procedure, or danger precautions, etc.
[0142] For example, the construction work details may be "Work name: Existing low-voltage switchboard work, Work time period: 8:00 - 11:00, Work location: Location A, Work procedure: 1. Confirm the work scope 2. Confirm the load measurement panel 3. Power outage operation 4. Add and modify the circuit breaker 5. Restoration of power supply, Danger precautions: After the power outage, confirm the power outage of the switchgear."
[0143] Similar to Embodiment 1, Server 1 performs information acquisition processing on the image of the construction site, and extraction processing of case data and reference data based on the information in the image. Server 1 generates an input prompt including the acquired construction work details for the current day, a preset query, the sentences included in the extracted case data and reference data, and the information related to the image of the construction site.
[0144] As an example, the generated input prompt may be "Today's work is the addition and modification of the circuit breaker in the existing low-voltage switchboard. The work time period is from 8:00 to 11:00. The work location is Location A. The work procedure is 1. Confirm the work scope 2. Confirm the load side of the existing switchboard 3. Power outage operation 4. Addition and modification of the circuit breaker 5. Restoration of power supply. The danger precaution is to confirm the power outage of the switchgear after the power outage. The door of the low-voltage switchboard is open. Regarding danger prediction, please refer to the following information and answer. {Path of the searched file (case data D-*****}···"
[0145] Server 1 inputs the generated input prompt into message output model 151 and outputs a danger prediction message at the construction site. Similar to the processing in Embodiment 1, Server 1 transmits the danger prediction message at the construction site output from message output model 151 to unit 2 attached to the helmet worn by the operator.
[0146] Figure 20 is a flowchart showing the processing procedure of a subroutine for generating an input prompt in Embodiment 3. Regarding the content overlapping with that in Figure 11, the same reference numerals are used and the description is omitted.
[0147] After executing the processing in step S16, control unit 11 of Server 1 acquires the construction content of the operator on the current day (such as work name, work time period, work location, work procedure, or danger precautions, etc.) from construction content DB156 in mass storage unit 15 based on the date of the current day and the operator ID (step S26).
[0148] Control unit 11 generates an input prompt including the acquired construction content of the current day and an image of the construction site (step S27). Specifically, control unit 11 generates an input prompt including the acquired construction content of the current day, a preset query, the extracted text, and information regarding the image of the construction site. Control unit 11 ends the subroutine for the input prompt generation processing and returns.
[0149] According to this embodiment, based on the construction content of the current day and an image of the construction site, it is possible to output a danger prediction message at the construction site using message output model 151.
[0150] (Embodiment 4) Embodiment 4 relates to a form in which advice regarding a standard is output by message output model 151 based on an image of a construction site and standard data related to the image. Regarding the content overlapping with that in Embodiments 1 to 3, the description is omitted.
[0151] When workers at a construction site carry out construction using incorrect implementation methods, based on the images of the construction site and the reference data related to the images, by using the message output model 151, advice regarding the standards can be transmitted to the unit 2 attached to the helmet worn by the workers. The advice regarding the standards includes the pointed-out content in the construction, the accuracy of the implementation method, or areas for improvement, etc.
[0152] For example, when more than a predetermined number of electric wires are inserted into an electric wire pipe, there is a high possibility of an overheating and burning accident occurring. In this case, by inputting an input prompt including an image of the construction site including the wiring of the electric wire pipe and reference data such as internal wiring regulations into the message output model 151, advice regarding the wiring of the predetermined number can be output. This makes it possible to prevent accidents.
[0153] The server 1 acquires an image of the construction site from the unit 2 attached to the helmet worn by the worker at the construction site. The server 1 acquires information regarding the acquired image of the construction site from the acquired image of the construction site. The server 1 may extract the feature amount of the image of the construction site, for example, using A-KAZE.
[0154] The server 1 acquires an image information vector of the information regarding the acquired image by vectorizing the information regarding the acquired image. The server 1 searches (extracts) reference data corresponding to a query (for example, "advice regarding the electric wire pipe") from the reference data DB155 based on the acquired image information vector.
[0155] Specifically, the server 1 calculates the similarity between the acquired image information vector and the vectors of the reference data (laws, regulations, rules, or guidelines, etc.) stored in the reference data DB155 respectively. The server 1 compares each calculated similarity with a predetermined similarity threshold. The server 1 extracts from the reference data DB155 the reference data corresponding to a similarity that is equal to or greater than the predetermined similarity threshold, or the file path of the document describing the reference data, etc.
[0156] Server 1 generates an input prompt that includes a query, information about the acquired image, the extracted reference data, or the file path of a document describing the reference data. Server 1 inputs the generated input prompt into the message output model 151 and outputs advice regarding the criteria. For example, the output advice regarding the criteria may be "Please put four in the wire pipe to prevent overheating and burning accidents."
[0157] Server 1 transmits the advice regarding the criteria output from the message output model 151 to unit 2 attached to the helmet worn by the operator. Unit 2 receives the advice regarding the criteria transmitted from Server 1. Unit 2 outputs the received advice regarding the criteria through speaker 27.
[0158] Figure 21 is a flowchart showing the processing procedure when transmitting advice regarding the criteria to unit 2. Regarding the content overlapping with Figure 9, the same reference numerals are used and the description is omitted. After executing the processing of step S103, the control unit 11 of Server 1 outputs the advice regarding the criteria from the message output model 151 (step S131). The address regarding the criteria may include link information (URL or file path of a document describing the reference data, etc.) for accessing the reference data.
[0159] The control unit 11 stores the advice regarding the criteria output from the message output model 151 in the danger prediction result DB153 of the mass storage unit 15 in association with the operator ID (step S132). Specifically, the control unit 11 stores the construction date, the image data of the construction site, the advice regarding the criteria, and the output time as one record in the danger prediction result DB153 in association with the operator ID.
[0160] The control unit 11 transmits advice regarding the standard to unit 2 via the communication unit 13 based on the unit ID (step S133). The control unit 21 of unit 2 receives the advice regarding the standard transmitted from server 1 via the communication unit 23 (step S231). The control unit 21 outputs the received advice regarding the standard via the speaker 27 (step S232). The control unit 21 ends the process.
[0161] According to the present embodiment, based on the image of the construction site and the reference data related to the image, it is possible to output advice regarding the standard by using the message output model 151.
[0162] According to the present embodiment, by outputting advice regarding the standard, it is possible to prevent accidents at the construction site in order to carry out the construction in the correct implementation method.
[0163] (Embodiment 5) Embodiment 5 relates to a form in which a danger prediction message for an operator is output to the administrator's computer. Note that descriptions of the contents overlapping with Embodiments 1 to 4 are omitted.
[0164] In the present embodiment, an administrator terminal (the administrator's computer) (not shown) is included. The administrator terminal is a terminal device that receives and displays information including a danger prediction message for an operator and an operator ID, etc. The administrator terminal is an information processing device such as a smartphone, mobile phone, tablet, or personal computer terminal.
[0165] FIG. 22 is a flowchart showing a processing procedure when outputting a danger prediction message and an operator ID to the administrator terminal. The control unit 11 of the server 1 acquires the operator IDs of a plurality of operators, the images of the construction site acquired from the unit 2 attached to the helmets worn by each operator, and the danger prediction messages for each operator from the danger prediction result DB 153 of the mass storage unit 15 (step S141). The control unit 11 transmits the acquired operator IDs of the plurality of operators, the images of the construction site, and the danger prediction messages to the administrator terminal via the communication unit 13 (step S142).
[0166] The administrator terminal receives the operator IDs of the plurality of operators, the images of the construction site, and the danger prediction messages transmitted from the server 1 (step S341). The administrator terminal displays the received operator IDs of the plurality of operators, the images of the construction site, and the danger prediction messages on the screen (step S342). The administrator terminal ends the process.
[0167] Note that, in the present embodiment, an example in which the danger prediction message for each operator is transmitted to the administrator terminal has been described, but the present invention is not limited thereto. For example, the server 1 may transmit the danger prediction message for the leader of each grouped group to the administrator terminal.
[0168] Specifically, the server 1 extracts the operator ID of the leader of each group from the operator DB 152 based on the group ID and the responsibilities of the operator (for example, management). The server 1 acquires the operator ID of the leader of each extracted group and the danger prediction message for the leader of each group from the danger prediction result DB 153. The server 1 transmits the acquired operator ID of the leader of each group and the danger prediction message for the leader of each group to the administrator terminal.
[0169] In addition, the administrator can determine the validity of the danger prediction message for the worker. For example, when the danger prediction message for a novice worker output from the message output model 151 is "You are an experienced worker...", the administrator determines that the danger prediction message for the worker is not valid.
[0170] In this case, the administrator terminal receives the determination result of the validity of the danger prediction message by the administrator. The administrator terminal transmits the danger prediction message for the worker to the server 1 together with the received determination result of the validity. The server 1 receives the determination result transmitted from the administrator terminal and the danger prediction message for the worker.
[0171] In addition, when it is determined by the administrator that the danger prediction message for the worker is not valid, the administrator terminal receives the correction of the danger prediction message by the administrator. For example, when it is determined by the administrator that the danger prediction message "Please connect wiring A to wiring B." is not valid, the administrator terminal may receive the corrected danger prediction message "Please connect wiring A to wiring C." by the administrator.
[0172] The administrator terminal transmits the received corrected danger prediction message to the server 1 together with the determination result of the validity. The server 1 receives the determination result and the corrected danger prediction message transmitted from the administrator terminal. The server 1 transmits the received determination result and the corrected danger prediction message to the unit 2 mounted on the helmet worn by the worker.
[0173] Note that the server 1 may also transmit the determination result and the corrected danger prediction message to the leader of the group to which the worker belongs, or to the unit 2 mounted on the helmet worn by other workers in the group.
[0174] Furthermore, by fine-tuning the message output model 151 using the corrected danger prediction message, the prediction accuracy of the danger prediction message can be improved. Specifically, the server 1 uses information on the image of the construction site corresponding to an inappropriate danger prediction message, data retrieved from each database (such as case data and reference data) based on the information on the image, and the corrected danger prediction message (correct answer) in a dataset (learning data) to relearn the message output model 151, enabling the output of a danger prediction message with high prediction accuracy.
[0175] According to this embodiment, by outputting a danger prediction message for the worker to the administrator terminal, the administrator can check the danger prediction message at the construction site at any time.
[0176] (Embodiment 6) Embodiment 6 relates to a form in which information on an image of a construction site and an input prompt including a worker's comment are input to the message output model 151 to output a danger prediction message. Note that descriptions of content overlapping with Embodiments 1 to 5 are omitted.
[0177] Comments from workers at the construction site can be received together with the image of the construction site. The worker's comment includes voice data based on the voice spoken by the worker, or text data input by the worker, etc.
[0178] For example, the unit 2 receives an image of the construction site and obtains information on the received image of the construction site (for example, information related to the danger of electric shock). Note that the process of obtaining information on the image is the same as that in Embodiment 1, so the description is omitted. The unit 2 receives the input of the worker's comment via the operation unit 26 or the microphone 28. For example, the received comment may be "Tell me about dangers other than electric shock", etc.
[0179] The unit 2 transmits information related to the acquired image of the construction site and the received comments of the worker to the server 1. The server 1 receives the information related to the image of the construction site and the comments of the worker transmitted from the unit 2. As in the processing in the first embodiment, the server 1 generates an input prompt including the information related to the received image of the construction site and the comments of the worker. The server 1 inputs the generated input prompt to the message output model 151 and outputs a hazard prediction message related to hazard prediction at the construction site.
[0180] In this way, by inputting an input prompt including both information about the image of the construction site and the worker's comments into the message output model 151, it is possible to predict not only the risk of electric shock, but also risks other than electric shock, such as being pinched, entangled, or falling.
[0181] Fig. 23 is a flowchart showing the processing procedure when a danger prediction message is transmitted to the unit 2 in the sixth embodiment. Note that the same reference numerals are used for the contents that overlap with Fig. 9, and the description thereof will be omitted.
[0182] After executing the process of step S201, the control unit 21 of the unit 2 acquires comments from the worker at the construction site (step S251). For example, the control unit 21 receives voice data based on the voice uttered by the worker via the microphone 28. Alternatively, the control unit 21 receives input of text data by receiving an input operation from the worker via the operation unit 26. The comments from the worker may include voice data, text data, or both voice data and text data.
[0183] The control unit 21 transmits the acquired images of the construction site and the comments of the worker to the server 1 via the communication unit 23 (step S252). The control unit 11 of the server 1 receives the images of the construction site and the comments of the worker transmitted from the unit 2 via the communication unit 13 (step S151). The control unit 11 executes the process of step S102.
[0184] FIG. 24 is a flowchart showing the processing procedure of a subroutine for generating an input prompt in Embodiment 6. Note that the same reference numerals are given to the contents overlapping with FIG. 11, and the description thereof is omitted.
[0185] After executing the process of step S16, the control unit 11 of the server 1 acquires the received operator's comment (step S31). The control unit 11 determines whether or not the operator's comment includes voice data (step S32). When the operator's comment does not include voice data (NO in step S32), the control unit 11 transitions to step S34 described later.
[0186] When the operator's comment includes voice data (YES in step S32), the control unit 11 analyzes it using a voice recognition engine or the like, and converts the voice data into text (step S33). For the conversion of voice data into text, for example, a voice text conversion tool such as Azure Speech API (Application Programming Interface) may be used. Alternatively, the control unit 11 may convert the voice into text using a learned model by machine learning such as AI (artificial intelligence). Note that the voice data text conversion process may be executed on the unit 2 side.
[0187] The control unit 11 extracts an article related to the operator's comment from the case data DB154 and the reference data DB155 in the mass storage unit 15 (step S34). Specifically, the control unit 11 calculates the similarity between the acquired operator's comment and the case data (such as danger, countermeasures, or similar cases) stored in the case data DB154, respectively. Based on each calculated similarity, the control unit 11 extracts from the case data DB154 a similarity that is equal to or higher than a predetermined similarity threshold, or case data corresponding to the highest similarity, or a file path of a document describing the case data.
[0188] Further, the control unit 11 calculates the similarity between the obtained operator's comment and the reference data (laws, regulations, rules, guidelines, etc.) stored in the reference data DB 155, respectively. Based on each calculated similarity, the control unit 11 extracts from the reference data DB 155 the similarity that is equal to or higher than a predetermined similarity threshold, or the reference data corresponding to the highest similarity, or the file path of the document in which the reference data is described, etc.
[0189] The control unit 11 extracts the sentences included in the extracted case data or reference data as sentences related to the operator's comment. When the control unit 11 extracts the file path of a document, it reads the data stored in each document. The control unit 11 may extract the sentences included in the read data.
[0190] In FIG. 24, the sentence extraction process based on the information related to the image and the sentence extraction process based on the operator's comment are respectively executed, but the present invention is not limited thereto. For example, the control unit 11 may execute a sentence extraction process based on both the information related to the image and the operator's comment.
[0191] Specifically, the control unit 11 calculates the similarity between the image information vector and the operator's comment vector obtained by the vectorization process and the vector of the case data stored in the case data DB 154, respectively. Based on each calculated similarity, the control unit 11 extracts the case data from the case data DB 154. Further, the control unit 11 calculates the similarity between the image information vector and the comment vector and the vector of the reference data stored in the reference data DB 155, respectively. Based on each calculated similarity, the control unit 11 extracts the reference data from the reference data DB 155.
[0192] The control unit 11 generates an input prompt including a preset query (e.g., "Today's work is to check the status of the high-voltage input panel of the UPS."), the extracted text, information regarding the image of the construction site, the operator's comment (e.g., "Tell me about the dangers other than electric shock."), and the text related to the comment (step S35). The control unit 11 ends the subroutine of the input prompt generation process and returns.
[0193] According to the present embodiment, it is possible to input an input prompt including information regarding the image of the construction site and the operator's comment into the message output model 151 to output a danger prediction message.
[0194] According to the present embodiment, by considering the operator's comment, it is possible to prevent accidents caused by various dangers at the construction site and improve the safety of the construction site.
[0195] The embodiments disclosed this time should be considered as illustrative in all respects and not restrictive. The scope of the present invention is shown not by the above meaning but by the claims, and it is intended that all modifications within the meaning and scope equivalent to the claims are included.
[0196] The matters described in each embodiment can be combined with each other. Also, the independent claims and dependent claims described in the claims can be combined with each other in all possible combinations regardless of the citation format. Furthermore, although the claims use a format (multi-claim format) for describing claims that cite two or more other claims, it is not limited thereto. A format for describing a multi-claim (multi-multi-claim) that cites at least one multi-claim may be used.
Explanation of Reference Numerals
[0197] 1 Information processing device (server) 11 Control unit 12 Storage unit 13 Communication unit 14 Reading Unit 15 Large-capacity Memory Unit 151 Message Output Model (Language Generation Model) 152 Operator DB 153 Danger Prediction Result DB 154 Case Data DB 155 Reference Data DB 156 Construction Content DB 1a Portable Memory Medium 1b Semiconductor Memory 1P Control Program 2 Unit 21 Control Unit 22 Memory Unit 23 Communication Unit 24 Camera 25 Camera Control Unit 26 Operation Unit 27 Speaker 28 Microphone 29 Display Unit 2P Control Program
Claims
1. Obtain an image of the construction site, Input an input prompt including information about the obtained image of the construction site into a language generation model for predicting danger related to danger prediction information at the construction site, and output a danger prediction message related to danger prediction at the construction site. Information processing method.
2. Input the input prompt into the language generation model, and output the predicted danger and countermeasures at the construction site as the danger prediction message. The information processing method according to claim 1.
3. Input the input prompt into the language generation model, and output similar cases associated with the danger prediction message. The information processing method according to claim 1 or 2.
4. Output link information for accessing the similar cases. The information processing method according to claim 3.
5. Obtain the attributes of the workers at the construction site, Identify the text corresponding to the obtained attributes, Input the input prompt including the identified text and information about the image into the language generation model, and output a danger prediction message according to the attributes. The information processing method according to claim 1 or 2.
6. Obtain the construction content of the day at the construction site, Input the input prompt including the obtained construction content of the day and information about the image into the language generation model, and output the danger prediction message. The information processing method according to claim 1 or 2.
7. Identify case data related to the obtained image from a database storing a plurality of case data including dangers, countermeasures, and similar cases at the construction site, Input the input prompt including the sentences included in the specified case data and the information about the image into the language generation model to output the risk prediction message. The information processing method according to claim 1 or 2.
8. The database stores a plurality of reference data regarding laws, regulations, rules or guidelines at the construction site. Refer to the database to identify the reference data related to the acquired image. Input the input prompt including the sentences included in the specified reference data and the information about the image into the language generation model to output advice regarding the standard. The information processing method according to claim 7.
9. Obtain the comments of the worker including the voice data or text data of the worker at the construction site. Input the input prompt including the information about the image of the construction site and the obtained comments of the worker into the language generation model to output a risk prediction message regarding risk prediction at the construction site. The information processing method according to claim 1 or 2.
10. A unit having a camera and a speaker mounted on a helmet or a chest pocket, and a computer communicating with the unit. The computer includes a control unit. The control unit Acquire an image of the construction site through the camera. Input the input prompt including the information about the acquired image of the construction site into the language generation model using the risk prediction information regarding risk prediction at the construction site, and output the risk prediction message regarding risk prediction at the construction site to the unit. The speaker of the unit Output the risk prediction message. Information processing system.
11. The camera has a camera-off mode The information processing system according to claim 10.
12. The unit is associated with the operator ID of the operator, Output the danger prediction message for the operator and the operator ID to the administrator's computer The information processing system according to claim 10 or 11.
13. A plurality of operators including a first operator are grouped, Output the danger prediction message for the first operator to the units of the other plurality of grouped operators The information processing system according to claim 10 or 11.
14. Acquire an image of the construction site, Input an input prompt including information about the acquired image of the construction site into a language generation model that uses danger prediction information related to danger prediction at the construction site, and output a danger prediction message related to danger prediction at the construction site A program for causing a computer to execute the processing.
Citation Information
Patent Citations
Risk prediction activity support system
JP2021033404A
Cited By
Work assisting device, work assisting system and work assisting method
JP2025134549A
Support system, support device, support method, and program
JP7793855B1