Program, computer, system and information processing method
An AI camera system addresses excessive communication traffic and administrative burden by analyzing video data to generate and transmit work-in-progress information, enhancing management efficiency in manufacturing environments.
Patent Information
- Application Number
- JP2025082044
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2025-05-15
- Publication Date
- 2025-10-31
- Estimated Expiration
- 2045-05-15
AI Technical Summary
Existing surveillance camera systems fail to provide comprehensive management and result in excessive communication traffic, burdening administrators with image data management and analysis.
Implementing an AI camera system that functions as a sensing, analysis, and transmission means to generate and transmit work-in-progress information by analyzing video data using image recognition models, reducing data volume through edge computing.
Reduces administrative burden and communication traffic by generating actionable work information efficiently, enabling effective management and monitoring in manufacturing environments.
Smart Images

Figure 0007762928000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to a program, a computer, a system, and an information processing method. [Background technology]
[0002] Currently, there is a serious labor shortage, especially in small and medium-sized manufacturing companies, and even if machinery and equipment is available, it may not be able to be operated. This labor force includes both workers and managers. Managers are expected to perform tasks such as quality control, education, and manufacturing operations. Of these, the burden on managers for monitoring work can be reduced by using surveillance cameras.
[0003] For example, Patent Document 1 discloses a surveillance camera system that can obtain an optimal image from among images captured by a plurality of surveillance cameras and can improve the communication efficiency of image data. [Prior art documents] [Patent documents]
[0004] [Patent Document 1] Japanese Patent Application Laid-Open No. 2006-86991 Summary of the Invention [Problem to be solved by the invention]
[0005] However, in Patent Document 1, image data to be sent to an external device is merely determined based on priority, and comprehensive management is not realized. Furthermore, in Patent Document 1, the amount of communication traffic required to send image data to a host computer is already large, so there is a demand for further reduction in the amount of data.
[0006] The present disclosure has been made in consideration of these points, and aims to provide a novel program, computer, system, and information processing method that can reduce the burden on administrators and reduce communication volume. [Means for solving the problem]
[0007] The program of the present disclosure is A program that causes an AI camera to function as a sensing means, an analysis means, and a transmission means, The AI camera acquires image data of the section of the work area where the worker works, the sensing means analyzes the video data using an image recognition model to generate numerical data corresponding to the shape or movement during the work; The analysis means analyzes the numerical data to generate work-in-progress information, The transmitting means transmits an output instruction signal to a server so that the work information is output in association with the section of the work place. It is characterized by:
[0008] The AI camera may include a sensor unit, The sensor unit generates auxiliary numerical data corresponding to the shape or movement of the section during work; In the program of the present disclosure, The analysis means can analyze the numerical data and the auxiliary numerical data to generate work-in-progress information.
[0009] In the program of the present disclosure, The analysis means may analyze the generated numerical data and generate, as work information, numerical information corresponding to a shape or movement during work.
[0010] The program of the present disclosure is The AI camera further functions as a comparison means, the comparison means or the server compares the generated numerical information with a predetermined threshold value and generates evaluation information from the comparison result; The transmitting means may transmit the evaluation information generated by the comparing means to the server.
[0011] In the program of the present disclosure, The analysis means or the server may use a trained regression model to analyze the numerical information to generate the probability that a specified task is being performed and the amount of the task as task information.
[0012] In the program of the present disclosure, The analysis means or the server may use a trained classification model to analyze the numerical information and generate, as work information, a classification result indicating whether a specified work content is being performed.
[0013] In the program of the present disclosure, The analysis means or the server may generate the presence or absence of an abnormality as work information by analyzing the numerical information using a learned anomaly detection model.
[0014] In the program of the present disclosure, The numerical information may be information for explaining the work content in numerical terms.
[0015] In the program of the present disclosure, The sensing means may select an existing image recognition model including a general-purpose image recognition model as the image recognition model or may notify the need for a new image recognition model at the time of initial setup before analysis of video data.
[0016] In the program of the present disclosure, The server receiving an image relating to work information to be input to the existing image recognition model for the selection or notification; outputting a display screen for adding annotation information to image data in the video; inputting the image data into the existing image recognition model to verify whether the annotation information is output, thereby determining whether to select or notify; The annotation information and the image data or video may be output to explain the work information.
[0017] The program of the present disclosure is The AI camera further functions as a reception unit and a model setting unit, When the reception means or the server receives attendance management information including work content, The model setting means or the server may set a trained model in accordance with the work content.
[0018] In the program of the present disclosure, The work information may be information that identifies any one of the work date and time, the individual worker, the section, the coordinates, and the work content.
[0019] In the program of the present disclosure, The server may receive attendance management information including work content and associate it with the work information.
[0020] In the program of the present disclosure, The server may calculate the progress of work as work information based on numerical information analyzed from the video data.
[0021] In the program of the present disclosure, The server may calculate the work accuracy as work information based on the numerical information analyzed from the video data.
[0022] In the program of the present disclosure, the server has a calculation unit that generates answer information corresponding to question information related to work information, The calculation unit may generate, as the response information, any one of text, graphs, reports, and images that explain the status of the work information.
[0023] In the program of the present disclosure, The server may allow an administrator to search for work information analyzed from the video data.
[0024] In the program of the present disclosure, The transmission means transmits the acquired video data to a manager terminal operated by a manager of the workplace, The server may identify video data corresponding to question information regarding work information from the manager terminal.
[0025] The computer of the present disclosure includes: A computer that functions as a sensing means, an analyzing means, and a transmitting means by executing a program, The computer acquires image data of a section of a work area where a worker performs work, the sensing means analyzes the video data using an image recognition model to generate numerical data corresponding to the shape or movement during the work; The analysis means analyzes the numerical data to generate work-in-progress information, The transmitting means transmits an output instruction signal to a server so that the work information is output in association with the section of the work place.
[0026] The system of the present disclosure comprises: An AI camera that captures video data of the work area where workers are working, a server that associates the work area sections with work information; A system comprising: The AI camera acquires video data of the section of the workplace where the worker works, the sensing means analyzes the video data using an image recognition model to generate numerical data corresponding to the shape or movement during the work; The analysis means analyzes the numerical data to generate work-in-progress information, The transmitting means transmits an output instruction signal to the server so that the work information is output in association with the section of the work area.
[0027] The information processing method of the present disclosure includes: An information processing method executed by a computer having a control unit, a step in which the control unit uses an image recognition model to analyze video data of a section of a work area where a worker performs work, thereby generating numerical data corresponding to the shape or movement during work; a step in which the control unit analyzes the numerical data and generates work information during the work; a step in which the control unit transmits an output instruction signal to a server so that the work information is output in association with a section of the work site; The present invention is characterized by the following. [Effects of the Invention]
[0028] According to the present disclosure, it is possible to provide a novel program, computer, system, and information processing method that can reduce the burden on the administrator and reduce the amount of communication traffic. [Brief explanation of the drawings]
[0029] [Figure 1] 1 is a diagram schematically illustrating a configuration of an exemplary information processing system according to an embodiment of the present disclosure. [Figure 2] 2 is a diagram showing an exemplary flow of information processing by the information processing system shown in FIG. 1. [Figure 3] 1. FIG. 4 is a diagram showing another exemplary flow of information processing by the information processing system shown in FIG. [Figure 4] FIG. 10 is a diagram illustrating an example of input and output of information for creating a procedure manual. [Figure 5] 10A and 10B are diagrams illustrating examples of input and output of information for creating a poster. [Figure 6] FIG. 10 is a diagram illustrating an example of a terminal screen for displaying generated work information. [Figure 7] FIG. 10 is a diagram illustrating an example of search conditions for generated work information. [Figure 8] FIG. 1 is a diagram illustrating an exemplary flow of information processing in an information processing method according to an embodiment of the present disclosure. [Figure 9] FIG. 10 is a diagram showing another exemplary flow of information processing in the information processing method according to the embodiment of the present disclosure. [Figure 10]FIG. 10 is a diagram showing another exemplary flow of information processing in the information processing method according to the embodiment of the present disclosure. [Figure 11] FIG. 10 is a diagram showing another exemplary flow of information processing in the information processing method according to the embodiment of the present disclosure. [Figure 12] FIG. 10 is a diagram showing another exemplary flow of information processing in the information processing method according to the embodiment of the present disclosure. [Figure 13] FIG. 10 is a diagram showing another exemplary flow of information processing in the information processing method according to the embodiment of the present disclosure. [Figure 14] FIG. 10 is a diagram showing an example of an image to be referenced when setting (selecting and generating) an image recognition model. [Figure 15] FIG. 10 is a diagram showing an example of a terminal screen for displaying and determining work information in real time. DETAILED DESCRIPTION OF THE INVENTION
[0030] Hereinafter, embodiments of the present disclosure will be described with reference to the drawings, but the invention according to the present disclosure is not limited thereto. Figures 1 to 15 are diagrams showing an information processing system 1 and an information processing method according to the present embodiment.
[0031] [Information Processing System 1] An information processing system 1 according to the present disclosure, shown in FIG. 1, includes multiple AI cameras (computers) 10 and a server (computing unit) 34. In FIG. 1, AI cameras 10A, 10B, and 10C are installed in sections A to C of a workplace X or Y, respectively, and acquire video data for each section. In the example shown in FIG. 1, sections A and B are installed in workplace X, and section C is installed in workplace Y. Sections A and B are connected by a passageway, which may be further divided into sections a to c. An AI camera 10 may be installed in each of sections a to c, or one AI camera 10 may be installed in each passageway (not shown). The workplaces, passageways, sections A to C, and sections a to c are not limited. The information processing system 1 may further include an administrator terminal (i, ...) and / or worker terminals (a to c, ...) (collectively referred to as "terminals"). The AI camera 10 and the server 34 are communicably connected to these terminals via a communication network (not shown) such as the Internet. Each of the components of the information processing system 1 will be described below.
[0032] <AIカメラ(コンピュータ)10(10A、10B、10C、…)> Figure 1 shows a schematic configuration of an AI camera 10 according to the present disclosure. The AI camera 10 can be configured from an industrial computer, a personal computer, or the like, and as shown in Figure 1, includes a control unit 12, an imaging unit 28, a sensor unit 29, a storage unit 30, and a communication unit 32. The control unit 12, storage unit 30, and communication unit 32 may be included in the housing (main body) of the AI camera 10 as an edge device, to which the imaging unit 28 and sensor unit 29 may be connected for wired or wireless communication.
[0033] (Control unit 12) The control unit 12 is composed of a CPU (Central Processing Unit), a GPU (Image Processing Unit), an AI inference device, etc., and controls the operation of each element of the AI camera 10. Specifically, the control unit 12 executes programs stored in a storage unit 30 (described later) to function as a reception unit 14, a sensing unit 16, a model setting unit 18, an analysis unit 20, a comparison unit 22, a transmission unit 24, etc. (see FIG. 1). Each of these units will be described later.
[0034] (Image capture unit 28) The imaging unit 28 acquires video data of the section of the work area where the worker is working. As such, there are no particular limitations as long as it can capture images of the work situation.
[0035] The imaging unit 28 may be installed anywhere in each section of the workplace, but is preferably installed in a position where it can capture images of the worker's work status. For example, depending on the type of work, it may be installed in a position where it can capture images of the worker's entire body, arms, fingertips, hands, etc. Furthermore, along with the worker, images of items related to the work (e.g., tools, machinery, equipment, finished products, parts, etc.) may also be captured (people and items are collectively referred to as "objects"). When installed in an aisle or section a thereof, it is preferable to install the imaging unit 28 in a position where it can track (identify) the worker.
[0036] (Sensor part 29) The sensor unit 29 is intended to detect objects, and generates auxiliary numerical data corresponding to the shape or movement of the worker in the area where the worker is working. The "auxiliary numerical data" may be coordinate data of the object (combined type), such as the numerical data (generated from video data / image data acquired by the imaging unit 28) described below, or may be numerical data that corresponds to the shape or movement of the object during work independently of the imaging unit 28 (additional type). An example of the combined type is data that detects the presence (shape or movement) of an object, and an example of the additional type is data that detects the shape or movement based on the five senses, such as sound or pressure (for example, event data for a switch, etc.).
[0037] By using such auxiliary numerical data, the trained model can be made lightweight (reducing the amount of learning) and data that cannot be obtained by the imaging unit 28 can be acquired as an auxiliary data. Furthermore, the AI camera 10 can use the auxiliary numerical data to start analysis from video data or compare it with a predetermined threshold, which can be used to determine the start of a specific task or the progress of the task (which can also have a power-saving effect). Due to the characteristics of the product, the imaging unit 28, such as a camera, is often installed in a position that allows the user to view the hand, but since it is difficult to obtain information from the lower body, this can be assisted by the sensor unit 29. In this way, the installation position of the sensor unit 29 can be appropriately determined separately from the imaging unit 28 depending on the type, purpose, etc.
[0038] Examples of the sensor unit 29 include a microphone sensor (microphone), a temperature sensor, a motion detection sensor (to detect objects), a gyro sensor, an acceleration sensor, a pressure sensor, a switch sensor (to detect the start of work), a variable resistance sensor, a gas sensor, a lidar sensor, a ranging sensor, etc.
[0039] As a specific example, pressure sensors can be used to confirm whether a worker is seated, check work progress, and measure concentration. By placing multiple pressure sensors on the seat like cushions, it is possible to obtain numerical data even when working while sitting and moving little (for example, measuring the progress and concentration of assembly, painting, and inspection work that involves sitting for long periods of time). Concentration can be estimated from changes in the position of the center of gravity on the seat, and it can be estimated that the longer a worker remains seated during continuous sitting work, the lower their concentration will be. Furthermore, in assembly line work on a conveyor belt or simple tasks, the movements of the work are cyclical, and this can be estimated from the movement of the worker's center of gravity on the seat.
[0040] Another example of the sensor unit 29 is a microphone that can detect the start of work (movement) from sound. It can also ascertain not only the presence of an object but also its position based on the intensity and pitch of the sound. Furthermore, for example, it can estimate the amount of paint to be applied from the sound of the paint being applied and the duration of the sound (quantity and quality of the sound), which can complement and enrich work progress information as work information, and can also obtain information on the remaining amount of consumables.
[0041] Another example of the sensor unit 29 is thermography, which can detect drowsiness during work (before work stops). This can prevent oversights and accidents due to dozing during night shifts when there is less supervision. An anomaly detection model trained on normal thermography data can detect drowsiness based on changes in facial surface temperature (especially the nose). Drowsiness can also be detected by the imaging unit 28 through image recognition and time series analysis (e.g., classification using a time series analysis model such as an LSTM model, or work detection using a classification model trained on periodic features using Fourier transform). Blink detection can also be considered to detect drowsiness. However, because these methods have low accuracy depending on the position of the imaging unit 28, thermography can replace the imaging unit 28 and image recognition model that detects blinking. Furthermore, when there are ethical issues with using a camera to capture images, non-visual object detection methods (sensor unit 29) are particularly effective.
[0042] Another example of the sensor unit 29 is an IMU (inertial sensor, etc.) that can prevent forgetting work procedures (poka-yoke). An IMU can easily quantify the movements of tools and other items that are constantly used in work. For example, in assembly work, it is possible to prevent omissions in work processes by detecting abnormalities in the movement of tools. This is because tool movements are cyclical if the work content is the same. While sensors cannot be attached to human bodies or parts, the sensor unit 29 can be installed relatively easily on tools, containers, and machines. By acquiring and analyzing auxiliary numerical data from these, it is possible to analyze the progress and accuracy of work. More specifically, parts and human bodies can be detected using an image recognition model, and tools can be digitized by the sensor unit 29, making the image recognition model a lightweight model.
[0043] Another example of the sensor unit 29 is a tactile switch, which can detect the start of a task. When the start of a task is difficult to determine through time series analysis, such a switch can easily identify the start of the task. For example, in a hand-washing detection system, a tactile switch can be installed on a hand soap container to determine when hand washing has begun and whether the "hand soap" has been used. The use of such an item can be detected, and the image capture unit 28 can be activated upon detection. This also saves power for the image capture unit 28. Similarly, in a roller-applying detection system, a tactile switch can be installed on the hook that hangs the roller to detect the start of the task. It can also be used as a trigger for the start of a task in tasks that involve a person "pushing" an object. Detecting the start of a task and analyzing the progress of the task sometimes require separate models, and the use of a tactile switch or similar can eliminate the need for the former model.
[0044] Another example of the sensor unit 29 is work progress detection using a load cell or gas sensor. By adding detection based on the weight or smell of the actual product or component to the visual (visual) work progress detection seen by the imaging unit 28, detection accuracy can be improved. For example, measuring the weight of a finished product during assembly can prevent components from being forgotten to be installed. Furthermore, the stable weight of an item untouched by human hands can be detected and used to grasp the progress of assembly. Another example is the use of a VOC (volatile organic compound) gas sensor to detect soldering work and confirm its progress. Gas sensors can be used when work steps involve handling alcohol-based solvents that emit specific gases, such as VOCs. While soldering irons can be detected using an image recognition model, the weight of the model can be reduced by using a gas sensor. Furthermore, this can handle cases where hand movements in the video data appear normal but actual work, such as installation or soldering, is not being performed.
[0045] (Storage unit 30) The storage unit 30 is configured, for example, with a hard disk drive (HDD), random access memory (RAM), read-only memory (ROM), solid state drive (SSD), etc. Furthermore, the storage unit 30 is not limited to being built into the AI camera 10, but may also be a storage medium (for example, a USB memory) that can be detachably attached to the AI camera 10.
[0046] Examples of information and data stored in the memory unit 30 of the AI camera 10 include video data captured by the imaging unit 28 and its image data, programs executed by the AI camera 10 (control unit 12) and various information acquired (numerical data, numerical information, etc.), registration information such as accounts corresponding to the manager terminal i operated by the workplace manager (unless otherwise specified, simply referred to as the "manager") and the worker terminal a operated by the worker, trained models, etc. Note that these programs, registration information, etc. may be stored in other storage means (such as a cloud server) rather than in the memory unit 30.
[0047] The acquisition and processing methods will be described later, but in this specification, "data" and "information" are used as different concepts even if they indicate the same content, unless otherwise specified.
[0048] First, "data" is a collection of numbers and symbols that cannot be understood by a person such as a manager. Examples of "data" include video data, image data, and numerical data, and these data are processed in this specification in this order. Of these, "numerical data" is necessary for analyzing work information extracted from video, sensor, and trained model analysis results, but is not directly presented to a manager. Specific examples of numerical data include coordinate data of objects in video (image) data acquired as work status, distances between object coordinates, positional relationship vectors, the probability that work is in progress, labels of classification models, output results of regression models, and reliability. While "image data" can constitute part of video data, it cannot constitute work information by itself.
[0049] On the other hand, "information," like "data," is information that can be understood by managers and other personnel (meaningful information), even if it is a collection of numbers and symbols. Such "information" is characterized by its small amount of information, particularly compared to video data and image data. Typically, when multiple cameras are installed in a factory, video data transmission must be limited to ensure communication stability and reduce impact on other operations. In contrast, the information processing system 1 disclosed herein processes video using the AI camera 10 to extract meaningful information and reduce data volume. For example, one second of full HD (1920x1080 pixels, 30 FPS) video can reduce the data volume from tens of megabytes to several to hundreds of bytes. In this case, the AI camera 10, acting as an edge device in the workplace, further improves information processing efficiency for work management (details will be described later).
[0050] An example of "information" is work information. "Work information" specifies information such as "when, who, where, and what" of a work. In other words, work information can be said to be information that specifies the date and time of work, the individual worker, the area, coordinates (within the video data), or the work content. Work information can also include work evaluation information, work progress information, work accuracy information, history, etc. In the information processing system 1, such work information can be output to each terminal as a digital twin (GUI).
[0051] Here, the "task date and time" may be the time information when the AI camera 10 analyzes the video data, or it may be the time information included in the video data. The "worker," "section," and "task content" can be obtained from the AI camera 10's identification information associated with facial recognition, shape information such as markers, or attendance management information (including worker and workplace information). The "section" may be the section (e.g., A-C) where the AI camera 10 is installed, or an arbitrary section set within the workplace or image. The "task content" varies depending on the workplace and is not limited to this. However, the work content can be appropriately categorized and subdivided based on the movements, posture, shape, and process of the work (job, task, etc.) such as hand washing and rolling. Examples of work content include actions such as holding, carrying, assembling, and releasing parts. Specific examples include the order of rotating, moving, and looking at products by hand during inspection work; changing paint and applying each paint during painting work; machine startup and shutdown procedures; workpiece installation, tool installation, tool installation position, and adjustment of the push amount during machining work. Such work content (work information) serves as a guideline for the output method of the GUI for work management, as well as for the training, configuration, and selection of various trained models, which will be described later.
[0052] In this specification, the term "workplace" is not limited to, and may include factories, hospitals, restaurants, etc., as long as the AI camera 10 of the present disclosure performs work (hand washing, inspection, etc.).
[0053] "Attendance management information" is information used by managers to manage worker attendance and workplace management. Attendance management information may include work content and worker information, and worker information may include the work content of the day (plans, schedules, shifts, and break times). Attendance management information may also include "work information (when, who, where, and what)," which is past (historical) information obtained through analysis by the AI camera 10 or server 34. In this way, it may be possible to automate the evaluation of work attitudes based on work information (including information on procedure compliance, work progress, and work accuracy) in the database (storage unit) of the AI camera 10 or server 34.
[0054] In this specification, "numerical information" refers to information expressed as meaningful numerical values analyzed from numerical data such as coordinates, and these numerical values are used to explain and evaluate the work content (work information). Numerical information is included in work information, but numerical information may also be verbalized to become work information. Numerical information can be generated by analyzing the aforementioned numerical data using a rule base (including a correspondence table) or a trained model (machine learning method). Examples of numerical information include the distance of an object's coordinates, a positional relationship vector, the probability that work is in progress, the label of a classification model, the output result of a regression model, and reliability. Furthermore, time-series information such as the position history of these items is also included.
[0055] Furthermore, "shape information" that represents the shape of a person or object can be generated from numerical data and / or numerical information. Examples of shape information include the shape, position, and angle of a person (whole body or part) or an object (including parts). If such shape information, worker information, and work content information are available, the amount of calculation required can be reduced.
[0056] The following is an example of information processing from numerical data to task information. First, when performing "hand skeleton detection" using sensing, coordinate data for 21 points on one hand can be obtained as numerical data. This coordinate data is converted into shape information (meaningful numerical information) such as "hand shape classification" and "finger angle." Then, by referring to this "hand shape classification" and previous task information (thresholds, etc.), task information such as "what task is being performed (task content)" and "task accuracy (evaluation information)" is generated, which can also be verbalized for explanation.
[0057] (Communication Department 32) The communication unit 32 connects the control unit 12 to an external device (e.g., a server 34, an operator terminal i, etc.) wirelessly or via a wired connection so that the control unit 12 can communicate with the external device. The control unit 12 can send and receive signals to and from the external device via the communication unit 32.
[0058] (others) Additionally, the AI camera 10 may be provided with an operation unit (input unit) such as a numeric keypad that allows a manager or worker to issue various commands to the control unit 12, and a display unit such as a monitor that displays various screens in response to display command signals from the control unit 12. In some embodiments, a display operation unit such as a touch panel that integrates the operation unit and display unit may be used. That is, in some embodiments, the AI camera 10 may also function as the manager terminal i or the worker terminal a.
[0059] Furthermore, as will be explained below, the AI camera 10 installed in the workplace may be provided with a warning unit that issues a warning of an abnormality using light, sound, or other means when an abnormality is detected (Figs. 2 and 3). Here, "abnormalities" include accidents, drowsiness, and slow work progress. Alternatively, a warning may be issued using a specific light or sound when work progress is too fast, or a specific warning may be issued as "advice" when specific work information is obtained.
[0060] (Details of control unit 12) (Method of reception 14) The reception means 14 receives various signals including attendance management information from an external device. For example, the reception means 14 can receive attendance management information including work content from the manager terminal i or the server 34. The reception means 14 may also receive instructions regarding the trained model to be used from the server 34.
[0061] (Sensing means 16) The sensing means 16 generates numerical data corresponding to the shape or movement during work by analyzing the video data using an image recognition model. As described above, the numerical data includes coordinate data. To obtain such coordinate data, the image recognition model may be used, or the coordinate data may be generated using data analyzed (feature extracted) by the image recognition model.
[0062] The sensing means 16 may use an image recognition model to analyze image data included in the video data and generate numerical data corresponding to the shape or movement of the worker during work. Furthermore, the sensing means 16 may continuously perform analysis (sensing) in conjunction with the imaging of the imaging unit 28. In this way, not only can motion information be acquired by viewing multiple numerical data (coordinates) in a time series (time series model), but also shape information that can be constructed from a single numerical data (coordinate) in the image data (position feature model). In the case of a time series model, numerical data may be generated continuously or every 1 to 180 seconds. For example, when assessing byte terrorism, a "position feature model" that simply processes images as images is used as an image recognition model, while when assessing work progress, a "time series model" that processes images as video is used (e.g., 2DCNN and 3DCNN).
[0063] Examples of "image recognition models" include skeleton detection models and object detection models. Depending on the work content and the detection (image recognition) method, one or both of the skeleton detection model and the object detection model can be used. Furthermore, these models may be publicly known general-purpose image recognition models or custom-made image recognition models.
[0064] In this way, the sensing means 16 may use an existing (specialized) image recognition model, including a general-purpose image recognition model, as the image recognition model used for sensing during initial setup before analyzing the video data. It can also notify (to the administrator terminal i, etc.) the need for a new image recognition model or selection of a new image recognition model. In other words, if the target detection target can be detected using either a known general-purpose image recognition model or a specialized image recognition model corresponding to existing work information, it can be automatically found and used (by rule-based or machine learning; the latter is referred to here as sensing AI). The selection and trial shooting (inference test) of an existing image recognition model can be performed by either the AI camera 10 or the server 34, and learning of a new image recognition model is performed by the server 34.
[0065] The selection of an existing image recognition model and the determination of whether a new image recognition model is necessary can be performed together with the output of work information normally required in each workplace. The following description will be given with reference to FIGS. 4 and 5. Examples of output modes for work information (including annotation information and image data entered by a user) include, but are not limited to, paper output such as a work procedure manual (FIG. 4(f)) or a notice (FIG. 5(e)), and digital output such as a training video for new employees. FIG. 4 shows a GUI (display screen) for creating a procedure manual displayed on a manager terminal i, etc., and examples of input and output of each piece of information. FIG. 5 shows a GUI (display screen) for creating a notice displayed on a manager terminal i, etc., and examples of input and output of each piece of information.
[0066] First, the server 34 displays "Add a new poster," "Add a new procedure manual," "Add AI as blocks," and "Add AI as code" on the GUI (display screen) of the terminal as shown in Figures 4(a) and 5(a), and accepts the selection of "Add a new poster" or "Add a new procedure manual" from the user (administrator, etc.). Note that "Add AI as blocks" and "Add AI as code" are buttons used when creating a new image recognition model (AI) through programming.
[0067] Next, as shown in the examples of Figures 4(b) and 5(b), a "Select Work" button is displayed to allow the user to select the type of work, and the user selects a category appropriate for the work they wish to manage. Images (templates) of various work scenes and existing image recognition models that can recognize these images are stored in the server 34, and candidates for existing image recognition models that can be selected via this "Select Work" button are narrowed down.
[0068] After selecting the type of work, the user is prompted to select the video of the scene they want to recognize (manage) (as shown in Figure 4(c) and Figure 5(c)), or to input (upload) the video they want the image recognition model to recognize.
[0069] Next, as shown in Figure 4(d), the user is asked to divide the video into task steps (task division). Here, the total number of steps is input. This information (interval annotation) can be used for time series analysis.
[0070] Next, as shown in Figures 4(e) and 5(d), the user is asked to enter text to match each work step with a diagram (image). Specifically, the user is asked to enclose the object to be recognized in a bounding box and then to describe each object, its actions, and the work content in text. The objects enclosed by the bounding box are the targets for detection or learning by the image recognition model. The text is the same when read by a human and is output as a procedure manual or poster, but it also serves as ground truth data (labels) for the image recognition model. Here, the user is asked to explain the "parts" and "tools" used in each step, the "location" where the work is performed, the "names" of the "semi-finished products" and "finished products" assembled in that step, and important points to note. Each content is explained for each bounding box. For example, Figure 5(d) shows an example in which information such as "who," "where," "when (conditions)," and "what" is explained in text, and the user is also asked to distinguish people in the image by their helmets (or badges or other decorations) and determine the operating conditions by the presence or absence of warning lights.
[0071] This annotation information (bounding boxes and text) and video are then input into an existing image recognition model, and a rule-based or machine learning approach is used to determine whether each model can recognize certain classes (objects). For example, if a part called "heat sink A" can be detected with high accuracy using an existing image recognition model's pre-trained class called "Lego blocks," the "Lego blocks" class can be used as the "heat sink A" class (simply change the class name). For example, if the coordinates (object location) detected by the image recognition model are detected with the same confidence level of 90%, the rule-based approach can be determined to be highly accurate. In this way, the rule-based approach has a 1:1 correspondence, making it easy for humans to control and understand. Furthermore, in cases where "firm tofu" is also recognized as the "heat sink A" class in addition to the "Lego blocks" class, or in multivariate cases, label (class) selection and weighting are possible using machine learning (sensing AI) such as SVM, enabling improved accuracy (optimization).
[0072] On the other hand, if an existing image recognition model cannot detect annotated objects with high accuracy or cannot output annotation information, the system will notify the user of the need for a new image recognition model (proposing custom-made model training to the user).In this custom-made model training, the information obtained through the GUI (total number of steps, video, segmentation, bounding box, text, etc.) can also be used as training data.
[0073] Typically, custom-made image recognition models are required when skeletal detection is not possible due to workers wearing special clothing or gloves, when parts or products are special or require precise detection, or when detailed classification of assembly processes is required. By creating custom-made image recognition models, it is possible to create models dedicated to each task and / or worker. On the other hand, general-purpose image recognition models can be used (selected) when skeletal detection is possible, when the worker is simply wearing a mask or helmet, or when detecting general objects such as screws and solder.
[0074] (Model setting means 18) The model setting means 18 sets a trained model (described later) in accordance with the work content. Here, "setting" a trained model includes selecting (switching) a trained model for the purpose of optimizing the analysis results, as well as adjusting and combining parameters of trained models. This setting can be performed daily at predetermined intervals in accordance with attendance management information, and is preferably performed based on attendance management information including the work content, worker, location (section, booth, seat), etc.
[0075] The model setting means 18 can also set a trained model for analysis according to the work content or the worker. This enables more precise analysis tailored to the work content and the worker. In particular, it can prevent situations where the work analysis results, such as work classification and work accuracy, are poor due to individual differences in the worker's physique, habits, rhythm, etc., even if the actual work results are good. In this way, the model setting means 18 can change the threshold for generating work information by checking the work attendance management information.
[0076] The "setting" by the model setting means 18 can also be performed automatically (using a rule-based or machine learning model), just like the sensing AI described above. Furthermore, when learning a model for each individual, parameters that are expected to vary from person to person can be automatically optimized using a rule-based system, additional annotations by the administrator (selecting a video that serves as a model for the individual's work; Figure 4), and reinforcement learning (reinforcement learning through evaluation of product defect rates, work efficiency, etc.).
[0077] (Analysis method 20) The analysis means 20 analyzes the numerical data generated by the sensing means 16 to generate work information during work (when video data is acquired). The analysis means 20 may analyze the generated numerical data to generate numerical information corresponding to the shape or movement during work as work information. The numerical information (which is meaningful information as described above) can be used as work information as it is. For example, when the number of workers or the number of items is expressed as work information, the numerical data can be analyzed using a rule base or a trained model (machine learning method) to create work information based on the numbers.
[0078] As an example of generating shape information (meaningful numerical information) based on a rule base (such as a preset correspondence table), the analysis means 20 can convert skeletal coordinate data generated by a skeletal detection model into the angle of each joint and obtain numerical information on the angle (such as the angle of the index finger or the angle of the elbow). Furthermore, the skeletal coordinate data may be classified by pose and used as a label value representing the posture during the inspection work, the shape of the hand during the inspection work, etc. As another example, the analysis means 20 can detect the coordinates of a soldering iron and a circuit board from the coordinate data generated by an object detection model and obtain numerical information on their positional relationship. Furthermore, the analysis means 20 may detect the coordinates of a finger and an assembled product from such coordinate data and convert the positional relationship into a vector (numerical information). As an example of work information (work content), assembly consists of multiple steps, and after predicting and classifying each step, the accuracy of the entire process can be evaluated.
[0079] As a machine learning method for the trained model, supervised learning, unsupervised learning, reinforcement learning, deep learning, etc. can be adopted. In addition, there are no particular limitations on the method for generating the trained model, as long as "numerical data" is input and "numerical information" is output. Below, we will explain each of the suitable machine learning methods for generating work information.
[0080] The analysis means 20 may use a trained regression model to analyze the numerical information to generate the probability that a predetermined task (e.g., a predetermined shape or movement) is being performed and the amount of the task as task information. For example, the task information may include categorical information (also referred to as categorical data or class information) corresponding to the task and the probability (certainty) that the task is being performed. The analysis means 20 may also analyze the numerical data generated by the sensing means 16 to generate task information.
[0081] The training method for the trained regression model is not particularly limited; as mentioned above, it is sufficient to input "numerical information" such as coordinate data and movement vectors and output "task information" such as important quantities (movements) and probabilities of the work. For each piece of numerical information (coordinates, vectors, etc.) corresponding to the work content, the model can be assigned correct answer data (label) for the probability of that work being performed and a correct answer target (objective variable) for the amount of work performed. As an example of the "amount of work," when detecting the amount of cutting using a lathe (rather than detecting the amount of workpiece grinding using a fixed AI camera 10), the amount of cutting can be estimated by observing the worker's hand movements (skeleton, shape). For example, when obtaining the "feed rate" of machining using an old lathe without sensors as output, the skeletal coordinates of the hands, waist, and head of the upper body can be used as input to estimate the lathe's "feed rate." It is also possible to train a model to learn the movements specific to a particular tool, but this can also be applied to other tools with similar shapes and movements. Reinforcement learning can also be used, with additional information added as appropriate.
[0082] Furthermore, the analysis means 20 may use the trained classification model to analyze numerical information to generate, as work information, a classification result indicating whether a predetermined work content (for example, a predetermined shape or movement) is being performed. Here, too, the analysis means 20 may generate, as work information, for example, categorical information corresponding to the work content and the probability (certainty) that the work content is being performed as a classification result. Furthermore, the analysis means 20 may also analyze numerical data generated by the sensing means 16 to generate work information. This can also be used for anomaly detection.
[0083] Methods for training trained classification models include SVM, CNN, RNN, LSTM, and other rule-based methods (such as thresholds), classical analysis methods such as Fourier transforms and statistical methods (linear classification, k-nearest neighbor methods), and combinations of these, but there are no particular limitations; as mentioned above, it is sufficient if "numerical information" is input and "classification results (work information)" can be output.
[0084] For example, the analysis means 20 uses a trained classification model generated by machine learning using training data including task details and numerical information related to each task (coordinates and polygons of each object), to generate task details (classification results) corresponding to the numerical information generated by the sensing means 16 as task information. In this way, assuming that the skeleton of the hand is detected, the hand shape can be classified as "grabbing an object" or "releasing an object," etc.
[0085] The analysis means 20 may generate the presence or absence of an abnormality as work information by analyzing the numerical information using the trained anomaly detection model. Here, the analysis means 20 may also generate, for example, categorical information corresponding to the presence or absence of an abnormality as work information. The analysis means 20 may also analyze the numerical data generated by the sensing means 16 to generate the presence or absence of an abnormality. Examples of "abnormalities" include accidents, dozing, and part-time work terrorism, and may also include actions or behaviors unrelated to the set work information.
[0086] There are no particular limitations on the learning method of the trained anomaly detection model, and it can use, for example, unsupervised learning clustering, autoencoder, GAN, etc. It is sufficient if it can output "anomaly" if something is different from normal.
[0087] Furthermore, the analysis means 20 can analyze not only the numerical data but also the auxiliary numerical data to generate work information during the work. The coordinates from the imaging unit 28 and the coordinates from the sensor unit 29 do not necessarily need to be the same, and they can be processed without any problems as input for machine learning even if they are not in the same coordinate system. However, when adjusting the coordinate system to make the data understandable to humans, it is preferable to adjust the coordinates based on the fixed position of the tool on which the sensor unit 29 is attached (calibration).
[0088] (Comparative means 22) The comparison means 22 generates evaluation information from the comparison result obtained by comparing the generated numerical information with a predetermined threshold. The comparison means 22 can generate "evaluation information" as work information that evaluates each task (movement, etc.) from the difference between the numerical information and a threshold (which may include an upper threshold and a lower threshold; in other words, a predetermined range). The threshold, evaluation information, evaluation criteria, etc. can be set appropriately by the manager depending on the task content. The comparison means 22 may generate evaluation information only when the numerical information exceeds a predetermined threshold, or may generate evaluation information by classifying the difference between the numerical information and the threshold into numerical ranges of a predetermined width (for example, evaluation information may be generated as "excellent," "poor," or "abnormal" depending on the magnitude of the difference).
[0089] Furthermore, the comparison means 22 may compare the confidence level of the probability or classification result generated by the analysis means 20 with a predetermined threshold to determine whether or not to generate task information (a combination of a regression model or a classification model with a threshold comparison is also possible). In this way, the amount of calculation can be reduced. For example, in the task (action) of crossing fingers, there are cases where the fingers are crossed at the "tip" and cases where the fingers are crossed at the "base of the fingers." If the fingers are crossed somewhere in between these, the threshold cannot be compared with either the tip or base of the fingers, so the action can be determined to be "in progress" or task information can be not generated.
[0090] (Transmission means 24) The transmission means 24 transmits an output instruction signal to the server 34 so that work information is output in association with a section of the workplace. The "output instruction signal" is a signal that instructs the server 34 to output work information associated with (corresponding to) a specific section of the workplace in response to, for example, an operation or question from a manager. That is, the server 34 transmits and displays the work information, including the numerical information, received from the transmission means 24 (AI camera 10) to the manager terminal i, and when an abnormality is detected as described above, the server 34 may not only display information indicating an abnormality on the manager terminal i, but may also output sound or light.
[0091] The transmission means 24 can also transmit the generated numerical information and evaluation information as work information to the server 34. By storing the above information in association with each other (creating a database), the work information can be unified in the server 34.
[0092] Furthermore, the transmission means 24 (AI camera 10) can also transmit the video data acquired by the imaging unit 28 to a manager terminal i operated by the manager of the workplace.
[0093] <Server 34> The server 34 of the present disclosure can be configured from an industrial computer, a personal computer, or the like, and includes a control unit 36, a storage unit 38, a calculation unit 40, and a communication unit 42, as shown in FIG.
[0094] (Control unit 36) The control unit 36 is composed of a CPU, GPU, AI inference device, etc., and controls the operation of each element of the server 34. By executing a program similar to the program stored in the memory unit 30 of the AI camera 10, the control unit 36 can perform all or part of the functions of the aforementioned reception means 14, model setting means 18, analysis means 20, comparison means 22, transmission means 24, etc. (see Figures 2 and 3). In other words, the AI camera 10 and the server 34 can perform similar functions, and it is possible to set which device will perform which function depending on the processing capabilities of the AI camera 10 and the server 34. Note that various trained models may be generated in the server 34 using the above-mentioned learning method.
[0095] One of the AI camera 10 and the server 34 may be a cheaper device with lower computing power, while the other may be a device with higher computing power. For example, the AI camera 10 may generate numerical information, and the server 34 may generate work information from the numerical information. As an example, the AI camera 10 may first use a trained regression model to predict the coordinates to which an object will move 10 seconds from the coordinates of the detected worker, converting this information into "numerical information," and then transmitting this numerical information to the server 34. Once the server 34 receives the numerical information, it may compare the coordinates of the no-entry area defined by the user with the predicted coordinates of the "numerical information," and generate the comparison result as "work information" indicating dangerous work.
[0096] In addition, the server 34 (control unit 36) can also calculate the work progress and work accuracy as work information based on the numerical information analyzed from the video data (numerical information generated by the analysis means 20 or the server 34) (see the information processing method described below for specific examples of these).
[0097] (Storage unit 38) The storage unit 38 is configured with, for example, an HDD, RAM, ROM, SSD, etc. Furthermore, the storage unit 38 is not limited to being built into the server 34, but may be a storage medium (for example, a USB memory) that can be detachably attached to the server 34. A plurality of storage units 38 may be provided.
[0098] The server 34 can receive various data and information from multiple AI cameras 10 and store them in the memory unit 38. The memory unit 38 of the server 34 then associates the work information with the sections of the workplace where each AI camera 10 is installed (database creation). The memory unit 38 may also associate work attendance management information, including work details received from the manager terminal i, with the work information, etc. Additionally, the memory unit 38 can also store various programs, trained models, training data, etc.
[0099] (Computation unit 40) The calculation unit 40 generates answer information corresponding to question information related to the work information. Specifically, the calculation unit 40 has a function of calculating and outputting answer information in response to question information from the manager terminal i and worker terminals a to c in accordance with the program according to the present disclosure. Furthermore, when outputting answer information, the calculation unit 40 may be caused by the RAG to search for information from a database (such as work information in the storage unit 38).
[0100] The calculation unit 40 can generate any of the following information as "answer information": text, graphs, reports, or video (including videos and images) explaining the status of the work information (e.g., Figures 6 and 7). This information can be generated in a predetermined format, for example, by providing instructions in JSON format. Note that the "video" here is different from video data captured by the AI camera 10; it refers to video for confirmation created based on the work information in the received question information (e.g., a short video or image). After reviewing this video, it is preferable to generate a report, graph, etc. If the administrator believes that final review should be done by the administrator himself or herself, or if reinforcement learning is enabled, it is particularly preferable to review the video. After reviewing the video, if there are no problems, the work information is saved, and any problematic work information, such as errors (false detection or misrecognition by each trained model), may be deleted. Once this review is complete, the information can be used as a reward for further reinforcement learning of the model.
[0101] The server 34 can also allow the administrator to search for work information analyzed from the video data (see FIG. 7). For example, the server 34 may send an output instruction signal via a predetermined program (such as an app) to display a search screen and search conditions on the terminal screen of the administrator terminal i or the like (see Information Processing Method 7). Then, when the calculation unit 40 receives question information as search conditions, it can generate answer information related to the work information included in the question information. For example, the search may be based on the work date and time, the section, the AI camera 10, the worker, the work content, the abnormality level, the abnormality content, and attendance management information. The administrator or the like can confirm such search results while viewing the video from the AI camera 10A or the like. Based on the search confirmation results as described above, the administrator or the like may choose to save or delete the work information.
[0102] The example shown in Figure 7 shows a terminal screen for searching and extracting work steps that are suspected to be abnormal by specifying the AI camera 10 (i.e., section, employee) and date and time, selecting the work content (step) using a specified word, or selecting the alert level (confidence level) and step elapsed time (upper and lower limit range) from a predetermined (all or part) trained model for that work step. Selection using a specified word means searching for verbalized work information, while selection by specifying a confidence level or range means searching for non-verbalized numerical information. Furthermore, a search using a high alert level presupposes trust in the confidence level of the trained model, and the number of search results can be adjusted by adjusting the left and right movement of the slide bar as shown. Furthermore, setting (selecting) a large upper limit for the step elapsed time means that delayed work can be searched for, and setting (selecting) a small lower limit for the step elapsed time means that work that is completed quickly can be searched for.
[0103] The calculation unit 40 preferably includes a learning model, such as a large-scale language model (LLM), an acoustic model, or a WaveNet model, that receives question information as input and answer information as output. As described above, task information is information that identifies the task date and time, the worker, the task content, etc., and can be verbalized or written down to be easily understandable by the user. The learning method of the calculation unit 40 is not particularly limited, as long as it can accept, for example, numerical information or task information that identifies the worker or the task content as "input (question information)" and generate a sentence explaining the corresponding task information (task content) as "output (answer information)." Examples of LLMs include BERT, XLNeT, T5, and GPT. Multimodal trained models are particularly preferred for the calculation unit 40. Fine-tuning can be performed on these models using task information. Furthermore, if an administrator terminal i or other device notifies the calculation unit 40 of an error in the task information, the model may be correctly labeled and re-trained.
[0104] The calculation unit 40 may also output voice information as a response. In this case, natural language processing, voice recognition, etc. may be performed to analyze the question, and then voice information of a response according to the question may be calculated and output. Other methods may also be employed.
[0105] The calculation unit 40 does not necessarily have to be provided within the management server 10, but may be included in the information processing system 1 as an external device of the management server 10. In one embodiment, a language model server may be provided in the information processing system 1 as the calculation unit 40, and may be communicatively connected to the management server 10 via a communication network.
[0106] (Communications Department 42) The communication unit 42 connects the control unit 36 to an external device (e.g., the AI camera 10, the worker terminal i, etc.) wirelessly or via a wired connection so that the control unit 36 can communicate with the external device. The control unit 36 can send and receive signals to and from the external device via the communication unit 42.
[0107] (others) In addition, the server 34 may be provided with an operation unit (input unit) such as a keyboard that allows an administrator or operator to give various commands to the control unit 36, and a display unit such as a monitor that displays various screens in response to a display command signal from the control unit 36. In one embodiment, a display operation unit such as a touch panel that integrates the operation unit and display unit may be used.
[0108] As described above, the server 34 can also determine whether to select or notify an image recognition model. For example, the server 34 receives, from the administrator terminal i, a video of work information to be input to an existing image recognition model for selection or notification in a simple, standard GUI format (e.g., FIG. 4 ). The video may include a video or an image. The server 34 then transmits to the administrator terminal i an output instruction signal to output a display screen (e.g., FIG. 4 ) for adding annotation information to image data in the video. In this way, upon receiving image data and annotation information that mark a division of work in the video, the server 34 can input the image data into an existing image recognition model and verify whether annotation information has been output from the image recognition model, thereby determining whether to select or notify.
[0109] The server 34 can also output the annotation information and image data or video input / accepted for the selection or notification decision as described above to explain the associated work information. That is, simultaneously with the selection of an image recognition model and the notification decision, the server 34 can output (e.g., print) the work information as a design image for a procedure manual or poster, as shown in FIGS. 4(f) and 5(e), using the processing (annotation information) up to FIGS. 4(e) and 5(d). It can also output a video (instructional video) containing image data and annotation information (not shown). The use of a GUI reduces the burden on the user. The contents of the procedure manual for Step 1, Step 2, etc. in FIG. 4(f) (omitted from the drawing) correspond to the categories set in FIG. 4(d). It is also possible to output information about prohibited acts and warnings, as shown in FIG. 5, and this information can also be used to generate anomaly detection models.
[0110] Traditionally, annotation has been performed solely for the purpose of training image recognition models, etc., and has been a specialized, technical task performed by specialized engineers. However, this annotation process is burdensome and costly. In contrast, the format shown in Figures 4(e) and 5(d) allows annotation to be performed in the same way as creating work instructions and notices, making it a natural and familiar task for managers without the burden or cost involved. Furthermore, the format allows annotation to be performed by managers, and only managers with a deep understanding of the objects being detected and the tasks being analyzed can accurately annotate key points in the analysis of the tasks. This reduces the burden of the work required to introduce novel equipment such as AI cameras on managers at small and medium-sized enterprises, thereby reducing the barriers to introducing AI cameras to the workplace.
[0111] Furthermore, by inputting image annotation data, which was previously created simply as data for training image recognition models, in the format of work procedure manuals and posters, it becomes possible to add context and meaning (analysis of how multiple objects work together) that was not possible with annotations from conventional "image recognition models," and it becomes possible to use annotation information to analyze work content.
[0112] Furthermore, conventional annotations were performed on individual objects in image recognition, and creating a program to combine these to analyze tasks required work such as identifying individual relationships and selecting analysis methods. In contrast, this disclosure allows annotation information, such as "the procedure for a certain task," to be formatted as a work procedure manual or notice, thereby creating a predetermined combination of objects and tasks (shapes and actions) (in a predetermined format). The predetermined format can be annotated with who, where, when (conditions), and what task (what) is being performed (e.g., Figures 4 and 5). In other words, annotations for the "image recognition model" required for the entire project at each workplace, as well as video segmentation annotations for the "time series analysis model" and "rule-based" required for generating and verbalizing task information, can be performed simultaneously, significantly reducing the labor and cost required for creating task analysis programs.
[0113] <Administrator terminal i... and worker terminals a-c...> The manager terminal i and the worker terminals a to c are terminals operated by the respective operators. The manager terminal i and the worker terminals a to c include industrial computers, personal computers, tablet terminals, mobile terminals, etc., but are not particularly limited as long as they are capable of inputting and outputting the above-mentioned information and transmitting and receiving warning signals, attendance management information, etc.
[0114] The administrator terminal i may be configured to be able to separately display the video data transmitted from the AI camera 10 without going through a GUI with the server 34.
[0115] [Information processing method 1] Next, the information processing method 1 in the information processing system 1 described above will be described with reference to FIG. 8, but the information processing method according to the present disclosure is not limited to this. In the following description, components with the same reference numerals are the same as the components described above, and duplicated descriptions will be omitted as appropriate. The processing described below is performed by the control unit 12 executing a program stored in the memory unit 30. It is assumed that video data of the section of the workplace where the worker works, acquired by the AI camera 10 (imaging unit 28), has already been acquired (the same applies to the example of the information processing method described below).
[0116] First, the control unit 12 (sensing means 16) detects the coordinates of the hand and the coordinates (numeric data) of the product from the video data (image data) using a skeleton detection model and an object detection model as image recognition models (step S10).
[0117] Next, the control unit 12 (analysis means 20) analyzes (rule-based) the coordinates of the hand and the coordinates of the product, and calculates the "distance between the coordinates" and the "positional relationship vector" (numerical information) (step S20).
[0118] Furthermore, the control unit 12 (analysis means 20) quantifies the "probability of being in work" and the amount of work from the "distance between coordinates" and "vector of positional relationship" (numerical information) calculated in step S20 using a "trained regression model" that has learned the "time-series work pattern" using the "coordinate distance," "positional relationship vector," and "label of correct work" as training data (i.e., numerical information is obtained) (step S30). Note that in this specification, there is no limit to the number of operations on numerical data to calculate numerical information, and there is also no limit to the number of operations on numerical information.
[0119] Next, the control unit 12 (transmission means 24) transmits the "probability of being in work" and the amount of work (work information) of up to several tens of bytes together with an output instruction signal to the server 34 so that the work information ("probability of being in work") is output in association with the work area section (step S40).
[0120] According to information processing method 1, the information transmitted is numerical information, so the communication charges between the AI camera 10 and the server 34 can be reduced from several MB of data to approximately one hundred thousandth of the amount required when transmitting video data. For example, even if information is sent from 100 cameras, if there is a worker who needs attention, such as someone who is dozing off while working, and who is unlikely to be there, it is possible to identify and notify the manager.
[0121] [Information processing method 2] Next, information processing method 2 in the above-described information processing system 1 will be described with reference to FIG. 9. Note that explanations that overlap with information processing method 1 will be omitted as appropriate (same below). Information processing method 2 is an example in which a machine learning classification model in the AI camera 10 device is used to confirm hand washing steps as an example of a manual task and check whether the hand washing was done thoroughly. Here, it is assumed that images of each hand washing step have been machine-learned using the classification model.
[0122] First, the hand washing video data is input to an image recognition model (step S110). This image recognition model detects the reliability (numerical data) of the class corresponding to the hand shape at each point in time.
[0123] Next, the control unit 12 (analysis means 20) uses the trained classification model to receive the coordinates of the hand shape as input, classify this numerical data into which step it is among all the hand washing steps, and generate a class and a reliability (step S120).
[0124] Next, the control unit 12 (comparison means 22) compares the time (numerical information) of each classified procedure with a threshold value to determine whether it has been performed for a sufficient number of seconds (step S130).
[0125] Next, the control unit 12 (transmission means 24) transmits the comparison result (evaluation information) to the server 34 (step S140).
[0126] In this information processing method 2, as in information processing method 1, the amount of communication traffic can be reduced by transmitting numerical information. Furthermore, the administrator can identify the work content, date and time, location, etc. of any work that may have been performed that is unsanitary, based on the database in server 34 (see FIG. 7).
[0127] [Information processing method 3] Next, information processing method 3 in the above-described information processing system 1 will be described with reference to Fig. 10. Information processing method 3 is an example in which classification and anomaly detection are performed using a machine learning model within the AI camera 10 device.
[0128] First, the control unit 12 (sensing means 16) detects "each coordinate of the whole body" using the skeleton detection model and "each coordinate of the product" using the object detection model from the video data (step S210).
[0129] Next, the control unit 12 (analysis means 20) calculates the "joint angles" of the whole body and the "distance" (numerical information) between the product and the worker from the numerical data such as "each coordinate of the whole body" and "each coordinate of the product" (step S220).
[0130] Next, the control unit 12 (analysis means 20) classifies the work content from each image data based on the "joint angle" and "distance" using the trained classification model (step S230).
[0131] Next, the control unit 12 (analysis means 20) uses the trained anomaly detection model to determine whether or not an abnormality exists for each classified work content, using the "whole body coordinates" and "product coordinates" (polygon) as input (step S240). This anomaly detection model is an example, and is a model that detects abnormal behavior such as part-time worker terrorism by detecting any action such as putting something into a finished product when no product has been detected as an abnormality.
[0132] Next, the control unit 12 (transmission means 24) transmits the results of the trained anomaly detection model (presence or absence of anomaly; categorical information) to the server 34 (step S250). If an anomaly is detected, the AI camera 10 simultaneously generates an alarm to alert the worker (step S260).
[0133] Like information processing method 1, information processing method 3 can also reduce communication volume by transmitting numerical information. Furthermore, even if information is sent from 100 cameras, if there is an abnormality such as a byte terrorism that mixes garbage, the administrator can be notified.
[0134] [Information processing method 4] Next, information processing method 4 in the above-described information processing system 1 will be described with reference to Fig. 11. Information processing method 4 is an example in which classification (time-series segmentation, work progress analysis) and object detection (hazard monitoring) using a machine learning model are performed simultaneously within the AI camera 10 device.
[0135] <Work progress analysis> First, the control unit 12 (sensing means 16) detects the coordinates of the hand using the skeleton detection model and the coordinates of multiple parts using the object detection model from the video data (step S310).
[0136] Next, the control unit 12 (analysis means 20) calculates the "joint angles" of the hand and the "distances" between the parts and each finger (numerical information) from these numerical data (step S320).
[0137] Next, the control unit 12 (analysis means 20) classifies the work content by work progress based on the "joint angles" and "distances" using the learned classification model (step S330). Note that important features of each step of the work are learned from work procedure manuals, etc., and the achievement level and elapsed time for each step can be used to determine progress.
[0138] Next, the control unit 12 (analysis means 20) calculates how much time each task took (step S340; see FIGS. 4(d) and 4(e)).
[0139] Next, the control unit 12 (transmission means 24) transmits the work progress information to the server 34 (step S350).
[0140] <Danger monitoring> Simultaneously with the above <Work progress analysis>, the control unit 12 (sensing means 16) detects the coordinates of the whole body using the skeleton detection model and the coordinates of the object (a forklift different from the above part) using the object detection model from the video data (step S410).
[0141] Next, the control unit 12 (analysis means 20) calculates the "distance" and "vector" between the coordinates of the whole body and the coordinates of the forklift from the numerical data (rule-based) (step S420).
[0142] Next, the control unit 12 (analysis means 20) calculates the probability of a collision between a person and a forklift from the "distance" and the "vector" using a trained regression model (step S430).
[0143] Next, the control unit 12 (transmission means 24) transmits the "distance," "vector," and collision probability to the server 34 (step S440). If the collision probability is higher than a predetermined threshold, the AI camera 10 issues an audio warning (step S450).
[0144] In this type of information processing method 4, if two AI applications (work progress analysis and hazard monitoring) are constantly running within the AI camera 10 device, the burden on the AI camera 10 may be so great that it may not be able to perform calculations. However, by using the entire information processing system 1 (AI camera 10 and server 34) and processing work progress analysis and hazard monitoring independently and alternately within the AI camera 10 as separate APIs, multiple monitoring functions can be realized even on an edge device (AI camera 10) that has many limitations in memory and computing units.
[0145] [Information processing method 5] Next, information processing method 5 in the above-described information processing system 1 will be described with reference to Fig. 12. Information processing method 5 is another example of analyzing work progress.
[0146] First, the control unit 12 (sensing means 16) detects the face of the worker (step S510). In this way, the individual worker is identified, and the subsequent analysis is made to address individual variations. Here, the individual may be identified using an image recognition model or the sensor unit 29 (for example, a fingerprint sensor) based on a marker attached to the worker, an IC card, or attendance management information.
[0147] Next, the control unit 12 (sensing means 16) detects the "coordinates of the whole body and fingers" from the video data using a skeleton detection model (step S520).
[0148] Furthermore, the control unit 12 (sensing means 16) detects the "coordinates of parts and products" including tools used in the object detection model (step S530).
[0149] Next, the control unit 12 (analysis means 20) digitizes the position, shape, and angle of the hands and body during work from the "whole body and finger coordinates" to generate "skeletal coordinates" (numeric data) (step S540).
[0150] Furthermore, the control unit 12 (analysis means 20) quantifies the degree of completion, class, and position of the product from the "coordinates of the parts and products" using the trained regression model (step S550).
[0151] Next, the control unit 12 (analysis means 20) clusters the skeleton coordinates (numerical data) generated in step S540 using MLP (step S560).
[0152] Next, the control unit 12 (analysis means 20) performs inference and classification of the numerical information generated in steps S550 to S560 in time series using the trained classification model, and outputs the class and the confidence level (step S570).
[0153] Next, the control unit 12 (comparison means 22) evaluates the confidence level of each class in step S570 by comparing it with a threshold value for each individual identified in step S510 (step S580).
[0154] Next, if the confidence level is equal to or greater than the threshold value in step S580, the control unit 12 (transmission means 24) transmits information that "the work is completed" to the server 34 (step S590).
[0155] [Information processing method 6] Next, information processing method 6 in the above-mentioned information processing system 1 will be described with reference to Fig. 13. Information processing method 6 is an example of analyzing work accuracy. In this example, tool coordinates (auxiliary numerical data) are detected by sensor unit 29 (inertial sensor), numerical information is generated by AI camera 10, and work accuracy (work information) is generated by server 34.
[0156] First, the control unit 12 (sensing means 16) detects a person (step S610). In this way, variations among individuals can be addressed in the subsequent analysis, but it is sufficient to identify the person, and the person may be detected by an image recognition model or the sensor unit 29 (thermography, etc.).
[0157] Next, the control unit 12 (sensing means 16) detects the "coordinates of the whole body and fingers" from the video data using a skeleton detection model (step S620).
[0158] Furthermore, the control unit 12 (sensing means 16) detects the "coordinates of the parts and products" using the object detection model (step S630).
[0159] Next, the control unit 12 (analysis means 20) digitizes the position, shape, and angle of the hands and body during work from the "whole body and finger coordinates" to generate "skeletal coordinates" (numeric data) (step S640).
[0160] Furthermore, the control unit 12 (analysis means 20) quantifies the degree of completion, class, and position of the product from the "coordinates of the parts and products" (step S650).
[0161] In addition, the control unit 12 (analysis means 20) uses the learned regression model to analyze the coordinates (auxiliary numerical data) of the tool detected by the sensor unit 29 and detects the position (numerical information) of the tool at each time (step S660).
[0162] Next, the control unit 12 (analysis means 20) uses the trained classification model to classify the numerical information generated in steps S640 to S650, and outputs the class and the confidence level (step S670).
[0163] Next, the control unit 12 (transmission means 24) transmits the personal information detected in step S610 and the numerical information obtained in steps S660 and S670 to the server 34 (step S680).
[0164] Next, the control unit 36 (server 34) quantifies the deviation of the numerical information generated in steps S640 to S660 from the values during normal work based on the information on the individual identified in step S610 and the information on the class classified in step S670, and generates the work accuracy for the work (step S700). The classified classes are matched on a rule basis.
[0165] Next, the control unit 36 (server 34) transmits information "Work completed: work accuracy XX%" in response to the question information from the terminal (step S710).
[0166] [Information Processing Method 7] Next, information processing method 7 in the above-described information processing system 1 will be described with reference to Figures 6, 7, 14, and 15. Information processing method 7 is an example in which work information performed after the information processing methods exemplified above is presented to a manager or worker for retrieval, etc. Here, it is assumed that a PC application is launched on a terminal to access server 34.
[0167] <Server initial settings> When the PC application is launched on the terminal, in the initial setting, the server 34 prompts the terminal to enter the number and location of the AI cameras 10, the SSID, the password, and select the start button, and then prompts the terminal to set up the camera connection (this setting information is saved and will be auto-filled from the next time).
[0168] <Camera connection settings> Based on the setting information, the server 34 displays the locations of the AI cameras 10 and workers in the workplace where the terminal belongs on a map (Fig. 6). Based on the user's settings, the icons of workers may be color-coded depending on whether they require confirmation (warning, etc.).
[0169] The server 34 allows the user to select and set an image recognition model (AI application). The user selects the icon of the AI camera 10 at the position they want to check, and displays available AI applications. Once the AI application is selected, the PC application setup is completed.
[0170] The server 34 can also adjust the AI for each worker. When the "AI Adjustment" button in the lower right of Figure 6 is selected, an order can be placed for fine-tuning of the image recognition model (including the trained model for analysis). AI adjustment for each worker may also include the generation of the custom-made image recognition model described above, as well as adjustment of parameters and thresholds for numerical information analysis.
[0171] In addition, the AI can be adjusted for each task (Fig. 14). As shown in Fig. 14, parameters and thresholds can be adjusted while viewing a graph showing numerical information (Fig. 14(a)) or video of video data (Fig. 14(b)). (Note that this function may be the same as the conditional search GUI shown in Fig. 7.)
[0172] Once the setup is complete, the server 34 automatically establishes a wireless connection with the AI camera 10 via a pre-registered IP address or Bluetooth. If the automatic connection is unstable, the server 34 may establish a connection by having the user read a QR code (registered trademark) displayed on the terminal screen with the AI camera 10.
[0173] <Work management support> The server 34 can use a PC app to check the status of the entire workplace hourly based on work information. As shown in the lower part of Figure 6, the time selection slider can be set to snap (automatically stop) during periods requiring attention when abnormal numerical information is detected. This eliminates the need to concentrate on checking the entire period (preventing oversight). For example, Figure 6 shows the log information indicating "13:59, camera ID: 10A (section) detected an error in the work procedure by a specific worker." The icon of this worker is also highlighted in section A on the map screen. From a work management perspective, it is preferable to build a database that associates attendance management information with work information so that each position on the time selection slider indicates "what the employee was doing at that time and in that location." It is also preferable to be able to check work information using video and text. As described above, the server 34 (the calculation unit 40) can automatically generate reports in a specified format (supporting the administrator).
[0174] The server 34 may allow the user to select a worker on the map screen whose work status the user wishes to view. Selecting a worker icon preferably displays the worker's work history and video data, such as "what time, where, and what the worker was doing." Information from all AI cameras 10 may also be aggregated on a single map, or information from all AI cameras 10 may be displayed in a list.
[0175] <Digital Twin> The server 34 may display the work information acquired by the AI camera 10 in real time (while the work is being done) as a digital twin. The icon display of the worker can be different depending on whether they are working or slacking. According to the present disclosure, the amount of communication traffic is reduced, so the AI's assessment results can also be reflected in real time. For example, a weekly work attitude score can be calculated in real time.
[0176] <Conditional search> The server 34 (calculation unit 40) can display only the images that should be viewed from the video data based on the judgment of the natural language processing AI, thereby reducing the burden of checking on the administrator. The server 34 (calculation unit 40) preferably allows the user to search by conditions (FIG. 7). Examples of search conditions include the AI camera 10, selection of the date and time to check, "threshold and range specification" related to the AI camera 10's judgment results, area, image recognition model, and attendance management information. It is also preferable to allow the user to adjust the conditions while checking the video data and work information as shown in FIG. 14. At this time, the user may be prompted to fine-tune the image recognition model.
[0177] <Providing information to workers> According to the present disclosure, in order to reduce the management burden, numerical information such as shape information and motion information is generated from numerical data such as coordinate data and is used to evaluate and determine work information. Such information can also be applied to work support, training, and self-improvement for workers.
[0178] Figure 15 shows an example of determining hand washing as one of the tasks. Hand washing can also be classified into multiple actions and steps, and the similarity to a model (correct answer data) for each step can be calculated and evaluated.
[0179] In the program, computer (AI camera 10), information processing system 1, and information processing method according to the present disclosure configured as described above, the AI camera 10 is provided to acquire video data of a section of a workplace where a worker performs work, and the AI camera 10 functions as sensing means 16, analysis means 20, and transmission means 24. The sensing means 16 analyzes the video data using an image recognition model to generate numerical data corresponding to the shape or movement during work. The analysis means 20 analyzes the numerical data to generate work information during work. The transmission means 24 transmits an output instruction signal to the server 34 so that the work information is output in association with the section of the workplace. The server 34 associates the work information with the section of the workplace.
[0180] As described above, the program, computer (AI camera 10), information processing system 1, and information processing method disclosed herein aggregate and output work information necessary for work management, rather than video data. This reduces the burden on managers and reduces communication traffic. In particular, even if the information processing system 1 has hundreds of AI cameras 10, the server 34 can communicate with these AI cameras 10 with minimal communication traffic, significantly benefiting from the aggregation of work information and reduced communication traffic. Furthermore, it reduces the need for reviewing surveillance camera (AI camera 10) footage, patrolling the site, re-educating employees on important points, and ensuring compliance with rules (considering training methods, devising communication methods with employees, reporting to superiors, meetings, and sample preparation). Because managers are essentially key workers on the front lines, performing the tasks of machine engineering and production management, this disclosure frees managers from these tasks, leading not only to labor savings but also to improved machine utilization rates, more efficient production activities, and overall factory productivity improvements.
[0181] Furthermore, in the program, computer (AI camera 10), information processing system 1, and information processing method according to the present disclosure, the AI camera 10 may include a sensor unit 29 that generates auxiliary numerical data corresponding to the shape or movement of the area during work, and the analysis means 20 analyzes the numerical data and auxiliary numerical data to generate work information during work. In this way, it is possible to generate work information that cannot be analyzed using only the numerical data obtained by the imaging unit 28, or to generate auxiliary numerical data that complements the numerical data, thereby improving the accuracy of the work information. For example, while it is possible to detect drowsiness as an abnormal shape or movement based on the work information from the imaging unit 28 (numerical information and its original numerical data), it is also possible to detect signs of drowsiness or a cold by also using body temperature information from a thermograph (sensor unit 29) in the analysis, as described above.
[0182] Furthermore, in the program, computer (AI camera 10), information processing system 1, and information processing method according to the present disclosure, the analysis means 20 can analyze the generated numerical data and generate, as work information, numerical information corresponding to the shape or movement during work. If the numerical information is generated as meaningful information in this way, it can be easily converted into words (work information).
[0183] Furthermore, in the program, computer (AI camera 10), information processing system 1, and information processing method according to the present disclosure, the AI camera 10 may further function as a comparison means 22, which generates evaluation information from the comparison result obtained by comparing the generated numerical information with a predetermined threshold, and the transmission means 24 transmits the evaluation information generated by the comparison means 22 to the server 34. Alternatively, the server 34 may generate the evaluation information from the comparison result obtained by comparing the transmitted numerical information with a predetermined threshold. Generating the evaluation information as work information in this way also makes it more efficient to search for potentially problematic work information, leading to a reduction in the burden of work management.
[0184] In addition, in the program, computer (AI camera 10), information processing system 1, and information processing method according to the present disclosure, the analysis means 20 or server 34 may use a trained regression model to analyze numerical information to generate task information indicating the probability that a specific task is being performed and the amount of task. This allows for fuzzy numerical values to be used to obtain ambiguous judgments, enabling flexible responses. This also allows for situations where the AI camera 10 is required to detect the "amount of operation."
[0185] Furthermore, in the program, computer (AI camera 10), information processing system 1, and information processing method according to the present disclosure, the analysis means 20 or server 34 may use a trained classification model to analyze numerical information to generate, as task information, a classification result indicating whether a specific task is being performed. Utilizing such a classification result makes it easier to understand and to verbalize.
[0186] In addition, in the program, computer (AI camera 10), information processing system 1, and information processing method according to the present disclosure, the analysis means 20 or server 34 may use a trained anomaly detection model to analyze numerical information to generate work information indicating the presence or absence of an anomaly, thereby reducing the work management burden.
[0187] In addition, in the program, computer (AI camera 10), information processing system 1, and information processing method according to the present disclosure, the numerical information can be information for explaining the work content numerically. By using such meaningful numerical information, it becomes easier to verbalize and evaluate the corresponding work information.
[0188] Furthermore, in the program, computer (AI camera 10), information processing system 1, and information processing method according to the present disclosure, the sensing means 16 may select an existing image recognition model, including a general-purpose image recognition model, as the image recognition model during initial configuration before analyzing the video data, or may notify the need for a new image recognition model. In this way, an existing image recognition model may be reused, or a custom-made image recognition model may be generated for highly accurate image recognition.
[0189] In addition, in the program, computer (AI camera 10), information processing system 1, and information processing method according to the present disclosure, the server 34 may receive video of work information to be input to an existing image recognition model, output a display screen for adding annotation information to image data in the video, and input the image data into the existing image recognition model to determine whether annotation information has been output, thereby determining whether to select or notify. The server 34 may also output the annotation information and image data or video to explain the work information. In this way, the video and annotation information used to select an existing image recognition model or generate a new image recognition model can be used to create work procedures, notices, and the like normally required in workplaces, thereby reducing the workload of managers.
[0190] Furthermore, in the program, computer (AI camera 10), information processing system 1, and information processing method according to the present disclosure, the AI camera 10 may further function as a reception means 14 and a model setting means 18. When the reception means 14 receives attendance management information including work content, the model setting means 18 sets a trained model according to the work content. The server 34 may also receive attendance management information including work content and set a trained model according to the work content. By taking attendance management information into account in this way, more accurate work information can be generated.
[0191] Furthermore, in the program, computer (AI camera 10), information processing system 1, and information processing method according to the present disclosure, the work information can be information that identifies the work date and time, the individual worker, the area, the coordinates, or the work content. Such work information facilitates work management and can be easily associated with attendance management information.
[0192] In addition, in the program, computer (AI camera 10), information processing system 1, and information processing method according to the present disclosure, the server 34 can receive attendance management information including work content and associate it with the work information. In this way, in combination with the attendance management information, more accurate work management can be achieved.
[0193] Furthermore, in the program, computer (AI camera 10), information processing system 1, and information processing method according to the present disclosure, the server 34 can calculate work progress and work accuracy as work information based on numerical information analyzed from video data. Providing work progress and work accuracy leads to more efficient work management.
[0194] Furthermore, in the program, computer (AI camera 10), information processing system 1, and information processing method according to the present disclosure, the server 34 can allow the administrator to search for work information analyzed from video data. For example, conditional searches can facilitate work management. In particular, if work information can be searched using more conditions, the video data can be narrowed down more precisely, further facilitating work management.
[0195] In the program, computer (AI camera 10), information processing system 1, and information processing method according to the present disclosure, the server 34 may include a calculation unit 40 that generates answer information corresponding to question information about the work information, and the calculation unit 40 can generate any of text, graphs, reports, and images that explain the status of the work information as the answer information. The use of an AI agent enables rapid work management support and a wide variety of support.
[0196] Furthermore, in the program, computer (AI camera 10), information processing system 1, and information processing method according to the present disclosure, the transmission means 24 transmits the acquired video data to a manager terminal i operated by a workplace manager, and the server 34 can also identify video data corresponding to question information about work information from the manager terminal i. By viewing the search results from the server 34 in this way (FIG. 14), the convenience of work management can be improved when the manager actually checks the video data via the manager terminal i.
[0197] The program, computer (AI camera 10), information processing system 1, and information processing method according to the present disclosure are not limited to the above-described aspects and combinations, and various modifications can be made.
[0198] The above describes an example in which the AI camera 10 converts numerical data into numerical information or work information, but the numerical data may also be sent from the AI camera 10 to the server 34, and the server 34 may convert the numerical data into numerical information or work information. [Explanation of symbols]
[0199] 1. Information Processing Systems 10(10a, 10b, 10c…) AI Camera (Computer) 12 Control Unit 14. Reception methods 16 Sensing Methods 18 Model Setting Method 20 Analysis methods 22 Means of comparison 24 Transmission Method 28 Imaging unit 29 Sensor section 30 Storage section 32 Communications Department 34 servers 36 Control Unit 38 Memory section 40 Arithmetic section 42 Communications Department i, ... Administrator terminal a, b, c... Worker terminal A, B, C, ... Sections X, Y, … Workshop
Claims
1. A program that causes an AI camera to function as a sensing means, an analysis means, and a transmission means, The AI camera acquires image data of a section of the work area where the worker performs work, the sensing means analyzes the video data using an image recognition model to generate numerical data corresponding to the shape or movement during the work; the analysis means analyzes the numerical data on a rule basis or using a trained model to generate numerical information corresponding to a shape or movement during work as work information during work; The transmitting means transmits an output instruction signal to a server so that the work information is output in association with the section of the work area.
2. The AI camera includes a sensor unit, The sensor unit generates auxiliary numerical data corresponding to the shape or movement of the section during work; 2. The program according to claim 1, wherein said analyzing means analyzes said numerical data and said auxiliary numerical data to generate work-in-progress information.
3. The AI camera further functions as a comparison means, the comparison means or the server compares the generated numerical information with a predetermined threshold value and generates evaluation information from the comparison result; The program according to claim 1 , wherein the transmitting means transmits the evaluation information generated by the comparing means to the server.
4. 2. The program according to claim 1, wherein the analysis means or the server uses a trained regression model to analyze the numerical information to generate the probability that a specified task is being performed and the amount of that task as task information.
5. The program according to claim 1 , wherein the analysis means or the server uses a trained classification model to analyze the numerical information and generate, as work information, a classification result indicating whether a specified work content is being performed.
6. The program according to claim 1 , wherein the analysis means or the server generates the presence or absence of an abnormality as work information by analyzing the numerical information using a learned anomaly detection model.
7. The program according to claim 1 , wherein the numerical information is information for explaining the work content in numerical terms.
8. 2. The program according to claim 1, wherein the sensing means notifies the user of the need for a new image recognition model or a selection of an existing image recognition model including a general-purpose image recognition model as the image recognition model during initial configuration before analysis of the video data.
9. The server receiving an image relating to work information to be input to the existing image recognition model for the selection or notification; outputting a display screen for adding annotation information to image data in the video; inputting the image data into the existing image recognition model, and verifying whether the annotation information is output, thereby determining whether to select or notify; 9. The program according to claim 8, wherein the annotation information and the image data or video are output for explaining the work information.
10. The AI camera further functions as a reception means and a model setting means, When the reception means or the server receives attendance management information including work content, The program according to any one of claims 4 to 6, wherein the model setting means or the server sets the trained model according to the work content.
11. 5. The program according to claim 1, wherein the work information is information that identifies any one of the work date and time, the individual worker, the section, the coordinates, and the work content.
12. The program according to claim 11 , wherein the server receives attendance management information including work content and associates the information with the work information.
13. The program according to claim 1 , wherein the server calculates work progress as work information based on numerical information analyzed from video data.
14. The program according to claim 1 , wherein the server calculates work accuracy as work information based on numerical information analyzed from the video data.
15. the server has a calculation unit that generates answer information corresponding to question information related to work information, The program according to claim 11 , wherein the calculation unit generates, as the response information, any one of text, graph, report, and video that explains the status of the work information.
16. The program according to claim 11 , wherein the server causes an administrator to search for work information analyzed from video data.
17. The transmission means transmits the acquired video data to a manager terminal operated by a manager of the workplace, The program according to claim 16 , wherein the server identifies video data corresponding to question information related to work information from the manager terminal.
18. A computer that functions as a sensing means, an analyzing means, and a transmitting means by executing a program, The computer acquires image data of a section of a work area where a worker performs work, the sensing means analyzes the video data using an image recognition model to generate numerical data corresponding to the shape or movement during the work; the analysis means analyzes the numerical data on a rule basis or using a trained model to generate numerical information corresponding to a shape or movement during work as work information during work; The transmitting means transmits an output instruction signal to a server so that the work information is output in association with the section of the work area.
19. An AI camera that acquires video data of the work area where the worker works; a server that associates the work area sections with work information; A system comprising: The AI camera functions as a sensing means, an analyzing means, and a transmitting means by executing a program, The AI camera acquires image data of the section of the workplace where the worker performs work, the sensing means analyzes the video data using an image recognition model to generate numerical data corresponding to the shape or movement during the work; the analysis means analyzes the numerical data on a rule basis or using a trained model to generate numerical information corresponding to a shape or movement during work as work information during work; The transmission means transmits an output instruction signal to the server so that the work information is output in association with the section of the workplace.
20. An information processing method executed by a computer having a control unit, a step in which the control unit uses an image recognition model to analyze video data of a section of a work area where a worker performs work, thereby generating numerical data corresponding to the shape or movement during work; The control unit analyzes the numerical data using a rule-based or learned model to generate numerical information corresponding to a shape or movement during work as work information during work; a step in which the control unit transmits an output instruction signal to a server so that the work information is output in association with a section of the work site; An information processing method comprising:
Citation Information
Patent Citations
Step change machine-learned model switching system, edge device, step change machine-learned model switching method, and program
WO2019229983A1
Monitoring camera system, apparatus and method for controlling monitoring camera system
JP2006086991A